China Merchants Bank Cuts AI Inference Costs 60% on Kubernetes
Shanghai – September 10, 2026 -- China Merchants Bank has won the Cloud Native Computing Foundation's (CNCF) End User Case Study Contest after lifting average accelerator utilization from 35% to more than 60% and cutting inference cost per 1 million tokens by over 60%.
Unified control plane pools nearly 10,000 accelerator cards
The bank's AI infrastructure team built a single Kubernetes control plane combining Kueue, KEDA, Prometheus, HAMi, and Fluid, allowing training, fine-tuning, and online inference workloads to share close to 10,000 heterogeneous accelerator cards. Working with its data center and other internal teams, the bank brought 99% of its accelerator compute under this framework.