Shanghai – September 10, 2026 -- China Merchants Bank has won the Cloud Native Computing Foundation's (CNCF) End User Case Study Contest after lifting average accelerator utilization from 35% to more than 60% and cutting inference cost per 1 million tokens by over 60%.
Unified control plane pools nearly 10,000 accelerator cards
The bank's AI infrastructure team built a single Kubernetes control plane combining Kueue, KEDA, Prometheus, HAMi, and Fluid, allowing training, fine-tuning, and online inference workloads to share close to 10,000 heterogeneous accelerator cards. Working with its data center and other internal teams, the bank brought 99% of its accelerator compute under this framework.
Efficiency gains hold under comparable model and service conditions
Under matched model and service conditions, the unified architecture cut the cost of processing 1 million combined input-and-output tokens by more than 60%. Kueue manages training admission and quotas to prevent idle capacity reservation, KEDA and Prometheus scale inference from live traffic signals, HAMi allocates shared accelerator capacity in fine-grained units, and Fluid speeds access to datasets, model weights, and checkpoints.
Twinkle framework lifts fine-tuning density fivefold
China Merchants Bank's in-house Twinkle training framework lets five LoRA tenants share one base-model instance by default. Reducing base-model replicas from five to one cut accelerator resource usage for that setup by 80% while increasing training density fivefold.
CNCF cites financial-sector infrastructure as a scalability model
"China Merchants Bank's approach to unifying AI training and inference is a clear example of how composable, vendor neutral infrastructure powered by CNCF projects delivers measurable efficiency at scale,