China’s AI Infrastructure Surge: How Rapid Model Upgrades Are Driving Cloud Capex
According to a Bank of America Merrill Lynch monthly monitoring report summarized by 36 Kr, August delivered simultaneous model upgrades from Tencent, Alibaba, DeepSeek, and Zhipu AI, tightening…
Owen Garfield·updated September 02, 2026

China's AI cloud buildout just got expensive. According to a Bank of America Merrill Lynch monthly monitoring report summarized by 36 Kr, August delivered simultaneous model upgrades from Tencent, Alibaba, DeepSeek, and Zhipu AI, tightening compute demand, and a sharp rise in cloud capex — a combination the bank frames as the industry accelerating from model competition toward commercial monetization. For anyone running inference at scale, the relevant question isn't who's "winning" the leaderboard. It's whether the compute pipeline behind those usage spikes holds up under your SLA.
Capex as the leading indicator
The numbers I'd watch first aren't the benchmark scores. They're the capex lines. Tencent and Alibaba reported Q2 capex of RMB 530 billion and RMB 680 billion respectively, with cloud revenue growth of just over 20% YoY for Tencent and 45% for Alibaba. BofA estimates the top four Chinese cloud platforms will grow total capex 98% YoY in 2026 and 44% in 2027. Over the past 30 days, consensus expectations for Tencent and Alibaba's next-12-month capex were revised up 25% and 29%.
That's not "AI enthusiasm." That's a fleet purchase. Someone is building out GPU clusters and committing balance sheet to do it. The market is repricing the upside.
Throughput beats benchmarks
Leaderboards are theater until you measure tokens served. OpenRouter data as of August 17 puts DeepSeek V4 Flash at the top with 31.6 trillion accumulated tokens for the month, Tencent Hy3 at 26.2 trillion, and Xiaomi MiMo-V2.5 at 19.1 trillion. DeepSeek's weekly token usage share hit 22%, up 4.7 percentage points from July. On Vercel AI, DeepSeek's average monthly token share rose to 29.7% from 26.0%, with Anthropic second at 24.7%.
On capability, Artificial Analysis has Claude Opus 5 first on the intelligence index at 63, followed by Claude Fable 5 (62) and GPT-5.6 Sol (61). Moonshot's Kimi K3 sits fifth at 60 — 95% of Opus 5's score — and Zhipu's GLM-5.3 is sixth, also at 60. On the Agentic Index, GLM-5.3 and Claude Opus 5 tie at 59; Qwen 3.8 Max is fifth at 58. In Code Arena WebDev, Opus 5 leads at 1691, with Kimi K3 at 1674 and Qwen 3.8 Max at 1669; Tencent Hy4 preview is sixth at 1633.
The capability gap is closing fast. The gap that actually matters — cost-per-token at production scale — is wider, and some vendors are raising token pricing during this same upgrade wave. Read the price sheets as carefully as the leaderboards.
Deploy or discard
The filter I use: p99 latency against your SLA, token cost against your throughput target, and whether the provider's inference fleet can absorb a 20%+ monthly token jump without OOM'ing your batch jobs. The OpenRouter and Vercel numbers say DeepSeek's serving stack is already absorbing pressure at scale. That's a stronger signal than any benchmark.
Near-term: test DeepSeek V4 Flash and the Qwen 3.8 line against your current default on real workload traces, not synthetic eval sets. Longer-term: BofA's capex trajectory implies a 2026–2027 window where GPU supply loosens at the margin. Plan capacity accordingly. Don't sign multi-year commits yet — technical debt from a locked-in inference provider is the kind of bottleneck you don't notice until your migration window has already closed.