
OpenRouter's latest weekly report puts the Chinese model market share in plain numbers. For the week of September 7–13, total global LLM traffic on the API aggregation platform reached 127 trillion tokens: Chinese models accounted for roughly 61.2 trillion, up 7.9% week on week, while US models logged 21.8 trillion, up 31.6%. It is the 20th consecutive week that Chinese models have led the United States in weekly call volume.
Chinese models hold at least seven of the top ten spots, a combined share above 70% — against just two a year ago. Tencent's Hunyuan Hy4 preview ranked second at 16.8 trillion tokens, Z.AI's GLM 5.3 Flash third at 11.9 trillion, DeepSeek's V4 Flash fourth at 11.6 trillion. DeepSeek's V4.1 Flash, released September 10, reached sixth place with 4.94 trillion tokens in three days; OpenRouter says the model processed about one trillion tokens within 24 hours of release, with roughly 90% of traffic coming from cache reads at a cost of about $0.006 per million tokens on market pricing — some 80% lower than comparable models such as GLM 5.3 Flash.
The "price butcher" label is not the whole explanation. Xiong Wei, an internet industry analyst at UBS Securities China, told Yicai that as AI scales into real business processes, attention is shifting from capability to return on investment, with "Token ROI" becoming a key term in the second half of 2026. The first half was about "token-maxxing" — big companies pushing employees to use as much AI as possible — until bills ballooned and value became hard to measure. From mid-year, enterprises moved to "token optimization," holding AI spending to a stricter ROI standard. A developer of the open-source project xiaobei told Yicai he switched his entire workflow to domestic models in May: "Chinese models can already handle most application-layer tasks; there is no need to pay tens of times the cost for top flagship models."
The capability leaderboard tells a different story. Per Artificial Analysis, the top five on its intelligence index are all overseas models: Anthropic's Claude Fable 5.1 first, OpenAI's GPT-6 Astra second, Meta's Muse Spark 1.3 fourth, xAI's Grok 4.6 eighth; the best-placed Chinese model is Z.AI's GLM-5.3 at seventh with a score of 45. Chinese models are not ahead on capability — they are winning traffic in a market where developers vote with their workflows.
Cost structure supports the migration. A UBS survey found Chinese model vendors' API gross margins mostly in the 20–40% range, meaning they are not subsidizing adoption but sustaining cost-effectiveness through architecture innovation and training and inference efficiency. OpenRouter's aggregate numbers amplify the trend: weekly traffic grew from 4.92 trillion tokens in September 2025 to 127 trillion in September 2026, roughly a 25-fold increase in a year, with DeepSeek first at 25.4% share and Chinese vendors above 45% against about 43% for US vendors.
Scope should be stated: OpenRouter counts only traffic routed through its platform. Vendor-owned APIs, cloud platforms, private deployments, and direct-to-consumer apps fall outside it, so the numbers are not global market share. The direction, though, is hard to dispute — low price is not a moat and Chinese vendors cannot yet claim victory; but in a market increasingly obsessed with input-output ratios, they have taken the seats, and are competing for the default entry point.
[1][2]