On September 22, the daily ranking on OpenRouter, the global API aggregation platform, showed Zhipu AI's GLM-5.3-Flash recording 4.36 trillion tokens in call volume for the day, up 73% from the previous period and ranking first among models on the platform, ahead of DeepSeek V4.1 Flash (2.43T) and Tencent's Hy4 preview (1.88T). The numbers come from real-time aggregation of developer request traffic, not self-reported benchmark scores.
[1][2]The nut graf: OpenRouter's daily chart is a real-time thermometer of developers voting with their keyboards — 4.36 trillion tokens says not just that the model is good, but that the price-capability combination of an open Chinese model has claimed a slot in the default-model choice of a large mass of small and mid-sized developers. The chart also shows the runner-up, DeepSeek V4.1 Flash (2.43T), down 2% period over period, and third-place Tencent Hy4 preview (1.88T), down 4%, while GLM-5.3-Flash jumped 73% in a single day — the increment came mostly to it.
Break the 73% jump into two verifiable pieces of context. First, GLM-5.3-Flash is a natively multimodal model positioned for efficient coding and long contexts — exactly the "get-work-done" category where Flash-tier models compete, and where the shift from "let me try it" to "set it as my default" shows up directly in the call-volume curve. Second, OpenRouter's ranking logic aggregates real developer traffic, with no vendor inflation or self-evaluation, so its signal about which open models are actually being used is closer to market truth than any benchmark leaderboard. Note that at the weekly scale, GLM 5.3 Flash's cumulative tokens already stand at 16.9T, which suggests sustained adoption rather than a single-day spike.
Attribution matters. The 4.36 trillion figure is the daily call volume on one aggregation platform, not Zhipu's channel-wide total; the +73% baseline is "the previous statistical period," whose exact window the platform defines; and the ranking reflects traffic rather than quality — a model can be heavily called because it is cheap or has high cache hit rates, which is not the same as being the most capable. Zoom out, though, and the numbers align with the half-year trend: the capability gap among open models is narrowing while the price gap widens, and call volume is concentrating on tiers that are "good enough and cheap." GLM-5.3-Flash's rise to the top reads as a signal that in the next phase of open-source competition, the metric will increasingly be how many tokens developers actually burned, not where they rank on a few leaderboards.
The wider read is about where the open-source economy sits. A year ago the top of OpenRouter's daily chart belonged to frontier US labs' flash tiers; the fact that two of the top three slots now belong to Chinese open-weight families — and that the gap between first and second is nearly two trillion tokens — is a structural shift in how much of the world's commodity inference is sourced. It also matters for the companies themselves: call volume on aggregators like OpenRouter is becoming a funding-valuation input, because investors now ask "who is actually being used" rather than "who has the best demo." None of this settles the quality question — traffic is not capability — but it does settle the market question: developers have chosen, and the choice is price-performance first, brand second. The next benchmark to watch is whether GLM-5.3-Flash can hold the slot after its promotional window and when the next DeepSeek or Tencent flash tier ships.
[1][2]