On September 23, iFlytek officially launched Spark-ASR-2.0, its latest speech recognition large model. Per the company, the model keeps high recognition accuracy while focusing on the fluency and normalization of output text, marking a shift from "transcribing content" to "writing fluent text"; it starts rolling into the iFlytek Input Method tomorrow, with open-platform API access, then into products like iFlytek AI Glasses, smart business pads, and iFlytek Tingjian.

[1][2]

The nut graf: speech recognition is one of the few moat-type tracks in the Chinese AI market, and the technical route here deserves attention — Spark-ASR-2.0 is not a brute-force parameter dump but uses "non-autoregressive plus LLM-enhanced autoregressive" cooperation to split the two goals of recognition: the non-autoregressive side is fast, the LLM-enhanced side understands. The selling point has also moved from "character accuracy" to "well-formed text," and with inference cost reportedly up only about 10% (per media citing official figures), the bundle reads less like another accuracy-leaderboard round and more like paving the way for speech to scale into every device.

The rollout cadence is dense. Starting tomorrow it steps into the iFlytek Input Method, API access goes live on the open platform in parallel, and then hardware and SaaS lines — AI Glasses, smart business pads, iFlytek Tingjian — plug in successively. That means Spark-ASR-2.0 is the base of a multi-device product matrix from day one, not a single demo. On the technical side, it extends the Spark-Audio-1.0-Preview speech base model, with key techniques including joint Chinese-English mixed text and acoustic enhancement and dynamic context injection, covering general recognition (mixed Chinese-English, dialects, technical terms) and complex acoustic scenes (high noise and similar). The more substantive change for developers is the API and SDK channel: iFlytek's open platform dictation and speech-model services have long been among the most-used transcription channels in the Chinese ecosystem, and if the new model's fluent-text capability is stable, the most direct beneficiaries are downstream applications like meeting transcription, subtitling, and customer-service QA.

Attribution works on two levels. The specific accuracy numbers and the "about 10% inference cost increase" come from vendor release materials; no independent third party has yet re-tested on a unified corpus, and "fluent text" is a subjective quality whose evaluation criteria (human scoring or automatic metrics) the official release did not detail. Stepping back: over the past two years, speech recognition was somewhat overshadowed by the "end-to-end" narrative of LLMs, but dialects, technical terms, noise, and colloquialism in Chinese remain the real battleground of accuracy. Spark-ASR-2.0's value is not that it refreshed some leaderboard; it is that it shows recognition and understanding can cooperate in layers within one model chain — possibly another step in speech interaction moving from usable to good.

The competitive context matters too. Chinese speech-to-text is a market where the incumbent faces open-weight and cloud giants on one side and specialized ASR startups on the other; what distinguishes this release is less raw accuracy claims and more the delivery system — input-method reach, an open platform with SDKs, and hardware lines already shipping. That distribution advantage is harder to copy than a benchmark number. Watch whether the fluent-text mode holds up in noisy real-world audio and dialects, because that is where the official claims meet the actual users.

[1][2]
Morning speech lab, an engineer's back wearing monitoring headphones at a mixing console, a large screen showing an audio waveform being smoothed into a clean curve with fine light points at its edges, acoustic foam on the walls; one hand holds the headphone, the other pushes a fader on the mixer, desk lamp and screen light mixed, morning city beyond the window. No text or characters.
Smoothing speech into text, AI-generated illustration, not a news photo