On September 23, Alibaba's Qwen released the Qwen-Audio-3.1 series of speech large models: speech recognition (ASR), speech synthesis (TTS), and real-time speech interaction (Realtime) all upgraded, plus two new audio creation models, Qwen-Audio-3.1-TTS-Next and Qwen-Audio-3.1-ASR-Next — five models whose APIs are already listed on the Qwen AI platform. At the same time, prices across the Qwen-Audio lineup were cut: TTS by about 70%, Realtime by about 85%, and ASR by as much as 95%.

[1][2]

The nut graf: putting five models and a lineup-wide price cut in the same release, Alibaba's message is clear — audio is no longer an accessory feature of LLMs but an independent capability stack laid out as understand-generate-interact-create, and this stack will compete the way text models did: cut prices until developers use it without thinking, then win on ecosystem scale. After a 95% ASR cut, speech recognition becomes close to a free resource, overturning the pricing premise that recognition is expensive.

The numbers make it concrete. Per official pricing as cited by Sina Finance: Qwen-Audio-3.1-ASR-Flash is 0.8 yuan per million input tokens and 2.7 yuan per million output tokens; TTS-Flash is 1.5 yuan input and 12 yuan output. The official claims are about 70% off for TTS, 85% for Realtime, and 95% for ASR. The two "Next" creation models are the new faces here: they push audio generation from "reading a script" to "making content" — text, voice, sound effects, and ambient sound in one pass — aimed at audiobooks, podcasts, and short-video dubbing. For developers the substantive change is that a product that speaks used to mean picking recognition, synthesis, and interaction separately with separate billing; now one Qwen-Audio-3.1 stack covers all of it, at lower prices.

Attribution: the cut percentages and unit prices are vendor-released figures with no third-party comparison yet; the capability boundaries of the audio-creation models (multi-track mixing, sound-effect controllability, rights handling) were not detailed in the release. In industry context, this release forms a contrast with iFlytek's Spark-ASR-2.0 the same day: one is an established speech vendor upgrading recognition toward fluent text, the other a cloud giant bundling the whole audio stack and cutting prices — the competition in speech has moved from "whose recognition is more accurate" to "who makes sound cheaper and easier to turn into products." For small developers, this may be the first time making a product speak costs little enough to ignore.

The strategic read goes beyond pricing. Bundling five models into one stack changes the developer default: instead of evaluating each vendor per capability, a team can adopt the whole audio layer from one provider and swap later. That is how platform lock-in starts in an API market — not through exclusivity, but through convenience and a price curve that makes switching feel wasteful. Alibaba has played this game in text; Qwen-Audio-3.1 is the same playbook applied to sound, and the two audio-creation models are the differentiation bet, aimed at the content-production boom rather than the commoditized transcription lane.

[1][2]
Afternoon sound studio, a recording engineer's back at a five-channel mixing console, a large screen showing five differently colored waveform tracks laid side by side, several smooth and regular, one with fine light points; monitor speakers standing in the room, one hand holding headphones, the other adjusting a fader, acoustic foam on walls, afternoon city beyond the window. No text or numerals.
Five tracks, one audio stack, AI-generated illustration, not a news photo