On September 18, Alibaba's Qwen team released Qwen3.8-Omni-Flash, a natively omni-modal model that accepts text, image, audio and video input inside a single 1M-token context window and returns text. The launch post calls it Qwen's first omni-modal model built around agentic capabilities: understand the content, plan the task, call tools, deliver the result. The model is available only as a hosted API on the Qianwen AI Platform and Alibaba Cloud Model Studio; no open weights were announced at launch. The architecture traces back to Qwen3.8-Flash-Next, a multimodal MoE released with open weights in August, which the company positions as the architectural preview of Qwen4.
[1][2]The headline of this release is not multimodality itself, which is no longer unusual. It is the cost structure. Qwen reports that the API price per hour of audio input falls by more than 98% against Qwen3.5-Omni-Plus, and the price per hour of audio-visual input by more than 93%; on a per-token basis, audio input is as cheap as 0.8 yuan per million tokens, with implicit cache hits at 0.016 dollars per million tokens on the international listing. If those numbers hold, an hour of meeting audio stops being an expensive experiment and becomes a routine API call, which puts direct pricing pressure on Western competitors selling audio-video understanding.
On quality, the vendor reports an average improvement of more than 25% across 29 evaluations against Qwen3.5-Omni-Plus, a 36.5-point gain on WildClawBench-MM, 69.6 on UniClawBench and a 22.3-point gain on AgenticVBench. On OmniVideoBench, an agentic perception approach — coarse scans followed by focused passes on relevant segments — raises accuracy from 63.4 to 67.8 while cutting token use by about 45.7%. None of these figures have been independently reproduced as of publication; they are all vendor-reported launch numbers, a caveat the company's own materials do not hide.
A few product details matter. The model outputs text only; Qwen positions it as an agent's brain rather than its mouth, and points speech-synthesis needs at Qwen3.5-Omni. Thinking is on by default through a reasoning_effort parameter that runs from low to xhigh. It supports 2-channel and 4-channel spatial audio, 113 input languages, and video files up to two hours at 2 GB via URL. Two companion projects are open-sourced alongside: Qwen-MM-Plugins under Apache-2.0, which gives text-only agent harnesses like Claude Code, Codex and Gemini CLI audio-video input through Skills and optional MCP servers, with capabilities such as turning a tutorial video into an illustrated PDF or building audio-visual memory from long videos; and Qwen-Live Harness for real-time interaction. A low-latency variant, Qwen3.8-Omni-Flash-Realtime, reportedly answers a 20-second audio prompt in about 981 ms over WebSocket or WebRTC. The model runs in six regions: Beijing, Singapore, Hong Kong, Tokyo, Frankfurt and Virginia.
The competitive signal is direct. Qwen says audio-video performance approaches Gemini 3.8 Flash, and the platform changelog shows the model arriving with aggressive pricing on the underlying multimodal pipeline. That is a pricing attack at the intersection of multimodality and agents, aimed squarely at the workflows where incumbents have been charging a premium for media input. The real test comes in independent evals and in real production workflows, where the cost per hour of media processing will decide whether the price cuts survive contact with actual usage patterns. But on price, Qwen has put its cards on the table.
[1][2]