A 26% improvement across 30 evaluations is not a number you see often from an incremental model update. Alibaba’s Qwen team has released Qwen3.8-Omni-Flash, a native omni-modal model that handles text, images, audio, and video inside a single workflow, and the benchmark gains over its predecessor, Qwen3.5-Omni-Plus, are significant enough to pay attention to.
The model supports a 1-million-token context window, which puts it in the same tier as Google’s Gemini 1.5 Pro and well ahead of what most compact multimodal models offer. That context length matters practically: Qwen3.8-Omni-Flash is built for long-video analysis, meeting summaries, video research, and multimodal tool use. These are real workflows that enterprises and developers are already trying to run, and most existing models either cap out on context or require splitting modalities across separate pipelines.
Alibaba reported gains specifically in audio-video agents, coding, long-context tasks, and real-time multimodal interaction. To support those use cases, the team also released two companion tools: Qwen-MM-Plugins for long-running workflows and Qwen-Live Harness for real-time applications. That’s a more complete release than just dropping a model checkpoint and letting developers figure out the rest.
The pricing is where things get interesting for anyone evaluating this against Western alternatives. API input is priced as low as RMB 0.8 per million tokens, which converts to roughly $0.11 at current rates. That’s cheaper than GPT-4o mini and competitive with Gemini Flash pricing, and it’s coming from a model that claims stronger multimodal benchmarks than either. Cost-sensitive applications, particularly in Asia-Pacific markets, will find this hard to ignore.
This release fits into a broader pattern from Alibaba this year. The team launched Qwen3.8 with 2.4 trillion parameters in August, then previewed the Qwen4 architecture shortly after. Qwen3.8-Omni-Flash looks like the production-ready, efficiency-optimized layer sitting between those larger research pushes and what developers actually deploy. The model is available now through the Qwen AI platform.
- 1 million token context window across text, image, audio, and video
- 26%+ average score improvement over Qwen3.5-Omni-Plus across 30 evaluations
- Qwen-MM-Plugins for long-running multimodal workflows
- Qwen-Live Harness for real-time interaction use cases
- API input pricing starting at RMB 0.8 per million tokens
For developers building multimodal pipelines, this is worth testing. The combination of native omni-modal processing, long context, and low API cost is a meaningful stack. The benchmark claims still need real-world validation, but Qwen has earned enough credibility over the past year that the numbers aren’t easy to dismiss.



