Alibaba Releases Qwen3.8-Omni-Flash for Tool-Using Media Tasks
The new hosted model is meant to turn long recordings into inputs for AI workflows. Its low listed API rates are clear; its effectiveness on complex, multi-step media jobs remains to be tested.
Listen to this story
The audio brief
Story brief
3 key pointsAlibaba is positioning Qwen3.8-Omni-Flash as a hosted multimodal agent backend, with Alibaba Cloud Model Studio handling inference alongside custom tools, web search, caching, and adjustable reasoning. Its one-million-token context could support long audio and video sessions, while listed API rates—$0.15 per million input tokens and $0.47 per million output tokens—undercut the comparison Gemini figures. The...
- 01
Supports text, image, audio, and video input, but its documented output is text.
- 02
The service lists a one-million-token context window and up to 131,072 output tokens.
- 03
Alibaba estimates audio input below $0.01 per hour and 720p video at roughly $0.20.
Alibaba has released Qwen3.8-Omni-Flash, a model built for AI agents that can accept text, images, audio and video, then call custom tools during a task. The release gives developers a hosted route to work through lengthy media while keeping tool use and reasoning in the same workflow.
Alibaba Cloud Model Studio provides inference for the model. Its documentation lists a one-million-token context window, up to 131,072 output tokens, automatic context caching, adjustable reasoning effort and web search. Those features give an agent room to retain a large amount of media and its working history in one session.
From media input to an agent workflow
Qwen describes Omni-Flash as its first multimodal model built for AI agents. The company says the system can process audio and video together, reach conclusions and use tools for tasks such as editing vlogs, translating short videos and summarizing films. The documented model output is text, so any finished edit or other artifact still depends on the connected software around it.
A cheaper API is not the same as a cheaper job
Qwen’s published rates are lower than the Gemini figures listed above, but input price is only one part of a media workflow. Qwen estimates audio input at under $0.01 an hour and 720p video with audio at one frame per second at about $0.20, excluding response costs. The total bill can still change with the amount of output and the tools a workflow invokes.
What developers get with the launch
- A service that accepts text, images, audio and video, with text as its documented output.
- Custom tool calling, web search and context caching in the documented service.
- Open-source Qwen-MM-Plugins for media-oriented workflows in tools including Claude Code, Gemini CLI and Qwen Code.
The handoff is the real test
The plugin layer matters because it gives the model a route into existing agent environments rather than isolating it as a media chatbot. But its practical value will rest on how reliably an agent identifies relevant material, calls the right tool and hands work off without losing context. The launch establishes the interface; real-world media workflows will determine whether that interface is dependable.
Sources
- alibabacloud.comqwen3.8-omni-flash - Alibaba Cloud
- the-decoder.comQwen3.8-Omni-Flash undercuts Google's Gemini Flash pricing while matching its multimodal benchmarks
Loading discussion...
Reader comments
Newest comments first. Replies stay oldest first.