Alibaba Releases Qwen3.8-Omni-Flash for Tool-Using Media Tasks

The new hosted model is meant to turn long recordings into inputs for AI workflows. Its low listed API rates are clear; its effectiveness on complex, multi-step media jobs remains to be tested.

By 2 min read
Alibaba Releases Qwen3.8-Omni-Flash for Tool-Using Media Tasks
Alibaba Releases Qwen3.8-Omni-Flash for Tool-Using Media Tasks

Listen to this story

The audio brief

About 1:30
0:001:30
Read transcript
Alibaba has released Qwen3.8-Omni-Flash, a hosted model designed to turn long audio and video sessions into inputs for AI workflows. It accepts text, images, audio, and video, then can call custom tools while it reasons through a task. The service runs through Alibaba Cloud Model Studio, with web search, automatic context caching, and adjustable reasoning effort built in. The headline capacity is a one-million-token context window, with up to 131,072 output tokens. In practical terms, that gives an agent room to keep a substantial recording and its working history in one session. Alibaba points to use cases including vlog editing, short-video translation, and film summarization. But there is an important boundary: the documented output is text. Any finished edit or other media artifact still has to be produced by connected software. The listed API pricing is fifteen cents per million input tokens and forty-seven cents per million output tokens. Alibaba estimates audio input at under one cent per hour, and 720p video with audio at about twenty cents, excluding response costs. Those rates are below the listed comparison figures for Gemini 3.8 Flash, but the total bill also depends on output and tool calls. Qwen-MM-Plugins connects the model to Claude Code, Gemini CLI, and Qwen Code. The key question now is whether agents can reliably find the right material, choose the right tool, and preserve context across those handoffs.

Story brief

3 key points

Alibaba is positioning Qwen3.8-Omni-Flash as a hosted multimodal agent backend, with Alibaba Cloud Model Studio handling inference alongside custom tools, web search, caching, and adjustable reasoning. Its one-million-token context could support long audio and video sessions, while listed API rates—$0.15 per million input tokens and $0.47 per million output tokens—undercut the comparison Gemini figures. The...

  1. 01

    Supports text, image, audio, and video input, but its documented output is text.

  2. 02

    The service lists a one-million-token context window and up to 131,072 output tokens.

  3. 03

    Alibaba estimates audio input below $0.01 per hour and 720p video at roughly $0.20.

Alibaba has released Qwen3.8-Omni-Flash, a model built for AI agents that can accept text, images, audio and video, then call custom tools during a task. The release gives developers a hosted route to work through lengthy media while keeping tool use and reasoning in the same workflow.

Alibaba Cloud Model Studio provides inference for the model. Its documentation lists a one-million-token context window, up to 131,072 output tokens, automatic context caching, adjustable reasoning effort and web search. Those features give an agent room to retain a large amount of media and its working history in one session.

From media input to an agent workflow

Qwen describes Omni-Flash as its first multimodal model built for AI agents. The company says the system can process audio and video together, reach conclusions and use tools for tasks such as editing vlogs, translating short videos and summarizing films. The documented model output is text, so any finished edit or other artifact still depends on the connected software around it.

A cheaper API is not the same as a cheaper job

Qwen’s published rates are lower than the Gemini figures listed above, but input price is only one part of a media workflow. Qwen estimates audio input at under $0.01 an hour and 720p video with audio at one frame per second at about $0.20, excluding response costs. The total bill can still change with the amount of output and the tools a workflow invokes.

What developers get with the launch

  • A service that accepts text, images, audio and video, with text as its documented output.
  • Custom tool calling, web search and context caching in the documented service.
  • Open-source Qwen-MM-Plugins for media-oriented workflows in tools including Claude Code, Gemini CLI and Qwen Code.

The handoff is the real test

The plugin layer matters because it gives the model a route into existing agent environments rather than isolating it as a media chatbot. But its practical value will rest on how reliably an agent identifies relevant material, calls the right tool and hands work off without losing context. The launch establishes the interface; real-world media workflows will determine whether that interface is dependable.

Sources

  1. alibabacloud.comqwen3.8-omni-flash - Alibaba Cloud
  2. the-decoder.comQwen3.8-Omni-Flash undercuts Google's Gemini Flash pricing while matching its multimodal benchmarks

Loading discussion...

YOUR READING SPACE

Notifications

Alibaba Releases Qwen3.8-Omni-Flash for Tool-Using Media Tasks | Superpower Daily