Qwen3.8-Omni-Flash matches Gemini Flash multimodal benchmarks at a fraction of the price

Qwen's new multimodal model for AI agents processes audio and video, matches Gemini 3.8 Flash on benchmarks, and costs roughly one-fifth the API price.

Teams building multimodal AI agents now have a budget option that holds its own on standard benchmarks. Qwen3.8-Omni-Flash, Qwen’s first multimodal model built for agents, processes audio and video together, draws conclusions from them, and uses tools on its own. On audio-video tasks, Qwen reports performance close to Gemini 3.8 Flash, while its API pricing sits at $0.15 per million input tokens and $0.47 per million output tokens.

What Qwen3.8-Omni-Flash does

The model handles audio and video as a single stream, reasons across both, and can call tools without human prompting. The practical examples Qwen lists include editing vlogs, translating short videos, and summarizing movies. The context window spans one million tokens, which lets an agent hold a long video and the surrounding instructions in working memory at the same time.

How it compares with Gemini 3.8 Flash on benchmarks

Qwen says Qwen3.8-Omni-Flash performs on par with Gemini 3.8 Flash across multimodal benchmarks. The full benchmark table Qwen published shows matched scores on audio and audio-video tasks, with Gemini 3.8 Flash holding small leads on a few categories. Both models are aimed at the same agent-building use case, but Qwen’s model is priced well below it.

How much it costs to run

Qwen’s API pricing works out to $0.15 per million input tokens and $0.47 per million output tokens. Qwen estimates audio input at under $0.01 per hour, and 720p video with audio sampled at one frame per second runs about $0.20 per hour, not counting response costs. For comparison, Gemini 3.8 Flash charges $0.75 for input and $3.75 for output per million tokens at its introductory rate, and those prices are scheduled to double on January 1, 2027. That puts Qwen at roughly one-fifth of Gemini’s input price and about one-eighth of its output price today, with the gap set to widen further once the Gemini increase lands.

Where the model is available

The model is available through Qwen Studio, Qwen Cloud, and the Alibaba Cloud Model Studio API. The open-source Qwen-MM-Plugins repo adds video editing, speaker recognition, PDF video notes, and reusable workflows to agents such as Claude Code, Gemini CLI, and Qwen Code. A separate open-source tool, Qwen-Live Harness, enables real-time interaction using a camera and microphone, so a developer can point the model at a live stream and let it react.

What the open-weight release unlocks

Because the model and its plugin stack are open-weight, teams can run Qwen3.8-Omni-Flash on their own hardware, fine-tune it on private video archives, or wire it into existing agent frameworks without sending audio or video through a third-party endpoint. The plugins handle the agent-side glue: speaker recognition for meeting recordings, video editing for short-form content, and PDF video notes for turning recorded presentations into searchable documents. For a team that already pays for a Gemini 3.8 Flash integration, the migration path is the same multimodal endpoints with a smaller line item at the end of the month.

What to watch before adopting it

Qwen’s published benchmark table covers audio and audio-video tasks. Teams that depend on long-context reasoning, code execution, or pure text benchmarks should test those cases directly rather than assume parity. Pricing for response tokens on video jobs is not bundled into the $0.20 per hour figure, so total cost will rise once the agent actually answers. Finally, the Gemini 3.8 Flash price increase in January 2027 makes the cost gap larger over time, but it also means any contract renegotiation with Google’s API will land in the same planning cycle.

FAQ

What is Qwen3.8-Omni-Flash?

Qwen3.8-Omni-Flash is Qwen’s first multimodal model built for AI agents. It processes audio and video together, draws conclusions, and uses tools on its own to edit vlogs, translate short videos, or summarize movies. The context window is one million tokens.

How does Qwen3.8-Omni-Flash compare with Gemini 3.8 Flash?

Qwen reports that Qwen3.8-Omni-Flash performs on par with Gemini 3.8 Flash on multimodal benchmarks, with Gemini holding small leads in a few categories.

How much does Qwen3.8-Omni-Flash cost?

The Qwen API charges $0.15 per million input tokens and $0.47 per million output tokens. Audio input is estimated at under $0.01 per hour, and 720p video with audio at one frame per second runs about $0.20 per hour, not counting response costs. Gemini 3.8 Flash charges $0.75 input and $3.75 output per million tokens at its introductory rate, with prices set to double on January 1, 2027.

Related coverage


This article summarizes reporting from the-decoder.com. See our editorial disclaimer for how our articles are produced.

🤖
Is your business visible to AI assistants?

Run a free scan to see your AI Visibility Score, SEO rating, and local citation accuracy.

Check Your Score →