Somewhere in your phone system there are fifty-odd hours of recorded calls from last month that nobody has listened to. Not because they do not matter. Because listening to them is fifty-odd hours of work, and nobody in the building has fifty-odd hours.
The short version: Alibaba’s Qwen team released Qwen3.8-Omni-Flash on September 18, a model that accepts audio, video, images and text. Every outlet reported the same price, $0.15 per million input tokens and $0.47 per million output. That pair is the cheapest cell on a rate card that has five of them. The rate governing audio, which is the whole reason a small business would reach for this model, is $3.81 per million tokens, and multimodal requests bill output at $3.06 rather than $0.47. Run a real month of phone calls through it and the bill lands near $12.82 instead of the $1.34 the headline implies. That is roughly ten times. It is also still under $13 a month, which is the part worth your attention.
What does an hour of audio actually cost?
Qwen tokenizes audio at 7 tokens per second, so an hour of recorded call is 25,200 tokens. At the audio input rate of $3.81 per million, that hour costs 9.6 cents.
Qwen’s own estimate, repeated by The Decoder and Startup Fortune, is that audio input runs “under one cent per hour.” Both numbers come from Qwen and differ by roughly ten times. The only difference is which rate cell you apply: under a cent at the $0.15 headline rate, 9.6 cents at the published audio rate.
Now scale it. Take a three-van home-service company that fields 45 calls on a normal day, averaging three and a half minutes. That is 57.75 hours of call audio across a 22-day working month, or 1,455,300 audio tokens. At the headline rate that is $0.22 for the month. At the published audio rate it is $5.54.
Five dollars is not a crisis. It is a rounding error against a phone bill. The reason to do the arithmetic anyway is that the same misreading applies to output, where the trap is better hidden.
Which rate card will you land on?
Two rate cards are circulating for one model, and which applies depends on the endpoint you sign up to rather than on anything you choose.
The widely reported card is $0.15 input, $0.47 output, $0.016 for cached input. The Singapore and international deployment card, as read by Aireiter, breaks the same model out by modality: $0.43 for text input, $3.81 for audio input, $0.78 for image or video input, $1.66 for output on a text-only request and $3.06 for output on a multimodal one.
Audio input costs 8.9 times what text input costs, and sending any audio raises your output rate by 84 percent because the request is now multimodal. The headline number is real, but it describes the cheapest possible request, and a business feeding it phone calls is never making that request.
We went looking for Alibaba’s own per-modality price list to settle it. Alibaba Cloud Model Studio’s model list, checked on September 20, 2026, carries qwen3.8-omni-flash under four categories, including speech recognition and omni, and prints no price for any of them. The prices live behind the console, which is worth knowing before you budget against a number from a news story, this one included.
The setting that quietly raises your output bill
A second detail does not appear in the coverage at all. Per MarkTechPost’s read of the model documentation, thinking is on by default, with reasoning_effort set to xhigh. Reasoning tokens bill as output tokens.
Take the same business, and say every call gets a short summary plus a few extracted details. Assume 2,000 reasoning tokens and 400 visible tokens per call, an assumption rather than a measurement and the one number here we did not read off a rate card. Across 990 calls that is 2,376,000 output tokens: $1.12 at the headline output rate, $7.27 at the multimodal one.
Add the audio and the month comes to $12.82 on the real card against $1.34 on the headline. Most of that gap is reasoning you did not ask for. If your workload does not need deliberation, turning that setting down is the single largest lever on the bill, and it is one line of configuration.
Does Qwen3.8-Omni-Flash have open weights?
No, and it is worth pinning down because at least one headline says otherwise. Startup Fortune’s piece is titled “Slashes Audio Pricing 98% and Drops Open Weights,” yet the body of that same article states that “the weights are not the front door.” The confusion is a near-identical name. Qwen3.8-Flash-Next has open weights on Hugging Face. Qwen3.8-Omni-Flash, the one that hears audio, does not.
The consequence: you cannot run this on your own hardware or inspect it, and your call recordings leave your building. For many small businesses that is fine. For anyone handling medical, legal or financial conversations it is the whole question, and this model does not give you the option. We covered the work you have quietly refused to upload to anyone else’s server when Meta shipped an open-weight model small enough to run offline.
What is this actually good for?
Not answering your phone. Somebody at your shop already does that, and they are better at it than a model is.
The opportunity is the pile. Fifty-eight hours of conversations already happened, and they already contain the quote you promised to follow up on, the customer moving in spring, the third caller this month asking about a service you do not offer yet. Nobody has the hours to go and get it. At thirteen dollars a month, going and getting it stops being a project and becomes a background process, and whoever answers the phone gets handed a follow-up list instead of an extra job.
Two cautions. It accepts audio in 113 languages and dialects, which matters in a market like Fresno, but MarkTechPost notes most tool harnesses cannot yet feed audio to the main model natively, so the plumbing may not exist in your product yet. And its benchmark claims sit close to Gemini 3.8 Flash rather than above it, a pattern we have seen in Alibaba’s own benchmark tables before.
The lesson keeps recurring this year. A headline model price is a marketing number, and the line that decides your bill is usually one nobody quoted. It was the cache read rate when Anthropic shipped Fable 5.1. This month it is which modality you send. Before you budget against any model, find the rate card with more than two numbers on it.
Frequently Asked Questions
How much does Qwen3.8-Omni-Flash cost per hour of audio?
At the published audio input rate of $3.81 per million tokens and Qwen’s tokenization of 7 audio tokens per second, one hour of recorded audio is 25,200 tokens and costs 9.6 cents, verified September 20, 2026. Qwen’s own estimate of “under one cent per hour” reflects the $0.15 headline input rate rather than the audio rate.
Why is the audio price different from the $0.15 headline price?
It bills by modality rather than at one flat rate. The international rate card lists $0.43 for text input, $3.81 for audio and $0.78 for image or video, so audio costs 8.9 times what text costs. Sending audio also shifts output from $1.66 to $3.06, because the request counts as multimodal.
Can I run Qwen3.8-Omni-Flash on my own hardware?
No. Despite headlines suggesting otherwise, it is API-only through QwenCloud, Alibaba Cloud Model Studio and Qwen Studio. The similarly named Qwen3.8-Flash-Next does have open weights on Hugging Face, which is the likely source of the confusion. If your calls contain medical, legal or financial detail that cannot leave your premises, this is not the right tool.
What would a month of phone calls actually cost my business?
For a three-van home-service company taking 45 calls a day at three and a half minutes each, roughly 58 hours of audio a month, expect about $12.82 at the real rate card against the $1.34 the headline price implies. That figure assumes 2,400 output tokens per call with reasoning enabled, and reasoning is on by default at the highest setting, so turning it down is the fastest way to cut the bill.
One thing we are curious about: if you could hand your team a list pulled from every call last month, what would you want it to find first?
