The short version: On August 13, OpenAI previewed a new API service tier called Ultrafast that runs its GPT-5.6 Sol model up to 14 times faster than standard, at up to 750 words-worth of output per second. The model itself does not get smarter. Only the waiting gets shorter. That sounds like a developer detail, but real-time AI for small business is the difference between a customer who stays on the line and one who hangs up, and it is the first time speed has been sold as a thing you buy separately from intelligence.
For two years the AI conversation has been a race about capability. Which model reasons better, codes better, scores higher. This week’s news is about a different axis entirely, and it is the one that quietly decides whether the AI you already pay for works in front of a customer.
What OpenAI actually announced
OpenAI previewed Ultrafast mode as a new tier in its API, running GPT-5.6 Sol at up to 14 times the speed of standard processing. The work is done on hardware from Cerebras, a chipmaker whose processors keep an entire model’s weights on the chip itself rather than shuttling them back and forth from memory. Cerebras reports up to 750 output tokens per second, and says the speedup comes with no loss of quality: same model, same answers, delivered faster.
Two caveats belong right next to those numbers. This is a limited preview, gated behind a waitlist, and OpenAI has not published a price or a general availability date for the tier. And the headline benchmarks come from Cerebras itself, which has an obvious interest in the result. Treat vendor benchmarks as a claim awaiting independent evidence, not as a settled fact.
So nothing here is something you can go buy on Monday. What it tells you is where the market is heading, and that is worth knowing before you sign anything for next year.
Why is speed suddenly sold on its own?
Because the industry has run into a tradeoff that customers can hear. Broadly, the models that reason best are the ones that take the longest to answer, since more thinking means more tokens generated before the first word reaches the user. Until now, a business choosing an AI tool made one decision that set both dials at once. Pick the smart model and accept the pause. Pick the fast model and accept the shallower answer.
A separate speed tier breaks those dials apart. That is a structural change, and it follows the same pattern we have watched all year, where vendors slice one product into several priced dimensions. OpenAI split GPT-5.6 into three models at different price points. It added a premium seat to its business plan. Google halved its workhorse model’s price with a date attached, which we covered in why the date beside the price now matters more than the price. Speed is simply the newest dial to get its own price tag.
What does real-time AI for small business actually change?
It changes the jobs where a human is waiting. Those are a specific and valuable set: the phone ringing at 7pm, the website chat at midnight, the quote a customer wants before they call the next company on the list. These are the leaks we mapped in our guide to AI customer service for small business, and they share one property. Nobody waits.
A two second pause in a written chat is tolerable. The same pause on a phone call is a dead line, and the caller starts saying “hello?” This is why AI that answers the phone has felt slightly wrong even when its answers were correct, a gap that showed up clearly when voice models learned to listen and talk at the same time. Conversation has a rhythm, and being late to your turn reads as not understanding.
Worth saying plainly: the value here is capturing work that currently falls on the floor. The calls nobody answers after hours are not somebody’s job. They are revenue you never hear about, and your team never sees the missed ones. Faster AI covers the hours and the overflow that no small business staffs for anyway, which is a different thing from replacing the person who answers during business hours.
Is your AI too slow, or just wrong?
Here is the part worth acting on today, and it costs nothing. Most small business owners who feel let down by a customer-facing AI tool describe the problem the same way: “it just isn’t smart enough yet.” Sometimes that is right. Often it is a misdiagnosis, and the two problems have opposite fixes.
Pull ten real transcripts or call recordings from your chat widget or phone assistant and read them cold, ignoring how they felt. Ask one question of each: was the answer correct?
If the answers were right but people abandoned the conversation anyway, you have a latency problem. Buying a bigger, smarter model will make it worse, because bigger models are generally slower. If the answers were wrong, you have a capability or setup problem, and more speed only delivers the wrong answer sooner. Owners routinely upgrade to a more expensive model to fix what was a timing failure, then wonder why the number did not move.
The part the 14x number does not cover
This is the detail most coverage skips. A 14x faster model does not produce a 14x faster phone call, because the model is only one link in the chain. A spoken exchange runs through speech-to-text, then the model, then text-to-speech, then the phone network, and each stage adds its own delay. Speeding up one link cannot shrink the total below what the others cost.
The practical version: if a vendor pitches you on model speed next year, ask what the end-to-end response time is on a real call, measured from the moment the customer stops speaking to the moment they hear a reply. That single number is the one your customers experience. Anything quoted about tokens per second is a component spec, not a customer experience, and the gap between the two is where a lot of disappointing demos live.
What to do with this now
Nothing urgent, which is the honest answer. There is no price and no launch date, so this is not a decision. It is a heads-up with three useful consequences.
First, run the transcript test above and find out which problem you actually have, since that is true regardless of what OpenAI ships. Second, if you are being sold an AI phone or chat system in the next few months, add end-to-end response time to the list of things you ask about, alongside price and accuracy. Third, expect speed to show up as a line item on your bill eventually. When a capability gets its own tier, it usually gets its own charge, and the pattern this year has been that introductory terms come with dates attached.
The broader shift is genuinely good for small operators. When AI competition moves from raw intelligence toward responsiveness and cost, the winners are the businesses running small, well-defined jobs, not the ones with research budgets. If you want the wider context for where these pieces fit, start with our guide to AI for small business.
Frequently Asked Questions
Can I use OpenAI’s Ultrafast mode right now?
Not generally. OpenAI announced it on August 13 as a limited preview available to a subset of API customers, with no published pricing and no stated general availability date. In practice this matters to you through the tools you already buy rather than directly, since most small businesses reach these models through a chat widget, phone system, or CRM built by somebody else. Your vendor will adopt it, or not, and you will see the result as a shorter pause.
Does a faster AI model give worse answers?
Not in this case, which is the notable part. Ultrafast runs the same GPT-5.6 Sol model on different hardware, so the speedup comes from the chips rather than from shrinking the model, and both companies state the output quality is unchanged. That differs from the usual way of getting speed, which is switching to a smaller and genuinely less capable model. It is worth knowing which of the two a vendor means when they tell you they made things faster.
How do I tell whether my AI tool is too slow or just not good enough?
Read ten real conversations and check only whether the answers were correct. If the answers were right but customers dropped off mid-conversation, the problem is timing and a smarter model will likely make it worse, since more capable models generally respond more slowly. If the answers were wrong, speed is not your issue and upgrading for performance will just deliver wrong answers faster. This one test separates two problems that feel identical from the owner’s chair.
Will this make AI cheaper for my business?
Probably not by itself, and it may do the opposite. Faster processing on specialized hardware is a premium capability, and the consistent pattern this year is that new capabilities arrive as new tiers with their own prices. The likelier benefit is not a smaller bill but a wider set of jobs AI can do acceptably well, particularly live conversations where the old response times simply did not work. Budget for capability, not for savings.
Here is what we are curious about: when you have called a business and an AI answered, what gave it away first, the wording of the answer or the length of the pause before it? Tell us in the comments.
