OpenAI released GPT-6 Astra on September 3 and, in the same announcement, said the model may represent artificial general intelligence. It is the largest claim the company has ever made about its own work. For anyone actually shopping AI software for small business use, it is also close to irrelevant this month, and the reason is sitting in plain sight on OpenAI’s public price list.
The short version: GPT-6 Astra is a real capability jump, and it costs $10 per million input tokens. The same price list still sells GPT-5-nano at $0.05 per million input tokens. That is a 200 to 1 spread inside a single catalog, and nearly every job a small business hands to AI today finishes correctly at the cheap end of it. What changed this week is not which model you should be paying for. It is the ceiling on what the software you already rent will be able to do next year.
What OpenAI actually shipped
The specifications are genuinely impressive, and worth separating from the framing. GPT-6 Astra carries a context window of 1,050,000 tokens with up to 128,000 tokens of output. On OpenAI’s eight-needle retrieval test, which measures whether a model can find specific facts buried inside a very long document, it scored 100 percent between 256K and 512K tokens and 96.3 percent between 512K and 1M. It reached 97.6 percent on FrontierMath Tier 4, against 83.0 percent for the previous generation.
The headline benchmark is ARC-AGI-3, a test built specifically to measure whether a model can learn tasks it has never seen before. Astra scored 99.9 percent. That result deserves an asterisk, and Simon Willison flagged it on the day: the score came through a custom adapter harness built by OpenAI, not the default evaluation setup. It is a real result produced under conditions the vendor chose. Those are not quite the same thing.
Access is staged. Enterprise customers already inside OpenAI’s Daybreak program get it first, with Plus, Pro, Business and Enterprise plans, the API and AWS following in the coming days. If you pay for ChatGPT, you will probably have it before you have decided whether you want it.
Why AGI-level AI does not change the AI software for small business
Here is the arithmetic the benchmark charts leave out. OpenAI’s published API pricing puts GPT-6 Astra at $10 per million input tokens and $50 per million output. GPT-5-mini sits at $0.25 and $2.00. GPT-5-nano sits at $0.05 and $0.40.
Now think about what you actually ask AI to do. Draft a reply to the customer who wants to reschedule. Summarize a forty minute site visit into five bullets. Sort thirty inbound emails into quote requests, complaints and noise. Rewrite a product description for the third time. These tasks were saturated two model generations ago. The cheap model gets them right, and the frontier model gets them right in a way you cannot perceive, at forty to two hundred times the cost.
This is the part that gets lost during AGI week. A more capable model does not make the cheap model worse. Everything nano did on Tuesday, it still does on Friday. A new tier is a ceiling being raised, not a floor being pulled out from under you.
There is a quieter point underneath it. You are mostly not the one choosing. When you buy scheduling software, or a CRM with an assistant bolted onto it, the vendor picks the model behind the feature and absorbs the bill. That decision gets made in a room you are not in, and it is where the real pricing story has been hiding all year.
The number that actually moved
One benchmark on this week’s list matters more to a small business than every reasoning score combined. On OSWorld 2.0, which measures whether a model can operate real desktop software the way a person does, clicking through menus and filling in fields, Astra scored 72.6 percent against 65.7 percent for the prior model, and finished in roughly 47 percent less time per task.
That is the capability that reaches small businesses as features rather than as models. It is the engine underneath the class of tools that reached general availability last month: software that can work inside an application with no integration and no API, which describes most of what a small business actually runs.
Read 72.6 percent honestly, though. Roughly one attempt in four still fails. That is a useful assistant working under supervision, and it is genuinely useful. It is not something you hand your banking login and walk away from, and any vendor implying otherwise this quarter is selling ahead of the evidence.
Two labs, one price, two different scores
The most revealing detail of the launch is one nobody put in a headline. Astra arrived at $10 and $50 per million tokens. That is the identical headline price Anthropic charges for Claude Fable 5.1, which shipped two days earlier. On Artificial Analysis’s general intelligence index, Astra comes in at 61 and Fable 5.1 at 66.
Two frontier labs, the same sticker, different scores. The headline price at the top of the market has stopped being a differentiator, which means the competition has moved to the lines underneath it: cache reads, speed, and the cost to finish a task rather than the cost of a token. For a small business, that means the vendor choosing a model on your behalf is now making a judgment call that price alone can no longer settle for them. That is a fair thing to ask about at your next renewal.
What should a small business do about GPT-6 Astra?
Nothing urgent, and that is a legitimate answer rather than a dodge. Four things are worth doing over the next month.
Start by checking what you are being charged for. If a tool you already pay for advertises that it is powered by the latest model, ask which one, and whether your plan actually includes it. Second, resist upgrading a workflow that already works. If your inbox sorting has run correctly for six months on a cheap model, a smarter model buys you nothing there. Third, revisit the job you gave up on. Most owners have one task they tried with AI a year ago, found unreliable, and quietly dropped; the OSWorld numbers say that is exactly the class of task that changed. Fourth, if you or a contractor build anything directly on the API, price the task rather than the token. A model at double the rate that succeeds in a third of the attempts is the cheaper model.
The honest summary of AGI week is that the ceiling went up and the floor stayed exactly where it was. For a small business that is good news, because the floor is where nearly all of the useful work happens, and it keeps getting cheaper to stand on. If you want the ground-level view of what that work looks like, we walked through it industry by industry here.
Frequently Asked Questions
Do I need to pay for GPT-6 Astra?
Almost certainly not as a separate purchase. If you subscribe to ChatGPT Plus, Business or Enterprise, access is rolling out to the plan you already hold at no extra cost. If you use AI inside other software, your vendor decides which model runs behind the feature and absorbs the API bill, so the choice is not yours to make. The only case where the sticker price touches you directly is if you or a contractor build on the API, and there the cheaper GPT-5-mini and GPT-5-nano tiers handle most small business tasks correctly at a small fraction of the cost.
Is GPT-6 Astra really AGI?
OpenAI has said the model may represent artificial general intelligence, and its ARC-AGI-3 score of 99.9 percent is the strongest evidence offered for that. The score was produced using a custom adapter harness built by OpenAI rather than the default evaluation setup, which makes it a real result under conditions the vendor selected. Independent testing over the coming weeks will settle how well it generalizes. For practical purposes the label changes nothing about what the model costs or what it can be trusted to do without supervision.
Will GPT-6 Astra replace jobs in my business?
The benchmark that moved most is OSWorld 2.0, which measures operating ordinary desktop software, and the score went from 65.7 percent to 72.6 percent. That means roughly one attempt in four still fails, which is not a standard anyone would accept from a person working unsupervised. What the model does well is the tedious middle of a job: filling in forms, moving data between systems that do not talk to each other, working steadily through a queue. That hands time back to the people you already employ rather than removing the reason you employ them.
Which AI model should a small business actually use?
Match the model to the task rather than to the headline. Routine text work such as drafting replies, summarizing notes, sorting inbound messages and rewriting descriptions runs correctly on the cheapest tier available, currently GPT-5-nano at $0.05 per million input tokens. Reserve the expensive frontier tier for work that genuinely fails on the cheap one, which in most small businesses means long document analysis or multi-step tasks that operate other software. Test the cheap model first and move up only when you can point at a specific failure.
So here is the question we keep coming back to: which task did you give up on last year? That is the one worth trying again this month, and we would genuinely like to hear what it was.
