The short version: Google announced Gemini 4 Argon on September 30, and the number worth an owner’s attention is not the benchmark tie with OpenAI’s GPT-6 Astra. It is what Argon does when it does not know. On Artificial Analysis’s AA-Omniscience test, published the same day, Argon gave a wrong answer in 15 percent of the cases where it lacked the right one. GPT-6 Astra did so in 51 percent. Run the two published rates over 100 hard questions and Astra gets 13 more right, but makes roughly 19 confident mistakes to Argon’s 7 or 8. Argon would rather say “I don’t know.” You cannot use it yet; Google has opened it only to vetted cyber defenders so far.
That trade, fewer answers but far fewer invented ones, is the most useful thing a frontier model has offered a small business in a while. Here is what the figures mean, how the two models compare question by question, and which jobs in your week favor which kind of assistant.
What did Google actually release?
Gemini 4 Argon is Google DeepMind’s first Gemini 4 model, aimed at long software projects, legal and financial work, and cyber defense. Google’s launch post (read October 1) highlights a jump in output length to 1 million tokens, up from 64,000, and scores of 77.9 percent on the DeepSWE coding test and 51.3 percent on AutomationBench.
Independent testing landed the same day. Artificial Analysis scored Argon 53 on its Intelligence Index, level with GPT-6 Astra at its maximum setting and one point ahead of GPT-6.1 Sol. A tie at the top is ordinary news in a year of monthly releases. What sits underneath the tie is not.
What does a 15 percent hallucination rate mean?
It does not mean Argon is wrong 15 percent of the time. AA-Omniscience asks 6,000 questions across business, law, health, software, science and the humanities. Artificial Analysis’s method page defines the hallucination rate as wrong answers divided by every response that was not fully correct: the wrong ones, the partial ones and the ones the model declined to attempt. In plain terms, it measures what a model does once it is out of its depth. Does it guess, or does it stop?
Argon stops far more often. Its overall accuracy, though, is lower: 50 percent of questions answered correctly, against 63 percent for GPT-6 Astra. Both figures matter, so we combined them.
Gemini 4 Argon vs GPT-6 Astra on the same 100 questions
Applying each model’s published accuracy and hallucination rate to 100 questions gives this split. The arithmetic is ours; the two input rates are Artificial Analysis’s, published September 30.
| Per 100 hard questions | Gemini 4 Argon | GPT-6 Astra (max) |
|---|---|---|
| Answered correctly | 50 | 63 |
| Confident wrong answer | 7.5 | 18.9 |
| Declined or partial | 42.5 | 18.1 |
| Right answers per wrong one | 6.7 | 3.3 |
Read across the rows and the two models arrive at nearly the same place by opposite roads. Artificial Analysis’s overall Omniscience score, which rewards right answers, penalizes wrong ones and ignores refusals, puts Argon at 42 and Astra at 43. Astra crosses more of the gap. It also walks off the edge about two and a half times as often. One model builds the whole bridge whether or not the middle is there; the other stops building where its knowledge runs out.
When Argon gives you a definite answer, it is right about 6.7 times for every miss. Astra manages about 3.3. For anyone who has had an assistant state a fake fact in a perfectly confident sentence, that ratio is the headline.
Which matters more for a small business: more answers or fewer wrong ones?
It depends on what happens after the answer, and that is a question about your workflow, not about Google.
Where the honest “I don’t know” wins: anything that leaves the building or gets acted on without a second look. A reply to a customer about warranty terms. A summary of what a supplier contract allows. A first pass at whether a job needs a permit in your county. A confident wrong answer here costs a refund, an argument or a redo, while “I can’t confirm that, check the county page” costs you two minutes. Take a three-person office that drafts 40 customer replies a week. Any reply that turns on a fact the model has to recall, rather than read from your own documents, is a hard question in the benchmark’s sense, and on those the gap points to roughly two and a half times as many invented details from the bolder model.
Where the bolder model wins: brainstorming, rough drafts and anything you were going to check line by line anyway. Ad headlines, a first outline of a training checklist, ten ways to word a price increase. Here a wrong suggestion costs nothing, and 13 extra hits per hundred are worth having.
The rule that has always kept small businesses safe with AI still holds: AI drafts, you approve. What Argon changes is how much the approving has to catch.
When can you use Gemini 4 Argon?
Not yet. Google says Argon is rolling out first to “trusted cyber defenders” through what it calls the Fairwind Program, and that it is taking part in the US government’s voluntary pre-release review. 9to5Google reports paid API customers and Google AI Ultra subscribers come next. Google gives no dates for either.
Pricing is set in advance: $2 per million input tokens and $10 per million output tokens during an introductory period, then $4 and $20 (Google’s launch post, read October 1). Google has not said when the introductory period ends. One caution for when it does arrive: Artificial Analysis found Argon used about 62,000 output tokens per test task against Astra’s 27,000. It thinks at length before it answers, and the meter runs while it thinks.
If you use Gemini today, through the Gemini app or inside Google Workspace, nothing changes this week, and that includes newer features such as Gemini’s Call for Me. Google has not said when Argon will reach either.
What you can do with the assistant you already have
You do not need Argon to see how your current tool behaves when it is out of its depth. Ask it five questions you know the answers to and five it cannot possibly know, such as your own return policy before you have shared it, or last month’s revenue. Count how many of the second five it answers anyway. That is your assistant’s personal hallucination test, and it takes ten minutes.
Then tell it, in its saved instructions, that “I don’t know” is an acceptable answer and that it should point you to where the answer lives. Give it the documents it needs, such as your price list, policies and service area, so it has less to guess about. If you are still choosing an assistant, our look at what the free Claude plan now does covers what one no-cost option includes.
The bigger signal is the direction. For two years the race between AI labs has been scored on how many questions a model gets right. Argon ties for the lead while declining to bluff, and an independent tester has put a number on it. If that becomes something labs compete on, the assistants small businesses rent will get easier to trust.
Frequently Asked Questions
What is Gemini 4 Argon?
Gemini 4 Argon is Google DeepMind’s newest frontier AI model, announced September 30, 2026. It is built for long, multi-step work such as software engineering, legal and financial analysis and cyber defense, can produce up to 1 million tokens of output, and tied OpenAI’s GPT-6 Astra on Artificial Analysis’s Intelligence Index with a score of 53.
Can a small business use Gemini 4 Argon yet?
Not yet. As of October 1, 2026, Google has released Argon only to vetted cyber defenders through its Fairwind Program. Paid API customers and Google AI Ultra subscribers are reported to be next in line, but Google has not published dates, and it has not said when Argon will reach the Gemini app or Google Workspace.
What does Gemini 4 Argon’s 15 percent hallucination rate mean?
It measures how often Argon gives a wrong answer when it does not have the right one, out of all its responses that were not fully correct. On Artificial Analysis’s AA-Omniscience test, Argon guessed wrong in 15 percent of those cases and declined or gave partial answers in the rest. GPT-6 Astra guessed wrong in 51 percent. Argon’s overall accuracy is lower, 50 percent against Astra’s 63 percent.
How much will Gemini 4 Argon cost?
Google lists an introductory price of $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 after the introductory period, with no end date announced. Cached input is 95 percent off. Argon also reasons at length before answering, and Artificial Analysis counted that reasoning among its output tokens, so real costs depend on how much it thinks per task.
Would you trade a few right answers for a lot fewer made-up ones? Tell us in the comments which jobs in your business you would hand to the model that says “I don’t know.”
