The short version: Mistral Large 4 went into public preview on October 6, and the number it offers for office work is 59.9 percent on AutomationBench, a test of 657 business workflows that Zapier built from real automations in Gmail, Sheets, Slack, Salesforce and HubSpot. Mistral compares that score only with other open models, and its post does not say which of the test’s two scoring methods produced it. The two public leaderboards for that test, read today, do not list Mistral Large 4 yet. What they do show is more useful than any one model’s score: the best AI on the board completes 93 percent of the work in those tasks and still loses about a sixth of its grade, because one broken business rule, like texting a job candidate who opted out, zeroes the whole task.
The models are already good at the work. What separates a safe automation from an embarrassing one is how often they break a rule you never wrote down.
What is Mistral Large 4?
Mistral, the Paris lab that publishes many of its models as downloadable open weights, announced Mistral Large 4 as a mixture-of-experts model with 1 trillion parameters, of which 49 billion are active on any one request. It reads text and images and writes text. The preview runs through the Mistral Studio API today. Mistral says it “will release the weights by the end of the month,” but the announcement does not name the license those weights will carry. Its predecessor, Mistral Large 3, shipped under Apache 2.0, a permissive license; whether Large 4 matches it is unknown until the license text appears.
Price is where it gets attention. Mistral lists Large 4 at $1.36 per million input tokens and $4.18 per million output tokens, and its model page shows half that during the preview: $0.68 and $2.09 (read October 6, 2026). For comparison, Claude Sonnet 5.5 costs $2 and $10. At preview rates, Large 4’s input costs about a third of Sonnet’s and its output about a fifth. One caution from Artificial Analysis, which ranks it 64th of 225 models on its general intelligence index: the model is very verbose, and a model that writes more words per task gives back some of what it saves per word.
What does AutomationBench actually test?
It is the benchmark that looks most like a small business’s week. Zapier says the 657 tasks come from patterns in “2B+ monthly tasks across 3.7M companies,” across six functions: sales, marketing, operations, support, finance and HR. Each task drops the AI into a simulated business with CRM records, inbox threads, calendars and spreadsheets, gives it one trigger message, and leaves it alone. The grade is the state of those systems when the AI is done, not anyone’s opinion.
The examples Artificial Analysis publishes read like a Tuesday at any office. Send interview reminders by text, but only to candidates who agreed to texts, and not to anyone whose interview was canceled. Route a contract for signature, outside signers first, and for a deal under $500,000 the VP of Sales signs, not the CEO. Flag the grant that went over budget, and do not falsely report one that did not.
Each of those has two kinds of check. Objectives are the work: the reminders sent, the record updated. Guardrails are the rules that were true before the AI started and must still be true after: the opted-out candidate was never texted. Break a guardrail and the task scores zero, no matter how much of the work was done.
How does Mistral Large 4 compare with Claude, Gemini and GPT?
Nobody can say yet. Mistral’s post puts its 59.9 percent ahead of Kimi K3, MiMo-V2.6-Pro and DeepSeek V4 Pro, three other open models, and sets it beside no closed model. The test is scored two ways, and the two produce very different numbers for the same models. Here is how the leaders stand on each board, read October 6, 2026:
| Model | Work done | Score after rules | Tasks fully correct |
|---|---|---|---|
| Gemini 4 Argon (High) | 93% | 77.5% | 51.29% |
| Claude Sonnet 5.5 (Max) | 91% | 71.8% | 44.75% |
| Claude Opus 5.5 (Max) | 90% | 69.5% | 42.47% |
| Mistral Large 4 | Not listed | Not listed | Not listed |
The first two columns come from Artificial Analysis’s independent run, which gives partial credit for the share of work done and zero for any task with a broken rule. The last column is Zapier’s own leaderboard (version 1.0.6), which counts a task only if every check passes. Mistral’s 59.9 percent sits below the closed leaders on the first scale and above all of them on the second, so which scale it came from is the whole story. Until Mistral says, or one of the boards runs it, treat 59.9 as a claim in search of a column.
Why do the rules matter more than the work?
Do the subtraction on the top three. Gemini 4 Argon completes 93 percent of the work and scores 77.5, so 15.5 points of work it actually did are wiped out by broken rules. Sonnet 5.5 loses 19.2 points the same way, and Opus 5.5 loses 20.5. For the best models in the world, unfinished work is now the smaller problem. The bigger one is the confident, finished task that texted the wrong person or let the wrong executive sign.
That changes what to ask of any AI that will act inside your business, whatever model runs it. Three habits follow:
Write your rules where the AI can read them. The test’s guardrails are things a new hire would learn by asking: who opted out, who can approve what, which accounts are off limits. An AI cannot ask. If the rule lives only in your head, it is not a rule to the software.
Keep a person on the actions that cannot be taken back. Texts, emails, payments and signatures are where a broken rule leaves the building. Drafting, sorting and tidying your own records are where a cheaper model saves the most hours with the least exposure. We walked through how approval levels work in practice in Meta’s Muse for Small Business.
Test on your own week, not the leaderboard. Take twenty tasks your team actually did last month, run them through the tool you are considering, and count two things separately: work done, and rules broken. The second number is the one to buy on.
There is genuinely good news in all of this. When Artificial Analysis first ran this test in July, the best model completed 73 percent of the objectives. Three months later the leader completes 93 percent. And Mistral says a model at a fraction of the leaders’ price will be downloadable by the end of the month, which, as we wrote when Mistral raised €3 billion last month, keeps every vendor’s prices honest. If you are new to putting AI to work, start with our practical guide to AI for small business, and see how Gemini 4 Argon handles questions it cannot answer.
Frequently Asked Questions
How much does Mistral Large 4 cost?
Mistral lists it at $1.36 per million input tokens and $4.18 per million output tokens, and its model page showed half those rates, $0.68 and $2.09, during the public preview that began October 6, 2026. Claude Sonnet 5.5 costs $2 and $10, so at preview rates Mistral Large 4 is roughly a third of the input price and a fifth of the output price.
Is Mistral Large 4 open source?
Not yet. Mistral says it will release the weights by the end of October 2026, but the announcement does not name the license. Its predecessor, Mistral Large 3, used Apache 2.0, a permissive license that allows commercial use. Check the license text when the weights arrive before assuming the same terms.
What is AutomationBench?
It is a test of 657 business workflows built by Zapier from real automation patterns, covering sales, marketing, operations, support, finance and HR in simulated apps such as Gmail, Google Sheets, Slack, Salesforce and HubSpot. It checks both the work done and whether the AI broke any business rule along the way, and a broken rule zeroes the task.
Which AI model is best at small business office work?
On both public AutomationBench leaderboards read October 6, 2026, Gemini 4 Argon leads, followed by Claude Sonnet 5.5 and Claude Opus 5.5. Mistral Large 4 is not on either board yet. The more useful test is your own: run twenty real tasks from last month and count rules broken separately from work completed.
What is one rule in your business that a new hire always gets wrong in the first week? It is probably the first one to write down for your AI, too.
