The short version: On a real small-business job, xAI’s new Grok 4.7 is now the cheapest flagship model you can rent. The month of 880 customer inquiries we priced last week for a five-person appliance repair shop costs $3.52 on Grok 4.7, against $4.40 on OpenAI’s GPT-6 Sol and $8.45 on Anthropic’s Claude Opus 5.5. That lead is 88 cents, and it rests on one setting. Grok 4.7 always thinks before it answers, xAI bills that thinking as output, the default setting is high, and about 170 extra thinking tokens per reply is all it takes to erase the difference.
So the price is real, and so is the catch. Both come straight from xAI’s own documentation, read on September 27, 2026.
How much does Grok 4.7 cost?
xAI released Grok 4.7 on September 21 and calls it its most capable model for code and knowledge work. According to xAI’s pricing page, Grok 4.7 pricing is $2 per million input tokens, $0.50 per million cached input tokens and $6 per million output tokens, for any request whose prompt stays under 200,000 tokens. Those are the same rates Grok 4.6 carried, so a business already running 4.6 moves to the newer model at no extra cost per token.
Here are the three flagship rate cards side by side, per million tokens, checked on September 27, 2026.
| Model | Input | Cached input | Output |
|---|---|---|---|
| Grok 4.7 | $2.00 | $0.50 | $6.00 |
| GPT-6 Sol | $2.00 | $0.20 | $10.00 |
| Claude Opus 5.5 | $4.00 | $0.20 | $20.00 |
Is Grok 4.7 really half the price of its rivals?
xAI’s launch announcement describes the model as “half the price of comparable models.” On the day it was published, that held up. OpenAI’s flagship that Monday was GPT-5.6 Sol at $4 input and $20 output, and Grok 4.7 came in at half the input price and under a third of the output price.
The claim had a shelf life of one day. On Tuesday OpenAI replaced that model with GPT-6 Sol at $2 and $10, which we priced against Claude Opus 5.5 on a real workload. Today, Grok 4.7 matches GPT-6 Sol on fresh input, undercuts it by 40 percent on output, and costs two and a half times more on cached input. Against Opus 5.5 the launch line still holds.
What does a real month cost on Grok 4.7?
Same shop, same job as last week: a five-person appliance repair business answering 40 inquiries a day, 22 working days a month, each with a drafted reply the owner approves before it goes out.
Each request carries about 2,000 tokens of cached standing context (price list, service area, hours, tone), about 300 tokens of customer message and about 400 tokens of reply. For the month, that is 1.76 million cached tokens, 264,000 fresh input tokens and 352,000 output tokens.
| Where the month lands | Cached input | Fresh input | Output | Total |
|---|---|---|---|---|
| Grok 4.7 | $0.88 | $0.53 | $2.11 | $3.52 |
| GPT-6 Sol | $0.35 | $0.53 | $3.52 | $4.40 |
| Claude Opus 5.5 | $0.35 | $1.06 | $7.04 | $8.45 |
Read across the Grok row. It loses to GPT-6 Sol on the cached price list by 53 cents and wins on the replies by $1.41. Replies are the expensive part of this job, so the cheaper output wins, and Grok 4.7 finishes the month 88 cents cheaper than Sol and $4.93 cheaper than Opus. Every one of those totals assumes the model writes only the 400 tokens of reply the customer sees, and that assumption is exactly where the ranking can flip.
Can you turn off reasoning in Grok 4.7?
No. xAI’s reasoning documentation says it in four words: “Reasoning cannot be disabled.” Grok 4.7 thinks before every answer. You can set how hard it thinks, at low, medium, high or xhigh, and the default is high, one step below the maximum. The same page says reasoning tokens are billed as part of your total consumption, and like every token the model generates, they count as output at $6 per million.
Now the arithmetic the launch coverage did not do. Grok 4.7’s lead over GPT-6 Sol on this job is 88 cents. Divide that by 880 replies at $6 per million output tokens and you get about 167 tokens per reply. If Grok 4.7 spends roughly 170 more thinking tokens per reply than GPT-6 Sol does, the cheaper model is no longer cheaper. Against Opus 5.5 the cushion is much wider, about 930 extra tokens per reply before the order changes.
We have not measured how much Grok 4.7 thinks at each setting on a job like this, and xAI does not publish a figure. The documentation does show which way the default leans. A drafted reply to “can you come Thursday to look at my dryer” does not need the second-hardest thinking setting a frontier lab offers.
The test is cheap and settles it for your workload, not ours. If Grok 4.7 runs your routine drafting, run one week at low effort and one at the default, then compare the bills and a sample of replies side by side. If customers cannot tell the two weeks apart, keep the cheaper one.
What happens above 200,000 tokens?
Grok 4.7 accepts prompts of up to 500,000 tokens, which invites you to paste in everything. The pricing page sets the terms in one sentence: “Requests whose prompt reaches the listed token threshold are billed at the higher rate for all tokens in the request.” All of them, not just the overflow.
Here is what that does at the line. A 199,000-token prompt costs about 40 cents of input. A 200,000-token prompt costs 80 cents, and every token of the answer bills at $12 per million instead of $6. One thousand more tokens in, and the whole request doubles.
At roughly three quarters of a word per token, 200,000 tokens is about 150,000 words: a full bid package with specifications, or a commercial lease with every exhibit attached. Near the line, splitting the job into two requests keeps both halves at the lower rate. OpenAI also lists higher long-context rates for GPT-6 Sol, so this is not an xAI quirk; xAI just states the rule more plainly.
Should a small business use Grok 4.7?
Start with the good news, because it is genuinely good. Three frontier labs now draft a whole month of customer replies for less than five dollars in model time, and the newest got there by competing on price. Every inquiry that used to wait until somebody got back off a job can get a first draft in seconds, with the owner still deciding what goes out.
Most owners will not choose a model directly. You buy the tool, and your vendor buys the tokens, which is why the median small business pays about $28 a month for AI software rather than a token bill. That makes two questions worth asking whoever supplies your AI features: which model sits underneath, and what reasoning setting it runs at. If the answer is Grok 4.7 at the default, you have found a free saving someone forgot to take.
Frequently Asked Questions
How much does Grok 4.7 cost?
Grok 4.7 costs $2 per million input tokens, $0.50 per million cached input tokens and $6 per million output tokens for prompts under 200,000 tokens, read from xAI’s pricing page on September 27, 2026. Once a prompt reaches 200,000 tokens, the entire request is billed at $4, $1 and $12. These are the same rates Grok 4.6 carried.
Is Grok 4.7 cheaper than GPT-6 Sol?
On a typical month of 880 drafted customer replies, Grok 4.7 costs about $3.52 in model time against $4.40 for GPT-6 Sol, because its output price is 40 percent lower. That lead is only 88 cents, though, and it disappears if Grok 4.7 spends about 170 more reasoning tokens per reply than GPT-6 Sol, since xAI bills reasoning as output.
Can you turn off reasoning in Grok 4.7?
No. xAI’s documentation states that reasoning cannot be disabled on its reasoning models. You can choose low, medium, high or xhigh effort, and the default is high. For routine drafting, testing the low setting for a week is the simplest way to see whether it lowers your bill without hurting the replies.
What is Grok 4.7 Fast?
Grok 4.7 Fast is a version of the model that generates output at roughly twice the speed, at double the standard price: $4 input and $12 output per million tokens below 200,000 tokens. It is available only through the Cursor coding editor and xAI’s Grok Build, so it matters to software developers rather than to most small businesses.
If you have run the same everyday job at low and high reasoning effort, on any model, how different was the bill, and could your customers tell the replies apart? We would genuinely like to know.
