Two of the three largest AI providers cut their prices this month. Both cuts have an expiry date printed on them, and one of those dates falls in the first week of January.
The short version: OpenAI now charges $4 per million input tokens and $20 per million output tokens for GPT-5.6 Sol, and its own documentation describes that as promotional pricing available “at least through November 21, 2026.” Google launched Gemini 3.7 Flash on August 13 at $0.75 and $3.75 per million tokens, half its standard rate, which returns to $1.50 and $7.50 on January 1, 2027. The underlying cost of running AI for small business tools keeps falling, and that trend is real. But the direct saving to you is probably measured in cents, because you do not buy tokens. You buy software. The discount landed on your vendor’s cost line, not yours, and that is where the interesting question lives.
What actually changed in August
On August 21, OpenAI reduced the API price of GPT-5.6 Sol, its frontier tier. Input fell to $4 per million tokens and output to $20, from a previously reported $5 and $30. The output side is the bigger move, down by a third. The company’s published pricing applies the new rate to the pay-as-you-go API, Codex credits, and eligible ChatGPT Work plans. Consumer Pro and Plus subscriptions were not part of it.
Eight days earlier, Google had shipped Gemini 3.7 Flash. VentureBeat reported the introductory rate at $0.75 input and $3.75 output per million tokens, with context caching at $0.075, all of it doubling on January 1, 2027. Google described the model as its “most intelligent workhorse model yet for coding and agents.” The benchmark gains it published were not marginal: production code quality moved from 34.4 percent to 43.6 percent on FrontierCode 1.1 Main, and long-horizon software engineering tasks from 49.0 percent to 65.3 percent on DeepSWE v1.1.
So the same fortnight brought a cheaper frontier model and a much cheaper workhorse model that got noticeably better at the long, multi-step tasks that agents actually run. Both discounts are temporary and both companies said so in writing.
Why doesn’t a 20 percent price cut show up on my bill?
Because almost no small business buys tokens directly. You buy a scheduling tool, an inbox assistant, a bookkeeping feature, a website builder. Those products buy tokens, mark them up, and sell you an outcome. A price cut at the model layer reaches you second-hand, if it reaches you at all.
The arithmetic is worth doing once, because it reframes the whole thing. Suppose a tool drafts 500 customer email replies a month for you. A reply might read 2,000 tokens of context and write 400 tokens back. At Gemini 3.7 Flash’s current rate that is about a third of a cent per reply, or roughly $1.50 a month in raw model cost. When the introductory pricing ends and the rate doubles, it becomes about $3.00.
Now compare that to what you actually pay. The median small business spending on AI software runs around $28 a month, a figure we broke down in our guide to what to actually pay for in AI software for small business. The tokens are a rounding error inside your subscription. Nearly everything you pay is the product wrapped around the model: the interface, the integrations, the support, the margin.
That is the non-obvious part. Your bill did not move because your bill was never mostly tokens. Your vendor’s cost of goods just fell, and their exposure in January just got scheduled.
The risk nobody is describing
When a vendor’s input costs drop for a fixed window and then snap back, they have three options. They can absorb the increase and keep your price flat. They can raise your price. Or they can quietly move the feature to a cheaper model and keep both the price and the margin.
The third one is the one to watch, because it is invisible from your side. There is no email announcing it. The feature still works. It just gets a little worse at the hard cases, which are exactly the cases you adopted it for. If your quoting assistant starts fumbling the unusual jobs in February, a model swap is at least as likely an explanation as anything you changed.
This is not a reason to distrust vendors, most of whom will handle it sensibly. It is a reason to ask a question you can ask cheaply right now: which model is behind this feature, and will that change when introductory pricing ends? A vendor with a good answer will give you one. A vendor who cannot say what is under the hood is telling you something too.
What falling inference prices actually make possible
Set the expiry dates aside for a moment, because the direction of travel matters more than any one promotion. Work that was too expensive to automate two years ago is now trivially cheap to run. Reading every receipt. Following up on every missed call. Drafting a first-pass reply to every enquiry that arrives at 9pm.
None of that is work anyone was doing. It is work that was falling through the cracks because there were not enough hours in the week to catch it. We made the same point when DeepSeek collapsed the price of AI document processing to a fraction of a cent per page. Cheap inference does not take a job off your team. It picks up the things your team never got to.
The honest caveat is that cost was rarely the real blocker anyway. Setup time, trust, and knowing which task to point it at are harder problems than price, and none of them got cheaper this month. A model that costs half as much and still needs a fortnight of your attention to configure is not half as easy to adopt.
What should a small business do before the discounts expire?
Spend the cheap window on learning, not on locking in. Concretely, that means three things.
Run the experiment you have been putting off. If there is a task you were unsure was worth automating, the next four months are the cheapest they will be to find out. Test it, measure whether it saved real hours, and keep the result.
Ask your vendors the model question. One email. Which model powers this, and does your pricing change in January? File the answers.
Be careful about annual contracts priced off temporary input costs. A vendor offering an unusually aggressive twelve-month deal right now is either confident about their margins or hoping you will not notice the renewal. Both are worth a conversation before you sign. The same logic applies to the way speed and intelligence are now sold as separate line items, a pricing structure that rewards owners who read the invoice closely.
The broader picture is straightforwardly good. AI capability is getting cheaper faster than most people planned for, and the businesses that benefit are the ones with a clear idea of which hour of their week they want back. That part does not expire in January.
Frequently Asked Questions
Did AI actually get cheaper in August 2026?
Yes, at the model layer. OpenAI cut GPT-5.6 Sol to $4 per million input tokens and $20 per million output tokens on August 21, down from a reported $5 and $30. Google launched Gemini 3.7 Flash on August 13 at $0.75 and $3.75, half its standard rate. Both companies published these as promotional or introductory prices with stated end dates rather than permanent reductions.
Will my AI subscription price go down because of this?
Probably not, and that is normal rather than a sign of gouging. Token costs are a small fraction of what you pay for AI software; the rest is the product built around the model, including the interface, integrations, and support. A model price cut improves your vendor’s margin first, and only reaches customers later, if competition pushes it through.
What happens when the introductory pricing expires?
OpenAI’s promotional rate for GPT-5.6 Sol runs at least through November 21, 2026, and Gemini 3.7 Flash returns to $1.50 and $7.50 per million tokens on January 1, 2027. Vendors facing that increase can absorb it, raise prices, or move features to a cheaper model. The third option is the hardest to detect, since the feature keeps working but may handle difficult cases less well.
Should I wait for prices to fall further before adopting AI?
Waiting on price is usually the wrong optimisation for a small business, because cost is rarely the real barrier. Setup time, deciding which task to automate, and building trust in the output take far longer than the money involved, and none of those get cheaper by waiting. Running a small, cheap experiment now tells you something waiting cannot.
Have you asked a software vendor which model sits behind one of their AI features, and did you get a straight answer? We would like to hear how that conversation went.
