The short version: AI document processing for a small business used to need a business case. It does not any more. On August 21 DeepSeek released a multimodal model that bills each image at up to 384 tokens, which on its published rates works out to roughly five hundredths of a cent to read a receipt. Amazon’s purpose-built receipt service charges a full cent per page. The capability was never the hard part; the price of using it casually was. That has now gone, and what is left standing is the awkward question of where your photos actually go and who checks the answer.
Here is the release itself. DeepSeek shipped deepseek-v4-flash-vision-exp, an experimental model that takes mixed text and image input by base64, external URL, or its Files API. The company says it matches the text-only V4-Flash on reasoning, agents, and world knowledge, and that on multimodal agent benchmarks it makes what it calls a major leap, bringing performance close to Opus-4.8. The line that matters for a business owner is buried in the billing note: images are tokenized at up to 384 tokens each, charged at ordinary V4-Flash rates.
What does AI document processing cost now?
Run the numbers yourself, because they are the whole story. DeepSeek’s published pricing is 44 cents per million input tokens and $1.32 per million output tokens at peak, half that off-peak. Peak runs 01:00 to 10:00 UTC, which is evening and overnight in the United States, so an American business doing this during its own working day pays the lower rate by default.
Take the expensive end anyway. One photo at 384 input tokens costs about $0.00017. Ask for a structured answer back, say 250 tokens of vendor, date, line items, and total, and that costs about $0.00033. Call it five hundredths of a cent per receipt, all in. A thousand receipts costs about fifty cents.
For comparison, Amazon’s Textract Analyze Expense API, built specifically to pull data off invoices and receipts, lists at $10.00 per thousand pages. That is a cent a page, twenty times more, and it does one thing. The general model reads the receipt, notices the date is smudged, and tells you so in the same call.
This is not a DeepSeek phenomenon and you should not read it as an endorsement of one vendor. Google cut Gemini 3.7 Flash to 75 cents per million input tokens through the end of December, and it reads images too. The direction is industry-wide. DeepSeek just put the clearest number on it.
Why does cheap reading matter more than cheap writing?
Most small businesses met AI as a writing tool, and writing was never their bottleneck. Paperwork was. Research from American Express and Small Business Saturday UK, surveying a thousand UK business owners, found they spend an average of eleven hours a week on administrative or finance tasks, roughly twice what they spend on sales and business development. It is a UK sample, so treat the exact figure as directional rather than as your number. The shape of it will be familiar to anyone reading this in Fresno or anywhere else.
Almost none of those eleven hours are spent thinking. They are spent transcribing: a delivery note into a spreadsheet, a photographed meter reading into a log, a stack of fuel receipts into accounting software, a handwritten job ticket into the invoicing system. The information already exists. It is simply trapped in an image, and getting it out cost either an employee’s afternoon or a per-page fee that made you think twice.
Here is the part worth sitting with. When a capability gets cheap enough, the cost of deciding whether to use it exceeds the cost of using it. At five hundredths of a cent, evaluating whether a photo is worth processing is more expensive than processing it. That inverts how you plan. You stop asking which documents justify automation and start pointing it at everything, because the losing bets cost nothing.
What does this actually change about the work?
It moves a job rather than removing one. The person who was retyping receipts was never being paid for their typing speed; they were being paid because someone had to look at a smudged total and decide what it said. That decision is still theirs. What changes is that they now start from a filled-in form instead of a blank one, and they spend their day on the twenty items the model flagged as uncertain rather than the eight hundred it read cleanly.
That is also where the real cost hides, and why the near-zero price tag is only half a plan. Reading was never the expensive part of a photo-to-data workflow. Exceptions were. The receipt with two totals, the delivery note signed by someone who no longer works there, the invoice in a currency you do not normally handle. Cheap reading floods your exception queue faster than before, and if nobody owns that queue, you have simply built a faster way to generate uncertainty. Decide who checks the flags before you switch anything on. Our guide to what to automate first covers the sequencing.
Practical places to point it, in rough order of how forgiving they are: fuel and materials receipts, packing slips and delivery notes, meter and equipment readings, business cards from a trade show, photographs of a job site before and after, handwritten counts from a stock take. Places to be slow about: anything where a misread number becomes a legal filing, a payroll figure, or a quote a customer will hold you to. Those need a human signature, not a confidence score.
The catch the release note does not lead with
Two of them, actually. The first is that this vision model is experimental and API-only. The text-only base, DeepSeek-V4-Flash-0731, was open-sourced under the MIT licence at the end of July, which means it can run on hardware you control. The vision variant was not. So the eyes stay in someone else’s building, and that building is in China. For photographs of a driveway that is a shrug. For photographs of customer records, medical intake forms, or anything with a social security number on it, it is a decision you should make deliberately rather than by default. If keeping documents in-house is the requirement, the open-weight models that run offline are the honest starting point, even where they are weaker.
The second catch is the word experimental, sitting right there in the model name. Do not wire a business process directly to a string ending in -exp. Whatever tool or script you build, build it so the model is a swappable part, because prices and model names in this market have the shelf life of milk. We wrote recently about reading the expiry date next to the price rather than the price itself, and that applies here with force. The interesting claim is not that DeepSeek is cheap this month. It is that reading an image has stopped being a line item anywhere, and that is unlikely to reverse.
Frequently Asked Questions
What is AI document processing, in plain terms?
It means handing a computer a photograph or scan of a document and getting structured data back, so a picture of a receipt becomes a vendor name, a date, a total, and a list of line items you can put in a spreadsheet. Older systems did this with optical character recognition, which read the characters but did not understand them and broke on anything unusual. Current AI models read the image and reason about it, so they can tell you that a total looks wrong or that a field is illegible instead of silently guessing.
Is it safe to send my business paperwork to a model like this?
It depends entirely on what is in the document, and the honest answer is that most owners have not thought about it. Sending photographs of fuel receipts or job sites to an overseas API is a low-stakes decision. Sending customer records, employee information, medical forms, or anything covered by a contract or a privacy regulation is a different decision, and DeepSeek’s vision model is API-only and hosted in China. Read the terms, decide per document type rather than in general, and use a self-hosted open-weight model when the answer needs to be no.
Do I need a developer to use any of this?
To call this specific model directly, yes, because it is a raw API and not an app. But you almost certainly do not need to. The pricing shift described here flows into the tools you already pay for within months, and your accounting software, field service app, or inventory system very likely already has a scan-a-receipt feature that got quietly better and cheaper for the same reason. Check what you already own before you commission anything custom.
Will this replace my bookkeeper or admin person?
It replaces the transcription, which is the part of their job nobody enjoys and nobody was really paying for. The judgment stays: deciding what a smudged figure says, spotting the invoice that does not match the delivery, knowing which supplier always bills late. What actually happens in practice is that the same person handles several times the volume and spends their time on the exceptions, which is both more useful to you and considerably less dull for them. The risk worth watching is the opposite one, where cheap automation produces more flagged items than anyone has time to check.
One thing we are genuinely curious about: if reading a photograph costs nothing, what is the pile of paper in your business that you have never bothered to digitise because it was never worth the effort? Tell us what is in it.
