Introduction
Why test models on payments and fintech tasks
There are already hundreds of AI models to choose from, and new ones arrive every few weeks as developers compete on capability and price. Any business building on AI should expect many more options within the next year.
Choosing a model is a question of AI economics. The model needs enough capability for the task in hand, at a price that works at the business's volumes, and it has to clear the business's other requirements, such as where the model is hosted, where the data goes and who the supplier is. At scale, paying for an advanced model on a basic task has a material effect on cost.
Public benchmarks and leaderboards test AI models in general, but few of them look at the AI economics of fintech tasks. This paper takes one payments task, reading card acquiring statements, measures accuracy in money, and weighs it against price and the other points that decide the choice, such as reliability, hosting and where the data goes.
This paper sets out the findings. The model-by-model results and the full test details are in a technical annex at the end of the online version of this paper, at scalepointpartners.com/labs/ai-economics-card-statements.
Executive summary
With good instructions, a model at half a cent matched one at 25 cents
Objective
The test set out to measure the AI economics of one fintech task: how accurately each model understands the contents of a document, and what that accuracy costs. I gave 34 models one task: read UK card acquiring statements, find every fee and spot any errors. The test weighed accuracy against price and the other points that decide the choice, such as reliability, hosting, developer location, where the data goes and what it might be used for.
Hypothesis
I went in with one view to test and one open question. The view, commonly held, is that the leading US frontier models are clearly the best. The question was how much difference better instructions would make, and whether they would let a cheaper model match the most expensive frontier model.
Test method
- I built synthetic statements from two invented acquirers with different layouts, since real layouts differ from one acquirer to the next, and filled them with traps and errors, ranging from easy to hard, such as a fee shown only in a footnote and pricing labelled pass-through that is really one flat rate. I knew the correct amount and type of every fee, so I could measure the cost of each model's mistakes.
- Each model received the text of the statement, taken from the PDF, with OCR used to turn the scanned pages into text. The models did not see the page itself, so this test measured how well they understood and classified a statement's contents. Reading the page image directly is a different skill, which this test did not cover (Section 10).
- I ran every model through OpenRouter, which connects to more than 400 models through one API, and picked 34 of them from the main developers in the US, Europe and China.
- In the first round, all 34 models read 16 statements with a plain prompt: nine basic rules and the names of the 18 fee types, with no definitions and no payments guidance.
- In the second round, seven finalists read 22 statements, adding six harder ones, at three prompt levels: plain, middle (plus a definition of each fee type) and expert (plus a briefing of about 1,500 words on UK statements).
- Every answer was scored against the correct answer for each fee, in pounds sterling. Section 2 sets out the method in brief, and the technical annex at the end of the online version has every model's results and the full test details.
Results
Round 1: all 34 models, plain prompt, 16 statements
1. Which model is best on its own? With the plain prompt, Claude Opus 5.5 and Grok 4.7 got almost no fees wrong, £4 across all 16 statements. Gemini 3.8 Flash came next at 0.37% of fees wrong, then GLM 5.3, Kimi K3 and Mistral Large 3 at about 0.5%, or about £40 on a typical statement with £8,000 of fees. GPT-5.5 got 0.94% wrong, about £75 on the same statement. See Section 3.
2. Which models were not usable, and why? 25 of the 34 models cleared a usable bar set before the test. The nine that failed returned broken answers, invented fee lines, or missed traps. See Section 3 and annex A2.
3. Where did the models go wrong? Every usable model found the fee lines and read the amounts. The mistakes came from putting fees under the wrong type, mostly surcharges on non-UK and commercial cards filed as penalties. See Section 3.
Round 2: seven finalists, three prompt levels, 22 statements
4. Can a cheaper model get the fees right? With the expert prompt, all seven finalists got less than 0.1% of fees wrong, from DeepSeek V4.1 Flash at $0.005 a statement to GPT-5.5 at $0.25. See Section 4.
5. Did that hold on the hardest tasks? On fee accuracy it did. Spotting the two faulty statements was harder. Claude Opus 5.5 and Grok 4.7 flagged both at every prompt level, the cheaper finalists needed the expert prompt, and Mistral Large 3 and GPT-5.5 missed at least one even with it. See Section 5.
6. What does it cost? With the expert prompt, the most expensive finalist cost 50 times as much per statement as the cheapest, for the same accuracy. However, even the most expensive cost only about 25 cents a statement, so the gap only adds up for a business that reads statements in bulk. See Section 6.
7. What still separates the finalists? Reliability, speed, where the data goes and who the supplier is. See Section 7.
Conclusions
The main finding is that once the prompt was right, all seven finalists were about equally accurate. The choice of model then comes down to price and other points such as reliability, hosting and where the data goes.
The chart below shows that finding for the seven round 2 finalists. Each arrow runs from a model's result with the plain prompt (grey dot) to its result with the expert prompt (coloured dot). Further right means fewer fees wrong, and higher means cheaper per statement. With the expert prompt, all seven ended up to the right of the 0.1% line, so the choice between them comes down to height: DeepSeek V4.1 Flash at half a cent a statement near the top, and GPT-5.5 at 25 cents at the bottom. Claude Opus 5.5 and Grok 4.7 barely move, because they were already accurate with the plain prompt.
Seven finalists, round 2. Each arrow runs from the plain prompt to the expert prompt. The fees-wrong scale is stretched to the right of the 0.1% line.
The most surprising result was GPT-5.5, OpenAI's flagship. With the plain prompt it got 0.94% of fees wrong, about £75 on a typical statement, while its frontier peers Claude Opus 5.5 and Grok 4.7 got almost none wrong. It also finished behind cheaper models from China and France, such as GLM 5.3 and Mistral Large 3. It did catch up with the expert prompt, but just like the cheaper models it needed that help, and even then it failed to flag the faulty subtotal. GPT-5.5's showing is all the more surprising because, at $0.25 a statement, it was the most expensive finalist. Hence the common view that the leading US frontier models are clearly the best held for Opus and Grok, and did not hold for GPT-5.5. What the higher price of Claude Opus 5.5 and Grok 4.7 bought was good results with basic instructions, the hardest judgement calls caught without being told what to look for, and no failed answers.
AI is moving fast, and over the next 12 months we can expect to see more and more companies running several models side by side, each matched to the task in hand. This one test challenged the commonly held view that the leading US frontier models all outperform other models by a long way. It also showed that better instructions can let a cheaper model match the most expensive frontier model. Hence any management team using AI should check how its own model choices hold up, starting with the questions below. Section 9 sets out how I expect the market to change over the next 12 months, what good practice will look like by then, and where routers fit. Section 10 sets out the questions this test leaves open, including a harder task, reading page images directly and a whole workflow run by an AI agent.
Questions for management
The findings point to five questions for any management team using AI models across its business:
- Which model does each of our AI use cases run on, and when did we last test that choice against the alternatives?
- Do we measure accuracy in money, on a test set built from our own data with known answers?
- How much of our own knowledge is written into our prompts, and who owns and maintains them?
- Where does each model run, under what retention terms, and which companies in the chain, including any router, see our data? Section 7 sets out where each finalist ran in this test.
- Which of our tasks need an advanced model, and what would a cheaper model that passes the same test save us at our volumes?
Part 1, The test
1Objective and hypothesis
The test set out to measure the AI economics of one fintech task: how accurately each model understands the contents of a document, and what that accuracy costs. Card acquiring statements suit that question. Every fee on a statement is a fixed amount worked out from the acquirer's price list, so a model's figure for it is either right or wrong, to the penny.
I went in with one view to test and one open question:
- The leading US frontier models are clearly the best. This is a commonly held view. I tested it by giving every model the same plain prompt, so that each model had to work on its own.
- How much difference do better instructions make? I tested this by giving the finalists two richer prompts and comparing their accuracy and cost with the plain prompt.
Alongside accuracy and price, the test recorded the other points that decide the choice, such as reliability, hosting, developer location and where the data goes. Two of these can rule a model out before price comes into it. The first is where the developer is based. Some Chinese developers face sanctions or government bans in the US and Europe, and a regulated firm's own supplier policy may rule them out even when the model runs on a Western host. The second is zero data retention. Under many standard terms, the company running the model keeps a copy of each request for a period, for example to check for abuse. Zero data retention means it keeps no copy once it has replied, which for a regulated firm is the safer starting point before sending real customer documents.
2Method
The statements and the calculated answers
I used AI to help write a program that built each statement from scratch. It made up a month of card sales for an invented merchant, applied an invented acquirer's price list to work out every fee, and produced the statement as a PDF file, with the scanned ones produced as page images. The same program wrote out the calculated answers: the list of every fee on that statement with its amount, currency and fee type. No real merchant or acquirer data was used, and because the statement and its answers come from the same figures, every fee has a known right answer. The statements come from two invented acquirers, one with an interchange-plus-plus layout modelled on the large UK acquirers and one with simple blended pricing. The typical statement is for a merchant taking about £1 million a month on cards and paying about £8,000 in fees. Every page is marked as a synthetic test document, so the files can't be mistaken for real statements if they are shared. Annex A3 notes the small risk that this marking changed how a model read them.
| Tier | Statements | What it contains |
|---|---|---|
| Standard | 12 | Interchange-plus-plus pricing with 30 to 40 fee lines |
| Traps | 4 | A fee only in a footnote, fees shown twice, fees taken out of the settlement, charges in two currencies |
| Hard | 6 | Three outlets on one statement, scanned pages, a subtotal that doesn't add up, pass-through pricing that is really one flat rate, a currency conversion fee |
For every statement the model had to find every fee line, read its amount and currency, put it under the right fee type and flag anything wrong with the statement itself.
The three prompt levels
| Level | What it contains |
|---|---|
| Plain | Nine rules, such as record each fee once, keep currencies separate and flag problems with the statement, plus the names of the 18 fee types |
| Middle | The plain prompt, plus a one-line definition of each fee type |
| Expert | The middle prompt, plus a briefing of about 1,500 words on how UK acquirers lay out statements, where fees hide and how to classify the awkward ones, much like the briefing you would give a new analyst |
The plain prompt answers "which model is best on its own?" The middle and expert prompts answer the open question: how much difference better instructions make, and whether they let a cheaper model match the most expensive frontier model. A business running statements in production would give the model detailed instructions, so the expert prompt is in my mind the realistic buying question. I wrote the expert briefing myself with an AI drafting tool, without access to the test statements, so it could not be tuned to the answers.
Choosing the 34 models
OpenRouter offers more than 400 models. I picked 34 to cover the current flagship and a cheaper model from each of the main developers in the US, Canada, Europe and China, with both open and closed weights, plus some older versions still in wide use. Testing all 400 would have cost several times my $50 budget, and many of them are older versions or variants of the same model. Each statement went through OpenRouter to a company hosting the model, such as Microsoft Azure for GPT-5.5, the model behind ChatGPT. The test called every model through its API, the way a business builds a model into its own systems, and didn't use any chat apps. Microsoft Copilot isn't in the test. It is an assistant built into Microsoft's products, and it can't be sent a statement through OpenRouter. Copilot is largely built on OpenAI's GPT models, so GPT-5.5 is the closest guide here, but Microsoft adds its own instructions and settings, so Copilot's results could differ. The same applies to the ChatGPT app compared with GPT-5.5.
The two rounds
- Round 1. All 34 models read 16 statements, the Standard and Traps tiers, once each with the plain prompt.
- Making the cut. A model first had to clear a usable bar, set before the test: its fee lines had to add up to the true total on at least 95% of the Standard statements, it could invent no fee lines, and it had to catch at least three of the four traps. 25 models cleared it. From those I picked six across the range of prices: the two most accurate at any price (Claude Opus 5.5 and Grok 4.7), three accurate models at the low end (Gemini 3.8 Flash, Mistral Large 3 and GLM 5.3) and one of the very cheapest (DeepSeek V4.1 Flash). The choice within each price band was my judgement. I added GPT-5.5 afterwards as a seventh finalist, because it is the best-known OpenAI model and readers would look for it.
- Round 2. The seven finalists read 22 statements, the same 16 plus the six Hard ones, at all three prompt levels. At the plain prompt the first-round answers to the 16 repeated statements were reused.
Sending and scoring
Every call went through OpenRouter[1] with three settings: the model received only the text of the statement, never the page itself, with the text extracted from the PDF once and the scanned pages turned into text by Tesseract, a standard optical character recognition (OCR) tool, so every model got exactly the same input; OpenRouter only used hosts that keep no copy of the data (zero data retention); and temperature was set to zero, so the same statement gets the same answer as far as the model allows.
The headline measure is fees wrong. Take a statement with £5,000 of fees. If a model misses a £45 fee and counts a £60 fee twice, it has £105 of fees wrong, or 2.1%. A fee put under the wrong type or read at the wrong amount counts the same way. Fees wrong is the total value of those mistakes as a share of the true fees on the statement, averaged across the statements.
I also ran a template parser without AI, the way statement reading was automated before AI models, as a baseline. The whole test cost about $62 in model fees.
The technical annex, at the end of the online version of this paper, has the fee types, the nine traps, the scoring calls, the calls and costs, and the test's limitations (annex A3).
A note on terms. "AI model" here means a large language model, the kind of model behind chat assistants. A model has "open weights" when anyone can download it and run it on their own servers. Models are priced per "token", roughly three-quarters of a word. Interchange-plus-plus (IC++) pricing shows the card issuer's interchange, the card scheme's fees and the acquirer's own charge separately. Blended pricing rolls them into one rate. "Fine-tuning" means training a model further on a business's own examples, which changes the model itself. A model with open weights can be fine-tuned by anyone who downloads it. A closed model can be fine-tuned only where its developer offers it, on the developer's terms.
Part 2, The results
Section 3 covers round 1, where all 34 models read 16 statements with the plain prompt. Sections 4 to 7 cover round 2, where the seven finalists read 22 statements at three prompt levels.
3Round 1: Which model is best on its own?
With the plain prompt in the first round, Claude Opus 5.5 and Grok 4.7 got almost no fees wrong, £4 across the 16 statements. Gemini 3.8 Flash was third at 0.37%.
Higher and further right is better. 16 Standard and Traps statements, one run each. Every model is listed in annex A1. Off the chart above 10%: Command A Plus (63%), Nova Pro (143%).
- After Opus, Grok and Gemini came NVIDIA's Nemotron 3 Ultra and GLM 5.3 at 0.48%, Kimi K3 at 0.49% and Mistral Large 3 at 0.56%. On a typical statement with £8,000 of fees, that is between £38 and £45 put in the wrong place.
- Other US models did worse than the leaders from China and France, such as GPT-5.5 at 0.94% (about £75 on a typical statement) and the previous Claude Opus 4.8 at 2.16% (about £173).
- Price was a poor guide to accuracy. Mistral Medium 3.5 cost $0.17 a statement and got 1.94% wrong, while Mistral Large 3 from the same company cost $0.012 and got 0.56% wrong.
- Nine of the 34 models failed the usable bar, mostly older or smaller models from developers whose newer models did clear it. They returned broken answers, invented fee lines or missed traps. Annex A2 lists them.
Where the models went wrong. Every usable model found at least 99% of the fee lines and read almost every amount correctly. The mistakes came from putting a fee under the wrong type, which is a question of payments knowledge. For example, GPT-5.5 put £979 in the wrong place across the 16 statements, which carried about £128,000 of fees. £676 of that was non-UK card surcharges filed as penalties, and £300 was two fees taken out of the settlement recorded as credits. A payments analyst learns these calls on the job, so I wrote them into the expert prompt.
The baseline without AI. The template parser got every fee right on the 12 Standard statements, at no cost. It missed two of the four traps, and on the six Hard statements it got 34% of fees wrong, mostly on the scans. It was written knowing both layouts, so this is its best case. An acquirer it has not seen would need a new template.
4Round 2: Can a cheaper model get the fees right?
With the expert prompt, all seven finalists got less than 0.1% of fees wrong. The cheaper models needed the expert prompt to get there, and so did GPT-5.5. Opus and Grok did almost as well without it.
22 Standard, Traps and Hard statements. Cost is per statement with the expert prompt, from actual usage.
| Model | Plain | Middle | Expert | Cost per statement, expert |
|---|---|---|---|---|
| Claude Opus 5.5 | 0.04% | 0.03% | 0.03% | $0.187 |
| Grok 4.7 | 0.04% | 0.11% | 0.03% | $0.049 |
| Gemini 3.8 Flash | 0.31% | 0.12% | 0.06% | $0.029 |
| Mistral Large 3 | 0.56% | 0.19% | 0.04% | $0.013 |
| GLM 5.3 | 0.61% | 0.27% | 0.03% | $0.008 |
| GPT-5.5 | 0.86% | 0.07% | 0.03% | $0.246 |
| DeepSeek V4.1 Flash | 0.99% | 2.31% | 0.07% | $0.005 |
- With the plain prompt, the order from round 1 held, and the four cheaper finalists got between 8 and 25 times as many fees wrong as Opus.
- The middle prompt cut that by more than half for Gemini 3.8 Flash, Mistral Large 3 and GLM 5.3, and took GPT-5.5 from 0.86% to 0.07%. The expert prompt brought all seven under 0.1%.
- DeepSeek V4.1 Flash did worse with the middle prompt because of one answer: on one Hard statement it kept writing its working-out, hit its length limit and returned nothing, which counts as every fee missed.
What the prompt levels show:
- The prompt carried the payments knowledge. I wrote into the briefing, for example, that a surcharge on commercial cards is charged because of the card, which puts it in a different category from a penalty for keyed or late transactions. With that written down, the other models classified the awkward lines as well as Opus.
- The briefing belongs to the business. It works with any of the seven models, so it stays with the business if it changes model or provider.
- On scans, the OCR set the limit. On one scanned statement the OCR read £55.17 as £56.17 and an £86.99 credit as £26.99. Every model reported what the text said, so none could get those two lines right. That accounts for the 0.03% that no model beat.
5Round 2: Did that hold on the hardest tasks?
On fee accuracy, the harder statements made little difference. With the expert prompt, all seven finalists got between 0.12% and 0.17% of fees wrong on the six Hard statements, and most of that was the two misread scanned figures.
The hardest task was spotting the two faulty statements among the Hard ones: a subtotal that doesn't add up, and a statement that claims to pass interchange through at cost but charges one flat rate. To catch either, the model has to check figures across many rows.
- Claude Opus 5.5 and Grok 4.7 flagged both at every prompt level.
- Gemini 3.8 Flash, GLM 5.3 and DeepSeek V4.1 Flash mostly missed the fake pass-through with the middle prompt, and flagged both every time with the expert prompt.
- Mistral Large 3 spotted the faulty subtotal, with the exact figures, but filed it under "other" in place of the label the prompt asked for. The score counts only the right label, because any automated routing to a person works from the label.
- GPT-5.5 never flagged the faulty subtotal, at any level, and caught the fake pass-through pricing only with the expert prompt.
Annex A4 has the run-by-run results.
An example: fake pass-through pricing
One of the Hard statements is for an invented pub group, Juniper Lane Taverns. The front page says "Pricing: Interchange Plus Plus", and the charges summary lists "Interchange (passed through at cost)" of £1,492.57. These are some of the rows from the detail page:
| Card type on the statement | Sales (£) | Interchange rate shown | What interchange should look like |
|---|---|---|---|
| Visa consumer debit, UK | 80,748.52 | 0.7100% | Capped at 0.2% |
| Mastercard consumer credit, UK | 29,746.00 | 0.7100% | Capped at 0.3% |
| Visa commercial credit, UK | 6,490.75 | 0.7099% | Not capped, so usually higher than consumer cards |
| Visa consumer credit, international | 3,153.36 | 0.7100% | Usually higher than UK cards |
Every row shows interchange at about 0.71%, whatever the card. Real interchange changes with the card type, so the same rate on every row means the acquirer is charging one blended rate and calling it pass-through. On the UK consumer debit row alone, interchange at the 0.2% cap would have been about £161, against the £573 shown. On this invented statement, the merchant is paying a blended rate while being told its costs are passed through at cost.
The trap is deliberately blatant. It tests whether a model checks the figures or accepts what the statement says about itself.
6Round 2: What does it cost?
With the expert prompt, reading a statement cost between $0.005 with DeepSeek V4.1 Flash and $0.246 with GPT-5.5, for the same accuracy. Gemini 3.8 Flash came out as the best value at $0.029.
| Model | Fees wrong, expert | Cost per statement | Cost compared with the cheapest |
|---|---|---|---|
| DeepSeek V4.1 Flash | 0.07% | $0.005 | 1× |
| GLM 5.3 | 0.03% | $0.008 | 1.6× |
| Mistral Large 3 | 0.04% | $0.013 | 2.7× |
| Gemini 3.8 Flash | 0.06% | $0.029 | 6× |
| Grok 4.7 | 0.03% | $0.049 | 10× |
| Claude Opus 5.5 | 0.03% | $0.187 | 38× |
| GPT-5.5 | 0.03% | $0.246 | 50× |
The difference comes from two things: first, each model's price per token, and second, how many tokens it uses when writing its answer. Some models "think" before they answer. They write out working notes first, which the user doesn't normally see but still pays for, so they use more tokens. For example, with the expert prompt DeepSeek V4.1 Flash wrote about 6,200 tokens of working-out per statement and Claude Opus 5.5 about 500. DeepSeek was still the cheapest, because its price per token is so low. Annex A5 discusses this in more detail and sets out what drove each model's cost.
Cost at volume
For a merchant reading its own statement once a month, the model's price hardly matters: even GPT-5.5, the most expensive finalist, would cost about $3 a year. The cost only adds up for a business that reads statements in bulk, such as an acquirer. At 10,000 statements a month, GPT-5.5 costs about $2,460 a month, Claude Opus 5.5 about $1,870 and Gemini 3.8 Flash about $290.
Actual cost per statement in this test, September and October 2026 prices. Pale bar: Gemini 3.8 Flash after its introductory price ends on 31 December 2026.
Prices and versions change quickly
Gemini 3.8 Flash, which Google launched on 2 September 2026, is on an introductory price until 31 December 2026, after which its price is set to double[2]. It was also released three weeks after Gemini 3.7 Flash, and Claude Opus moved from 4.8 to 5.5 in four months. The hosts for GLM 5.3 and DeepSeek V4.1 Flash even halved their prices between the morning and the afternoon of one test day. Hence best practice is to lock in the exact model version used in production, and to rerun the test whenever a new version or a price change comes along.
7Round 2: What still separates the finalists?
Once accuracy was level, the finalists still differed on reliability, on spotting faulty statements, on speed and on where the data goes. For a regulated business, where the data goes can rule a model out before anyone compares prices.
The seven finalists compared, expert prompt
| Model | Fees wrong, expert | Cost per statement | Flags caught (of 9) | Median time | Failed answers | Host in this test |
|---|---|---|---|---|---|---|
| Claude Opus 5.5 | 0.03% | $0.187 | 9 | 46 s | None | Google Vertex |
| Grok 4.7 | 0.03% | $0.049 | 9 | 73 s | None | xAI |
| Gemini 3.8 Flash | 0.06% | $0.029 | 9 | 32 s | 4 cut off by the host, rerun | |
| Mistral Large 3 | 0.04% | $0.013 | 7 | 56 s | None | Mistral, EU |
| GLM 5.3 | 0.03% | $0.008 | 9 | 79 s | None | Reka |
| GPT-5.5 | 0.03% | $0.246 | 8 | 23 s | None | Microsoft Azure |
| DeepSeek V4.1 Flash | 0.07% | $0.005 | 9 | 171 s | 1 runaway answer, 1 cut off | Wafer, InferenceNet |
Notes: 22 statements, expert prompt. Flags are the nine traps across the Traps and Hard tiers, counted only when the model used the right label. Failed answers are counted across all three prompt levels.
Reliability. Five of the seven returned a clean answer every time. The host cut off four of Gemini 3.8 Flash's 110 answers and one of DeepSeek V4.1 Flash's. That means the company running the model stopped sending the answer part-way through, for example because its servers were busy or a time limit was reached, so the answer arrived incomplete. Each was fine when sent again. DeepSeek also had one answer that never finished, because it kept writing its working-out until it reached its length limit. Around 100 answers each is too few to prove one service more reliable than another, so a business should check reliability over a longer trial, and anyone using these models in production needs automatic retries and a person checking the log.
Speed. Median time per statement ranged from 23 seconds for GPT-5.5 to 171 seconds for DeepSeek. An overnight batch can absorb that, but a customer or an analyst waiting on the answer would notice the slower models.
Where the data goes. Zero data retention means the company running the model keeps no copy of the document or the answer once it has replied[3]. With it required, OpenRouter would not send requests to Anthropic's, OpenAI's or DeepSeek's own servers, so Claude ran through Google Vertex, GPT through Microsoft Azure and DeepSeek on third-party hosts. Two Alibaba models could not run with zero retention at all. The open-weight finalists ran on hosts most compliance teams will not have reviewed, GLM 5.3 on Reka and DeepSeek on Wafer and InferenceNet. Of the 34 models, 15 have open weights, including three of the finalists: Mistral Large 3, GLM 5.3 and DeepSeek V4.1 Flash. The alternative for any of them is to run the weights on the business's own servers, so its documents never leave the business.
Who the supplier is. Two finalists come from Chinese developers. Zhipu AI, which makes GLM 5.3, has been on the US Entity List since January 2025[4]. DeepSeek's app was blocked in Italy in January 2025, Berlin's data protection authority asked Apple and Google to remove it in June 2025, and the US has ordered it off Defense Department systems[5,6,7]. The UK has not banned it[8]. The router is a supplier too. Using OpenRouter, now part of Stripe, adds one more company to the data path, which under DORA in the EU and the PRA's outsourcing rules in the UK needs the same third-party review as any other provider[9,10].
Part 3, Conclusions
8Conclusions
The main finding is that once the prompt was right, all seven finalists were about equally accurate. The choice of model then comes down to price and other points such as reliability, hosting and where the data goes. The price gap is at most about 25 cents a statement, so in money terms the gap only adds up to a material amount for a business that reads documents in bulk.
Seven finalists, round 2. Each arrow runs from the plain prompt to the expert prompt. The fees-wrong scale is stretched to the right of the 0.1% line. All runs used zero data retention (ZDR). Direct means the developer's own service qualified.
Summary of findings
- The common view held for Opus and Grok, and not for GPT-5.5. In round 1, with the plain prompt, Claude Opus 5.5 and Grok 4.7 were the most accurate. GPT-5.5, OpenAI's flagship, got 0.94% of fees wrong, behind cheaper models from China and France, and in round 2 it needed the expert prompt as much as the cheaper models did.
- The models found the fees, and their mistakes were in classifying them. In round 1, every usable model found at least 99% of the fee lines. The errors came from lines that need payments knowledge to file correctly.
- A good prompt closed the gap. In round 2, the expert prompt took the five other finalists from between 0.3% and 1.0% of fees wrong to under 0.1%. That prompt is the business's own knowledge, written down, and it works with any of the seven models.
- Price was a poor guide to accuracy. In round 1, Mistral Medium 3.5 cost fifteen times as much as Mistral Large 3 and got more fees wrong. In round 2, with the expert prompt, the cheapest finalist matched the most expensive for a fiftieth of the cost.
- The higher price of Opus and Grok bought judgement and reliability. They did well with basic instructions in both rounds. In round 2 they caught the faulty statements without being told what to look for, and never failed.
- Where the data goes can decide the choice. Zero data retention changed which hosts could be used for Claude, GPT and DeepSeek, and ruled out two Alibaba models. Two finalists come from developers that will come up in sanctions screening.
Discussion
Who the price gap matters for
For a single merchant, the model's price makes little difference. Reading one statement a month costs about $3 a year with GPT-5.5 and about 6 cents with DeepSeek V4.1 Flash. What matters at that scale is accuracy: about £40 a statement between the best models and the rest in round 1, and about £3 in round 2 with the expert prompt.
The price gap matters for a business that reads statements in bulk, such as an acquirer checking its merchants' statements, a lender reading card statements to underwrite merchant cash advances, or a fintech building this into a product. At 10,000 statements a month, GPT-5.5 costs about $29,500 a year yet DeepSeek about $600, and GLM 5.3 both matched GPT-5.5's accuracy for about $940.
The skills the task needed
The task needed three different skills, and the models were not equally good at each. The first row draws on both rounds, and the other two on round 2.
| Skill | What it involves | How the models compared |
|---|---|---|
| Reading the statement | Finding every fee line and reading its amount and currency | Every usable model found at least 99% of the fee lines. On scanned pages, the conversion to text set the limit for all of them. |
| Payments knowledge | Putting each fee under the right type, such as a card surcharge versus a penalty | With a plain prompt, Claude Opus 5.5 and Grok 4.7 were best. With the expert prompt, which supplies the knowledge, all seven finalists were equal. |
| Reasoning across the statement | Comparing many rows to spot a pattern that should not be there, such as the fake pass-through pricing | Claude Opus 5.5 and Grok 4.7 did it at every prompt level. Gemini 3.8 Flash, GLM 5.3 and DeepSeek V4.1 Flash did it reliably only with the expert prompt. GPT-5.5 and Mistral Large 3 missed at least one. |
Hence a good prompt can supply payments knowledge, and the price premium for Opus and Grok mostly buys reasoning without being told what to look for.
Some of the models "think", writing out working notes before they answer. On this task the amount of thinking made no visible difference to accuracy, and it added cost. Annex A6 sets out the figures.
Was the task hard enough?
This task may not have been hard enough to separate the models. With good instructions, reading and classifying the fees on a card statement was within reach of every finalist, and the gaps showed only on the hardest judgement calls. A harder task would show whether this conclusion still holds and where a difference in capability starts to show (Section 10).
9What this means for your AI strategy
The model market is moving fast. During this test alone, Google released Gemini 3.8 Flash three weeks after 3.7 Flash, on an introductory price that doubles at the end of the year. Claude Opus moved from 4.8 to 5.5 in four months, and the hosts for two finalists halved their prices within a single day. With new models and new prices arriving this often, no single model is likely to stay the best choice for every task for long. Hence over the next 12 months we can expect to see more and more companies running several models side by side, each matched to the task in hand, and changing them as better or cheaper ones arrive. I also expect the routers that connect a business to those models to move on from keeping a service running to choosing models on price and accuracy.
Open-weight models are already part of that shift. A September 2026 survey by Enterprise Technology Research found that they handle 34% of enterprise AI token usage, up from 23% a year earlier, and that 42% of respondents run at least one in production[29]. The Financial Times has reported Tinder moving some work from frontier models to open-weight ones to control cost, with PNC, CH Robinson and Siemens also talking about using them[30]. That points to a tiered set of models: open-weight models for high-volume routine tasks, and the leading closed models for the hard judgement calls.
This test shows why that matters. With the expert prompt, reading and classifying the fees was within reach of a model costing half a cent a statement. The higher price of Claude Opus 5.5 and Grok 4.7 paid off on the hardest judgement calls, such as spotting a faulty statement without being told what to look for. Hence the useful question for a business is which of its tasks need an advanced model, and which are paying for capability they don't use.
That choice can be set out like any other investment. The costs are the model's price per task, the human checking its errors need, and the test that shows which models are good enough. The benefit is accuracy, measured in money. On this task, at 10,000 statements a month, the model alone costs between about $600 and $29,500 a year depending on the choice, and the test that showed which models were good enough cost about $62.
In 12 months I expect good practice to look like this:
- A set of models matched to tasks. The business runs a small set of models, each matched to a task, and keeps the most capable ones for the hard judgement calls.
- A test set for each important task. It holds a test set with known answers, built from its own data, and measures accuracy in money.
- Its own knowledge in its own prompts. It owns and maintains the prompts, so it can move between models without starting again.
- Locked versions and regular reruns. It locks in the exact model version it runs in production, and reruns the test whenever a new version or a price change comes along.
- Data rules and supplier reviews up front. It sets its rules on where data can go, and reviews its model hosts and any router as suppliers, before testing starts.
Routers make switching between models practical, and there are several kinds. Model marketplaces such as OpenRouter reach hundreds of models through one connection. AI gateways such as LiteLLM, Portkey and Cloudflare AI Gateway sit in front of a business's own model accounts. The cloud platforms offer their own catalogues of models, through Amazon Bedrock, Google Vertex AI and Microsoft Azure AI Foundry. Today routers are used mainly for resilience, moving to a backup model when one fails, and for one connection and one bill across many models. Where they cut cost, it is mostly by sending a model's traffic to the cheapest host for the same model. Matching each task to the cheapest model that is good enough still needs a test on the business's own data.
The same kind of test works on any fintech task, and the answer can differ from one task to the next.
Questions for management
The findings point to five questions for any management team using AI models across its business:
- Which model does each of our AI use cases run on, and when did we last test that choice against the alternatives?
- Do we measure accuracy in money, on a test set built from our own data with known answers?
- How much of our own knowledge is written into our prompts, and who owns and maintains them?
- Where does each model run, under what retention terms, and which companies in the chain, including any router, see our data?
- Which of our tasks need an advanced model, and what would a cheaper model that passes the same test save us at our volumes?
Download the five questions as a one-page scorecard (PDF)
10Questions this test leaves open
This test leaves four questions open.
- Page images in place of text. In this test every model received the statement as text. The scanned pages were turned into text first by Tesseract, a standard open-source OCR tool that works separately from the AI models. I chose text so that all 34 models started from the same input, since not all of them can read images, and because images use more tokens and would have cut the number of models I could test within the budget. Reading a page image is a different skill from understanding its contents, and many current models can now do both. A future test could send the page images to the finalists that can read them, to see whether reading the page directly beats OCR on scanned statements, and what it adds to the cost.
- A harder task. More judgement calls that need reasoning across a whole document, to see where a difference in capability shows.
- A whole workflow, run by an agent. Test a model working as an agent inside a business's own tools, taking several steps on its own. For example, it could fetch statements from a shared drive, read and check them, and write the results to a spreadsheet, using connectors such as MCP (the Model Context Protocol). The test would show whether accuracy and cost hold up from end to end, and where a person still needs to check the work.
- A fine-tuned open-weight model. No model in this test was fine-tuned. A business could train an open-weight model such as Mistral Large 3, GLM 5.3 or Llama 4 Maverick on its own statements, to see whether a small model matches the expert-prompt results without the 1,500-word briefing, and whether that lowers the cost further.
Happy to talk through the method with anyone planning a test like this on their own tasks or workflows, and to hear which of these questions matters most to you. Let me know at martin@scalepointpartners.com.
ATechnical annex
This annex holds the detail behind the findings: every model's results, the models that weren't usable, the full test details, the flags run by run, what drove each model's cost, whether thinking helped, and the developers behind the models.
A1Model by model
This part gives a verdict on each model, based on its scores, with supplier and data points in a separate column. The seven finalists ran at all three prompt levels. The other 27 ran with the plain prompt only, so their verdicts rest on the first round alone. The nine that failed the bar are listed in annex A2.
The seven finalists
| Model | Plain | Expert | Cost, expert | Verdict |
|---|---|---|---|---|
| Claude Opus 5.5Anthropic, US | 0.04% | 0.03% | $0.187 | Best accuracy, high cost. Right even with a plain prompt, caught every flag, no failed answersClosed weights. Ran on Google Vertex, because Anthropic's own service did not qualify for zero data retention (see Section 7) |
| Grok 4.7xAI, US | 0.04% | 0.03% | $0.049 | Strong. As accurate as Opus at a quarter of the price, caught every flag, no failed answersClosed weights. Launch page silent on where data is processed. EU and Ofcom investigations into the consumer Grok service |
| Gemini 3.8 FlashGoogle, US | 0.31% | 0.06% | $0.029 | Best value. Caught every flag with the expert promptClosed weights. Google enterprise terms. Price doubles in January 2027; versions change often |
| Mistral Large 3Mistral AI, France | 0.56% | 0.04% | $0.013 | Strong on fees, best EU option. Weak on the faulty-statement flagsOpen weights, Apache 2.0. Ran on Mistral's own EU hosting |
| GLM 5.3Zhipu AI, China | 0.61% | 0.03% | $0.008 | Strong. Matched Opus with the expert prompt for a twenty-fourth of the priceOpen weights. Zhipu is on the US Entity List. Ran on Reka, a third-party host |
| GPT-5.5OpenAI, US | 0.86% | 0.03% | $0.246 | Accurate with help, most expensive. Fastest finalist; never flagged the faulty subtotalClosed weights. Ran on Microsoft Azure, because OpenAI's own service did not qualify for zero data retention (see Section 7) |
| DeepSeek V4.1 FlashDeepSeek, China | 0.99% | 0.07% | $0.005 | Cheapest, less reliable. Accurate with the expert prompt, but one runaway answer and slow at the tailOpen weights, MIT. Its own service did not qualify for zero retention. App blocked in Italy; ordered off US defence systems |
Claude Opus 5.5. Anthropic is a San Francisco company founded in 2021 by former OpenAI staff and backed by Amazon and Google[11]. Opus 5.5, released on 22 September 2026, is its flagship model[12]. It needed the least help from the prompt: 0.04% wrong with the plain prompt and 0.03% with the expert prompt, with no failed answers and a median of 46 seconds a statement. In February 2026 the US Defense Department labelled Anthropic a supply-chain risk after a dispute over the limits Anthropic puts on military use, and a federal court blocked that in March 2026[11].
Grok 4.7. Elon Musk founded xAI in 2023. SpaceX bought it in February 2026 and renamed it SpaceXAI, and SpaceX has been listed on Nasdaq since June 2026[13]. Grok 4.7 was released on 21 September 2026[14]. It matched Opus with both the plain and the expert prompt, caught every flag and never failed, at $0.049 a statement, though it was slower at 73 seconds. The European Commission and Ofcom opened investigations in January 2026 into sexualised deepfakes made with the consumer Grok service on X[15,16]. Those cases concern the consumer product, and a management team will still want to know about them.
Gemini 3.8 Flash. Google DeepMind builds Gemini, and Google's parent Alphabet is listed on Nasdaq. Gemini 3.8 Flash was released on 2 September 2026[2]. It improved at every prompt level, from 0.31% to 0.06%, caught all nine flags with the expert prompt and took 32 seconds a statement. The host cut off four of its 110 answers, all fine when rerun.
Mistral Large 3. Mistral AI was founded in Paris in 2023 and is private, with the Dutch chip-equipment maker ASML as its largest investor[17]. Mistral Large 3 was released in December 2025 under the Apache 2.0 licence, so a business can run it on its own servers[18]. It went from 0.56% with the plain prompt to 0.04% with the expert prompt, with no failed answers. It is the only finalist from a European developer, and in this test it ran on Mistral's own EU hosting.
GLM 5.3. Zhipu AI is a Beijing company spun out of Tsinghua University in 2019 and listed in Hong Kong since January 2026, with Alibaba, Tencent and Shanghai state funds among its backers[19]. The GLM 5.3 weights were released in August 2026 under Zhipu's own licence, which requires very large hosts to pass a Zhipu security review[20]. It went from 0.61% to 0.03%, caught every flag with the expert prompt and never failed, at $0.008 a statement.
GPT-5.5. OpenAI is a San Francisco company in which Microsoft holds about 27%[23]. GPT-5.5 is its flagship model. It went from 0.86% with the plain prompt to 0.07% with the middle prompt and 0.03% with the expert prompt, never failed, and was the fastest finalist at 23 seconds a statement. At $0.246 a statement it was also the most expensive. It caught the fake pass-through pricing only with the expert prompt and never flagged the faulty subtotal. It ran on Microsoft Azure, because OpenAI's own service did not qualify for zero retention.
DeepSeek V4.1 Flash. DeepSeek is a Hangzhou company founded in 2023 by Liang Wenfeng and funded by his hedge fund High-Flyer. In its first outside funding round in mid-2026, a Chinese state AI fund took a stake with voting rights[21]. V4.1 Flash was released on 10 September 2026 under the MIT licence[22]. It made the biggest improvement of the seven, from 0.99% to 0.07%, and was the cheapest at $0.005, but it had one runaway answer and was the slowest at 171 seconds.
The other 18 usable models
| Model | Fees wrong, plain prompt | Cost per statement | Verdict, supplier and data |
|---|---|---|---|
| Nemotron 3 UltraNVIDIA, US | 0.48% | $0.016 | Strong, not tested with the expert promptOpen weights. Ran on DeepInfra |
| Kimi K3Moonshot AI, China | 0.49% | $0.065 | Usable, better options exist. As accurate as GLM 5.3 at nearly five times the costOpen weights, with a separate agreement needed for large providers. Ran on Wafer |
| Claude Sonnet 5Anthropic, US | 0.83% | $0.086 | Usable, better options existClosed weights. Google Vertex |
| MiniMax M3MiniMax, China | 0.89% | $0.006 | Strong for the price, not tested with the expert promptOpen weights. Ran on Together |
| Claude Sonnet 4.6Anthropic, US | 0.95% | $0.103 | Superseded by Sonnet 5Closed weights. Google Vertex |
| GPT-5.2OpenAI, US | 1.17% | $0.071 | Superseded by GPT-5.5Closed weights. Microsoft Azure |
| GPT-5.4 miniOpenAI, US | 1.18% | $0.023 | Usable, better options existClosed weights. Microsoft Azure |
| GPT-4.1OpenAI, US | 1.27% | $0.052 | SupersededClosed weights. Microsoft Azure |
| Gemini 3.1 ProGoogle, US | 1.32% | $0.096 | Usable, better options exist. Gemini 3.8 Flash is more accurate at under a third of the costClosed weights. Google |
| Gemini 3.5 Flash LiteGoogle, US | 1.40% | $0.019 | Usable, better options existClosed weights. Google |
| Qwen 3.5 397BAlibaba, China | 1.52% | $0.046 | Usable, better options existOpen weights. Ran on DeepInfra |
| GLM 5.3 FlashZhipu AI, China | 1.78% | $0.003 | Usable, the cheapest in the testOpen weights. Zhipu is on the US Entity List |
| Mistral Medium 3.5Mistral AI, France | 1.94% | $0.174 | Usable, better options exist. Fifteen times the cost of Mistral Large 3 and less accurateOpen weights. Mistral, EU |
| Claude Opus 4.8Anthropic, US | 2.16% | $0.184 | Superseded by Opus 5.5Closed weights. Google Vertex |
| Llama 4 MaverickMeta, US | 3.01% | $0.006 | Usable, better options existOpen weights. Ran on Parasail |
| Qwen 3.8 FlashAlibaba, China | 3.21% | $0.011 | Usable, better options existOnly on Alibaba's servers, no zero retention |
| Qwen 3.8 27BAlibaba, China | 3.25% | $0.012 | Usable, better options existOpen weights. Ran on Reka |
| Claude Haiku 4.5Anthropic, US | 3.40% | $0.044 | Usable, better options existClosed weights. Google Vertex and Amazon Bedrock |
A2The nine models that weren't usable
A model counted as usable if it cleared the bar set before the test: its fee lines added up to the true total on at least 95% of the Standard statements, it invented no fee lines, and it caught at least three of the four traps. 25 of the 34 models cleared it and 9 did not. The finalists were chosen from the 25.
| Model | Fees wrong, plain prompt | Why it failed |
|---|---|---|
| GPT-4.1 miniOpenAI, US | 1.52% | Missed traps and failed to add up |
| Qwen 3.8 MaxAlibaba, China | 1.86% | Invented fee lines. Also only on Alibaba's servers, with no zero retention |
| Gemini 2.5 ProGoogle, US | 1.91% | Invented fee lines |
| DeepSeek V3.2DeepSeek, China | 1.95% | Fee lines failed to add up on too many statements |
| gpt-oss-120bOpenAI, US | 7.15% | Invented a fee line, missed traps and failed to add up |
| Mistral Small 4Mistral AI, France | 8.04% | Missed traps |
| Gemini 2.5 FlashGoogle, US | 8.72% | Invented fee lines |
| Command A PlusCohere, Canada | 63% | Broken or empty answers on most statements |
| Nova ProAmazon, US | 143% | Broken answers on over a third of statements, and the fee lines it invented were worth more than the true fees |
Six of the nine are older or smaller models from developers whose newer or larger models did clear the bar.
A3Test details
The statements
I used AI to help write a program that built 30 card statements from scratch. It made up a month of card sales for each invented merchant, applied an invented acquirer's price list to work out every fee, and produced each statement as a PDF file, with the scanned ones produced as page images. The same program wrote out the calculated answers for every statement. No real merchant or acquirer data was used. The statements come from two invented acquirers. Kestrel Merchant Services uses an interchange-plus-plus layout modelled on the large UK acquirers, and Brightwater Payments uses simple blended pricing. Besides the three tiers used in the test, there was a Clean tier of 8 statements: blended pricing on one page, with three to six fee lines. It was not used in either round.
Each model received only the text of the statement, never the page itself. The text was extracted from each PDF once, and the scanned pages were turned into text by Tesseract, a standard open-source OCR tool, so every model got exactly the same input. Each model returned every fee as a line with its description, amount, currency, fee type, card type and where on the statement it found the fee, plus any problems it saw with the statement. The middle prompt also added a rule that a fee shown as a deduction is still a positive amount.
Runs per finalist
In round 2 the four cheapest finalists, Gemini 3.8 Flash, Mistral Large 3, GLM 5.3 and DeepSeek V4.1 Flash, ran twice at the middle and expert levels, to check that their answers were consistent. The other three ran once at each level.
Scoring
A scoring program, written in Python for this test with AI help, compared each fee line a model returned with the calculated answers, matching on amount, currency and description.
Two scoring calls were settled during the test. An American Express line on an interchange-plus-plus statement can fairly be called blended or a service charge, and so can a correction to an earlier month's service charge, so both answers count as right. How these calls are set can move a model a long way: before they were settled, Claude Sonnet 5 scored 5.5% wrong, and afterwards 0.8%.
The template parser
Before AI models, statement reading was automated with OCR and a template written for each acquirer's layout. I wrote a template parser of that kind for the two layouts in the test and ran it on the same statements in both rounds. It was written knowing both layouts, so its results are its best case.
Fee types
The fee types and their definitions, as given to the models with the middle and expert prompts:
| Fee type | Definition |
|---|---|
| Interchange | The part of an IC++ charge paid to the card issuer, shown in its own column |
| Scheme fee | Charges levied by Visa, Mastercard or another card scheme, re-billed by the acquirer |
| Acquirer service charge | The acquirer's own margin on an IC++ row, on top of interchange and scheme fees |
| Blended MSC | A single all-in rate per card type that includes interchange, scheme fees and margin |
| Card-type uplift | An extra charge because of the card used: commercial, premium, or non-UK and international cards |
| Non-qualifying surcharge | A penalty because of how a transaction was processed: not secure, keyed, settled late or missing data |
| Authorisation | A charge per authorisation request |
| 3D Secure and fraud screening | 3D Secure, fraud screening and risk tools |
| Tokenisation and account updater | Network tokens, card-on-file tokens and account updater |
| Refund | The fee for processing a refund, not the refunded sale amount |
| Chargeback | The fee for handling a chargeback, not the disputed amount |
| Terminal or gateway | Terminal or card reader rental, gateway and online checkout fees |
| PCI | PCI DSS programme and non-compliance fees |
| Minimum monthly charge | A top-up charged when fees fall below an agreed monthly minimum |
| Account and admin | Account, statement and admin fees |
| FX and settlement | Currency conversion charges and fees for paying out or speeding up settlement |
| Rebate or credit | A rebate or credit of fees, including a credit that corrects an earlier period |
| Other | Anything that fits none of the types above |
The nine traps
| Trap | Where it appears | What the model has to do |
|---|---|---|
| Footnote fee | Traps tier | Find an annual fee mentioned only in a note |
| Double count | Traps tier | Record a fee once when it appears on the summary and the detail pages |
| Netted fees | Traps tier | Find fees taken out of the settlement and never invoiced, and flag them |
| Two currencies | Traps tier | Keep euro charges in euros and flag the second currency |
| Faulty subtotal | Hard tier | Report the printed figures and flag that they do not add up |
| Fake pass-through pricing | Hard tier | Flag interchange-plus-plus pricing that charges the same rate on every card type |
| Minimum charge | Hard tier | Find a top-up to a monthly minimum on a statement with three outlets |
| Rebate | Hard tier | Record a rebate as a credit |
| Prior period | Hard tier | Mark a correction to an earlier month as such |
Calls and costs
The test made about 1,100 model calls. OpenRouter charged about $27 for the first round, about $20 for the second round with the first six finalists, about $13 for GPT-5.5 and about $2 for set-up tests, about $62 in all against a budget of $50.
Each call sends about 5,000 words of statement and prompt, and the model writes back about as much again, because it lists every fee line with its details. Most of the cost is in what the model writes back. A single call costs between a fraction of a cent and about 55 cents. The first round was 34 models reading 16 statements each, about 560 calls. Five expensive models, Claude Opus 4.8, Mistral Medium 3.5, Claude Opus 5.5, GPT-5.5 and Gemini 2.5 Pro, cost about $12 of the $27 between them. The 16 cheapest models cost about $2 together. Taking all 34 models through the second round would have cost about $110 at first-round prices.
Limitations
- The statements are synthetic. Real statements come in more layouts, and many arrive as poor scans. Any business should repeat the test on its own documents before relying on it.
- The sample is small. The finalists read 22 statements once or twice each, so gaps of a few hundredths of a percentage point at the top are noise.
- First-round answers were reused. In the second round, the plain-prompt results for the 16 repeated statements are the first-round answers, so only the six Hard statements were new at that level.
- The finalists were chosen by judgement. Only the usable bar was fixed before the test, and GPT-5.5 was added afterwards.
- Text only. Every model received the statement as text, with Tesseract used for the scanned pages, so the results on scans reflect the OCR as much as the model. On one scanned statement the OCR misread two figures, which no model could then get right. Reading the page image directly is one of the open questions in Section 10.
- No model was fine-tuned. Every model ran as released by its developer. The open-weight models could be fine-tuned on a business's own documents, which may improve their results further, and this test did not measure that. Llama 4 Maverick, for example, got 3.01% of fees wrong with the plain prompt and was not tested with the expert prompt or fine-tuned.
- Prices and versions move quickly. All prices are OpenRouter prices in September and October 2026, and the scorer reflects my judgement on how each fee should be classified.
- The models could see the statements were tests. Every page carries a line saying it is a synthetic test document, so the files cannot be mistaken for real statements. Mistral Large 3 mentioned this in 23 of its answers, and two other models once each. There is no sign it changed how any model read the fees, but a model could in principle take less care with a document it knows is a test.
A4The faulty-statement flags, run by run
The table shows how often each finalist flagged each of the two faulty statements with the right label, at each prompt level.
| Model | Faulty subtotal | Fake pass-through | ||||
|---|---|---|---|---|---|---|
| Plain | Middle | Expert | Plain | Middle | Expert | |
| Claude Opus 5.5 | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 |
| Grok 4.7 | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 |
| Gemini 3.8 Flash | 1/1 | 2/2 | 2/2 | 1/1 | 1/2 | 2/2 |
| GLM 5.3 | 1/1 | 2/2 | 2/2 | 1/1 | 0/2 | 2/2 |
| DeepSeek V4.1 Flash | 1/1 | 1/2 | 2/2 | 1/1 | 0/2 | 2/2 |
| Mistral Large 3 | 0/1 | 0/2 | 0/2 | 0/1 | 0/2 | 0/2 |
| GPT-5.5 | 0/1 | 0/1 | 0/1 | 0/1 | 0/1 | 1/1 |
The number of runs differs because the four cheapest finalists ran twice at the middle and expert levels.
- Claude Opus 5.5 and Grok 4.7 flagged both at every prompt level.
- Gemini 3.8 Flash, GLM 5.3 and DeepSeek V4.1 Flash flagged both in their single plain-prompt run, mostly missed the fake pass-through with the middle prompt, and flagged both every time with the expert prompt.
- Mistral Large 3 spotted the faulty subtotal in its runs, with the exact figures, but filed it under "other" in place of the label the prompt asked for. It noticed the fake pass-through pricing once in five runs and filed that under "other" too. The score counts only the right label, because any automated routing to a person works from the label.
- GPT-5.5 never flagged the faulty subtotal, at any level. It caught the fake pass-through pricing only with the expert prompt.
A5What drove the cost
A model is charged per token for the text sent in and the text it writes back. The statement and the prompt were the same for every model, so the difference comes from each model's price list and from how much it writes. Models that think before answering write more. With the expert prompt, DeepSeek wrote about 6,200 tokens of working-out per statement, Grok about 3,200 and Opus about 500. DeepSeek was still the cheapest, because its price per token is so low. GPT-5.5 did almost no working-out, but its price per token and its long answers made it the most expensive.
The longer expert prompt added little. It raised Opus's cost per statement by about 9%, Gemini 3.8 Flash's by under 1%, and made no measurable difference for Mistral Large 3.
A6Did thinking help?
Some models "think" before they answer. They write out working notes first and then give the final answer. Those notes are counted and billed as reasoning tokens, and the user does not normally see them.
Whether a model thinks is a design choice by its developer. Several developers, including xAI, DeepSeek and Anthropic, now ship models that think by default or let the user switch it on. Mistral Large 3 answers straight away, and Mistral offers its thinking models as a separate line, which this test did not include. Some hosts also do not report thinking, so a model that shows none may still think without saying so.
In this test, the amount of thinking made no visible difference to accuracy:
- In the first round, the 15 usable models that reported thinking had a median of 1.17% of fees wrong, and the 10 that did not had 1.08%.
- The heaviest thinkers did not do best in round 1. Qwen 3.8 Flash wrote about 16,000 tokens of working-out per statement and got 3.2% wrong. Gemini 3.8 Flash reported none and got 0.37%.
- The two most accurate models in round 1 differed too. Grok 4.7 wrote about 3,700 tokens of working-out and Claude Opus 5.5 about 200, and both got almost no fees wrong.
- On the faulty statements in round 2 the picture was mixed. Claude Opus 5.5 and Grok 4.7 think and caught both every time, DeepSeek V4.1 Flash and GLM 5.3 think and needed the expert prompt, and Gemini 3.8 Flash caught both with the expert prompt without reporting any thinking. GPT-5.5 reported almost no thinking and missed the faulty subtotal.
Thinking does add cost, because every token of working-out is paid for, and it caused the one answer DeepSeek V4.1 Flash never finished. Hence on this task thinking did not buy accuracy. It may matter more on harder tasks that need reasoning across a whole document.
A7The developers
The 34 models came from 14 developers, nine based in the US, Canada or France and five in China. 15 of the models have open weights, and two more are closely related to open-weight models.
| Developer | Ownership, September 2026 | Models tested |
|---|---|---|
| Anthropic[11]San Francisco, US | Private, backed by Amazon and Google. Filed confidentially for an IPO in June 2026 | Claude Opus 5.5, Opus 4.8, Sonnet 5, Sonnet 4.6, Haiku 4.5 |
| OpenAI[23]San Francisco, US | Private. Microsoft holds about 27%. IPO filing confirmed June 2026 | GPT-5.5, GPT-5.2, GPT-5.4 mini, GPT-4.1, GPT-4.1 mini, gpt-oss-120b |
| Google[24]Mountain View, US | Part of Alphabet, listed on Nasdaq | Gemini 3.8 Flash, 3.1 Pro, 3.5 Flash Lite, 2.5 Pro, 2.5 Flash |
| xAI, now SpaceXAI[13]US | Bought by SpaceX in February 2026. SpaceX listed on Nasdaq in June 2026 | Grok 4.7 |
| NVIDIASanta Clara, US | Listed on Nasdaq | Nemotron 3 Ultra |
| MetaMenlo Park, US | Listed on Nasdaq | Llama 4 Maverick |
| AmazonSeattle, US | Listed on Nasdaq | Nova Pro |
| Mistral AI[17]Paris, France | Private. ASML holds about 11% | Mistral Large 3, Medium 3.5, Small 4 |
| Cohere[25]Toronto, Canada | Private | Command A Plus |
| Zhipu AI (Z.ai)[19]Beijing, China | Listed in Hong Kong since January 2026. On the US Entity List since January 2025 | GLM 5.3, GLM 5.3 Flash |
| Moonshot AI[26]Beijing, China | Private, backed by Alibaba and Tencent | Kimi K3 |
| DeepSeek[21]Hangzhou, China | Private. Founded and funded by the hedge fund High-Flyer. A Chinese state AI fund invested in 2026 | DeepSeek V4.1 Flash, V3.2 |
| Alibaba[27]Hangzhou, China | Listed in New York and Hong Kong | Qwen 3.8 Max, 3.8 Flash, 3.8 27B, 3.5 397B |
| MiniMax[28]Shanghai, China | Listed in Hong Kong since January 2026 | MiniMax M3 |
Sources
Sources 11 to 28 are cited in the technical annex.
- Stripe, "Stripe agrees to acquire OpenRouter", press release, August 2026. OpenRouter connects to more than 400 models from over 80 providers; Stripe's reported price was $7.5bn (stripe.com)
- Google, "Gemini 3.8 Flash", launch post and pricing, September 2026 (blog.google)
- OpenRouter, Zero Data Retention guide, checked September 2026 (openrouter.ai)
- US Department of Commerce, additions to the Entity List, Federal Register, 16 January 2025 (federalregister.gov)
- Euronews, "DeepSeek AI blocked by Italian authorities as other member states open probes", 31 January 2025 (euronews.com)
- BleepingComputer, "Germany asks Google, Apple to remove DeepSeek AI from app stores", June 2025 (bleepingcomputer.com)
- US National Defense Authorization Act for FY2026, section 1532, as summarised by ETO AGORA (agora.eto.tech)
- Insurance Journal, reporting on the UK government's position on DeepSeek, February 2025 (insurancejournal.com)
- Regulation (EU) 2022/2554 on digital operational resilience for the financial sector (DORA) (eur-lex.europa.eu)
- Prudential Regulation Authority, SS2/21 Outsourcing and third party risk management, March 2021 (bankofengland.co.uk)
- Anthropic, company background and funding, checked September 2026 (wikipedia.org)
- Anthropic, Claude Opus 5.5 launch page, September 2026 (anthropic.com)
- SpaceX, pricing of initial public offering, June 2026; SpaceXAI background (ir.spacex.com)
- xAI, Grok 4.7 launch post, September 2026 (x.ai)
- PBS NewsHour, "EU investigates Musk's AI chatbot Grok over sexual deepfakes", January 2026 (pbs.org)
- Ofcom, investigation into X over Grok sexualised imagery, January 2026 (ofcom.org.uk)
- Mistral AI, company background and ASML investment, checked September 2026 (wikipedia.org)
- Mistral AI, Mistral Large 3 model card, December 2025 (docs.mistral.ai)
- Caixin Global, "China's Zhipu AI jumps in Hong Kong debut", 8 January 2026 (caixinglobal.com)
- The New Stack, on the GLM-5.3 weights and licence, August 2026 (thenewstack.io)
- Forbes, "DeepSeek just raised $7.4 billion. Here's the catch", 17 June 2026 (forbes.com)
- DeepSeek, DeepSeek V4.1 Flash model card, September 2026 (huggingface.co)
- OpenAI, company background and ownership, checked September 2026 (wikipedia.org)
- Alphabet Inc., company background, checked September 2026 (wikipedia.org)
- Cohere, company background, checked September 2026 (wikipedia.org)
- Moonshot AI, company background, checked September 2026 (wikipedia.org)
- Alibaba Group, company background, checked September 2026 (wikipedia.org)
- MiniMax Group, company background, checked September 2026 (wikipedia.org)
- Techstrong.ai, "Open-Weight AI Models Expected to Capture 41% of Enterprise Token Usage", reporting Enterprise Technology Research's September 2026 survey, 22 September 2026 (techstrong.ai)
- Financial Times, report on US companies moving work to open-weight models, 27 September 2026 (ft.com)
This paper reflects the author's own views. The test statements are synthetic, and prices are OpenRouter prices in September and October 2026; results on real documents and at other prices will differ. Model names and trademarks belong to their owners.