We asked 15 AI models to quote a price 3,595 times
By HelpStack
Before any of it, the thing you should know while reading: we sell an AI helpdesk, the shop whose price list this runs on is a paying customer of ours, and the human-approval step this post concludes you need is our own product's core workflow. Read the numbers rather than our framing of them.
In August we published a small test: eight AI models, one real shop, two questions. It carried an honest caveat, that it was a single generation per question per model, so the ordering was directional at best.
Two people on r/AI_Agents went straight for that caveat. One of them wrote:
have you thought about running like 20-30 generations per model per question? pricing errors especially can be intermittent, and one correct pass doesnt tell you much about reliability at volume
Another reframed the whole metric in a way we liked enough to steal. At volume the question stops being which model is right once, and becomes how many wrong quotes per thousand replies reach a customer.
We said we would rerun it properly. This is the rerun: 15 models, 8 cases, 30 generations each, 3,595 graded replies, two languages.
Two findings from the first post did not survive it, and both were ours.
What we tested#
The shop is a photo printing business. Its prices run on a quantity ladder: from 10 prints the unit price is 0.32 EUR, from 100 it drops to 0.30, from 150 to 0.28. Simple, and exactly the shape that catches models out.
We asked eight questions. Six in Slovenian, the shop's own language, so the numbers stay comparable with round one. Two in English on the most informative quantities, to answer something the first run could not: when a model misreads a ladder, is that a reasoning failure or a language one?
The quantities are clustered around the price breaks rather than spread evenly. The important pair is 99 and 100. One unit apart, different unit price, and reading one row too far down the ladder is precisely the failure we saw the first time.
One thing we did not control. We never set temperature or top_p, so every model ran at its provider's default, and those defaults are not the same number across providers. That is a real limitation rather than a footnote: it means the thirty generations per case are not thirty draws from a distribution we chose, and it is the most likely explanation for how deterministic they turned out.
One deliberate change of scope. Round one called the shop's production API for real. Doing that three and a half thousand times would put real load on a live customer's system for no gain, so tool responses are replayed from a fixture that reproduces the documented ladder exactly. Every model sees byte identical, correct input. That makes this a test of the model rather than of the whole system, which is the right scope: round one's own finding was that the correct data was in the context window both times a model failed. Nothing was hallucinated. The ladder was misread.
Ten of fifteen were correct on every single reply#
Every case, both languages, including the boundary trap. An eleventh, Kimi K2.5, never quoted a wrong price either; its one miss out of 240 was a reply that gave no price at all. Two years ago that would have been the headline. Now it means the metric is spent. If two thirds of the field is perfect, correctness no longer tells you which model to buy.
Look at the price range instead. Ling 3.0 Flash was correct on all 236 replies that completed, Claude Haiku on all 240, and Haiku costs 89 times more per reply. On accuracy that premium bought nothing, though it did buy the second fastest 95th percentile in the test, which turns out to matter. Notice too that the five failures are scattered through the price range rather than bunched at the cheap end.
The five models that did fail are more interesting, because they failed in three completely different ways.
The failure that never varies#
GPT-4o mini got 180 of 240 right. All sixty of its errors landed on one case.
At 99 photos, where the correct answer is 31.68 EUR, it quoted 29.70 every single time. That is 99 multiplied by 0.30, the unit price from the band that starts above it. Thirty for thirty in Slovenian. Thirty for thirty in English.
One unit lower and it is perfect. At exactly 100 it is perfect again. This is not an unreliable model. It misreads the ladder at one specific place, and it does so every time.
The English result answers the language question directly. This has nothing to do with Slovenian.
Why 29.70 is worse than a wildly wrong answer#
The correct total is 31.68. The model said 29.70. That is 6 percent off.
Our defence against a bad reply is a human approving the draft before it sends. That defence works beautifully against nonsense. If a model quotes 39 EUR for an order that costs 9.60, anybody catches it in two seconds, because it looks wrong. Nobody catches 29.70 against 31.68. It looks exactly like a price.
So we counted two kinds of error separately. A loud error is far enough off that a reviewer catches it. A quiet error is within 25 percent: plausible enough to approve, and therefore the only kind that actually reaches a customer.
Twenty-five percent is a judgement, not a measurement. We never tested where a real reviewer's threshold sits. It also turns out to barely matter: at a 10 percent band the split is 127 quiet to 17 loud, at 25 percent it is 131 to 13. Almost every wrong price this run produced was plausible enough to approve. That, rather than the threshold, is the finding.
GPT-4o mini produced 250 quiet errors per thousand replies and not one loud one. That rate is per thousand replies to our case mix, and our case mix is deliberately loaded with boundary quantities, so read it as a stress-test rate rather than a production one. On real traffic, where most orders sit nowhere near a price break, the same model would produce far fewer. Its 75 percent is worse in practice than the number suggests, because every failure is the kind that ships.
Two of the perfect models never show their working#
If arithmetic no longer separates the field, something else has to. So we graded the writing.
We sampled 120 replies, stripped every trace of which model wrote which, shuffled them, and scored them blind on five dimensions: whether the reply reads as written to a customer rather than a developer, whether it shows its working, whether it leaks anything internal, completeness, and tone. Only after grading did we match scores back to models.
Four of those five dimensions turned out not to separate anything.
Nothing internal leaked in almost any reply: twelve of fifteen models scored a flat 5, and the whole field fits in a quarter of a point. Completeness, tone and writing-for-a-customer are tight too. Whatever we were worried about when we designed the rubric, the models have collectively solved it.
Showing the working is the exception, and it is not close. It spans 2.00 to 4.63, a spread of 2.63, against 1.50 for the next widest and 0.25 for the tightest. Showing the working asks a simple question: can the customer check the number? A bare total cannot be checked by anyone. A reply that gives the unit price, the tier it came from and the total can be checked by the customer, by the agent approving it, and by anyone reading it later.
Two models sit at the floor on that dimension: Grok 4.3, and GPT-5.6 Luna, which is the model we had made our own production default after round one. Both are 100 percent correct on arithmetic. Both hand the customer a bare number with nothing to check it against.
One caveat we owe this finding: showing the working correlates with reply length at 0.85 across models, and our own system prompt asks for short replies. The two at the floor are also the two shortest in the field, 51 and 55 characters on average against 100 for the next. Some of what this dimension measures is verbosity, and it is not obvious a shop should punish a model for doing what it was told.
Seven of the ten perfect models do show their working, so this is not a rule about accuracy. It is a fact about two of them, and one of them was ours. We chose it on the strength of being correct and cheap, without ever looking at what it wrote.
Three replies in the sample contradicted themselves inside three lines. One stated 0.30 EUR per piece for 120 photos and then headlined a total of 9.60. The thing a customer would notice is the contradiction sitting in plain sight, and it is only visible because those replies showed their working. An unauditable reply cannot contradict itself. It can only be wrong.
Where the customer's message actually goes#
Everything above ranks models on accuracy, cost, speed and writing. None of those is the first question a European merchant should ask, and we nearly published without asking it.
A support conversation is not a benchmark prompt. It carries a customer's name, their email, what they ordered, where it is going and what they are unhappy about. Sending that to a model means sending it to whoever runs the model, wherever they run it.
Six of the twelve vendors are Chinese, four American, one Korean, one European. Fifteen models, twelve vendors: OpenAI supplied three of them and DeepSeek appears twice. The five best value models are all non-EU, four of them Chinese, and the only EU vendor in the test finished last on accuracy.
That is an uncomfortable result and we are not going to dress it up. Three things are worth separating, because they get conflated constantly.
Who made the model is not the same as who runs it. Ten of the fifteen were reached through OpenRouter and five called direct, so the weights may be served from infrastructure in a different country from the company that trained them.
That split also limits the cost axis, and our own table shows how. DeepSeek V4 Flash is in this test twice, the same weights on an almost identical token profile, at 0.044 and 0.392 dollars per thousand replies. The 9x gap is two pricing pages, not two models. Gateway rows are what OpenRouter actually billed, excluding their platform fee; direct rows are list price times tokens.
A gateway adds a hop. Routing through one means the message reaches a third party before it reaches any model vendor, and whether that hop can be pinned to a region is a question worth asking whoever you route through.
Open weights change the question entirely. Several of these models can be self hosted. A Chinese trained model running on a server in Frankfurt is a different data protection proposition from the same model called over an API in Hangzhou, and no benchmark can tell them apart.
We are not going to tell you any vendor here is untrustworthy, because we have not audited them and neither has this test. What we will say is that the ranking is not a shopping list. For a merchant bound by GDPR, jurisdiction is usually a gate rather than a trade off. It decides which rows you may consider before price enters the conversation.
Fast, accurate, cheap: pick two#
GLM 5.3 Flash scored 240 out of 240 and posted the best writing score in the field. Its 95th percentile response time is 42 seconds. Kimi K2.5 sits at 32. Neither can serve a customer watching a typing indicator.
Gemini 3.1 Flash Lite makes the point from the other side at 7 seconds, and Claude Haiku is the only other model under ten. Ling, which wins on nearly every other measure here, has a 95th percentile of 26 seconds and absorbed four of the run's five unrecoverable errors. All five were gateway rate-limit rejections rather than anything the models did.
Those numbers were measured with three requests in flight and retries on failure, so the absolute values are inflated and the inflation is not evenly spread. Read the ordering, not the seconds.
What this cost us#
We are re-examining our default. Luna is correct and cheap and writes replies nobody can check. That is not automatically wrong for a desk where an agent approves every draft, but we picked it without knowing, and that is the part that bothers us.
We are adding a prompt rule about upselling. Nine of the fifteen models pointed a customer at the next price tier when the customer was already past it, framing a more expensive order as a saving, in 140 of the 900 replies to our two mid-ladder quantities. We first saw it in the 120-reply writing sample and assumed it was a quirk of two models; counting the whole corpus says it is close to a default behaviour. One quoted 30 EUR for 100 photos and then suggested 150 at a total of 42. The instinct is right and the execution is mechanical. That guardrail is needed whichever model runs.
We found a bug in our own grader after drafting this. Every question names the print format, and the 15 in "10x15" survived our price extraction as a candidate number. Twenty-three replies that quoted no price at all were filed as wrong prices, and several genuine errors were recorded as quoting 15 when the model had actually said 8.00, or 9.60, or 210. No model's correctness moved, because an exact match on the right total is unambiguous. The split between loud and quiet errors did move, and that split is half the point of this post. We re-graded from the saved transcripts, which cost nothing, and every figure above is the corrected version.
Which is the real lesson, and it is not about models. Round one kept no transcripts. Its strangest result, two models producing no closing text at all, could never be settled: we suspected our own harness and had no way to check. This time the harness bug was real, we found it, and we could prove it in an afternoon without spending a cent.
What this does not tell you#
Eight cases on one price ladder. A model that reads this ladder correctly may still fail a different one, and the boundary case is deliberately the hardest thing here. Per-thousand rates imply more precision than eight prompts support, and our own data shows why. Repeated generations were close to deterministic: 101 of the 120 model-by-case cells returned the same verdict all thirty times, and mean agreement inside a cell is 97 percent. GPT-4o mini failing 30 out of 30 at one quantity and 0 out of 210 everywhere else is the clearest example. So the independent unit here is the case, not the reply. On eight cases a model with no observed errors has a 95 percent upper bound near 375 per thousand, not the 12 you would get by treating all 240 replies as independent.
Which is also an answer to the question that prompted this rerun. Thirty generations per case were asked for, and twenty-nine of them were mostly redundant. That is worth knowing before anyone else spends the money.
The writing sample is one reply per model per case, 120 in total and eight per model. That is the same sample size this post spends its opening criticising, which is exactly why the writing result is indicative and the accuracy is solid. The blind grades were produced by Claude, and Claude Haiku was one of the models under test. Identity was stripped and the order shuffled, which reduces the risk without removing it. Haiku placed ninth of fifteen on writing, which is not what favouritism looks like, but a reader is entitled to discount the dimension.
Prices move. Every figure was read from the providers on 2 September 2026. One vendor changed its rates on the morning of our first run.
If you run a shop and you are choosing a model, the useful version of this post is not our table. It is the method: take your own pricing data, ask your own question at a quantity that sits just below a price break, run it thirty times, and check the arithmetic yourself.
Frequently asked questions
- Which AI model is best for customer support pricing questions?
- Ten of the fifteen models we tested never quoted a wrong price across 240 replies each, so accuracy alone no longer separates them. Ling 3.0 Flash was both the cheapest and the most accurate, at roughly one eighty-ninth of Claude Haiku per reply. The better question is which model shows its working, because that decides whether anyone can catch an error before it reaches a customer.
- Do cheap AI models make more pricing mistakes?
- Not reliably. The cheapest model in our test was flawless across 236 replies, while a mid-priced one failed 25 percent of the time. Five generations never completed at all, but all five were rate-limit rejections from the gateway we called ten of the models through, on two routes. That is a property of how we reached those models, not of their price.
- What is a quiet pricing error?
- A wrong price close enough to the correct one that a human reviewer approves it without noticing. In our test GPT-4o mini quoted 29.70 EUR where the correct total was 31.68, a 6 percent error, in 30 out of 30 attempts at that quantity. An obviously wrong number gets caught in review. A plausible one gets sent to the customer.
- Does the language of the customer affect AI pricing accuracy?
- It can, but not in the way we expected. GPT-4o mini failed identically in Slovenian and English, so its error was about the price ladder rather than the language. Only one model showed a measurable gap on questions asked in both languages. Writing quality was a different story: every language error in our sample was Slovenian, and English produced none.
- Where is my customer data processed when an AI answers a support message?
- Wherever the model vendor runs inference, plus any gateway in between. The fifteen models come from twelve vendors: six Chinese, four American, one Korean and one European. For a business bound by GDPR that is usually a gate rather than a trade-off. Ask where inference physically executes, whether it can be contractually pinned, what is retained, and whether anything is used for training.