A cheaper token does not necessarily mean a cheaper completed task.

That is the reason I wanted to start this monthly pricing watch.

I have been building more of my work through agents, APIs, and model harnesses. The bill can move surprisingly quickly. The obvious response is to compare the published price per million tokens, but that is only the first number in a much more interesting equation.

A model can be cheaper per token and use more tokens. It can need more retries. It can make more tool calls. It can take a longer route through an agentic loop. It can produce an answer that still needs another model to check it. Most importantly, it can fail to complete the task.

So this first issue creates a baseline. It asks two separate questions:

  1. What do the major providers charge for direct API tokens today?
  2. What would a transparent, fixed reference workload cost at those rates?

It does not pretend those reference workloads are real average tasks. We will need observed runs for that, and I want to build those carefully later.

Three useful kinds of model

The model names change too quickly for a simple comparison to remain useful. I have therefore grouped them by the job I would expect them to do.

CategoryPlain-English meaningReference workload
QuickHigh-volume, lower-cost work such as extraction, classification, transformation, and narrow routine tasks.10,000 input + 1,000 output tokens
GeneralEveryday professional and agentic work: documents, analysis, planning, coding, and reliable multi-step tasks.50,000 input + 5,000 output tokens
FrontierDifficult reasoning, coding, research, and long-running work where model quality can materially change the outcome.250,000 input + 25,000 output tokens

I have selected no more than one current model from each provider for each category. Where there is no honest equivalent, I have left the slot empty. This is not a claim that every model in a row has equal quality. It is a purchasing map, not a benchmark leaderboard.

How I priced it

The figures below are standard, real-time, direct API prices. They exclude subscriptions, harness plans, VAT, enterprise discounts, batch processing, priority processing, tool charges, storage, and negotiated rates.

Provider prices are shown in US dollars because that is how all eight providers publish the selected prices. I converted them at £0.7414 per US dollar, the latest rate available in the Bank of England database when this snapshot was made, dated 16 July 2026. The Bank describes these as indicative spot rates rather than official rates.

That distinction will matter in future issues. If a dollar price stays still while the pound moves, the UK cost changes even though the provider did not alter its price. The dataset therefore preserves both currencies.

Three small-multiple charts comparing current GBP input and output prices per million tokens for quick, general, and frontier AI models.
Standard direct API prices converted to GBP. Each category has its own scale so the lower-cost models remain visible. Long-context surcharges are explained in the table.

The July 2026 price table

USD figures are provider-published prices per million tokens. GBP figures use the 16 July exchange-rate snapshot. Cached input is shown where the provider publishes a directly comparable rate.

Quick models
ProviderModelInput $ / £Cached $ / £Output $ / £Quick task £
OpenAIGPT-5.6 Luna$1.00 / £0.7414$0.10 / £0.0741$6.00 / £4.4484£0.0119
AnthropicClaude Haiku 4.5$1.00 / £0.7414$0.10 / £0.0741$5.00 / £3.7070£0.0111
GoogleGemini 3.1 Flash-Lite$0.25 / £0.1854$0.025 / £0.0185$1.50 / £1.1121£0.0030
MistralMistral Small 4$0.15 / £0.1112Not listed$0.60 / £0.4448£0.0016
DeepSeekDeepSeek V4 Flash$0.14 / £0.1038$0.0028 / £0.0021$0.28 / £0.2076£0.0012
Alibaba CloudQwen 3.6 Flash$0.165 / £0.1223Separate cache rules$0.99 / £0.7340£0.0020
General models
ProviderModelInput $ / £Cached $ / £Output $ / £General task £
OpenAIGPT-5.6 Terra$2.50 / £1.8535$0.25 / £0.1854$15.00 / £11.1210£0.1483
AnthropicClaude Sonnet 5$2.00 / £1.4828$0.20 / £0.1483$10.00 / £7.4140£0.1112
GoogleGemini 3.5 Flash$1.50 / £1.1121$0.15 / £0.1112$9.00 / £6.6726£0.0890
xAIGrok 4.3$1.25 / £0.9268$0.20 / £0.1483$2.50 / £1.8535£0.0556
MistralMistral Large 3$0.50 / £0.3707Not listed$1.50 / £1.1121£0.0241
Moonshot AIKimi K2.6$0.95 / £0.7043$0.16 / £0.1186$4.00 / £2.9656£0.0500
Alibaba CloudQwen 3.7 Plus$0.32 / £0.2372Separate cache rules$1.28 / £0.9490£0.0166
Frontier models
ProviderModelInput $ / £Cached $ / £Output $ / £Frontier task £
OpenAIGPT-5.6 Sol$5.00 / £3.7070$0.50 / £0.3707$30.00 / £22.2420£1.4828
AnthropicClaude Fable 5$10.00 / £7.4140$1.00 / £0.7414$50.00 / £37.0700£2.7803
GoogleGemini 3.1 Pro Preview$2.00 / £1.4828$0.20 / £0.1483$12.00 / £8.8968£1.0750*
xAIGrok 4.5$2.00 / £1.4828$0.30 / £0.2224$6.00 / £4.4484£0.9638*
DeepSeekDeepSeek V4 Pro$0.435 / £0.3225$0.003625 / £0.0027$0.87 / £0.6450£0.0968
Moonshot AIKimi K3$3.00 / £2.2242$0.30 / £0.2224$15.00 / £11.1210£0.8341
Alibaba CloudQwen 3.7 Max$1.65 / £1.2233Separate cache rules$4.951 / £3.6707£0.3976

* The Google and xAI frontier reference baskets cross their published 200K long-context threshold. Their task figures therefore use the higher long-context rates, even though the headline columns show the standard short-context rate.

Mistral's frontier slot is intentionally empty. Its current 256K-context candidates do not safely hold this issue's fixed 250K-input plus 25K-output reference basket. That does not make them poor models; it makes this particular normalized comparison unsuitable.

Promotions and qualifications matter

Claude Sonnet 5 is shown at its current introductory price of $2 input and $10 output per million tokens. Anthropic says that ends after 31 August 2026, when the standard price becomes $3 and $15.

Qwen 3.7 Plus is shown at Alibaba Cloud's current 20% international discount for prompts up to 256K tokens. The pricing page calls it limited-time but does not publish an end date. For Qwen 3.7 Max, I used a stable versioned global model price rather than a moving alias whose promotion changes by time of day.

Google counts thinking tokens in the output price. Anthropic says its newer tokenizer can produce approximately 30% more tokens for the same text than its earlier tokenizer. DeepSeek and Kimi publish unusually low cache-hit rates, but the reference workloads are deliberately uncached. Those details can matter more than a small difference in the headline rate.

What the reference work costs

Three small-multiple charts comparing the GBP cost of the quick, general, and frontier reference workloads.
These are normalized uncached token baskets, not observed average task costs or measures of quality.

The reference calculation is deliberately boring:

task cost = (input tokens x input rate) + (output tokens x output rate)

That lets us compare the rate cards without pretending all models consume the same number of tokens in real work.

The quick basket ranges from about one tenth of a penny to just over one penny. The general basket ranges from roughly 1.7 pence to 14.8 pence. The frontier basket ranges from about 9.7 pence to £2.78.

Those are enormous percentage differences. They are not proof that the lowest-price model is the best-value model. If the cheaper run fails, needs three retries, or produces substantially more tokens, the completed task may cost more.

The real price of an AI task

In practice, I would want to know:

  • How many input, cached, reasoning, and output tokens did the model use?
  • How many tool calls did the agent make?
  • How many loops, retries, and corrections were required?
  • How long did it take?
  • Did it complete the task?
  • Was the result good enough to use?

That is the next layer of this work. I intend to define versioned benchmark tasks and record the actual billed cost, latency, and success. I have not done that in this first issue because a bad benchmark can create more confidence than knowledge.

History starts here

A July 2026 baseline marker explaining that price trends begin only after a second verified monthly snapshot.
One observation is a baseline, not a trend.

Every monthly snapshot will preserve the provider's original currency as well as the GBP conversion. Retired models will remain in history but disappear from the current comparison. New providers will enter only when they have a relevant text model, direct public API access, and verifiable public pricing.

I am also keeping Meta Llama outside the direct-price table. Llama is important, but there is no single comparable Meta API list price. Its token cost depends on whether you self-host it or which inference provider you choose. That deserves its own hosting comparison rather than a made-up universal price.

For July, the lesson is already clear.

Do not buy a token price.

Choose a model and a route that can complete the work, then measure what the successful task actually cost.

Sources and notes

Prices were verified against the providers' official pages on 21 July 2026. The accessible tables above are the source behind the charts. The complete machine-readable snapshot is retained in the TonyWood.org source repository for monthly comparison.