A cheaper token does not necessarily mean a cheaper completed task.

That remains the point of this monthly pricing watch. August is the second issue, which means we can finally make a direct comparison with July instead of treating one rate card as a trend.

The first result is pleasantly small and useful. Across 20 matched provider, model, and category entries, one published provider-currency price card changed. Nineteen did not. The pound price for the unchanged cards still moved a little because the exchange rate moved.

That is exactly why I want to keep the original provider currency next to the GBP conversion. It stops us mistaking a currency wobble for a model-price revolution.

What changed in August

Claude Haiku 4.5 is the one material price-card change. Its published direct API price moved from $1.00 input and $5.00 output per million tokens in July to $0.50 input and $2.50 output in this August snapshot. The quick reference workload therefore falls from about 1.11 pence to 0.56 pence.

That is a real provider price change. It is not a claim that Haiku is automatically the best quick model for every task. Quality, latency, context, tool use, retries, cache behaviour, and task success still matter.

The other matched rate cards were unchanged in their published dollar price. The Bank of England quote used here moved from GBP 0.7414 per US dollar on 16 July to GBP 0.7426 per US dollar on 6 August. For a fixed dollar price, that makes the GBP figure around 0.16% higher. It is small, but it is not a provider price increase.

A July-to-August comparison showing 20 matched model entries, one provider-currency price change, 19 unchanged rate cards, and the separate small GBP-to-USD exchange-rate movement.
Two snapshots make a comparison, not a forecast. Provider-currency changes come before GBP conversion.

Three useful kinds of model

The names will keep moving, so I am grouping models by the work I would expect them to do.

CategoryPlain-English meaningReference workload
QuickHigh-volume, lower-cost work such as extraction, classification, transformation, and narrow routine tasks.10,000 input + 1,000 output tokens
GeneralEveryday professional and agentic work: documents, analysis, planning, coding, and reliable multi-step tasks.50,000 input + 5,000 output tokens
FrontierDifficult reasoning, coding, research, and long-running work where model quality can materially change the outcome.250,000 input + 25,000 output tokens

I selected no more than one current model from each provider for each category. Where there is no honest equivalent, I have left the slot empty. This is a purchasing map, not a quality leaderboard.

How I priced it

The figures below are direct, real-time, self-service API prices. They exclude subscriptions, harness plans, VAT, enterprise discounts, batch and priority processing, tool charges, storage, and negotiated rates.

All 20 selected provider prices are published in US dollars. I converted them at GBP 0.7426 per US dollar, the latest available daily Bank of England quote when this snapshot was taken, dated 6 August 2026. The Bank describes the rates as indicative spot rates rather than official rates.

The reference workloads are intentionally transparent and uncached:

  • Quick: 10,000 input and 1,000 output tokens.
  • General: 50,000 input and 5,000 output tokens.
  • Frontier: 250,000 input and 25,000 output tokens.

They are price baskets, not observed average tasks and not quality benchmarks.

Three small-multiple charts comparing current GBP input and output prices per million tokens for quick, general, and frontier AI models.
Standard direct API prices converted to GBP. Each category has its own scale so the lower-cost models remain visible.

The August 2026 price table

USD figures are provider-published prices per million tokens. GBP figures use the 6 August exchange-rate snapshot. Cached input is shown only where the provider publishes a directly comparable rate.

Quick models
ProviderModelInput $ / GBPCached $ / GBPOutput $ / GBPQuick task GBP
OpenAIGPT-5.6 Luna$1.00 / GBP 0.7426$0.10 / GBP 0.0743$6.00 / GBP 4.4556GBP 0.0119
AnthropicClaude Haiku 4.5$0.50 / GBP 0.3713$0.05 / GBP 0.0371$2.50 / GBP 1.8565GBP 0.0056
GoogleGemini 3.1 Flash-Lite$0.25 / GBP 0.1857$0.025 / GBP 0.0186$1.50 / GBP 1.1139GBP 0.0030
MistralMistral Small 4$0.15 / GBP 0.1114Not listed$0.60 / GBP 0.4456GBP 0.0016
DeepSeekDeepSeek V4 Flash$0.14 / GBP 0.1040$0.0028 / GBP 0.0021$0.28 / GBP 0.2079GBP 0.0012
Alibaba CloudQwen 3.6 Flash$0.165 / GBP 0.1225Separate cache rules$0.99 / GBP 0.7352GBP 0.0020
General models
ProviderModelInput $ / GBPCached $ / GBPOutput $ / GBPGeneral task GBP
OpenAIGPT-5.6 Terra$2.50 / GBP 1.8565$0.25 / GBP 0.1857$15.00 / GBP 11.1390GBP 0.1485
AnthropicClaude Sonnet 5$2.00 / GBP 1.4852$0.20 / GBP 0.1485$10.00 / GBP 7.4260GBP 0.1114
GoogleGemini 3.5 Flash$1.50 / GBP 1.1139$0.15 / GBP 0.1114$9.00 / GBP 6.6834GBP 0.0891
xAIGrok 4.3$1.25 / GBP 0.9283$0.20 / GBP 0.1485$2.50 / GBP 1.8565GBP 0.0557
MistralMistral Large 3$0.50 / GBP 0.3713Not listed$1.50 / GBP 1.1139GBP 0.0241
Moonshot AIKimi K2.6$0.95 / GBP 0.7055$0.16 / GBP 0.1188$4.00 / GBP 2.9704GBP 0.0501
Alibaba CloudQwen 3.7 Plus$0.32 / GBP 0.2376Separate cache rules$1.28 / GBP 0.9505GBP 0.0166
Frontier models
ProviderModelInput $ / GBPCached $ / GBPOutput $ / GBPFrontier task GBP
OpenAIGPT-5.6 Sol$5.00 / GBP 3.7130$0.50 / GBP 0.3713$30.00 / GBP 22.2780GBP 1.4852
AnthropicClaude Fable 5$10.00 / GBP 7.4260$1.00 / GBP 0.7426$50.00 / GBP 37.1300GBP 2.7848
GoogleGemini 3.1 Pro Preview$2.00 / GBP 1.4852$0.20 / GBP 0.1485$12.00 / GBP 8.9112GBP 1.0768*
xAIGrok 4.5$2.00 / GBP 1.4852$0.30 / GBP 0.2228$6.00 / GBP 4.4556GBP 0.9654*
DeepSeekDeepSeek V4 Pro$0.435 / GBP 0.3230$0.003625 / GBP 0.0027$0.87 / GBP 0.6461GBP 0.0969
Moonshot AIKimi K3$3.00 / GBP 2.2278$0.30 / GBP 0.2228$15.00 / GBP 11.1390GBP 0.8354
Alibaba CloudQwen 3.7 Max$1.65 / GBP 1.2253Separate cache rules$4.951 / GBP 3.6766GBP 0.3982

* The Google and xAI frontier reference baskets cross their published 200K long-context threshold. Their task figures therefore use the higher long-context rate, while the headline columns show the standard short-context rate.

Mistral's frontier slot is intentionally empty. Its current 256K-context candidates do not safely hold this issue's fixed 250K-input plus 25K-output reference basket. That does not make them poor models; it makes this particular normalized comparison unsuitable.

Promotions, announcements, and qualifications

Claude Sonnet 5 remains at its current introductory price of $2 input and $10 output per million tokens. Anthropic says that price ends after 31 August 2026, when the standard price becomes $3 and $15. That scheduled change is not yet entered as August's active price.

Qwen 3.7 Plus remains on Alibaba Cloud's 20% international promotional price for prompts up to 256K tokens. The pricing page calls it limited-time but does not state an end date. Qwen 3.7 Max uses a stable versioned global model price rather than a moving alias whose promotion changes by time of day.

DeepSeek's official page signals a future increase. It is not in this table because a proposed future rate is not a current rate. Google counts thinking tokens in output pricing. Cache rules, long-context tiers, regional access, and provider tokenizers can all move the bill away from the simple headline card.

What the reference work costs

Three small-multiple charts comparing the GBP cost of the quick, general, and frontier reference workloads.
Normalized uncached token baskets. They are not observed average task costs or measures of quality.

The calculation is deliberately simple:

task cost = (input tokens x input rate) + (output tokens x output rate)

The quick basket ranges from roughly 0.12 pence to 1.19 pence. The general basket ranges from roughly 1.66 pence to 14.85 pence. The frontier basket ranges from roughly 9.69 pence to GBP 2.78.

Those are large percentage differences. They are not proof that the lowest-price model is the best value. A cheaper run may use more tokens, call more tools, need more retries, or fail to complete the work.

What I want to measure next

In practice, I would want to know:

  • How many input, cached, reasoning, and output tokens did the model use?
  • How many tool calls, loops, retries, and corrections were required?
  • How long did it take?
  • Did it complete the task?
  • Was the result good enough to use?

That is the next layer of this work. A later benchmark will need versioned tasks, clear success criteria, and real billing capture. Until then, these reference workloads are deliberately only what they claim to be.

For now, the monthly lesson is simple: do not buy a token price. Choose a model and route that can complete the work, then measure what the successful task actually cost.

History starts to become useful

July was the baseline. August is the first comparison. The machine-readable snapshots preserve the provider's original currency alongside the GBP conversion, so future issues can distinguish an actual provider change from currency movement. Retired models will remain in history but disappear from the current comparison. New providers will enter only when they have a relevant text model, direct public API access, and verifiable public pricing.

Read the July 2026 baseline.

Sources and notes

Prices were verified against official provider pages on 9 August 2026. The accessible tables above are the source behind the charts; the month-by-month source data is retained in the TonyWood.org repository.