September. Another month, another model menu.

Astra is here, and naturally the first temptation is to select it for everything. The hardest architecture problem. The awkward migration. The tiny job that changes three words in a document. All of it.

But hang on. What are we actually paying for?

This is the third AI Model Pricing Watch: a UK view of direct API token prices. It follows July and August. This month I have expanded the shortlist to 26 model entries across eight providers, including Astra and newer alternatives. The categories are shopping baskets, not intelligence scores.

Checked on 8 September 2026. The tables show US-dollar rates as the providers publish them, alongside fixed workload costs in pounds. These are API prices, not the number of tokens included in a ChatGPT, Codex, Claude or Cursor subscription.

September in a minute

  • Astra enters at $10 input and $50 output per million tokens. Sol remains available at a lower promotional rate.
  • Sol's standard uncached reference basket is 25% cheaper in dollars than August. That is a genuine same-model price change.
  • Sonnet 5 stays at $2 input and $10 output. Its previously announced September increase has been cancelled.
  • Fable 5.1, Gemini 3.8 Flash, Grok 4.6, Kimi K2.7 Code and Qwen3.8 are on the menu. New models get new entries, not invented August prices.
  • DeepSeek now needs a time-of-day label. Peak and off-peak prices are different.

The supporting provider pages are linked in the tables and the notes below. My practical conclusion has not changed: buy the least expensive route to a result you can actually use. Sometimes that is the more capable model. Sometimes it absolutely is not.

Astra arrives. Sol gets cheaper.

OpenAI's 3 September API changelog records GPT-6 Astra's release. Its model page lists standard prices of $10 for input, $1 for cache hits and $50 for output per million tokens. Cache writes are separate, at $12.50 per million.

On our large basket of 250,000 uncached input tokens and 25,000 output tokens, that is $3.75, or about £2.77. Sol costs $1.50, or about £1.11, for the same token counts. That makes Astra 2.5 times the token bill in this particular uncached example. It does not mean it costs 2.5 times as much to finish the same real job.

The 21 August announcement reduced Sol to $4 input and $20 output. Its current model page says the promotional price is available at least through 21 November 2026. That is not a promise it ends the next day.

For our basket, Sol falls from $2 to $1.50 before currency conversion. I would test Astra on the jobs where wrong turns are expensive, then measure the whole run against Sol. I would not promote every background task just because the dropdown has a new name.

The September price tables

All input prices below are uncached standard text-input rates. Cache-hit prices are for eligible repeated input, not every prompt. Output includes billable reasoning where the provider charges for it. A token is a billing unit, not a universal quantity of useful work: different models can tokenise the same text differently.

Every basket uses the same token counts as previous editions. For the large basket, Gemini 3.1 Pro and Grok 4.6 cross their 200K input threshold, so the basket uses their more expensive long-context rates even though the table's rate columns show the starting tariff. Astra and Sol do not cross OpenAI's 272K input threshold: 250K input plus 25K output is not 275K input.

Quick: 10,000 input + 1,000 output tokens

Quick models: September standard rates, with stated promotions and DeepSeek peak pricing.
Model / sourceInput $/1MCache hit $/1MOutput $/1MBasket £
GPT-5.6 Luna0.200.021.200.00237
Claude Haiku 4.510.1050.01110
Gemini 3.5 Flash-Lite0.300.032.500.00407
Mistral Small 40.15Not tabulated0.600.00155
DeepSeek V4 Flash (peak)0.440.0141.320.00423
Qwen3.8 Flash0.113Not tabulated0.3820.00112

General: 50,000 input + 5,000 output tokens

General models: September standard rates, with stated promotions and DeepSeek peak pricing.
Model / sourceInput $/1MCache hit $/1MOutput $/1MBasket £
GPT-5.6 Terra20.20120.11838
Claude Sonnet 520.20100.11099
Gemini 3.8 Flash0.750.0753.750.04162
Gemini 3.7 Flash0.750.0753.750.04162
Grok 4.31.250.202.500.05549
Grok Build 0.110.2020.04439
Mistral Large 30.50Not tabulated1.500.02405
Mistral Medium 3.51.50Not tabulated7.500.08324
Kimi K2.60.950.1640.04994
Kimi K2.7 Code0.950.1940.04994
Qwen3.7 Plus0.32Not tabulated1.280.01657

Frontier: 250,000 input + 25,000 output tokens

Frontier models: September standard rates, with stated promotions and DeepSeek peak pricing.
Model / sourceInput $/1MCache hit $/1MOutput $/1MBasket £
GPT-6 Astra101502.77462
GPT-5.6 Sol40.40201.10985
Claude Fable 5.1100.25502.77462
Claude Opus 550.50251.38731
Gemini 3.1 Pro Preview20.20121.07285
Grok 4.620.5060.96187
DeepSeek V4 Pro (peak)1.320.0443.960.31742
Kimi K330.30150.83239
Qwen3.8 Max 09021.65Not tabulated4.9510.39679

Currency: £0.7399 per US dollar, from the Bank of England table dated 4 September 2026 returned during this check. It is an indicative spot rate, not a card-payment quote or a claim about today's executable exchange rate. Multiply any dollar rate by 0.7399 for the corresponding GBP rate. VAT, currency conversion fees, subscriptions, tools, cache storage, cache writes, hosting and retries are outside these uncached baskets.

“Not tabulated” does not mean free or unavailable. I have not assigned a model-specific cache-hit price where the checked card did not supply one clearly enough for this comparison. This is especially relevant to Mistral and Qwen, whose caching arrangements need their own check.

Three charts showing September GBP costs for fixed quick, general and frontier token baskets.
Compare within a category. Each category has a different workload and its own scale. These are prices, not quality scores.

The small print that changes the bill

Anthropic: a cancelled increase and cheaper cache reads

Anthropic's current pricing page says Sonnet 5's $2/$10 introductory input/output price is now standard. The planned increase to $3/$15 on 1 September will not happen. Fable 5.1 costs $10/$50, with cache reads at $0.25; five-minute cache writes cost $12.50 and one-hour writes $20. Those are different operations.

Fable 5.1 is a new version, not evidence that Fable 5's entire tariff fell. Mythos 5.1 has limited availability, so it is noted here rather than presented as an ordinary self-service choice. Opus 5 remains in the tables as a useful lower-priced comparator.

Google: put the promotion in your diary

Gemini 3.7 Flash arrived on 13 August and 3.8 Flash on 2 September. Their standard introductory tariff is $0.75 input, $0.075 cache hit and $3.75 output through 31 December 2026. Google lists $1.50/$0.15/$7.50 from 1 January 2027. Do not build a permanent budget around the temporary figure. Gemini 3.5 Flash-Lite is also newly included in this watch; that does not make it a September release.

Grok, Kimi and Mistral: more useful choices

Grok 4.6 was released on 12 August. At 200K input and above, its tariff rises to $4 input, $1 cached and $12 output. Grok Build 0.1 is included as a coding option. Its smaller context capacity is why it sits in the general basket.

Kimi's cards list K2.7 Code at the same $0.95/$4 input/output rate as K2.6, but with a different cache-hit rate. The coding model requires thinking mode; its high-speed variant is not the price shown here. K3 remains the larger comparator.

Mistral Medium 3.5 is added alongside Small 4 and Large 3. It is an existing model, not a new September launch. The API page also offers third-party GLM 5.2 at $1.40 input, $0.14 cache hit and $4.40 output. That is a Mistral-hosted Z.ai model, so I have kept it outside the direct model-maker comparison rather than quietly call it a Mistral model.

DeepSeek: what time are you running it?

DeepSeek's current card resolves Flash to version 0731 and Pro to 0813. The main tables use peak rates. Off-peak input/cache-hit/output prices are $0.22/$0.007/$0.66 for Flash and $0.66/$0.022/$1.98 for Pro, half the peak rates. Peak hours are Monday to Friday, 01:00–04:00 and 06:00–10:00 UTC; all other hours are off-peak.

That can matter for work which can wait. It does not mean every job is available at the off-peak headline price. The old snapshot did not establish the same version-and-time basis, so these are not plotted as simple like-for-like price increases.

Qwen: the endpoint is part of the price

Qwen's release list includes 3.8 Flash and 3.8 Max 0902. The rates here use Global deployment in Frankfurt for those two. Qwen3.7 Plus retains the Singapore International alias used in August: $0.40/$1.60 list less its published 20% promotion gives $0.32/$1.28. No expiry is specified there. Do not substitute a different region, dated endpoint or nightly discount and call it the same price.

August to September: the lines we can defend