September. Another month, another model menu.

Astra is here, and naturally the first temptation is to select it for everything. The hardest architecture problem. The awkward migration. The tiny job that changes three words in a document. All of it.

But hang on. What are we actually paying for?

This is the third AI Model Pricing Watch: a UK view of direct API token prices. It follows July and August. This month I have expanded the shortlist to 26 model entries across eight providers, including Astra and newer alternatives. The categories are shopping baskets, not intelligence scores.

Checked on 8 September 2026. The tables show US-dollar rates as the providers publish them, alongside fixed workload costs in pounds. These are API prices, not the number of tokens included in a ChatGPT, Codex, Claude or Cursor subscription.

September in a minute

  • Astra enters at $10 input and $50 output per million tokens. Sol remains available at a lower promotional rate.
  • Sol's standard uncached reference basket is 25% cheaper in dollars than August. That is a genuine same-model price change.
  • Sonnet 5 stays at $2 input and $10 output. Its previously announced September increase has been cancelled.
  • Fable 5.1, Gemini 3.8 Flash, Grok 4.6, Kimi K2.7 Code and Qwen3.8 are on the menu. New models get new entries, not invented August prices.
  • DeepSeek now needs a time-of-day label. Peak and off-peak prices are different.

The supporting provider pages are linked in the tables and the notes below. My practical conclusion has not changed: buy the least expensive route to a result you can actually use. Sometimes that is the more capable model. Sometimes it absolutely is not.

Astra arrives. Sol gets cheaper.

OpenAI's 3 September API changelog records GPT-6 Astra's release. Its model page lists standard prices of $10 for input, $1 for cache hits and $50 for output per million tokens. Cache writes are separate, at $12.50 per million.

On our large basket of 250,000 uncached input tokens and 25,000 output tokens, that is $3.75, or about £2.77. Sol costs $1.50, or about £1.11, for the same token counts. That makes Astra 2.5 times the token bill in this particular uncached example. It does not mean it costs 2.5 times as much to finish the same real job.

The 21 August announcement reduced Sol to $4 input and $20 output. Its current model page says the promotional price is available at least through 21 November 2026. That is not a promise it ends the next day.

For our basket, Sol falls from $2 to $1.50 before currency conversion. I would test Astra on the jobs where wrong turns are expensive, then measure the whole run against Sol. I would not promote every background task just because the dropdown has a new name.

The September price tables

All input prices below are uncached standard text-input rates. Cache-hit prices are for eligible repeated input, not every prompt. Output includes billable reasoning where the provider charges for it. A token is a billing unit, not a universal quantity of useful work: different models can tokenise the same text differently.

Every basket uses the same token counts as previous editions. For the large basket, Gemini 3.1 Pro and Grok 4.6 cross their 200K input threshold, so the basket uses their more expensive long-context rates even though the table's rate columns show the starting tariff. Astra and Sol do not cross OpenAI's 272K input threshold: 250K input plus 25K output is not 275K input.

Quick: 10,000 input + 1,000 output tokens

Quick models: September standard rates, with stated promotions and DeepSeek peak pricing.
Model / sourceInput $/1MCache hit $/1MOutput $/1MBasket £
GPT-5.6 Luna0.200.021.200.00237
Claude Haiku 4.510.1050.01110
Gemini 3.5 Flash-Lite0.300.032.500.00407
Mistral Small 40.15Not tabulated0.600.00155
DeepSeek V4 Flash (peak)0.440.0141.320.00423
Qwen3.8 Flash0.113Not tabulated0.3820.00112

General: 50,000 input + 5,000 output tokens

General models: September standard rates, with stated promotions and DeepSeek peak pricing.
Model / sourceInput $/1MCache hit $/1MOutput $/1MBasket £
GPT-5.6 Terra20.20120.11838
Claude Sonnet 520.20100.11099
Gemini 3.8 Flash0.750.0753.750.04162
Gemini 3.7 Flash0.750.0753.750.04162
Grok 4.31.250.202.500.05549
Grok Build 0.110.2020.04439
Mistral Large 30.50Not tabulated1.500.02405
Mistral Medium 3.51.50Not tabulated7.500.08324
Kimi K2.60.950.1640.04994
Kimi K2.7 Code0.950.1940.04994
Qwen3.7 Plus0.32Not tabulated1.280.01657

Frontier: 250,000 input + 25,000 output tokens

Frontier models: September standard rates, with stated promotions and DeepSeek peak pricing.
Model / sourceInput $/1MCache hit $/1MOutput $/1MBasket £
GPT-6 Astra101502.77462
GPT-5.6 Sol40.40201.10985
Claude Fable 5.1100.25502.77462
Claude Opus 550.50251.38731
Gemini 3.1 Pro Preview20.20121.07285
Grok 4.620.5060.96187
DeepSeek V4 Pro (peak)1.320.0443.960.31742
Kimi K330.30150.83239
Qwen3.8 Max 09021.65Not tabulated4.9510.39679

Currency: £0.7399 per US dollar, from the Bank of England table dated 4 September 2026 returned during this check. It is an indicative spot rate, not a card-payment quote or a claim about today's executable exchange rate. Multiply any dollar rate by 0.7399 for the corresponding GBP rate. VAT, currency conversion fees, subscriptions, tools, cache storage, cache writes, hosting and retries are outside these uncached baskets.

“Not tabulated” does not mean free or unavailable. I have not assigned a model-specific cache-hit price where the checked card did not supply one clearly enough for this comparison. This is especially relevant to Mistral and Qwen, whose caching arrangements need their own check.

Three charts showing September GBP costs for fixed quick, general and frontier token baskets.
Compare within a category. Each category has a different workload and its own scale. These are prices, not quality scores.

The small print that changes the bill

Anthropic: a cancelled increase and cheaper cache reads

Anthropic's current pricing page says Sonnet 5's $2/$10 introductory input/output price is now standard. The planned increase to $3/$15 on 1 September will not happen. Fable 5.1 costs $10/$50, with cache reads at $0.25; five-minute cache writes cost $12.50 and one-hour writes $20. Those are different operations.

Fable 5.1 is a new version, not evidence that Fable 5's entire tariff fell. Mythos 5.1 has limited availability, so it is noted here rather than presented as an ordinary self-service choice. Opus 5 remains in the tables as a useful lower-priced comparator.

Google: put the promotion in your diary

Gemini 3.7 Flash arrived on 13 August and 3.8 Flash on 2 September. Their standard introductory tariff is $0.75 input, $0.075 cache hit and $3.75 output through 31 December 2026. Google lists $1.50/$0.15/$7.50 from 1 January 2027. Do not build a permanent budget around the temporary figure. Gemini 3.5 Flash-Lite is also newly included in this watch; that does not make it a September release.

Grok, Kimi and Mistral: more useful choices

Grok 4.6 was released on 12 August. At 200K input and above, its tariff rises to $4 input, $1 cached and $12 output. Grok Build 0.1 is included as a coding option. Its smaller context capacity is why it sits in the general basket.

Kimi's cards list K2.7 Code at the same $0.95/$4 input/output rate as K2.6, but with a different cache-hit rate. The coding model requires thinking mode; its high-speed variant is not the price shown here. K3 remains the larger comparator.

Mistral Medium 3.5 is added alongside Small 4 and Large 3. It is an existing model, not a new September launch. The API page also offers third-party GLM 5.2 at $1.40 input, $0.14 cache hit and $4.40 output. That is a Mistral-hosted Z.ai model, so I have kept it outside the direct model-maker comparison rather than quietly call it a Mistral model.

DeepSeek: what time are you running it?

DeepSeek's current card resolves Flash to version 0731 and Pro to 0813. The main tables use peak rates. Off-peak input/cache-hit/output prices are $0.22/$0.007/$0.66 for Flash and $0.66/$0.022/$1.98 for Pro, half the peak rates. Peak hours are Monday to Friday, 01:00–04:00 and 06:00–10:00 UTC; all other hours are off-peak.

That can matter for work which can wait. It does not mean every job is available at the off-peak headline price. The old snapshot did not establish the same version-and-time basis, so these are not plotted as simple like-for-like price increases.

Qwen: the endpoint is part of the price

Qwen's release list includes 3.8 Flash and 3.8 Max 0902. The rates here use Global deployment in Frankfurt for those two. Qwen3.7 Plus retains the Singapore International alias used in August: $0.40/$1.60 list less its published 20% promotion gives $0.32/$1.28. No expiry is specified there. Do not substitute a different region, dated endpoint or nightly discount and call it the same price.

August to September: the lines we can defend

Among the 11 matched entries, one uncached reference basket changed; ten were unchanged between the recorded snapshots. Sol's basket is down 25% in dollars. An unchanged dollar basket is about 0.36% cheaper in pounds using these two FX observations. That small movement is not a provider price cut.

A correction to the previous watch: August treated Haiku 4.5's $0.50/$2.50 input/output figures as standard. They match Anthropic's batch tariff, while its launch announcement and current standard card show $1/$5. I cannot substantiate the earlier standard-price reduction claim. September uses $1/$5 and excludes Haiku from the historical lines; it would be wrong to manufacture a doubling from that baseline.

The original snapshots remain frozen for traceability. New models and newly covered siblings have no invented past. DeepSeek's two entries are excluded because the previous version/time basis is not sufficiently comparable. That leaves 11 matched entries, three explicit exclusions and 12 entries new to this watch.

August to September lines for eleven recorded model IDs and token baskets. Sol falls; ten other uncached dollar baskets are unchanged.
Each line follows the same recorded model ID and basket; aliases do not guarantee unchanged underlying weights. Separate logarithmic GBP scales keep small and large costs visible; this is not a forecast.
Matched uncached baskets. New versions and unsupported historical comparisons are excluded.
CategoryProviderModelAugust £September £USD basket change
QuickOpenAIGPT-5.6 Luna0.002380.00237Unchanged
QuickMistralMistral Small 40.001560.00155Unchanged
GeneralOpenAIGPT-5.6 Terra0.118820.11838Unchanged
GeneralAnthropicClaude Sonnet 50.111390.11099Unchanged
GeneralxAIGrok 4.30.055700.05549Unchanged
GeneralMistralMistral Large 30.024130.02405Unchanged
GeneralMoonshot AIKimi K2.60.050130.04994Unchanged
GeneralAlibaba CloudQwen3.7 Plus0.016630.01657Unchanged
FrontierOpenAIGPT-5.6 Sol1.485201.10985Down 25%
FrontierGoogleGemini 3.1 Pro Preview1.076771.07285Unchanged
FrontierMoonshot AIKimi K30.835420.83239Unchanged

What I would do with this information

First, look at the actual tasks. Extraction. Classification. Drafting. Coding. Architecture. Checking the work. Those are not all the same job, and they do not all deserve the same model.

I would try the quick group on bounded, repeatable work with clear checks. I would compare the general group on normal delivery tasks. I would test Astra, Fable, Sol and the other larger options when the cost of a bad answer or a failed run justifies it. These are starting points for an evaluation, not promises about which provider wins.

Then measure cost per accepted result. If Astra finishes a task correctly with fewer loops, a higher price per token may still be worth paying. If a smaller model produces the same acceptable result, spending more has not made you cleverer. It has just made the bill bigger.

A useful test prompt for your harness is:

Compare two suitable models on this task. Keep the requirements and acceptance tests the same. Record uncached input, cache hits, cache writes, billable output including reasoning, tool charges, retries, elapsed time and whether the result passed. Recommend the lower-cost reliable route, and explain the uncertainty.

I would also set a spending limit before letting an agent run, and review the model routing monthly. Promotions change. Models change. Yesterday's sensible default can become today's unnecessary expense.

That is the point of this series. Not to crown a winner. To make the choices visible enough that we can spend deliberately.

Astra is welcome. So is a smaller bill.

Related: August's AI Model Pricing Watch, No Astra Yet? Excellent. Another Reset. and The Future Is Either ERP Or AI

Sources and method

This is a dated observation of provider-published prices, not independently audited invoices or a complete catalogue. I checked the previous eight-provider universe and expanded the text/coding shortlist. Image, audio, embedding, specialist cyber, limited-access and high-speed products are not interchangeable with these text-token baskets. Self-hosting costs are also outside this comparison.

Prices were checked against current official rate cards and available release notes. The source links in the tables are the billing authorities for each row; the notes identify promotions, context thresholds and regional exceptions. Calculations use the unrounded stored rates and the stated FX observation. Displayed figures are rounded. The baskets count billable tokens, not equal-length documents, identical quality or guaranteed task completion.

Before committing meaningful spend, check your chosen endpoint, account access, tariff and provider's current terms. This monthly watch is a starting point for that decision, not a substitute for it.