price against intelligence

Which model delivers the most per dollar?

Fifteen models from eight providers, their price per million tokens and their score on an independent index. Choose your scale or enter your own usage, and everything recalculates live.

As of 5 September 2026. Prices and models change quickly; check the provider’s official page before you pay.

The comparison

Switch providers on and off, slide the limits for price and score, and choose a scale of one, ten or twenty million tokens in and out. Without JavaScript the whole table is there too, at the setting of one million.

Provider
What you weigh the price against
Scale: million tokens in and out
Or your own usage per month

15 of 15 models

cheapest in your selectionGLM-5.3-Flash$0.33 per month
highest index in your selectionClaude Fable 5.1index 57
most expensive versus cheapest185 timesas expensive, at this scale
Fifteen models with their index score, their price per million tokens and the cost at your usage, with source and check date per value
Checked
GLM-5.3-FlashZ.ai · on the frontier4647.5source for the index score of GLM-5.3-Flash, opens in a new tabinput per 1M$0.075the page notes a temporary discount of 50 percent heresource for the input price of GLM-5.3-Flash, opens in a new taboutput per 1M$0.25same temporary discountsource for the output price of GLM-5.3-Flash, opens in a new tab$0.33
GPT-5.6 LunaOpenAI · on the frontier43133.9setting maxsource for the index score of GPT-5.6 Luna, opens in a new tabinput per 1M$0.20source for the input price of GPT-5.6 Luna, opens in a new taboutput per 1M$1.20source for the output price of GPT-5.6 Luna, opens in a new tab$1.40
MiniMax-M3MiniMax · on the frontier3686.6source for the index score of MiniMax-M3, opens in a new tabinput per 1M$0.30up to 512k tokens per request; above that $0.60source for the input price of MiniMax-M3, opens in a new taboutput per 1M$1.20up to 512k tokens per request; above that $2.40source for the output price of MiniMax-M3, opens in a new tab$1.50
DeepSeek V4 ProDeepSeek · on the frontier4266.4setting max, version 0813source for the index score of DeepSeek V4 Pro, opens in a new tabinput per 1M$0.66off-peak hours without a cache hit; double in peak hourssource for the input price of DeepSeek V4 Pro, opens in a new taboutput per 1M$1.98off-peak hours; in peak hours $3.96source for the output price of DeepSeek V4 Pro, opens in a new tab$2.64
Gemini 3.5 Flash-LiteGoogle · on the frontier28370.6source for the index score of Gemini 3.5 Flash-Lite, opens in a new tabinput per 1M$0.30source for the input price of Gemini 3.5 Flash-Lite, opens in a new taboutput per 1M$2.50source for the output price of Gemini 3.5 Flash-Lite, opens in a new tab$2.80
Gemini 3.7 FlashGoogle · on the frontier45310.5setting highsource for the index score of Gemini 3.7 Flash, opens in a new tabinput per 1M$0.75up to and including 31 December 2026; after that $1.50source for the input price of Gemini 3.7 Flash, opens in a new taboutput per 1M$3.75up to and including 31 December 2026; after that $7.50source for the output price of Gemini 3.7 Flash, opens in a new tab$4.50
GLM-5.3Z.ai · on the frontier4980.0setting maxsource for the index score of GLM-5.3, opens in a new tabinput per 1M$1.40source for the input price of GLM-5.3, opens in a new taboutput per 1M$4.40source for the output price of GLM-5.3, opens in a new tab$5.80
Grok 4.6xAI · on the frontier5163.9setting highsource for the index score of Grok 4.6, opens in a new tabinput per 1M$2.00up to 200k tokens; above that $4.00source for the input price of Grok 4.6, opens in a new taboutput per 1M$6.00up to 200k tokens; above that $12.00source for the output price of Grok 4.6, opens in a new tab$8.00
GPT-5.6 TerraOpenAI · on the frontier47108.5setting maxsource for the index score of GPT-5.6 Terra, opens in a new tabinput per 1M$2.00source for the input price of GPT-5.6 Terra, opens in a new taboutput per 1M$12.00source for the output price of GPT-5.6 Terra, opens in a new tab$14.00
Kimi K3Moonshot AI · on the frontier5038.7setting maxsource for the index score of Kimi K3, opens in a new tabinput per 1M$3.00price without a cache hit; with a hit $0.30source for the input price of Kimi K3, opens in a new taboutput per 1M$15.00source for the output price of Kimi K3, opens in a new tab$18.00
GPT-5.6 SolOpenAI · on the frontier5180.9setting maxsource for the index score of GPT-5.6 Sol, opens in a new tabinput per 1M$4.00source for the input price of GPT-5.6 Sol, opens in a new taboutput per 1M$20.00source for the output price of GPT-5.6 Sol, opens in a new tab$24.00
Claude Opus 5Anthropic · on the frontier5457.5measured on the setting Adaptive Reasoning, Max Effortsource for the index score of Claude Opus 5, opens in a new tabinput per 1M$5.00source for the input price of Claude Opus 5, opens in a new taboutput per 1M$25.00source for the output price of Claude Opus 5, opens in a new tab$30.00
Claude Fable 5.1Anthropic · on the frontier5768.1setting Adaptive Reasoning, Max Effort, with standard fallbacksource for the index score of Claude Fable 5.1, opens in a new tabinput per 1M$10.00source for the input price of Claude Fable 5.1, opens in a new taboutput per 1M$50.00source for the output price of Claude Fable 5.1, opens in a new tab$60.00
Claude Fable 5Anthropic · on the frontier5370.0setting Max Effort with fallback to Opus 4.8source for the index score of Claude Fable 5, opens in a new tabinput per 1M$10.00source for the input price of Claude Fable 5, opens in a new taboutput per 1M$50.00source for the output price of Claude Fable 5, opens in a new tab$60.00
GPT-6 AstraOpenAI · on the frontier55–setting maxArtificial Analysis reports Speed N/A here: this model has not been measured for speed yetsource for the index score of GPT-6 Astra, opens in a new tabinput per 1M$10.00source for the input price of GPT-6 Astra, opens in a new taboutput per 1M$50.00source for the output price of GPT-6 Astra, opens in a new tab$60.00

The ring next to each model is the index score on the scale of this table, from 28 to 57. Point at a model or move into it with the keyboard, and the input and output prices slide open with their source. Your cost is our sum of those two prices times your usage; no provider publishes this amount in this form. A note on a price is not small print but a condition that changes the outcome.

What on the frontier means

Five of the fifteen carry that label at one million tokens in and out: GLM-5.3-Flash, GLM-5.3, Grok 4.6, Claude Opus 5, Claude Fable 5.1. For those five, there is no model in this list that both costs less and scores higher. That follows directly from the two numbers in the row, and it shifts when you change the ratio between input and output; that is why the table recalculates it live.

It does not say that such a model suits you. A model that scores four points lower but costs ten times less is often the better choice for the work you do. The label only tells you where you are not paying without getting something back for it.

Where the scores come from

The index comes from the Artificial Analysis Intelligence Index v4.2. A composite score from ten evaluations (AA-Briefcase, GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1), scaled from 0 to 100. Higher is better. Artificial Analysis is an independent measuring party and not the provider; for a benchmark score that is exactly the right source, because the score only exists at the party that measures it. On 5 September 2026 the index moved from v4.1.1 to v4.2: an evaluation was added and the weighting changed, so every model on this page comes out nine to eleven points lower than in August. The order of the thirteen models listed at the time did not change. A score from v4.1.1 and one from v4.2 are therefore not comparable.

The speed comes from the same party: Artificial Analysis Output Speed. Output tokens per second, measured on the provider’s own API. Higher is better. It is a median over repeated measurements and not a guarantee: what you get depends on your region, your prompt length and how busy it is at that moment. Speed and intelligence do not rise together, and that is exactly why they sit side by side here: the fastest model on this page is not the smartest, and the smartest is not the cheapest.

What is deliberately not in here is the Coding Agent Index from the same party. It does not rank models but agent variants: a combination of model, settings and the way the agent runs. Two lines in that list can be the same model with a different configuration, and then there is no price per million tokens to set beside it. It is here as a link, because for the question of which agent codes well it is the better source; as a column in this table it would compare apples with oranges.

Everywhere else on this site, a value has to come from the provider’s own page. For a benchmark score that rule is reversed, and that is the same rule and not an exception: a price exists with whoever charges it, a score exists with whoever measures it. Our check breaks the build if a score points to a provider or a price points to the measuring party.

What is deliberately not here

A row only goes in if both numbers could be checked at the source: the price on the provider’s page and the score on the measuring party’s. That is why there are fifteen models here and not everything we came across.

  • Qwen3.8-Max from AlibabaThe price is there ($2.00 input and $6.00 output per million, international rate on Alibaba Cloud’s own page), but the measuring party lists its scores under a different model name. We could not establish that link from a source, and this table does not guess.
  • Cheaper siblings such as Claude Sonnet 5, DeepSeek V4 Flash and Qwen3.8-FlashFor those we have the price from the source, but we did not get the score from the source. A row with half a box starts to look like a judgement, while it only says something about us.
  • No judgement of qualityThe index is one measurement by one party across ten evaluations. Whether a model does your work well, you only know once you let it do your work.
  • No converted amountsEverything is in dollars, as the providers show it. Converting would put today’s exchange rate into an amount that is no longer right tomorrow.
  • No subscriptionsThis page is about individual tokens through an API. Monthly subscriptions such as ChatGPT Plus or Claude Pro are on the subscription comparison.

Further reading

Three places that connect to what you are looking for here.

From first prompt to tax return.

basestep is not in this comparison and is not compared against anything here.