The comparison
Switch providers on and off, slide the limits for price and score, and choose a scale of one, ten or twenty million tokens in and out. Without JavaScript the whole table is there too, at the setting of one million.
15 of 15 models
No model fits within these limits. Widen them a little or click reset all.
The ring next to each model is the index score on the scale of this table, from 28 to 57. Point at a model or move into it with the keyboard, and the input and output prices slide open with their source. Your cost is our sum of those two prices times your usage; no provider publishes this amount in this form. A note on a price is not small print but a condition that changes the outcome.
What on the frontier means
Five of the fifteen carry that label at one million tokens in and out: GLM-5.3-Flash, GLM-5.3, Grok 4.6, Claude Opus 5, Claude Fable 5.1. For those five, there is no model in this list that both costs less and scores higher. That follows directly from the two numbers in the row, and it shifts when you change the ratio between input and output; that is why the table recalculates it live.
It does not say that such a model suits you. A model that scores four points lower but costs ten times less is often the better choice for the work you do. The label only tells you where you are not paying without getting something back for it.
Where the scores come from
The index comes from the Artificial Analysis Intelligence Index v4.2. A composite score from ten evaluations (AA-Briefcase, GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1), scaled from 0 to 100. Higher is better. Artificial Analysis is an independent measuring party and not the provider; for a benchmark score that is exactly the right source, because the score only exists at the party that measures it. On 5 September 2026 the index moved from v4.1.1 to v4.2: an evaluation was added and the weighting changed, so every model on this page comes out nine to eleven points lower than in August. The order of the thirteen models listed at the time did not change. A score from v4.1.1 and one from v4.2 are therefore not comparable.
The speed comes from the same party: Artificial Analysis Output Speed. Output tokens per second, measured on the provider’s own API. Higher is better. It is a median over repeated measurements and not a guarantee: what you get depends on your region, your prompt length and how busy it is at that moment. Speed and intelligence do not rise together, and that is exactly why they sit side by side here: the fastest model on this page is not the smartest, and the smartest is not the cheapest.
What is deliberately not in here is the Coding Agent Index from the same party. It does not rank models but agent variants: a combination of model, settings and the way the agent runs. Two lines in that list can be the same model with a different configuration, and then there is no price per million tokens to set beside it. It is here as a link, because for the question of which agent codes well it is the better source; as a column in this table it would compare apples with oranges.
Everywhere else on this site, a value has to come from the provider’s own page. For a benchmark score that rule is reversed, and that is the same rule and not an exception: a price exists with whoever charges it, a score exists with whoever measures it. Our check breaks the build if a score points to a provider or a price points to the measuring party.
What is deliberately not here
A row only goes in if both numbers could be checked at the source: the price on the provider’s page and the score on the measuring party’s. That is why there are fifteen models here and not everything we came across.
- Qwen3.8-Max from AlibabaThe price is there ($2.00 input and $6.00 output per million, international rate on Alibaba Cloud’s own page), but the measuring party lists its scores under a different model name. We could not establish that link from a source, and this table does not guess.
- Cheaper siblings such as Claude Sonnet 5, DeepSeek V4 Flash and Qwen3.8-FlashFor those we have the price from the source, but we did not get the score from the source. A row with half a box starts to look like a judgement, while it only says something about us.
- No judgement of qualityThe index is one measurement by one party across ten evaluations. Whether a model does your work well, you only know once you let it do your work.
- No converted amountsEverything is in dollars, as the providers show it. Converting would put today’s exchange rate into an amount that is no longer right tomorrow.
- No subscriptionsThis page is about individual tokens through an API. Monthly subscriptions such as ChatGPT Plus or Claude Pro are on the subscription comparison.
Further reading
Three places that connect to what you are looking for here.