Benchmark · USD per 1M tokens · 173 modelsas of 2026-08-22
Token Prices
posted price by model · input & output · USD per 1M tokens · 2026-08-22
input output
Rows carrying an n are served medians across that many providers.
The rest are the model owners' posted prices. Two constructions,
never pooled.
01
Frontier posted — first-party
from the owners' price lists
OA OpenAIstandard tier
Model
Input
Cached
Output
30d
Last reprice
gpt-5.6-sol
$4.00
$0.40
$20.00
2026-08-22 (-33%)
gpt-5.6-terra
$2.00
$0.20
$12.00
2026-07-30 (-20%)
gpt-5.6-luna
$0.20
$0.02
$1.20
2026-07-30 (-80%)
gpt-5.5
$5.00
$0.50
$30.00
none in record
gpt-5.5-pro
$30.00
—
$180.00
none in record
gpt-5.4-mini
$0.75
$0.07
$4.50
none in record
gpt-5.4-nano
$0.20
$0.02
$1.25
none in record
gpt-4o
$2.50
$1.25
$10.00
none in record
gpt-4o-mini
$0.15
$0.07
$0.60
none in record
o3-pro
$20.00
—
$80.00
none in record
o3
$2.00
$0.50
$8.00
none in record
o4-mini
$1.10
$0.28
$4.40
none in record
gpt-4-turbo-2024-04-09
$10.00
—
$30.00
none in record
gpt-3.5-turbo-instruct
$1.50
—
$2.00
none in record
AN Anthropicmodel table
Model
Input
Cached
Output
30d
Last reprice
Claude Fable 5
$10.00
$1.00
$50.00
none in record
Claude Mythos 5
$10.00
$1.00
$50.00
none in record
Claude Opus 5
$5.00
$0.50
$25.00
none in record
Claude Sonnet 5
$2.00
$0.20
$10.00
none in record
Claude Haiku 4.5
$1.00
$0.10
$5.00
none in record
GO Googlestandard tier
Model
Input
Cached
Output
30d
Last reprice
gemini-3.7-flash
$0.75
$0.07
$3.75
none in record
gemini-3.5-flash-lite
$0.30
$0.03
$2.50
none in record
gemini-3.1-pro-preview
$2.00
$0.20
$12.00
none in record
gemini-3-flash-preview
$0.50
$0.05
$3.00
none in record
gemini-2.5-pro
$1.25
$0.13
$10.00
none in record
gemini-2.5-flash-lite-preview
$0.10
$0.01
$0.40
none in record
gemini-robotics-er-2
$2.00
$0.20
$10.00
none in record
gemini-robotics-er-2-streaming
$2.00
—
$10.00
none in record
gemini-2.5-computer-use-preview-10-2025
$1.25
—
$10.00
none in record
DS DeepSeekstandard price
Model
Input
Cached
Output
30d
Last reprice
deepseek-v4-flash
$0.44
$0.01
$1.32
2026-08-17 (371%)
deepseek-v4-pro
$1.32
$0.04
$3.96
2026-08-17 (355%)
deepseek-v4-flash-vision-exp
$0.44
$0.01
$1.32
—
none in record
MS Moonshotper-model pages
Model
Input
Cached
Output
30d
Last reprice
kimi-k3
$3.00
$0.30
$15.00
none in record
kimi-k2.7-code
$0.95
$0.19
$4.00
none in record
kimi-k2.7-code-highspeed
$1.90
$0.38
$8.00
none in record
moonshot-v1-128k
$2.00
—
$5.00
none in record
moonshot-v1-128k-vision-preview
$2.00
—
$5.00
none in record
XA xAImodel catalogue
Model
Input
Cached
Output
30d
Last reprice
grok-4.20-0309-non-reasoning
$1.25
$0.20
$2.50
none in record
grok-4.20-multi-agent-0309
$1.25
$0.20
$2.50
none in record
grok-4.20-0309-reasoning
$1.25
$0.20
$2.50
none in record
grok-build-0.1
$1.00
$0.20
$2.00
none in record
grok-4.6
$2.00
$0.50
$6.00
none in record
ME Metastandard tier
Model
Input
Cached
Output
30d
Last reprice
muse-spark-1.2
$1.25
$0.15
$4.25
none in record
AL Alibabainternational storefront
Model
Input
Cached
Output
30d
Last reprice
qwen3.7-max
$2.50
—
$7.50
none in record
qwen3.7-max-preview
$2.50
—
$7.50
none in record
qwen3.7-plus
$0.40
—
$1.60
none in record
qwen-plus-latest
$0.40
—
$1.20
none in record
qwen-turbo
$0.05
—
$0.20
none in record
ZA Z.AImodel table
Model
Input
Cached
Output
30d
Last reprice
GLM-5.3
$1.40
$0.26
$4.40
none in record
GLM-5-Turbo
$1.20
$0.24
$4.00
none in record
GLM-4.7-FlashX
$0.07
$0.01
$0.40
none in record
GLM-4.5-X
$2.20
$0.45
$8.90
none in record
GLM-4.5-Air
$0.20
$0.03
$1.10
none in record
GLM-4.5-AirX
$1.10
$0.22
$4.50
none in record
GLM-4-32B-0414-128K
$0.10
—
$0.10
none in record
GLM-5V-Turbo
$1.20
$0.24
$4.00
none in record
GLM-4.6V
$0.30
$0.05
$0.90
none in record
GLM-OCR
$0.03
—
$0.03
none in record
GLM-4.6V-FlashX
$0.04
$0.00
$0.40
none in record
MM MiniMaxstandard tier
Model
Input
Cached
Output
30d
Last reprice
MiniMax-M3
$0.30
$0.06
$1.20
none in record
MiniMax-M2.7-highspeed
$0.60
$0.06
$2.40
none in record
TM Thinking Machinesserverless inference
Model
Input
Cached
Output
30d
Last reprice
Inkling-Small
$0.30
$0.06
$1.20
none in record
Inkling
$1.00
$0.17
$4.05
none in record
XI Xiaomioverseas price list
Model
Input
Cached
Output
30d
Last reprice
mimo-v2.5-pro
$0.43
$0.00
$0.87
none in record
mimo-v2.5
$0.14
$0.00
$0.28
none in record
Scope: these are standard self-serve rates, the price a buyer pays
without a contract. Batch, provisioned-throughput and committed-spend
tiers are separate prices and are not pooled into these figures, so a
large buyer's realized rate sits below them. Cached input is shown
wherever the owner posts it.
One vendor, one day, three tiers: on 2026-07-30 the posted output
price fell 80% on luna and 20% on terra while sol held.
02
Open-weight models — by serving breadth
cross-provider medians · n ≥ 3 · 2026-08-22
Model
Providers
Input median
Output median
Output range
MiniMax M3
6
$0.30
$1.20
$0.96 – $2.40
DeepSeek V4-Pro
5
$1.32
$3.20
$2.55 – $3.96
DeepSeek V4-Flash
5
$0.14
$0.28
$0.18 – $1.32
gemma-4-31B-it
4
$0.26
$0.65
$0.38 – $1.15
gpt-oss-120b
4
$0.10
$0.42
$0.17 – $0.60
kimi-k3
4
$3.00
$15.00
$14.25 – $15.00
GLM-5
4
$1.00
$3.20
$2.08 – $3.20
DeepSeek V3.1
4
$0.41
$1.32
$0.95 – $4.50
MiniMax-M2
4
$0.30
$1.20
$1.02 – $1.20
MiniMax M2.5
4
$0.30
$1.20
$1.15 – $1.20
MiniMax-M2.7
4
$0.30
$1.20
$1.00 – $2.40
DeepSeek V4-Flash-0731
4
$0.17
$0.47
$0.18 – $1.32
Kimi K2 Instruct
3
$0.57
$2.30
$2.00 – $2.50
Qwen3 Coder 480B A35B Instruct
3
$0.40
$1.60
$1.55 – $1.80
DeepSeek-V3.2
3
$0.27
$0.40
$0.38 – $4.50
qwen3.7-max Currently equivalent to qwen3.7-max-2026-05-20 context caching discount
3
$2.50
$7.50
$3.75 – $7.50
Each row is a median across independent providers of the same model,
n disclosed. A wide range is dispersion in the serving market.
Some models sit in both sections: the owner posts a price and
independent providers also serve the open weights. The two numbers are
different measurements, not a restatement. deepseek-v4-flash is
posted at $1.32 per 1M output tokens above,
and the median of independent providers here is
$0.28. What a model owner asks and what
the serving market charges are separate facts, and neither is pooled
into the other.
Method
USD per 1 million tokens, as posted. Record runs from 2026-07-03.
Posted rows come from the model owners' own price lists, captured
daily; a reprice appears on the day the list changed.
Served rows are medians across independent providers of the same
open-weight model, published at n ≥ 3.
The two are never pooled. A posted price and a cross-provider
median answer different questions.
Token prices are not evidence about GPU rental rates. The link
between them is serving efficiency, a vendor variable rather than a
market price. CCIR publishes the GPU-hour record separately.