Open-weight models have no single price. The same weights, served by a
different vendor, can cost several times more — and prompt caching moves the ranking
again. Pick a model, a job, and your cache-hit rate.
MODELJOBCACHE HIT RATE0%
COST PER JOB — RANGE ACROSS HOSTS
HOST
IN $/M
OUT $/M
CACHED IN $/M
COST / JOB
JOBS PER $100
VS CHEAPEST
What this comparison does not tell you
A price table that hides its own limits is worse than no table. Four things
to check before moving a workload on the strength of this page.
These are list prices. Any committed-spend or negotiated rate you hold directly
with a vendor won't appear here, and it can beat every number in the table.
Dedicated and self-hosted deployments aren't included. Above a certain sustained
volume, renting GPUs by the hour beats per-token pricing entirely — a different calculation.
Throughput and latency differ enormously between hosts for the same weights.
Speed specialists charge more per token and may still be the right answer for interactive
traffic; the cheapest host is not automatically the best host.
Quantisation, context limits and rate limits vary by host. Same model name does
not always mean an identical serving configuration — worth confirming before you switch.