The independent decision layer for AI model choice

Which model should run this task?

Your team defaults to the biggest model because nobody can police the choice. PerDollar answers it per task instead — the cheapest model that gets the work accepted, from the ones you can actually reach, inside your budget and inside your jurisdiction. Then it records what that decision saved.

Free, no sign-up. Independent of every model vendor — we do not sell inference.
ASKED: one coding-agent step · must stay in Germany 20,000 STEPS / MONTH
What most teams run
Frontier model, chosen once at launch
$3,205
per month
What the router returns
mistral-small @ IONOS · German data centre
$28
per month, same work, EU-resident
The difference is roughly a developer and a half a year, in salary terms, for identical work. Residency is applied before cost — ask for no US processing and a US-operated API never appears, however cheap it is.
Prices checked daily by a bot ·24 models + EU-resident hosts ·Residency: any · EU · Germany
The problem

Model choice is made a thousand times a day. Nobody is watching.

Every engineer picks the biggest model because it is right there and the cost is somebody else's line item. A policy telling them otherwise is a nag nobody enforces. The decision is too frequent and too small to govern by hand — so it has to be answered where it is made.

01

Governance by policy does not work

Asking engineers to justify a model per task is a tax on everyone. The choice has to be made for them, at the moment of use, or it will not be made at all.

02

Set once, never revisited

Most teams picked a model at launch. New models land every couple of days and prices move underneath you — the default stopped being the right answer months ago.

03

Compliance beats cost, and nobody prices it

For a European customer, "must stay in the EU" is binding before price is discussed. Other routers can send sensitive traffic to an internal model; none can tell you which EU vendor is cheapest, because that pricing is published nowhere.

A worked example · what the decision is worth

What one resolved ticket costs, across models.

A mid-size merchant resolves ~50,000 tickets a month with an AI first-line bot. Each resolution reads the ticket, customer history and policy (~2,500 tokens) and writes a reply (~450 tokens).

ModelCost / resolutionMonthly costVs. frontier default
GPT-5.5 (default)$0.0260$1,300
Claude Haiku 4.5$0.0048$240−$1,060 / mo
Gemini 3 Flash$0.0022$108−$1,192 / mo
+ prompt caching on policy text$0.0014$70−$1,230 / mo

Read straight down the column: the same 50,000 resolutions swing from $1,300 to under $110 a month by right-sizing the model, and lower again with caching — a ~$14,000/year line item on one workload. The caveat we always attach: a cheaper model that fails the task isn't cheaper, so every switch is validated on a 100-ticket sample before you commit — which is why the router records outcomes rather than assuming them.

Why us

Money decisions can't run on data that's right half the time.

Ask a chatbot what your workload costs across models and it will answer confidently and wrong — stale prices, invented rates, no accounting for answer length. PerDollar is built the opposite way.

Ask an LLM / read a price chart
Confidently wrong
  • Training-cutoff prices, often months out of date
  • Per-token numbers that ignore how much answer you actually get
  • No view of caching, batch, or promo-vs-permanent pricing
  • No idea whether the cheap model can do your job
PerDollar
Verified & job-costed
  • Prices checked against first-party sources, refreshed daily
  • Cost per finished job, adjusted for each model's answer length
  • Caching, batch and launch-promo pricing tracked as distinct, with lifespans
  • Every switch validated on your real traffic before you commit

And why not the horizontal spend tools?

Ramp, CloudZero and Vantage will all show you an AI cost dashboard. They're built to report spend across a whole company — not to tell a commerce team which model to run for a support resolution or a product description, or whether the cheap one can actually do the job.

Horizontal AI spend dashboards
Visibility, not decisions
  • Tell you what you spent, after the fact — not what to switch to
  • Route or rank on their own generic benchmarks, not your workloads
  • Built for the whole company; commerce jobs are one unlabelled slice
  • No independent, daily-checked price record you can audit
PerDollar
Commerce decisions, validated
  • Names the switch and the saving, per commerce workload
  • Costs your support / catalogue / search jobs on your real traffic
  • Every recommendation validated on a sample before you commit
  • An independent price record, checked by a bot every day, with history no dashboard keeps
What it does

Three questions, answered where the work happens.

Your agent asks before it starts and reports when it finishes. No proxy sits in the request path, so nothing slows down and no prompt leaves your infrastructure.

01Decide — with jurisdiction first

The cheapest model that will get this task accepted first time, from the models you can reach. Data residency is a hard filter applied before cost, so a Germany-only team is routed to German infrastructure rather than told to be careful.

02Stay inside a budget

An engineer no longer costs a salary. They cost a salary plus whatever their tools burn in inference, and that half moves weekly. Give it $500 per developer per month and it returns the most capable model that fits, not the cheapest — and as the month's spend accumulates the ceiling falls and the pick downgrades on its own.

03Count the work that stuck

Every routed task records what it cost, what your engineer's default would have cost, and whether it needed doing again — a retry or an escalation is an observable fact, not a judgement. Cheap output that gets redone twice is not cheap, and nothing else measures that.

04On prices you can check

Every price links to the provider's own page and says who confirmed it and when. Where we cannot source a number we say so rather than publish it. Corrections to our own errors are published as corrections.

Why not the others

Everyone routing your traffic has a stake in the answer.

Model routing is not new. What is missing is a router that is not also selling you the inference, locked inside one editor, or optimising its own margin.

Vendor and IDE routers
Optimising their margin
  • Free routers from Ramp and Cursor decide inside their own product, on their own definition of a quality bar, and report their own savings figure
  • A gateway that sells inference, or a card company that monetises spend, has a reason to prefer particular answers
  • Savings are vendor-stated from vendor A/B tests, not measured on your traffic
  • No concept of data residency, which is the binding constraint for European buyers
PerDollar
No inference to sell
  • Works wherever your work happens — agents, support, catalogue, search
  • Independent: we make nothing whichever model wins
  • Savings computed against the model your engineer would have chosen
  • Residency applied before cost — over EU-resident hosts whose prices no aggregator carries
Coverage

A new model ships roughly every ten hours.

Nobody can verify that. Sites tracking 300+ models are republishing whatever the aggregators say, which is fine until a number is wrong and no one can tell you who checked it. We track fewer models on purpose, and say who confirmed each price and when.

WHAT WE ADD

Models people actually route to

Current frontier and budget tiers from the major labs, plus every EU-resident host we can find. New releases enter a review queue, not the ledger.

WHAT WE SKIP

The long tail

Image, video and audio models, superseded versions, and anything we cannot source from the provider's own page. Breadth we cannot verify is breadth we would not defend.

WHAT YOU SEE

The label, always

Verified by a human on a date, confirmed by the daily agent, tracked from a published comparison, or source unconfirmed. Four states, never blurred into one.

Compliance

Where your data goes is a legal question, not a preference.

Two separate regimes bear on AI model choice, and they are often confused. PerDollar helps with one of them directly and gives you evidence for the other. It does not make you compliant with either — no tool can, and anyone claiming otherwise is selling you something.

Where PerDollar helps directly
Data transfers
  • Under GDPR Chapter V, sending personal data to a US-operated API is a restricted international transfer that needs a legal basis
  • PerDollar filters by where the operator actually processes: Germany only, EU, or no constraint
  • It carries EU-resident hosting — IONOS, Scaleway, OVHcloud, STACKIT, Aleph Alpha, Mistral, Nebius, Gcore — that no pricing aggregator lists
  • So "which vendor may we legally use, and what does it cost" becomes one answer instead of two separate investigations
Where it gives you evidence, not compliance
AI governance
  • The EU AI Act's Article 50 transparency duties applied from 2 August 2026 — disclosing AI interaction, marking AI-generated content, labelling deepfakes
  • Those are obligations on your product, not on your model choice, and PerDollar does not discharge them
  • What it does give you is a record: which model ran which task, why it was chosen, and what it cost — an audit trail rather than a recollection
  • Useful for internal governance and for answering auditors. Not a substitute for legal advice

The dates, plainly. Article 50 has applied since 2 August 2026; penalties reach €15 million or 3% of worldwide annual turnover, whichever is higher. A grace period runs to 2 December 2026 for machine-readable marking of generative systems already on the market before August. Annex III high-risk obligations were deferred to 2 December 2027. The Act reaches beyond Europe: it applies to providers and deployers whose AI outputs are used in the EU, wherever the company sits. We are not lawyers, and this is a summary, not advice.

Ask it a question right now.

No sign-up, nothing to install. Pick a task and a jurisdiction and see what comes back — the same answer your agent would get.

When you want it inside your workflow: one MCP server for coding agents, or a plain HTTP call from anything else.

IN THE BROWSER
Try the router
A form, not a terminal. Shows the equivalent API call as you change it.
IN YOUR AGENT
claude mcp add perdollar
Your agent asks before each task and reports what it cost afterwards.
Try the router →
1BEFORE THE TASK

Your agent asks

Task type, your available models, your budget and your jurisdiction go in. A model, a reason and a runner-up come back.

2AFTER THE TASK

It reports the real numbers

Actual tokens, actual cost, and the model your engineer would otherwise have used. The counterfactual is computed from both.

3END OF MONTH

You have a number, not a claim

Spent, would-have-spent, saved — from your own traffic. The kind of figure a finance team can check rather than take on trust.

Stop guessing which model to run.

Ask the router one question. If the answer is useful, wire it into your agent in a line. If your team spends enough for this to matter, come and talk.

Try the router
or book 30 minutes →