Your team defaults to the biggest model because nobody can police the choice. PerDollar answers it per task instead — the cheapest model that gets the work accepted, from the ones you can actually reach, inside your budget and inside your jurisdiction. Then it records what that decision saved.
Every engineer picks the biggest model because it is right there and the cost is somebody else's line item. A policy telling them otherwise is a nag nobody enforces. The decision is too frequent and too small to govern by hand — so it has to be answered where it is made.
Asking engineers to justify a model per task is a tax on everyone. The choice has to be made for them, at the moment of use, or it will not be made at all.
Most teams picked a model at launch. New models land every couple of days and prices move underneath you — the default stopped being the right answer months ago.
For a European customer, "must stay in the EU" is binding before price is discussed. Other routers can send sensitive traffic to an internal model; none can tell you which EU vendor is cheapest, because that pricing is published nowhere.
A mid-size merchant resolves ~50,000 tickets a month with an AI first-line bot. Each resolution reads the ticket, customer history and policy (~2,500 tokens) and writes a reply (~450 tokens).
Read straight down the column: the same 50,000 resolutions swing from $1,300 to under $110 a month by right-sizing the model, and lower again with caching — a ~$14,000/year line item on one workload. The caveat we always attach: a cheaper model that fails the task isn't cheaper, so every switch is validated on a 100-ticket sample before you commit — which is why the router records outcomes rather than assuming them.
Ask a chatbot what your workload costs across models and it will answer confidently and wrong — stale prices, invented rates, no accounting for answer length. PerDollar is built the opposite way.
Ramp, CloudZero and Vantage will all show you an AI cost dashboard. They're built to report spend across a whole company — not to tell a commerce team which model to run for a support resolution or a product description, or whether the cheap one can actually do the job.
Your agent asks before it starts and reports when it finishes. No proxy sits in the request path, so nothing slows down and no prompt leaves your infrastructure.
The cheapest model that will get this task accepted first time, from the models you can reach. Data residency is a hard filter applied before cost, so a Germany-only team is routed to German infrastructure rather than told to be careful.
An engineer no longer costs a salary. They cost a salary plus whatever their tools burn in inference, and that half moves weekly. Give it $500 per developer per month and it returns the most capable model that fits, not the cheapest — and as the month's spend accumulates the ceiling falls and the pick downgrades on its own.
Every routed task records what it cost, what your engineer's default would have cost, and whether it needed doing again — a retry or an escalation is an observable fact, not a judgement. Cheap output that gets redone twice is not cheap, and nothing else measures that.
Every price links to the provider's own page and says who confirmed it and when. Where we cannot source a number we say so rather than publish it. Corrections to our own errors are published as corrections.
Model routing is not new. What is missing is a router that is not also selling you the inference, locked inside one editor, or optimising its own margin.
Nobody can verify that. Sites tracking 300+ models are republishing whatever the aggregators say, which is fine until a number is wrong and no one can tell you who checked it. We track fewer models on purpose, and say who confirmed each price and when.
Current frontier and budget tiers from the major labs, plus every EU-resident host we can find. New releases enter a review queue, not the ledger.
Image, video and audio models, superseded versions, and anything we cannot source from the provider's own page. Breadth we cannot verify is breadth we would not defend.
Verified by a human on a date, confirmed by the daily agent, tracked from a published comparison, or source unconfirmed. Four states, never blurred into one.
Two separate regimes bear on AI model choice, and they are often confused. PerDollar helps with one of them directly and gives you evidence for the other. It does not make you compliant with either — no tool can, and anyone claiming otherwise is selling you something.
The dates, plainly. Article 50 has applied since 2 August 2026; penalties reach €15 million or 3% of worldwide annual turnover, whichever is higher. A grace period runs to 2 December 2026 for machine-readable marking of generative systems already on the market before August. Annex III high-risk obligations were deferred to 2 December 2027. The Act reaches beyond Europe: it applies to providers and deployers whose AI outputs are used in the EU, wherever the company sits. We are not lawyers, and this is a summary, not advice.
No sign-up, nothing to install. Pick a task and a jurisdiction and see what comes back — the same answer your agent would get.
When you want it inside your workflow: one MCP server for coding agents, or a plain HTTP call from anything else.
Task type, your available models, your budget and your jurisdiction go in. A model, a reason and a runner-up come back.
Actual tokens, actual cost, and the model your engineer would otherwise have used. The counterfactual is computed from both.
Spent, would-have-spent, saved — from your own traffic. The kind of figure a finance team can check rather than take on trust.
Ask the router one question. If the answer is useful, wire it into your agent in a line. If your team spends enough for this to matter, come and talk.
Try the router