Meta / Contents
The Cheap Model Was Fine. Proving It Was the Hard Part.
I built a gateway that sends each prompt to the cheapest Gemini model that can answer it, then benchmarked it on 300 prompts with a Pro-tier judge. The cheap model held up. My test set, my verifier, and my first table of numbers did not. Six lessons for anyone paying an AI bill.
Every team shipping something on top of a large language model eventually has the same argument. Someone wants the flagship model everywhere, because quality. Someone else points at the invoice. Nobody in the room has measured anything.
I've had that argument with myself. Last month I wrote about replacing a regex with a real model router inside my agent platform, and the design rested on a claim I had never tested: cheap models can handle most of the work, if something reliable decides which work. So I spent a week in the AI Labs building the smallest thing that could test it, and the results changed my mind about which half of that sentence is hard.
Short version: the cheap model was fine. Everything I built to prove it was fine broke at least once.
The 8x problem
Google's Gemini comes in three tiers that matter here. Flash-Lite is the cheap, fast one. Flash is the middle. Pro is the flagship that thinks before it answers.
| Tier | Input, per million tokens | Output, per million tokens |
|---|---|---|
| Flash-Lite | $0.25 | $1.50 |
| Flash | $0.50 | $3.00 |
| Pro | $2.00 | $12.00 |
Pro costs eight times what Flash-Lite does on the sticker. The real gap is worse, because Pro thinks, and Google bills thinking tokens at the output rate. On my benchmark the all-Pro bill came out 34 times the all-Flash-Lite bill for the same 300 prompts. Same questions, same number of requests.
So the temptation is obvious. Send the easy stuff to the cheap model, send the hard stuff to Pro, pay a fraction. The catch is the word decide. Something has to look at each prompt and pick a tier, and that something can be wrong in both directions. Route a hard question to Flash-Lite and you get a confident, wrong answer. Route everything to Pro "to be safe" and you're back where you started.
The question I wanted answered: can a gateway pick the tier per request, and can I prove with numbers that it picks well?
The experiment, in one breath
I spec'd the project on September 8 and had it live two days later, built in six phases by coding agents against written specs. The real benchmark ran on the 11th. The whole thing is open at github.com/jking-ai/thrifty-router, with a live report at thrifty-router.jking.ai.
Thrifty Router is a small web service on Cloud Run that sits in front of the three Gemini tiers. A request comes in, the gateway picks a tier using one of four strategies, calls the model, and returns the answer with the exact cost in the response headers. A daily budget ceiling and a cache sit in front of all of it.
The four strategies, in plain language:
- Fixed. Always use one tier. This is the control group: all-Lite, all-Flash, all-Pro.
- Semantic. Turn the prompt into an embedding (a list of numbers that captures its meaning) and compare it against a handful of example prompts for each tier. Closest match wins. No model call needed to decide.
- Classifier. Ask the cheapest model, "how hard is this?" and route on its answer. One small extra call per request.
- Cascade. Try the cheapest tier first. Check the answer. If it fails the check, escalate to the next tier and try again. This is the strategy the FrugalGPT paper made famous, and it was the one I expected to win.
To score them I needed a golden set: a fixed collection of prompts with grading rubrics, so every strategy answers the same questions. I wrote 300 of them across eight everyday task types: summarizing, classifying, answering questions about a document, explaining code, math, generating diagrams, parsing grading rubrics, and creative writing. Each one carries a label for the tier I expected it to need.
Then an LLM judge (Gemini Pro, run at zero temperature so it's as deterministic as it gets) scored every answer from 1 to 5 against the rubric, with a reference answer where one exists. Six strategies, 300 prompts, 1,800 answers, 1,800 judgments.
That's the setup. Here's what it taught me.
Lesson 1: My first test set was padding
I didn't find this by being careful. I found it by building a feature.
The report page said "300 curated prompts," and I wanted readers to be able to browse them instead of digging through a JSON file in the repo. So I had an agent build a little explorer with filters by category and tier. The moment it rendered, the problem was on the screen.
The set had 85 distinct prompts. The other 215 were copies with a suffix like "(Variation 2 focusing on specific context 7)" bolted on. Some base prompts appeared nine times. The rubrics repeated the same way. The original synthesis step had been asked for 300 items and had hit the number the lazy way.
Every published figure rested on that set. So I threw the set and the numbers out, hand-wrote 300 self-contained prompts (114 with reference answers), and reran the entire benchmark.
The lesson isn't "check your data," although, yes. It's that publishing the data is what forced the check. As long as the golden set was a file in a repo, nobody was ever going to sort it by category. Put it on a web page with filters and the padding is the first thing you see. If you're going to make a claim from a dataset, make the dataset as visible as the claim.
Lesson 2: The cheap model kept 98% of the quality at 3% of the cost
Here's the real table, from the rerun.
| Strategy | Judge score (1 to 5) | Cost per 1,000 requests | Cost vs all-Pro |
|---|---|---|---|
| Fixed: Flash-Lite | 4.82 | $0.50 | 3% |
| Fixed: Flash | 4.90 | $2.78 | 16% |
| Fixed: Pro | 4.90 | $16.98 | 100% |
| Semantic | 4.88 | $6.61 | 39% |
| Classifier | 4.84 | $8.30 | 49% |
| Cascade | 4.82 | $0.53 | 3% |
Flash-Lite alone scored 4.82 against Pro's 4.90. That's 98% of the flagship's judge score for 3% of its cost. Flash, the middle tier, matched Pro outright at 16% of the cost.
The honest headline, then: on these everyday tasks you can cut the bill by 61% to 84% with no measurable quality loss, or by 97% if you'll accept a 2% drop. That first version of the benchmark had claimed 77% savings for the cascade. The real number for the same quality is better, and it comes from a boring strategy: pick the middle tier and stop.
Two smaller surprises. Flash-Lite answered in about 2.5 seconds at the median. Pro took over 10. Speed came free with the savings.
The other: Pro lost to Flash-Lite, by a hair, on math, code explanation, and creative writing. My best read is that Pro writes long. It thinks at length, then answers at length, and ten of its answers ran into the output cap even after I raised it to 8,192 tokens. A short, correct answer scores a 5. A long, correct answer that gets cut off does not.
Lesson 3: Don't ask a model to grade its own work
The cascade was supposed to be the star, and its mechanism is elegant on paper.
Send the prompt to Flash-Lite with one extra instruction: after your answer, on the last line, write CONFIDENCE: followed by a number from 0 to 100. If the number is under 70, throw the answer away and escalate to Flash. If Flash is also unsure, escalate to Pro. No extra model calls, no separate judge on the hot path. The answerer grades itself.
Across 300 prompts, including formal mathematical proofs, Flash-Lite reported a confidence of 95 or higher every single time. The gate never fired. Escalation rate: zero. The cascade was Fixed Flash-Lite with a longer prompt, and it cost slightly more because of the longer prompt.
I don't think the model was lying. I think it had no idea. Self-reported confidence is a number the model generates the same way it generates everything else, by predicting what text should come next. Nothing in that process consults a ground truth.
A gate has to be able to fail. A JSON schema check can fail. A comparison against a reference answer can fail. A second model reading the answer cold can fail. The answerer's opinion of its own answer cannot, and any routing design that leans on it is not routing.
Lesson 4: The routers that thought about it paid to think
The semantic and classifier strategies did what they were built to do. They looked at each prompt and spread traffic across the tiers: roughly a third to Flash-Lite, close to half to Flash, and a fifth to a quarter to Pro.
That last slice is where the money went. Semantic routing cost 39% of the all-Pro bill and classifier 49%, and neither scored higher than plain Fixed Flash at 16%. They paid Pro prices on the prompts they judged hard, and on this set those prompts didn't need Pro.
There's a subtler point in the accuracy numbers. The classifier agreed with my hand-written tier labels 80% of the time. Semantic agreed 50%. Those sound like a good router and a bad one. They're not, because my labels were wrong too. I labeled 25% of the prompts as needing Pro, and the results say almost none of them did. A router that agrees with a human is measuring the human, not the task.
Routing overhead only pays where the cheap tier actually fails. Find those prompts first. Then route only those.
Lesson 5: The measuring instrument has bugs too
A benchmark is software. It fails the way software fails, and three of its failures nearly became findings.
The output cap was scoring truncation. The gateway shipped with a 1,024-token limit on model output. Gemini's Pro tier counts its thinking against that limit, so most Pro answers were being cut off mid-sentence, and the judge was dutifully marking them down. I raised the cap and replayed 166 responses. Without that fix, the "Pro loses to Flash-Lite" story would have been much bigger, and wrong.
The judge was grading from memory. The same semantic cache that saves money in production was also serving the judge, so a bad verdict on one answer could be replayed for a near-identical one. Judging now bypasses the cache entirely.
The backend was crashing my laptop. Every request built a fresh cloud client, which shelled out to a command-line tool to look up the project, which forked a process from inside a server full of live network threads. On macOS that trapped dozens of times a minute. It was invisible in production, where the metadata server answers the same question, and loud on the machine I ran the benchmark from.
None of these are exotic. All of them would have quietly shifted the numbers if I hadn't been watching the run instead of just its output.
Lesson 6: A test everyone aces ranks nothing
Look at the score column again. Every strategy landed between 4.82 and 4.90 out of 5. The judge was sitting at the ceiling.
That's a problem with the test, not a compliment to the models. When every tier scores an A, the benchmark can tell you that the cheap tier is fine, which is useful. It can't tell you where the cheap tier stops being fine, which is the thing a router needs to know.
So the follow-ups write themselves. A harder golden set, weighted toward the prompts where Flash-Lite actually stumbles. A verifier for the cascade that can reject. Both are on the list.
What I'd tell a team on Monday
If you're paying an AI bill and wondering whether routing would help, here is the order I'd do things in now.
Start on the cheapest tier and build the eval before the router. A golden set and a judge cost me about $22 in API spend for the whole experiment, generation and judging included. That is nothing next to a month of flagship pricing on traffic that didn't need it.
Measure where the cheap tier fails, then route only that. Don't build a router for the general case. Build it for the specific prompts your eval flagged.
Never gate on self-reported confidence. If the check can't fail, it isn't a check.
Publish the test set. Or at least put it somewhere your team will look at it with filters on. Mine was padding for a full day and nobody, including me, noticed until it was on a page.
Treat the benchmark as production code. Output caps, caches, and client lifecycles all bit me. They'll bite you.
The uncomfortable summary is that the interesting engineering wasn't the router. It was the measuring. The router turned out to be a config setting: use Flash, stop. The proof that this was safe took a hand-written dataset, a judge, three bug fixes, and a rerun.
Measure before you trust. Everything else was details.
The live report, the browsable golden set, and the code are all public, including the finding that the cascade never escalates. I'd rather publish the number than the story.
–Jeremy