
Premium models need failure-cost routing
Fable 5 reopened the price gap between frontier and routine inference — $10/$50 per million tokens. The router that earns that money asks one question the leaderboard never does: what does a wrong answer cost here?
Claude Fable 5 has reopened a gap the market had been quietly closing: the distance between frontier pricing and routine pricing. Anthropic lists it at $10 per million input tokens and $50 per million output — several multiples of the workhorse tier — plus a 1.1x multiplier if you need US-only inference. The newsletters are already producing when-to-use-it guides for GTM teams; one I read this week claims the model is twice as capable as Opus 4.8, a ratio I could not verify and which I suspect means nothing measurable anyway.
Because "how capable" is the wrong axis for the buying decision. The gap that matters is not between the models' benchmark ranks. It is between what their mistakes cost you — and that number is a property of your workflow, not of the model.
The question the leaderboard never asks
Take a GTM stack, since that is where the guides are aimed. Under one roof: lead classification (thousands a day, a mistake costs a mis-tagged row someone fixes next week), draft email copy (a human reads it before send, so failures are caught free), and the recommendation that goes in front of a client with your name on it (a bad one costs the deal, or worse, the relationship). I keep meeting versions of this stack where all three route to the same expensive model — usually because the team upgraded everything the day a launch impressed them, and nothing ever got downgraded.
The router that actually earns Fable-tier money asks three questions per task class, none of which appear on a leaderboard. What does a wrong outcome cost, at the margin? A mis-tag is pennies; a wrong client recommendation is the whole engagement. Is failure detectable before it does damage? Reviewed-by-a-human tasks can fail cheap, so the premium buys little; fire-and-forget tasks fail invisible, which is where quality is worth real money. Can a cheaper model refuse? If the workhorse can flag its own uncertain cases and escalate them upward, the premium model becomes the specialist you consult, not the generalist you salary.
Run those three and the routing table mostly writes itself: high failure cost + low detectability + rare volume routes premium; everything else routes cheap with an escalation path. Model selection stops being a leaderboard exercise and becomes error economics — which is what it always was, we just let the launches distract us.
At work I ran this exact exercise with a delivery team last quarter. Their pipeline sent every task to the priciest model available "to be safe". We tagged each task class with failure cost and detectability; roughly 80% of volume was cheap-to-fail work that a mid-tier model handled indistinguishably, 15% was human-reviewed anyway, and about 5% — the client-facing synthesis — genuinely justified the premium. The bill dropped by more than half, and the interesting part: quality complaints fell, because the escalation rule caught uncertain cases that the old everything-premium setup had been waving through unreviewed on the assumption the expensive model is always right.
My rule, stated once: pay for the model by the price of its mistakes, not the price of its tokens. Cheap failures deserve cheap models; expensive failures deserve the frontier plus a detection net.
Steal this: three columns on your task inventory — failure cost (money, honestly estimated), detectability (caught before damage: yes/no), monthly volume. One hour of tagging. Route premium only where column one is high AND column two is no. Everything else gets the workhorse and an escalation rule.
The frontier model is insurance, and nobody insures postcards at diamond rates — price the loss, then pick the courier.


