
Cheaper AI routing should be a policy you can inspect
A percentage fee creates an incentive to examine, not proof of bad routing. Own the quality threshold and inspect why each model was selected.
I would stop short of saying a router paid on spending cannot want your bill to shrink. That is too neat. A supplier can earn less on one request and still gain through retention, volume or a more competitive service. The sharper question is whether the customer can inspect the policy choosing the bill.
July reporting about Stripe and OpenRouter concerned acquisition talks, not a completed transaction. The proposed deal is an interesting commercial backdrop. It does not establish that routing decisions favour expensive models.
OpenRouter’s FAQ describes pass-through inference pricing with fees associated with purchasing credits, along with routing controls. That structure deserves scrutiny without turning an incentive into an accusation. The exact fee schedule can change; the ownership question remains.
Separate the incentive from the observed decision
Consider an illustrative fee of five percent on funded usage. A customer buying $100 of usage produces a $5 fee; buying $60 produces $3. That arithmetic describes a marginal incentive under the example’s assumptions. It tells us nothing by itself about which model the router actually selected or why.
A more expensive selection could be correct. The cheaper model may fail the task, exceed the latency limit, or force a second attempt that costs more overall. Equally, a premium model may be unnecessary for a routine transformation. “Cheapest” becomes meaningful only after someone specifies the work that must succeed.
My acceptance rule would put that specification on the customer’s side. Define the approved models, the quality threshold and the conditions that permit escalation. Then retain a record of the route actually taken. A dashboard showing a lower blended token price is helpful, but it cannot explain a particular escalation.
Reconstruct one billable choice
For a hypothetical support-summary workflow, I would compare two policies on the same permitted examples. One starts with the least expensive approved model and escalates after a defined failure. The other starts with a stronger model. Count accepted summaries, retries, elapsed time and total usage for each.
A summary that omits the promised next action should fail even if it is beautifully written. A successful cheaper route should remain successful when customer names, wording and message length change. Those are proposed checks, not measurements from a router trial I have performed.
The useful output is a decision someone can contest: this request used this model because that condition was met. If the reason was a provider outage, distinguish it from a quality escalation. If the customer pinned a model, preserve that fact too. Otherwise several different commercial and technical events collapse into an average.
I would also test whether the policy can be changed without changing the application. A customer who can describe the desired trade-off but cannot enforce it has purchased advice about routing, with limited control over the result.
There are limits to this exercise. Small test sets can miss difficult requests, and provider availability changes. Customer control also creates responsibility: a badly chosen quality threshold can make a perfectly obedient router produce poor work. Owning the policy means owning its mistakes as well as its savings.
This is different from asking what telemetry a hosted router retains. That remains a separate procurement question. Here the concern is narrower: can the buyer see and revise the rule that converts a task into model spending?
Judge a router by the cost of accepted work and the routing decisions you can inspect, not by an assumed motive attached to its fee.


