
Free LLM routers are paid for in telemetry
Ramp opened its internal LLM router to the public, free during beta. The routing layer's durable product is the cross-customer model-usage graph — price the telemetry as a trade, not a gift.
Ramp opened its internal LLM router to the public on July 20, 2026, and the launch is doing the rounds as "free AI infrastructure." The company's own numbers are worth quoting precisely: Router moves 2.75 trillion tokens a month and cut Ramp's internal AI costs 30%. Both figures come from the launch post — the company's own reporting, unaudited — so hold them as claims, not measurements.
The mechanics are genuinely convenient. It is an OpenAI-compatible endpoint, a one-line base-URL change, and it "picks the cheapest approved model that clears your quality bar, with automatic fallbacks." The models listed on the page are OpenAI and Gemini frontier models plus select open-source options including Kimi. Press summaries also claim Anthropic support; I could not confirm that on the page itself, so treat it as unverified.
The consensus read is: great, someone finally made model routing free. Read the page more slowly and "free" is doing two jobs. The routing layer is free during beta only. Token spend is billed at list price from day one, with $100 in promotional credits for the first 500 off the waitlist. You are not getting free inference; you are getting a free middleman, temporarily.
Ask what the middleman keeps
Here is my bet: the durable product is not the router, it is the usage graph. Ramp is a spend-management company. A hosted router operated by a spend-management company sees, across every customer at once, which models win at which price for which workload — the exact dataset that pricing negotiations, benchmark marketing, and procurement products are built from. Saved tokens are the pitch; procurement telemetry is the asset. Linas Beliūnas made a directionally similar argument in a paywalled deep-dive I have only seen the preview of, but the specific mechanism I am arguing is the spend data: the router operator learns the industry's real demand curve for models before the industry does.
Notice also who defines the quality bar the router optimises against. Ramp does. When your routing policy lives inside someone else's product, "cheapest model that clears the bar" quietly becomes "cheapest model that clears their bar."
None of this makes the trade a bad one. A 30% cost reduction, if it holds outside Ramp's own workloads, is real leverage — for many teams the telemetry is worth trading. My objection is that the launch prices it as a gift, and a trade priced as a gift is a trade you lose.
Where this bit me in a design review
At work, I went through exactly this fork with a client platform: hosted router versus a self-hosted routing layer. The hosted option was cheaper to stand up by weeks. What settled it was writing down, in one column, everything the hosted operator would observe: every prompt category, every model choice, every fallback event, the full cost curve of the client's AI usage — visible to the router's operator before the client's own CFO could assemble the same picture. For that client, in a regulated space, the column was disqualifying. For a startup burning runway, it might be a fine sale. The point is we priced it.
Steal this before you sign up for any hosted router, this one or another: put four questions in the vendor thread. What routing metadata do you retain, and for how long? Do your terms permit aggregated or derived use of my usage data? Can I export my full routing logs, so the decision history is mine? And who defines the quality bar — can I pin my own evals to it? If the answers are vague, that vagueness is the price tag.
A free router is a bill that arrives as telemetry — read what the operator learns about you before you admire the 30%.


