
Inference moats are financed in megawatts
Groq's $650M raise talks like a utility — 13 data centers, a 200 MW target by 2027. The tell for buyers: chip benchmarks decay in quarters; deployed capacity, financing, and placement decide what you can actually buy.
Groq announced $650 million in growth capital this week, and the language of the announcement is the story. The headline numbers: 13 operating data centers across North America, Europe, the Middle East and APAC; more than 5 million developers served; and a target of 200 megawatts of capacity by the end of 2027. All company-reported, and the capacity plan is a forward statement I can't verify — but notice what kind of numbers these are. Sites. Power. Global placement. Long-lived capital. Not a single one is a chip benchmark.
Three years ago, every inference-provider pitch led with tokens per second on a bar chart. This raise reads like a utility's expansion filing. That shift in vocabulary is worth more attention than the dollar figure, because it tells you where the durable competition in inference actually lives — and it is not where most enterprise buyers are still looking.
Why the benchmark decays and the megawatt doesn't
A silicon speed advantage is real but perishable. Competitors iterate, model architectures shift to favor different hardware, and — as Groq's own December licensing arrangement with NVIDIA illustrates — the boundaries between "chip company" and "cloud on someone's chips" blur within a couple of product cycles. Whatever the leaderboard says this quarter, the safe assumption is that it says something different in four.
Deployed capacity obeys a different clock. A data center is permits, land, transformers, grid interconnection queues, and construction crews — years, not sprints. Power purchase agreements and site financing are decade-shaped commitments. And geography is close to permanent: a customer with EU data-residency obligations does not care that your tokens are fast in Texas. The same week's other headline — Meta reportedly contracting Crusoe for major inference capacity — is the same pattern from the demand side: the hyperscale players are locking up supply, not admiring benchmarks.
rendering diagram…
The economics compound the point. An inference cloud is a high-fixed-cost, utilization-driven business — the unit margin lives or dies on keeping expensive capacity busy. That is a utility's profit model, and it produces utility behavior: sell long contracts, smooth demand, expand where power is cheap and customers are obligated to stay. The 200 MW target is not a technology claim. It is a statement about how much of that business Groq intends to finance into existence, and $650 million is the down payment.
The question set enterprise buyers keep skipping
At work, the platform-lead conversations I sit in still start — and too often end — with token speed and price per million. Both matter. Neither answers the questions that will actually bite during the contract term.
Where does the capacity sit? Residency, latency, and jurisdiction are properties of buildings, not models. If the provider's EU capacity is one site with a waiting list, your residency story is one incident away from an exception memo. How is it financed? A provider running on short-term capital with aggressive expansion targets faces a different set of temptations in a downturn than one on long-dated infrastructure financing — and capacity that gets mothballed mid-contract is a risk no benchmark surfaces. And who holds priority when demand spikes? This is the one that separates utility thinking from benchmark thinking. When a model launch or a seasonal surge saturates supply, somebody's workloads get throttled, and it is whoever's contract lacks a capacity commitment. Ask directly: is my throughput reserved, or am I buying from the spot pool with a nicer logo on it?
I watched a client discover the third question the hard way during a previous demand spike — their provider's benchmark numbers were unchanged and magnificent, and their batch jobs sat in a queue behind customers with committed-capacity contracts. The remedy cost nothing but negotiation: a reserved-throughput clause at renewal. The lesson cost a quarter's roadmap, and it would have been free if anyone had asked the surge question during procurement.
Reading providers like utilities
The practical upgrade is to move a third of your inference-vendor diligence from the model card to the infrastructure disclosures. Providers increasingly publish them — site counts, regions, power targets — precisely because the sophisticated buyers now ask. A provider talking megawatts and interconnects is telling you they intend to be constrained by physics and finance, not fashion; that is, on balance, whom you want to be locked in with.
Steal this clause list for your next inference contract: capacity commitment (reserved throughput, not best-effort), placement guarantee (named regions, with remedies), surge policy in writing (who gets throttled, in what order), and a financing question in the diligence call that you actually ask out loud — what is the capital structure behind the capacity I am renting? A provider that answers crisply is a utility. A provider that pivots to the benchmark slide is a bar chart with a burn rate.
Benchmarks are weather; megawatts are climate — buy inference the way you buy power, because that is what you are buying.


