← back to the archiveCover illustration for “Google Cloud's 82% is not a cloud-only signal”
POSTday 53·6 days ago·by Andy Padia

Google Cloud's 82% is not a cloud-only signal

Alphabet's Cloud segment now blends services with TPU system sales. Mid-market AI teams should answer with a three-layer stack: cloud, a local GPU rack, and AI-capable endpoints.

Alphabet reported on July 22 that Google Cloud revenue jumped 82% year over year in Q2 2026, from $13.624 billion to $24.768 billion. Its operating profit reached $8.814 billion, up from $2.826 billion. The clean market read is obvious: enterprise AI demand is flooding into the cloud.

The filing says something less tidy. Alphabet describes the service side as usage fees and subscriptions; product sales principally mean TPU systems. It does not split the $24.768 billion between rented compute, software subscriptions and systems shipped into customer-owned data centres.

That makes the quarter stronger, but the headline less useful as a buying signal. Google is not proving that every AI workload belongs in its cloud. It is proving that customers want compute in more than one place.

Hybrid now has three layers

Here is my bet: the sensible mid-market AI stack will become three-layered within 24 months. Not “cloud or on-prem”, and not a heroic private cluster pretending to be a hyperscaler. Cloud, a modest local GPU rack and AI-capable employee endpoints will each carry the work they are naturally good at.

  • Cloud gets frontier reasoning, bursty demand and workloads where managed operations matter more than the last rupee per call.
  • The local rack gets steady, sensitive, high-volume inference where open weights clear the eval and the hardware can stay utilised.
  • The endpoint gets small, repetitive, latency-sensitive work: transcription, redaction, classification, embeddings and first-pass document handling before anything leaves the device.

The endpoint layer is no longer theoretical. AMD says its Ryzen AI PRO 400 mobile processors provide up to 60 NPU TOPS for local AI acceleration. That is enough to make endpoint inference a serious benchmark candidate across a commercial laptop fleet.

But do not turn TOPS into procurement astrology. Peak TOPS is not tokens per second, and it says nothing by itself about model memory, runtime support, quantisation, output quality or the cost of managing the fleet. An AMD-backed laptop is not automatically cheaper than an API. It has simply earned a place in the test.

I have already seen the split pay

At work, I ran the numbers with a client on self-hosting Qwen for a high-volume internal workload versus keeping everything on a closed API. Raw token economics favoured self-hosting. Then we priced the idle GPU capacity, serving and upgrade time, and the eval work needed to prove the open model was good enough.

The all-in gap narrowed. The client kept the reasoning-heavy path on the API and moved bulk summarisation to self-hosted open weights. No named client, no victory-lap percentage — just a deployment that got better when we stopped forcing one economic model onto two different kinds of work.

I am adding a third column to that same spreadsheet now: endpoint. For many mid-market firms, the laptop fleet is already bought, distributed and powered. If a quantised model clears the task-specific eval on the NPU or integrated GPU, local execution may remove a recurring call, a network round trip and a data transfer in one move.

Price the workflow three times

My rule now is simple: no mid-market AI implementation gets approved until every inference step is priced in three places — cloud, rack and endpoint. Use the same eval set in all three. Include utilisation, engineering time, failure handling, security review and refresh cycles; the cheapest demo is often the most expensive production route.

Steal this for the next architecture review. Mark every step in one workflow C, R or E. Benchmark only the ambiguous steps before Friday. Route by measured quality and all-in cost, not by whichever vendor owns the loudest quarterly number.

Cloud won the quarter; mid-market implementers should win the workload by routing it across cloud, rack and endpoint.

#cloud#hybrid-ai#inference#amd#mid-market
← older drop
An isolated sandbox is a claim, not a property
newer drop →
The FCA put the model vendor inside the sandbox

related drops

explore all 78 drops →
← back to the archiveday 59