← back to the archiveCover illustration for “Long tool observations need a separate prefill budget”
POSTday 120·today·by Andy Padia

Long tool observations need a separate prefill budget

HySparse2 runs prefill through 25 of 49 layers, then uses the full model for decode. Measure long tool observations as a separate capacity workload.

An agent can emit a six-token tool call and still make the cluster ingest a 100,000-token observation before the next answer starts. Xiaomi's MiMo team designed HySparse2 around that imbalance. In its 49-layer model, the prefill node hosts only the first 25 layers; decode still needs the full network.

My rule is to budget long tool observations as a separate serving workload. Prompt ingestion and token generation may share a request ID, but they do not have to share the same capacity plan, placement or latency target.

HySparse2 turns prefill and decode into different deployments

Prefill processes the accumulated prompt and builds the key-value cache. Decode consumes that cache while generating new tokens. Tool-heavy agents make the first stage unusually lopsided: short actions return long search results, logs and documents, then repeat across turns.

The HySparse2 paper splits its model into a self-decoder and a cross-decoder. Its KV Bridging projections construct the cross-decoder's full-attention cache from self-decoder hidden states. That lets prefill exit after the self-decoder instead of running every layer. Under prefill-decode disaggregation, the authors place the first 25 of 49 layers and the bridging projections on the prefill node, then transfer the smaller projected cache to decode nodes.

That is more than a faster attention trick. It makes prefill-node model footprint, decode-node model footprint and the network between them separate capacity decisions.

The 1M-token numbers need their experimental frame

On the authors' 80B-A3B mixture-of-experts models, HySparse2 reports 5.02 times fewer prefill FLOPs than the Hybrid SWA baseline at one million tokens. Its KV cache is 2.69 GB, versus 12.09 GB for Hybrid SWA and 6.72 GB for HySparse.

Those are author-run analytical and experimental results from arXiv v1, not wall-clock production measurements I reproduced. The models were pretrained on roughly 500 billion tokens at 32k context and post-trained on another 100 billion tokens while extending context to 256k. HySparse2 also uses multi-query attention while the two baselines use grouped-query attention, so the cache comparison is not an isolated swap of one selector.

The quality trade is visible too. In the local-window ablation, the forced-window design that enables early-exit prefill beat gated sliding-window attention on RULER-v2 and GraphWalks, but trailed it by 5.08 points on GSM8K and 4.99 points on MRCR-v2. AgentPPL, one of the agent measures, is an internal 1,000-trajectory benchmark. I could not verify a public production deployment, wall-clock cluster benchmark or cost model.

Measure the observation lane before buying the architecture

The earlier context-window capacity argument budgets live cache per request. The agent-cost ledger keeps tool time, waiting and regenerated prefixes attached to the accepted task. HySparse2 adds another line: observation-ingest work can deserve its own pool.

For this article, I have not deployed HySparse2. This is labelled editorial judgment: before adopting any prefill-decode split, replay the actual agent traces and record four numbers for each stage—queue time, time spent computing, memory occupied and bytes transferred. Run long tool returns in bursts, not evenly spaced. Then compare the split design with the current stack at the same task quality and tail-latency target.

If prefill is the bottleneck, scale or place that pool independently. If cache transfer becomes the bottleneck, the split has only moved the queue. If the workload mostly has short observations and long generation, this architecture may be solving the wrong half.

What's in it for you

  • Separate time-to-first-token from output-token latency in agent dashboards.
  • Size prompt-ingest capacity from tool-return length and burstiness, not action length.
  • Charge cache transfer and any quality loss against the same accepted task.

A short agent action can hide a long ingestion job; give prefill its own budget before it becomes the queue nobody measured.

#agents#inference#prefill#kv-cache#capacity-planning
← older drop
A cache selector needs a random baseline

related drops

explore all 355 drops →
← back to the archiveday 120