
Price AI work by accepted outcome, not by token
Token spend is an input price. Useful AI unit economics joins model, tool, infrastructure, retry and review costs to an accepted unit of business work.
A token is a billable input. It is not a unit of business work.
That distinction sounds obvious, yet many AI cost reviews still begin and end with model spend: tokens by provider, tokens by team, perhaps cost per conversation. Those numbers help reconcile an invoice. They do not tell a bank whether an agent resolved a complaint, produced an accepted credit memo or created another item for a human queue.
Brillio's banking tokenomics article asks four useful questions about consumption, blended token cost, utilization and ownership. I would push the unit one step further: price accepted work, not attempted inference.
Name the business unit before measuring the model
The denominator should describe something the organisation values and can verify. “One million tokens” fails that test. “One complaint resolved without reopening within seven days” is closer. So is “one credit memo accepted after review” or “one code change merged without rollback.”
The unit does not have to be perfect on day one. It has to be concrete enough that finance, product and operations can disagree about the same outcome.
Once that unit exists, token cost becomes one component. A workload may use several models, retrieval calls, a browser, a sandbox, evaluation passes and human review. It may retry twice and discard the final output. Every one of those costs belongs to the attempted unit; only accepted work belongs in the success denominator.
rendering diagram…
Blended cost has to include waiting and control
The model invoice is easy to see because it arrives as a clean number. The surrounding system is fragmented across infrastructure and teams.
AgentSysBench makes the fragmentation visible. Its August 2026 paper reports 4,641 benchmark requests, 64,924 model calls and 118,274 tool calls across ten agent applications, plus 178,799 production sessions. In five of the ten benchmark applications, tools or environments dominated or co-dominated latency.
The paper's cost estimates are workload-specific, not a universal allocation. Their value is structural: an agent can hold a sandbox, wait for a model, call search, evaluate an answer and retry while the model-token dashboard captures only one layer.
I would join costs on a task ID and record state over time: model running, tool running, environment held, waiting for a person, retrying, accepted or abandoned. Idle state matters when resources remain reserved. Review matters when every output requires expert attention. Evaluation matters when a “successful” call produces work the business rejects.
Forecast ranges, not a fixed cost fantasy
AI workloads vary with context length, reasoning depth, provider routing, tool use and failure recovery. A fixed cost per request will look precise until the workload reaches production.
Use at least three ranges: normal, difficult and recovery. Normal represents the common accepted path. Difficult includes longer context, deeper reasoning or more tool use. Recovery includes retries, escalation and the work needed after a failed action.
Then connect each range to a quality floor and stop condition. A cheap answer that fails acceptance is not efficient. An agent that retries until it finds an agreeable judge can turn a quality problem into an unbounded cost loop. A human escalation may be the economically correct terminal action.
Ownership belongs here too. Every workload needs someone who can change model routing, narrow the scope, accept a higher review cost or stop the workload entirely. “Shared AI platform” cannot mean shared accountability.
Keep cost reduction tied to outcome quality
Model routing, caching, smaller contexts and cheaper providers can bend the cost curve. Each optimisation can also change the answer. Cost experiments therefore need paired acceptance results.
They also need versioned assumptions. Store the model and price effective date, routing policy, workflow version, evaluation version and exchange-rate basis where relevant. Without that receipt, a lower monthly bill can be misread as an architectural improvement when it came from a temporary price change, a quieter workload or a quality regression. Unit economics should be reproducible for the period in which the business decision was made.
For a bank, I would report a compact operating line: attempted units, accepted units, reopened or rolled-back units, total blended cost, human minutes, and cost per accepted unit. Segment it by workload and risk class. Do not average a low-risk drafting assistant with a decision-support system whose review is necessarily expensive.
This complements my earlier argument that agent bills are moving to the control plane. The control plane can observe the full task. The finance step is to join that trace to a business outcome someone accepts.
What's in it for you
- Replace cost per token with cost per accepted business unit.
- Put retries, discarded outputs, tools, sandboxes and review in the numerator.
- Forecast normal, difficult and recovery paths separately.
- Give one owner authority over scope, quality, spend and the stop condition.
If an AI invoice cannot be joined to accepted work, it is spend reporting—not unit economics.


