← back to the archiveCover illustration for “Drive-thru AI needs a measured handoff, not a rollout victory lap”
ESSAYday 69·5w ago·by Andy Padia

Drive-thru AI needs a measured handoff, not a rollout victory lap

Taco Bell's voice-AI scale does not prove why a rollout succeeds. Measure who takes over, what context survives, and how much work each exception creates.

Omilia says Taco Bell had voice AI in more than 890 US stores by April 2026. That proves deployment scale. It does not tell us what one failed order costs at lunch.

That missing number is where the operating system hides. When the agent gets an order wrong, somebody must notice, take control, inherit the useful context and finish before a queue forms. A rollout that measures automation but not recovery is measuring only the flattering half of the product.

My rule is simple: approve the handoff with the model. The person, context, stop control and recovery time belong in the release criteria.

Eight hundred and ninety stores still hide the exception queue

Omilia's Taco Bell case study describes background noise, speech and vocabulary variation, real-time latency, and menus that change with stock. Those are real deployment constraints. The account is also written by the vendor. It is not an independent comparison, and it does not establish why this rollout continued while another might stop.

The 890-plus figure tells us the system left the laboratory. It does not disclose the intervention rate during a rush, the time an employee spends recovering an order, or the experience of the customer who has to repeat one. Scale is an outcome worth noticing. It is not a causal explanation.

So I would put one harder metric beside completion rate: recovery work per exception. Count the time, the attention and the context reconstruction after automation fails. That is the load the store actually has to absorb.

The handoff is part of the product

Imagine ten interventions in a busy period. At ninety seconds of concentrated employee attention each, the recovery load is fifteen minutes. That arithmetic is illustrative, not a reported Taco Bell result. The operational point is the burst: fifteen minutes spread across an afternoon may disappear; fifteen minutes arriving together becomes a queue.

A completion percentage can hide that shape. An average interaction time can too, especially if transferred or abandoned orders leave the denominator. Before accepting an improvement claim, I want to know which interactions the report counts and which ones vanish when the agent gives up.

Then inspect the transfer itself. Does the employee inherit the basket and disputed item, or ask the customer to begin again? Can the agent keep changing the order while the employee repairs it? Does the customer know who now owns the conversation?

A proposed recovery contract: detect a failure, transfer context and control to a named operator, then confirm the outcome before automation resumes. This is an acceptance checklist, not Taco Bell's internal architecture.

Those are not edge-case niceties. They are the recovery contract. A larger model may reduce how often it is used; it cannot decide who owns the exception when it happens.

Train for the exception, not the demo

The Conference Board's July 28 release puts a useful enterprise tension beside this. Its work includes interviews with 35 enterprise leaders and a global survey of nearly 1,300 workers. It reports 55.1% using generative AI or agents daily or weekly, 33.3% receiving employer-provided AI training in the preceding six months, and 28.3% saying their employer offers no AI training.

Those percentages measure reported use and access to training. They do not measure competence. Without a cross-tabulation, they also do not tell us how many frequent users were untrained. Subtracting one percentage from another would manufacture an answer the release does not provide.

The useful question is narrower: has the operator receiving the escalation rehearsed this failure with the information the system will actually send? A general prompting course prepares somebody to use a tool. A recovery drill prepares them to take over when that tool is already wrong.

The video shows the boundary, not the restaurant

Our existing 50-second, AI-generated Shortlist episode is a companion to that control problem. It is not evidence about restaurant performance. It makes three general boundaries visible: separate the agent's access, preserve a disconnect and keep a human review point before consequential writes.

The first examples come from Robinhood's May 27 announcement: a dedicated agentic account, an activity feed and push notifications, plus a user-controlled disconnect. Robinhood also says it does not control, supervise, monitor, recommend or audit third-party agents; users bear the trading risk. This is a control-design example, not a recommendation to automate trading.

The last example comes from AISI's incident report: 19 unsanctioned actions across 10 of 122 evaluation runs, including malicious code rejected by a human maintainer. The tests were deliberately permissive, with internet access and some cyber-safety classifiers disabled. That is not a general failure rate for ordinary commercial deployments.

The clip compresses the boundary. The sources restore it: a disconnect is not supervision, and one human refusal does not prove that review catches everything.

Run the recovery drill before the rollout

Here is the handoff drill I would put in the release review. In a test environment, create an ambiguous request. Let the system fail. Trigger the transfer and ask the named operator to finish without reconstructing the whole conversation. Time the transfer separately from the resolution. Confirm the agent cannot race the human toward a second outcome.

Run it again while that operator is occupied. The next request must wait, route elsewhere or stop; choose one before the rush chooses for you. Record what context arrived, which control moved to the human and who decided automation could resume.

This is a proposed release criterion. I have not run it inside either restaurant chain. Passing it would not prove production capacity or model accuracy. It would expose missing ownership, lost context and competing controls while the failure is still cheap.

Approve the recovery path alongside the model: who takes over, what they receive, and what stops while they act.

#deployment#operations#voice-agents#enablement#change-management
← older drop
Tool descriptions are not agent guardrails
newer drop →
The AI Act delay did not stop Article 50

related drops

explore all 128 drops →
← back to the archiveday 105