← back to the archiveCover illustration for “Agent go-live starts the expensive phase”
ESSAYday 9·7 weeks ago·by Andy Padia

Agent go-live starts the expensive phase

Salesforce's lessons from 12,000+ Agentforce deployments say the quiet part: test before launch, monitor after. Budget most of an agent program for what comes after go-live — that is when the real debt surfaces.

Salesforce has been publishing what it learned shipping Agentforce at scale — the newsletter version circulating this week says 20,000 enterprise deployments; the accessible Salesforce account says more than 12,000 over a year. I could not reconcile the two numbers, so take the smaller one. Either way, it is the largest public corpus of production agent experience anyone has described, and the shape of the advice is more interesting than the case studies.

Two claims stand out. Salesforce reports a 30% engineering cycle-time improvement from its own internal use, and — the number that should reframe your budget — automatic remediation of 87% of detected incidents within 20 minutes. Both are first-party and unaudited, so hold them loosely. But notice what the second number implies: a mature agent operation is one that expects incidents continuously, detects them fast, and has invested so heavily in remediation machinery that most fixes need no human. That is not a pilot capability. That is a running cost, staffed and tooled, forever.

Which is the quiet message underneath all the operational guidance: test before deployment, monitor after it. Translated out of vendor-speak: go-live is not the finish line. It is the start of the expensive phase.

Why the debt only surfaces after launch

Enterprise agent programs fund themselves like theatre productions — months of rehearsal, a big opening night, and then the assumption that the show runs itself. The reason this fails is structural, not motivational: three kinds of debt are invisible until real users arrive, because only real users push the system off its designed paths.

Data debt. The pilot ran on the datasets someone curated for it. Production runs on the knowledge base as it actually is — the stale article, the two systems that disagree about a customer's status, the sync job that quietly broke. An agent surfaces every one of these as a confident wrong answer, at retail volume.

Policy debt. The pilot's permissions were whatever made the demo work. In production, an agent discovers every gap between what your policies say and what your systems enforce — the refund limit that lives in a manager's head, the escalation rule nobody wrote down. Each gap becomes either an incident or a new rule you now maintain.

Exception debt. Designed paths cover the common cases. Users bring the rest: the order that is half-cancelled, the account with two owners, the request that is reasonable but unanticipated. Every exception is a decision — handle, escalate, or refuse — and the backlog of those decisions is the real work of the first six months.

rendering diagram…

None of this debt is visible in the pilot, because the pilot was designed not to hit it. The launch does not create the debt. It reveals it — and it reveals it on a schedule you don't control, at whatever volume your users happen to bring on the day the sync job breaks.

There is also a fourth cost that isn't debt at all: drift. The business changes under a healthy agent — a new product line, a renamed policy, a reorganized team — and answers that were right in March are wrong in August with no incident, no error, and no alert. Only scheduled review catches it.

The plan I keep rejecting

At work, the client plan I see most often funds a polished pilot, a launch date, a communications push — and then assigns the agent to "the platform team" as one more thing they own. No rotation for reading traces. No owner for policy fixes. No budget line for keeping evals current as the business changes. No rehearsed rollback. The program's org chart simply stops at go-live.

When I ask who reads the traces in week three, the answer is usually "the dashboard will alert us". But dashboards alert on failures the builders anticipated; the debt above is by definition what they didn't. The only instrument that finds it is a human reading real transcripts on a schedule and filing what they find — wrong answers to the data team, permission surprises to the policy owner, new exception classes to the workflow backlog.

My rule for sizing this now, after watching a few programs through their first year: plan the pilot-to-launch effort as the smaller half. Whatever you spent getting to go-live, reserve at least as much again for the twelve months after — a trace-review rotation with real time carved out, an owner for policy repair with authority to change rules, eval maintenance wired to every model and prompt change, and a rollback path you have actually exercised. If the budget cannot fund the second half, ship a narrower agent. A small agent with a funded operations phase beats a broad one abandoned at launch — the broad one does not stay launched.

Salesforce's own 87%-in-20-minutes figure, whatever its audit status, is the strongest version of this argument: the most experienced agent operator on record responded to production reality by building an incident-remediation machine. They did not get to skip the expensive phase. They industrialized it.

Steal this

Before your next agent go-live, write the week-three rota on one page: who reads twenty traces a day, who owns policy fixes with what authority, who re-runs evals when anything upstream changes, who can roll back and how fast — with names, not team labels. If any line says "TBD", the launch date is fiction; you are scheduling an incident, not a release. And when the go-live retro happens, hold the applause for month six — that is when you will know whether you shipped an agent or an announcement.

The pilot proves the agent can work; the year after go-live decides whether it does — fund the year, not the launch party.

#agents#enterprise-ai#operations#deployment#governance#agentforce
← older drop
Agent loops are release artifacts
newer drop →
Launch buzz measures demoability

related drops

explore all 78 drops →
← back to the archiveday 59