
Klaviyo's L3 mandate needs an output ledger
Klaviyo told 2,300 people to reach L3 agent fluency and reported revenue per employee up 28%. A mandate starts behaviour; output and rework decide whether it worked.
SaaStr's August account embeds the original conversation with Klaviyo co-founder Andrew Bialecki. In it, Bialecki says the company had roughly 2,300 employees and told them to reach “L3” by the end of June. L1 meant using AI for search. L2 meant running an agent. L3 meant constantly running multiple sessions or a team of agents, decomposing work and validating the output.
Klaviyo's official Q2 results supplied a second number. Annualized revenue per employee was up 28% year over year. In the earnings call, co-founder Andrew Bialecki connected that increase to the company building agents for itself.
Those measurements will be collapsed into one success story. They should not be. L3 is a leading signal that a new behaviour has started. Revenue per employee is a broad result, but it cannot show which behaviour produced it. An enterprise needs both—and the conversion ledger between them.
The mandate does one useful job
A company-wide deadline is a forcing function. It creates a shared vocabulary, makes experimentation part of the job and denies every function the easy excuse that agents are only for engineering. SaaStr describes product managers, designers, sales and marketing working against the same ladder. That is a job redesign, not another optional software rollout.
Calling L3 a login target would be too crude. The published definition includes problem decomposition and output validation, which are materially harder than opening a tool. But the measurement is still about capability and behaviour. Neither Klaviyo's August 5 earnings release nor its call transcript reports the L3 pass method, the attainment rate across 2,300 people or a workflow-level before-and-after result tied to the mandate.
That does not make the deadline theatre. It puts the deadline in the right column. A capability ledger asks whether people can work this way. It should record who can decompose a representative task, run several agents without losing control and reject a bad result. It should not claim that the business shipped more simply because those behaviours were demonstrated.
My rule: an adoption mandate gets one quarter as a headline KPI. After that it becomes a capability requirement, not a business result.
Klaviyo published a number closer to output
The original Q2 release reported $370.6 million in revenue, up 26%, and annualized revenue per employee up 28%. In the official call transcript, Bialecki attributed the capacity gain to building agents internally.
That is management's attribution, not a controlled result. Revenue per employee rises whenever the numerator grows faster than headcount. It can reflect pricing, sales execution, customer expansion, hiring pace, product mix and automation well beyond the L3 programme. It does not reveal whether a campaign took less time, a feature needed less review or a support answer required less rework.
The number is still useful. It is far closer to business output than sessions, seats or self-reported fluency. Treat it as a company-wide throughput proxy, then demand the workflow evidence beneath it. Otherwise the executive dashboard jumps from “people used agents” to “the company got more efficient” without measuring the step where one was supposed to cause the other.

Pair fluency with conversion
The figure is the operating model. Keep the L1/L2/L3 ladder, because leading indicators matter. Pair it with a second ledger that starts at a named recurring workflow and ends at accepted work. The unit changes by function: an approved campaign, a resolved support case, a retained code change or a finance analysis that survives review. “Things generated” is not an accepted unit.
Run the pair for 30 days:
- Choose one recurring, high-volume workflow per function and name what counts as accepted output.
- Capture the pre-mandate baseline: volume, end-to-end cycle time, cost, rework and a quality outcome.
- Record L3 exposure and verified behaviour on that workflow, not a survey answer or a raw session count.
- Compare the same workflow at 30, 60 and 90 days; separate adoption from changes in team size or demand where the data allows.
- Scale the pattern only when accepted output improves without rework, defects or escalations getting worse.
This is the missing connection between two earlier arguments. AI training needs production telemetry because attendance cannot show whether work changed. AI spend per engineer needs an accepted-work denominator because input intensity is not value. An L3 mandate sits between them: it makes the behaviour non-optional, then needs both telemetry and a denominator to earn a productivity claim.
Do not make one denominator fit 2,300 jobs
The wrong response is a single company-wide “AI output” score. Code changes, campaigns and customer resolutions have different acceptance rules and failure costs. A uniform agent level can build common fluency; the outcome ledger has to stay close to the work.
I would approve L3 as an enablement target. I would not let it appear alone on a productivity slide. Put capability, accepted output and quality beside one another. Show the broad revenue-per-employee trend above them, not instead of them. If a function cannot name its accepted unit, the next job is instrumentation—not another autonomy level.
Klaviyo's mandate is interesting precisely because it is more serious than a login count. It asks people to decompose, orchestrate and validate. The stronger the behaviour standard, the more tempting it is to mistake compliance with it for value. Resist that shortcut. The mandate starts the operating change; the paired ledger decides whether the change deserves to remain.
Mandate fluency to start the behaviour; measure accepted work and rework to decide whether it worked.


