
The 3x agent multiple is a night-shift ledger
In short: OpenAI’s task success improved, but its 3.1× runtime measure also counts review agents. Connect accepted outcomes to the full cost before calling it productivity.
OpenAI reports 3.1 agent-workdays for every human workday by mid-August 2026. But the more useful number is further down its September 6 report: for tasks estimated at four to eight human hours, success without intervention rose from about 18% in January to 53% in July.
That is real evidence of improving autonomy. It still doesn't make machine runtime a productivity result. The 3.1 is a night-shift ledger. To value the shift, connect its hours to accepted work and the supervision that work consumed.
I went through the chart methods and extracted the published outcome values. They sharpen the argument in both directions: the autonomy gain deserves credit, and the runtime numerator includes something a productivity slide can easily hide.
The numerator includes agents checking agents
OpenAI's runtime calculation counts subagents separately. It also counts agents automatically reviewing other agents. The denominator assumes eight hours per research employee on each calendar day; it doesn't measure their actual hours worked. The chart uses a lagging 28-day window, excluding company holidays.
An extra review agent can therefore increase the multiple even when the team accepts exactly the same research result. That checking may prevent an expensive mistake. It belongs in the operating cost, and its benefit belongs in the quality record. Counting its runtime as another completed unit of research loses both distinctions.
Tomasz Tunguz's night-factory interpretation captures the capacity change: parallel work can continue beyond a human shift. But the public ratio doesn't tell us how much work happened overnight, how much supervision waited until morning, or how much useful research survived inspection.
My correction for a team using this metric: attach every worker, subagent and reviewer to the same task lineage. Count the accepted result once. Keep every attempt on the bill.
The intervention headline hides progress
The report says more than half of successful four-to-eight-hour tasks involved intervention. Its January–July outcome chart makes that concrete: approximately 43% succeeded without intervention, 45% succeeded with intervention, 10% failed and 2% hit tool errors.
Those are user-weighted shares after uncertain outcomes are excluded. The assisted share divided by the two success shares is about 51%. That is a seven-month aggregate, not July's intervention rate, and certainly not a 51% defect rate.
The separate monthly series shows zero-intervention success rising from 18% to 53%: 35 percentage points. June was 54%, so there is no evidence of a further July gain in this bucket. The broader improvement matters; freezing the pooled intervention figure into a permanent autonomy ceiling would misread the source.

OpenAI weights users equally, omits uncertain outcomes and suppresses sparse monthly cells. Late-July sessions lack full follow-up. The classifier was checked against 25 manually labelled tasks. Zero recorded corrections also does not establish zero human review. These are useful descriptive results across changing tasks, not a controlled measure of human hours saved.
Human support traffic also declined in one internal troubleshooting channel. That deserves a place beside the intervention numbers. A claim that agents merely manufacture more human work would ignore evidence in the same report.
Put four numbers on the shift ledger
I couldn't find public session records joining runtime, human correction time, spend and accepted outcomes. Multiplying the August runtime ratio by a January–July success share would mix windows and denominators. It would manufacture a productivity number.
For a deployment decision, I would use OpenAI's own investment guidance and track:
- Accepted outcomes per agent-workday, against a test specified before the run.
- Unattended yield, with uncertain outcomes reported separately.
- Human correction minutes and waiting time per accepted outcome, kept as separate measures.
- All-in cost per accepted outcome, including subagents, automated review, retries and human review.
An accepted research outcome can be a useful dead end. It needn't be a breakthrough. The denominator should reward reliable work, not pressure researchers to produce positive findings.
That extends the archive's rule that AI spend needs a work denominator. Intervention debt is the supervision queue I would watch for, not a backlog these public data prove exists at OpenAI.
Credit the autonomy gain. Book productivity when accepted work improves against the full cost of producing it.


