
AI-native teams still need eval ownership
Codex ships 10-12 surfaces with 2 PMs; Cursor runs 40 engineers on 1. The deleted layers were carrying decisions — acceptance, risk, launch evidence, rollback. Delete the layer, keep the decision, name its owner.
Product Growth published the org charts everyone in enterprise AI will be asked about by Friday. At OpenAI, the Codex group reportedly runs 10 to 12 product surfaces with 2 PMs, 1 designer, and about 40 engineers — surfaces that at a traditional company would each get a squad of fifteen. Cursor reportedly operates 40 engineers with a single PM. The numbers are unverified — reported headcounts from an inside-look piece, not filings — but the shape rings true, and the shape is what your executive team will want to copy.
The article's causal story is right, as far as it goes. When an agent builds a working version of a feature in under an hour, the sequence inverts: build first, evaluate second. Codex reportedly ships about two of every ten things it builds and discards the rest — being wrong got cheap, so the coordination machinery that existed to prevent expensive wrongness (sprints, PRDs, handoffs, dedicated QA) stops paying rent. Fewer translators between intent and code. Smaller teams. Faster loops.
Here is the part the copy-paste version misses, and it is the part that will hurt: the layers were carrying decisions, not just coordination. Delete the layer and the decision doesn't disappear — it goes unowned.
What the handoffs were quietly deciding
Walk through what a "bloated" product process actually adjudicated. The PRD forced someone to write down what counts as working — acceptance criteria. The QA gate forced someone to ask how does this break, and whom does it hurt — the harmful edge cases. The launch review forced someone to assemble evidence that it works before customers found out otherwise. And the release process meant somebody could answer how do we un-ship this at 2 a.m. — rollback.
Four decisions: acceptance, risk, evidence, reversal. In the fifteen-person squad they were smeared across roles so thoroughly that nobody noticed they were being made. That smearing was waste, mostly — three meetings to decide what one competent person could decide alone. The AI-native insight is that the smearing was the waste. The copy-paste failure is concluding that the decisions were.
rendering diagram…
Look at what the lean teams in the article actually have, on the reported description. Two of every ten prototypes ship — which means somebody is evaluating ten and choosing two, against some bar. That bar is eval ownership, operating at high cadence with no ritual around it. Cursor's one PM is not doing one-fifteenth of the old PM job; they are doing the concentrated decision core of it while agents and engineers absorb the rest. The layers are gone. The ownership is not — it is denser.
The flattening I keep getting called after
At work, the call I now get quarterly: a leader read a piece like this one, flattened the AI product team, velocity went up, everyone was thrilled — and then a launch went sideways and the post-mortem could not answer four questions. Who signed off that this was good enough to ship? Nobody; the demo looked great. Who owned the harmful edge cases? QA, except QA was deleted. What evidence existed at launch? A screen recording. Who could roll it back? Eventually, someone, after four hours of figuring out how.
None of that is an argument against the lean shape. It is the punch list for adopting it honestly. The fix in that engagement was one document and one habit: a decision-rights page naming a human owner for acceptance, risk, evidence, and reversal on each surface — and evals as the standing artifact that carries the first three. Prototypes stayed fast and disposable. The two-in-ten that shipped now crossed a bar someone owned. Velocity survived; the next incident had a name attached within minutes, in the good sense.
My rule for the flattening conversation: you may delete any layer whose decisions you can name and reassign — and no layer whose decisions you cannot name. If nobody can articulate what the launch review used to decide, you are not ready to delete the launch review, because you will be deleting it blind.
Steal this
Before copying anyone's org chart, run the four-question audit on each product surface, today, at current headcount: who owns what counts as working (and where are those criteria written — "the evals" is the right answer), who owns harmful edge cases, who owns launch evidence, who can reverse a ship and how fast. Write four names per surface on one page. Every blank is a decision currently being made by default — which is to say, by luck. Fill the blanks first; flatten second. The lean teams you are copying did it in that order, even if the article about them leads with the headcount.
The AI-native teams didn't delete the decisions — they concentrated them; copy the concentration, not just the deletion.


