← back to the archiveCover illustration for “More experiments need a false-discovery budget”
POSTday 79·3w ago·by Andy Padia

More experiments need a false-discovery budget

An experiment’s significance threshold does not tell us what share of declared winners are real. Faster AI-assisted experimentation makes that distinction more useful, not less.

If AI makes experiments cheaper to propose, I want the team to track how many declared wins survive confirmation. The count of experiments shipped is the easy part.

Ravi Mehta and Matthew Mamet's essay on the full-stack builder argues for judgment and curation as AI accelerates product work. It invokes the false-positive risk associated with significance thresholds. The qualification I would add is that a test's false-alarm rate is conditional on there being no real effect; it is not the proportion of declared winners that are false.

That distinction matters when deciding how much confidence to place in a busy experiment pipeline. The answer also depends on how often the ideas have real effects and how well the tests detect them.

I would not infer either of those quantities from the fact that AI helped generate the ideas. More volume alone does not prove a lower proportion of good hypotheses.

Put the denominator into a small example

Imagine 1,000 hypothetical experiments. Assume 100 test real effects and 900 test cases with no effect. Also assume 80 percent power and a correctly calibrated ten-percent false-positive rate, with fixed testing rules.

The expected counts are 80 detected real effects and 90 false positives. Among the 170 expected declared wins, roughly 53 percent would therefore be false positives. That is an illustrative calculation under stated assumptions, not a measurement of any product team's experiments.

Change the base rate or power and the result changes. That sensitivity is the point. “Ten percent significance means only one in ten winners is wrong” would answer a different question from the one the test actually controls.

Real programmes add other complications, including repeated looks at results and many metrics. Those choices need an appropriate analysis plan; the simple example is deliberately not a complete model of an experimentation platform.

I would use it to make one distinction visible before discussing whether a looser threshold is a sensible business trade.

Curation can improve the pipeline without blocking invention

A proposed experiment should explain the expected mechanism, the outcome that matters and the evidence that would change the decision. That is a useful review even when a prototype took only an afternoon to create.

For a hypothetical feature team, I would keep the initial test result separate from the decision to roll out permanently. A surprising win might justify confirmation or a limited extension, especially when the effect is expensive to reverse or the initial evidence is thin.

Track what happens afterward. Do the gains persist? Do later tests reproduce the direction? Are rejected ideas and null results retained so the team can learn from the whole pipeline rather than only its most exciting outputs?

The right process can vary with the cost of a mistake. A reversible presentation experiment and a consequential workflow change need not have identical gates. Speed can still be valuable when the team understands the evidence it is choosing to act on.

The essay's broader concern about judgment is worth taking seriously. I would make the statistical part more explicit so the argument supports better decisions rather than merely a preference for fewer experiments.

Evaluate an experiment pipeline by confirmed useful effects, and keep the false-positive rate distinct from the share of announced wins that may be false.

#experimentation#product-development#statistics#ai-workflows
← older drop
Weight updates within a context do not establish personal memory
newer drop →
Vendor-selection advice needs the same evidence check as vendors

related drops

explore all 243 drops →
← back to the archiveday 106