← back to the archiveCover illustration for “AI research needs an interest gate”
POSTday 125·today·Published ·by Andy Padia

AI research needs an interest gate

In short: BootLoops shows why exact answers still need an expert interest gate: verification proves the result, not that the question deserves the next experiment.

On 1 October 2026, Harvard physicist Matthew Schwartz reported 36 manuscripts across 18 fields in three months with Claude, BootLoops and 19 human coauthors. The uncomfortable result was not that the AI kept getting the mathematics wrong. It was that technically correct work could still leave domain experts unmoved.

AI research needs an interest gate after the correctness gate. Exact checks can establish that an answer follows. They cannot establish that the question is novel, consequential or worth the next experiment.

BootLoops verifies answers before claiming discovery

Schwartz built BootLoops as an open-source harness for precision quantitative science. Its public repository describes certified tools, patched computational engines and protocols that an AI agent can drive.

The fit is deliberate. In the October account, Schwartz says the semi-numerical physics problems were attractive because a second script could check a result to as many digits as needed. BootLoops reportedly completed 30 integrals: 15 reproductions of known results and 15 integrals not previously computed by that method.

Author-reported project scaleCount
Candidate problems consideredabout 400
Manuscripts produced36
Fields represented18
Human coauthors19
Project period3 months

These figures come from Schwartz's guest post; I did not reproduce the code or scientific results.

That is a strong verification story, not a blanket discovery claim. A calculation can be exact while its interpretation is ordinary, its question is stale or its consequence is already known by the field.

AI research needs an interest gate after verification

The ecology example makes the gap visible. Claude recognised a 20-year-old equation as BootLoops-shaped and solved it at scale. The resulting model said species composition in a well-studied forest changed 4.5 times faster than neutral theory allowed.

When Schwartz took the result to plant biologist James O'Dwyer, the technical feat survived. The scientific excitement did not. Ecologists already knew qualitatively that neutral theory lagged real forests. The expert then redirected the work toward a classification that separated species-level longevity, growth and recruitment from random fluctuation.

Population genetics followed a similar arc. A solved 30-year-old expression impressed an expert without compelling him. The useful redirect examined correlations between nearby mutations; the team then analysed 5.7 billion mutation pairs and reported evidence for gene conversion.

Those outcomes are author-reported and still need the normal scientific scrutiny. The operating lesson is narrower: expert review was not a ceremonial sign-off after the calculation. It changed the question being asked.

My practitioner rule: require a consequence receipt

This is labelled editorial judgment. I have not deployed BootLoops at Trigent or for a client, and I did not run its tools. If I were reviewing an AI research workflow, I would keep two separate acceptance fields.

The correctness receipt asks whether another operator can reproduce the computation, inspect the inputs and make the result fail under a targeted counter-check. That extends the archive's rule that a plausible explanation is not a root cause.

The consequence receipt names the domain owner and records what changed: a live hypothesis, a discriminating experiment, a decision, a dataset to collect or a method worth replacing. “The expert found it interesting” is too soft. The receipt should point to the next committed action.

Verification buys permission to believe the result. The interest gate buys permission to spend the next week on it. Neither can substitute for the other, and the same person should not quietly approve both.

This also complements scientific AI's need for an experiment trace. Lineage preserves how a result emerged. An interest gate decides whether that result deserves another loop.

What's in it for you

  • Add separate correctness and consequence fields to the research review.
  • Require a named domain owner and one changed next action before scaling the work.
  • Keep technically sound dead ends; they may become tools, but do not call them discoveries.

An exact answer earns belief; only a consequence earns the next experiment.

Sources

#scientific-ai#evaluation#ai-agents#research#human-in-loop
← older drop
Karpathy’s four ideas, turned into working examples

related drops

explore all 371 drops →
← back to the archiveday 125