
Cheap alignment search makes failure discovery more valuable
Anthropic's automated researchers improved measured alignment failures, including withheld tests. Cheaper optimisation increases the value of finding failures the tests still miss.
Anthropic reports that an automated researcher found improvements across ten categories of alignment failure. Its August 28 account includes gains on benchmarks withheld from the research loop and on Petri's adversarial scenarios. This is stronger than an agent merely memorising the test it can see. The reported experiment.
It still leaves a budget decision that the headline about automated research does not settle. If searching for fixes becomes cheaper, where should the human research effort go? My bet is that discovering and defining neglected failures becomes more valuable, not less.
A fast optimisation loop creates demand for good targets. The danger is allocating attention according to which problems already have convenient scores, then confusing rapid progress on those problems with broad coverage.
Transfer is real evidence with a boundary
The researchers describe a loop of literature search, proposed methods, training and testing. They rejected methods that harmed a predetermined set of general capabilities. They also acknowledge that the studied failures were narrower than production failures and that unmeasured capabilities could still regress. Method and limitations.
That is an important limit on what the result certifies. Passing additional withheld evaluations supports generalisation beyond the visible target. It does not establish that every important behaviour has been tested.
I would avoid the opposite overstatement too: that an automated researcher can only improve the exact benchmark somebody handed it. The withheld results contradict that narrow reading. The useful distinction is between a method's ability to transfer and our evidence that it transfers to the particular situations we care about.
Give discovery its own budget
Here is the allocation I would propose for a hypothetical customer-support agent. One track works on known failures such as unsupported refund promises. A second track deliberately looks for new combinations of circumstances that the current evaluation set does not represent.
The second track might inspect situations where an otherwise correct reply reaches the wrong account, or where two individually acceptable actions combine into an unacceptable outcome. These are illustrative failure searches, not incidents measured in the Anthropic study. Their value is that they can expose a missing test rather than improve an existing score.
I would report discoveries separately from fixes. A team that finds a serious blind spot should not appear to perform worse merely because its evaluation score falls after adding the new cases. Otherwise the reporting system rewards keeping the test set comfortable.
For each proposed fix, retain the failure definition, the examples used during optimisation and a separate acceptance set. Then ask someone outside the optimisation loop to suggest a nearby case that changes the meaning of success. That review can reveal when the model learned a convenient proxy instead of the intended behaviour.
This is a workflow proposal, not a reproduction of the research. Its purpose is to prevent cheap iteration from consuming the entire safety budget simply because the iteration count is easy to display. Discovery work is slower to score, but it determines whether the optimisation work is aimed at the right target.
Anthropic's comparison with human proposals also deserves its stated qualification: the humans could not iterate on their submissions in the same way. The result supports exploring a productive research workflow; it does not establish a universal replacement price for a human researcher. Comparison scope.
Before expanding an automated improvement loop, I would name the person or process responsible for challenging its test coverage. If nobody owns that work, the organisation can get very efficient at closing the gaps it already knows how to count.
As fixes get cheaper, protect the budget for discovering failures that do not yet have a score.


