← back to the archiveCover illustration for “Anthropic’s five-point safeguard gap already has a version problem”
POSTday 94·12d ago·by Andy Padia

Anthropic’s five-point safeguard gap already has a version problem

Fable 5.1 and Mythos 5.1 share a base model, but their published benchmark gap reflects earlier safeguards. A difference in the table is not a permanent price of safety.

Five point one percentage points looks like a clean price tag for safeguards. Anthropic’s own footnote makes it considerably less clean.

In its Fable 5.1 and Mythos 5.1 announcement, Anthropic reports Terminal-Bench 4.0 scores of 55.8% and 60.9%. It says these are the same underlying model with different safeguard levels. It also says the gap reflects earlier, less precise cyber safeguards and expects the difference to shrink with the changes announced at launch.

So I would not circulate “safety costs five points” as a stable product specification. The table is evidence of an intervention effect under a particular evaluation setup. The launch text says the setup itself was changing.

That is more interesting than a permanent tax. It means a benchmark result can become stale without a new set of model weights.

The version number needs to reach beyond the model

We already expect a model name to identify what was tested. This example requires a second question: which safeguards were active during the run?

Anthropic’s evaluation notes also distinguish what happened after an intervention. Some tasks received a zero; other interventions routed tasks to another model. Those outcomes enter the score differently. A refusal, a fallback that succeeds and a fallback that fails are not interchangeable events, even if a final aggregate makes them look like ordinary right and wrong answers.

This is a narrower issue than the procurement argument in “Less restrictive is now a model spec”. Specifying a restriction profile tells us what we intend to buy. Recording the profile used for a benchmark tells us whether the evidence describes that purchase.

My proposed addition to an evaluation record is a dated description of intervention behavior. Keep it beside the model, harness and test version. If the vendor changes a classifier or fallback policy, that should trigger a check of affected tasks even when the advertised model name stays unchanged.

Re-run the cases whose outcome can move

Suppose, hypothetically, I am choosing a coding assistant for an internal maintenance workflow. A few legitimate tasks trigger safeguards during our trial. I would keep those tasks in a small replay set with the permitted action and expected result documented.

After a safeguard update, I would run the same authorized cases again. A higher completion rate would be useful. I would still inspect whether the work was answered by the intended model, routed elsewhere or completed with a different set of tools.

I would also retain tests of actions that should remain disallowed. Fewer interruptions on legitimate work is desirable; a higher score alone does not tell me whether the boundary became more precise or simply moved.

That exercise does not measure a universal cost of safety. It measures the effect of a particular change on our workload. Its modest scope is an advantage: the team can repeat it and explain why an acceptance decision changed.

The published five-point difference remains informative historical evidence. It shows that access controls can materially affect reported task performance. What it cannot supply is a timeless exchange rate between protection and capability, or a promise that removing controls always buys the same improvement.

When safeguards change, the benchmark configuration changes—even if the model’s name and weights stay the same.

#anthropic#evaluations#safeguards#models
← older drop
A perfect ExploitBench score changes what the test can tell us
newer drop →
Perplexity’s growth curve needs units before a multiplier

related drops

explore all 329 drops →
← back to the archiveday 106