← back to the archiveCover illustration for “A perfect ExploitBench score changes what the test can tell us”
POSTday 94·12d ago·by Andy Padia

A perfect ExploitBench score changes what the test can tell us

Astra’s reported 100% result needs contamination and configuration context. Keep public tests for regression while demanding separate evidence for harder capability claims.

I would keep a benchmark after a model scores 100%. I would stop using that score to explain how much better the next model is.

OpenAI’s September 1 Astra assessment reports a perfect ExploitBench result, then describes a newer internal set of 20 high-severity V8 vulnerabilities because of contamination concerns. The company also describes expert-led assessments. Its Critical cybersecurity designation draws on that combined investigation, not just the twenty-item replacement.

That last distinction matters. A private test is harder for outsiders to reproduce, but it is inaccurate to say the whole classification rests on one hidden number. We should challenge the visibility of the evidence without deleting the evidence categories the lab actually reports.

My bet is that the next useful benchmark discussion will spend less time celebrating ceilings and more time explaining what each test still distinguishes.

Saturation and contamination ask different questions

A saturated test has little remaining room to separate strong systems. It may still catch a regression: the same system that previously passed can fail after a change to its tools, instructions or deployment.

Contamination is a different concern. If prior exposure could help solve the cases, a high score may be less informative about performance on unfamiliar problems. A newer set can reduce that particular concern without automatically becoming representative of every environment a defender cares about.

Neither issue makes the public test worthless. Neither is solved merely by making the replacement private. The useful response is to state the role of each set and the limits of the conclusions drawn from it.

For a capability comparison, I want enough hard, fresh cases to discriminate. For release regression, I want continuity with known cases. For a deployment decision, I need the configuration that will actually be available to the people doing the work.

That third requirement is especially relevant here. OpenAI says the reported Astra cyber results reflect Daybreak Blue access, rather than the default production configuration. A procurement team cannot treat a restricted-access result as a promise that its ordinary account will complete the same work.

Keep a public anchor and add a harder test

Imagine a hypothetical security team evaluating an assistant for authorized vulnerability triage. I would retain its existing known cases, even after most tools pass them. Then I would add a separate set drawn from newly available, authorized test environments and record when each case entered the evaluation.

The report would show the two results separately. It would also identify access tier, tools, effort allowance and whether a human supplied intermediate hints. Combining them into one impressive percentage would erase information we need when a later release behaves differently.

I would ask the vendor for independent assessment where available, a description of selection criteria and a clear account of what cannot be disclosed. Sensitive details can justify limited publication. They do not remove the need to distinguish a reported result from something my team has reproduced.

A small internal set can be useful evidence. It cannot tell us every failure mode, and it should not be asked to. The mistake is making a test carry a broader claim than its cases and configuration support.

The perfect score is an invitation to improve the measurement, not a reason to declare either universal capability or universal safety.

Keep the old benchmark as an anchor; require fresh cases and the actual access configuration for the next decision.

#evaluations#cybersecurity#openai#benchmarks
← older drop
Grok Bot roles share the authority of their user's computer
newer drop →
Anthropic’s five-point safeguard gap already has a version problem

related drops

explore all 329 drops →
← back to the archiveday 106