← back to the archiveCover illustration for “The AI safety index grades disclosure, not safety”
POSTday 52·1 week ago·by Andy Padia

The AI safety index grades disclosure, not safety

FLI's Summer 2026 index put Anthropic on top with a C+ and handed three labs an F. The scores largely measure what labs publish — reading them as safety measurements misreads the instrument.

The Future of Life Institute published its Summer 2026 AI Safety Index on July 7, and it circulated through my feeds this week under "best lab scored C+" headlines. The numbers, verified against the FLI primary: Anthropic tops the table at 2.66 (C+), then OpenAI 2.28 (C), Google DeepMind 2.01 (C), Meta 1.32 (D+), Z.ai 0.88 and Alibaba Cloud 0.87 (both D-), and xAI 0.65, DeepSeek 0.47 and Mistral 0.33 — three F grades spanning the US, China and Europe. Nine companies, 37 indicators, six domains, a seven-expert panel.

One newsletter framed this as an industry that voluntarily called in independent experts to grade it. The FLI primary does not support that. FLI is an independent nonprofit that grades labs whether they cooperate or not; being graded is not submitting to grading. The methodology: public materials collected until June 3, 2026, plus a "targeted company survey" — and the report does not say which of the nine companies answered the survey, which I could not determine either.

That methodology line is the whole story. An index built substantially on public disclosure cannot distinguish "unsafe" from "undocumented". Put the question at its sharpest: is Mistral eight times less safe than Anthropic, or eight times less published? The index cannot say. That is not a flaw FLI is hiding — grading on documentation is a defensible way to pressure labs toward transparency, and I think the index is useful for exactly that. The failure happens downstream, when a transparency score gets read as a safety measurement by someone with a procurement decision to make.

Where the misreading bites

I have watched this live. At work, in a vendor risk review this month, a model choice was defended with the lab's index grade — the letter did the arguing, and nobody in the room had read a line of the methodology. The mirror-image error was in the same meeting: skepticism toward a lab with thin public documentation, scored low, whose actual deployed controls nobody had examined. Both positions treated the grade as evidence about the model in front of us. It is evidence about the lab's publishing habits, collected before June 3, about systems we were not even discussing.

My rule, and the one I now put in writing for clients: third-party index grades go in the evidence appendix of a risk review, never in the controls section. A grade can prompt a question — "Meta scored D+, what does their documentation not cover?" — but it cannot answer one. The controls section gets filled by things you can verify against your own deployment: the eval you ran on your task, the data-handling terms in your contract, the incident-response commitment with your name on it, the access controls you tested. A lab's C+ does not stop your prompt-injection incident, and a lab's F does not cause one.

There is also a second-order effect worth betting on. My bet: within two grading cycles, labs will optimize for the index — publishing more frameworks, more policies, more whistleblowing pages — and scores will rise faster than underlying practice changes. Documentation is the cheapest indicator to move. When the Winter index shows broad improvement, remember that the instrument measures what it measures.

Steal this for your next review: take the 37 indicators, mark which ones your organization could verify independently for your actual deployment, and use only that subset in the decision. In my quick pass, most indicators are disclosure checks; the verifiable-by-you set is small. That small set is your real checklist — the rest is context.

A C+ measures what a lab publishes, not what your deployment risks — file index grades under evidence, and fill the controls column with checks you ran yourself.

#ai-safety#governance#vendor-risk#evals#procurement
← older drop
AI pricing changes now arrive as quota emails, not price lists
newer drop →
Stateless MCP moves the state problem onto your side

related drops

explore all 78 drops →
← back to the archiveday 59