
Vulnerability counts need a CVE ledger
Anthropic's Project Glasswing claims Mythos found 10,000+ high-severity vulnerabilities. An audit found one CVE explicitly attributed to it. The gap isn't fraud — it's a claim versus the public ledger that checks it.
The headline number is genuinely large: Anthropic says Project Glasswing partners used Claude Mythos to find more than 10,000 high- or critical-severity vulnerabilities across major operating systems and browsers, with roughly 50 partner organizations involved. A security audit of the public record, though, counts the vulnerabilities explicitly attributed to Glasswing itself in the CVE ledger at a strikingly smaller figure — as of the accounting I read, essentially one, CVE-2026-4747, a FreeBSD NFS remote-code-execution flaw (alongside a few dozen more attributed to Anthropic researchers broadly). Ten thousand claimed. A tiny number in the ledger. That gap is the whole post.
First, the fair part, because the gap is not evidence of fraud and I won't pretend it is. Coordinated disclosure lags discovery by design — several of the headline finds (a decades-old OpenBSD flaw, a long-lived FFmpeg bug, kernel privilege-escalation chains) reportedly sit under embargo while patches are prepared, and responsible researchers don't publish CVEs before fixes ship. The capability also looks real: independent assessment of a sample put the true-positive rate high, and exploit-generation success jumped from near-zero to meaningfully high. Something genuine is happening. A full public accounting is reportedly due shortly, and the right posture is to update against it.
But hold the two numbers side by side, because they measure different things and only one is checkable. "We found 10,000 vulnerabilities" is a first-party claim about internal counts — pre-triage, pre-dedup, pre-rejection, pre-disclosure. The CVE ledger is the public, adversarial, cross-checked record of vulnerabilities that have been validated, coordinated, and (usually) patched. The first number can include duplicates, false positives, low-impact findings, and issues no maintainer will ever accept. The second cannot — that's what the ledger is for. So the honest read of the gap is not "the capability is fake." It's "the 10,000 is a lab metric and the ledger is the audited one, and until they converge, only the ledger tells you what actually reduced risk."
This is the same distinction I keep coming back to on security tooling: discovery counts are the cheap supply side; validated, accepted, patched fixes are the constraint. Glasswing is just the largest, most-cited instance, which makes it the perfect case study for a discipline every security buyer now needs — because the "our model found N vulnerabilities" slide is about to appear in every vendor pitch, and N will always be the impressive pre-triage number.
The deployable check is one question: which N have CVE numbers, embargo records, or patched releases attached? Not "how many did the model find" — "how many made it to the ledger, or are on a dated, credible path to it." A vendor with a real pipeline answers with a number and a link. A vendor with a trophy count answers with a methodology slide and a promise. And notice the failure mode isn't even the vendor lying — it's the amplification layer: I watched one newsletter relay the 10,000 figure without the one-CVE audit that its own source publication had run separately. The claim travels; the accounting doesn't. Your job as the reader is to demand the accounting the retweet dropped.
At work, this is now my first response to any "AI found N flaws" claim, ours or a vendor's: show me the ledger slice. CVEs assigned, disclosures coordinated, patches shipped or dated. Everything above that line is capability demonstration — real, interesting, worth watching — but it is not risk reduced, and it should not be priced or reported as if it were.
Steal this for your next security-vendor review: when the vulnerability-count slide appears, ask for the same figure filtered to CVE-assigned-or-embargoed-with-a-date, and weight only that number in the decision. Then do the honest thing the vendor should model — commit to revisiting when the full public accounting lands, and actually change your view if the ledger catches up to the claim. It might. The point isn't that Glasswing is hollow; it's that you can't yet know, and "10,000" isn't allowed to stand in for knowing.
A vulnerability isn't found until it's in the ledger — count CVEs and patches, not model outputs, and make the amplifiers show the accounting they skipped.


