Mythos AI Severity Ratings: 14 of 27 CVEs Were Wrong

Udi Doron
Udi Doron
Oct 08, 2026 | 12 Minutes
Mythos AI Severity Ratings: 14 of 27 CVEs Were Wrong

Key takeaways

  • Only 1 of 8 findings rated Critical survived independent review, and 14 of the 27 assigned CVEs carried a severity mismatch in total.
  • The funnel's bottleneck was human review, not model accuracy - 90.8% of reviewed findings were valid, while roughly 21,000 were never reviewed at all.
  • One finding was rated too low, so a triage process tuned only to catch inflation will still miss things.
  • The findings skew heavily to C and heap overflows (67% and 37%), and the corpus composition has not been published - so no general claim about coverage is currently supportable.
  • Severity was already a weak proxy for exploitability - 10–20% of exploited vulnerabilities are rated only medium - and an upstream generator biased toward Critical makes sorting by score worse, not better.

Anthropic's Mythos Preview generated 23,019 candidate vulnerability findings. Twenty-seven became assigned CVEs - under 1% - and when those 27 were independently re-scored, 14 carried a severity mismatch. Thirteen were overstated and one understated. Of the eight findings the model rated Critical, exactly one survived independent review. If you are planning to feed AI-generated vulnerability reports into a prioritisation pipeline, that ratio is the number to plan around.

What happened to 23,019 findings?

The disclosure funnel is the most instructive artefact of the whole exercise, because almost all the attrition happens before a human ever looks:

Mythos Preview disclosure funnel: 23,019 candidate findings to 27 assigned CVEs
Stage Count
Candidate findings generated 23,019
Externally reviewed 1,900 (90.8% confirmed valid)
Confirmed 1,726
Reported to maintainers 467, plus 1,129 direct by Anthropic = 1,596
Acknowledged by maintainers 1,451
Patched upstream 97
Advisories published 88
CVEs assigned 27
Source: Echo, Mythos Readiness Report.

Source: Echo, Mythos Readiness Report.

‍

Two figures deserve emphasis. Roughly 21,000 findings were never reviewed at all - not rejected, simply never examined. And of the sample that was reviewed, 90.8% were confirmed valid. The model was not mostly wrong. It was mostly unread, which is a different and more awkward problem: the constraint was human review capacity, not model precision.

How badly were the severities wrong?

Echo pulled the full commit history for each of the 27 from cvelistV5 and re-derived every severity rating from its CVSS vector. The comparison:

Severity ratings for Mythos Preview's 27 assigned CVEs, before and after independent review
Severity Mythos initial rating After independent CVSS and maintainer review
Critical 8 1
High 15 16
Medium 4 8
Low 0 2
Source: Echo, Mythos Readiness Report.

14 of 27 carried a mismatch - 13 overstated, 1 understated. The understated one matters as much as the rest: jq (CVE-2026-32316) was rated too low, which is the failure mode no triage process catches, because nobody re-reviews the things that came in as minor.

Two worked examples show the scale of the drift:

  • Temporal Server (CVE-2026-5199) - rated Critical by the model; maintainers scored it 2.3, Low. Three rungs.
  • MinIO (GHSA-xh8f-g2qw-gcm7) - rated Critical by the model; MinIO scored it Medium. Two rungs.

One caveat, stated plainly: this finding is about Mythos Preview's published set of 27, not about AI models generally or about any vendor's current release. It is a sample of 27, from one model, in one disclosure programme.

What kinds of bugs did the model actually find?

The distribution is narrow in a way that should temper any general claim about AI vulnerability discovery:

  • By language: C accounts for 18 of 27 (67%), and all 12 memory-corruption findings. The other nine - spanning Go, Ruby, Java, JavaScript, PHP and Rust - involved no memory corruption at all.
  • By weakness: heap buffer overflow accounts for 10 of 27 (37%). That is notable because heap overflows rank only #16 in MITRE's 2025 CWE Top 25, and were absent from the 2024 list entirely.
  • By project: wolfSSL 9, FreeRDP 3, Mastodon 2, nginx 2, then one each across CraftCMS, gitoxide, ImageMagick, jq, junrar, libyang, MapServer, MinIO, Nomad, Temporal and Ghost.
  • By category: cryptography and TLS 9, servers and network protocols 5, media and file parsers 4, web apps and CMS 4, cloud infrastructure and orchestration 3, developer tooling 2.
  • By vintage: the 27 trace to code introduced between 2004 and 2026, including 2024–2026 - so this is not simply resurfacing decades-old code.

There is an unresolved confound here worth naming. Anthropic has not published the language breakdown of the 281 projects scanned, so it is impossible to separate benchmark composition from model strength. A 67% C skew could mean the model is good at memory corruption, or it could mean the corpus was mostly C.

Why this breaks severity-based prioritisation specifically

Most triage pipelines sort by severity. That works when severity is an independent assessment. It stops working when the rating and the finding come from the same automated source, because the queue then reflects the generator's calibration rather than your risk.

The underlying problem predates AI. Echo's research found that 10–20% of exploited vulnerabilities are only rated medium severity - precisely the band teams deprioritise - while the share of disclosed CVEs ever exploited has fallen from about 2.2% in 2021 to well under 1%. Severity has been a weak proxy for exploitability for years. An upstream generator biased toward Critical simply makes the proxy worse.

This is the argument for prioritising on exploitability and business impact rather than score alone, which we set out in EPSS vs CVSS: how to prioritise container CVEs, and it is why most container security tools fail at vulnerability prioritization.

How to handle AI-generated vulnerability reports

Practical rules that follow from the data above:

  • Re-derive severity from the CVSS vector, not the label. The vector is auditable; the label is an assertion. Echo re-scored all 27 this way, which is how the 14 mismatches surfaced.
  • Review the low-rated findings too. One of the 27 was understated. Processes tuned to catch inflation will never catch deflation.
  • Weight maintainer assessment above generator assessment. Maintainers know the reachability and the deployment context; the model does not.
  • Do not treat "acknowledged" as "patched." 1,451 findings were acknowledged; 97 were patched upstream. Acknowledgement is not remediation.
  • Keep a human in the loop for fixes. Among security leaders already using AI extensively for supply chain security, only one in eight would trust an AI-generated fix without human review - and that instinct is well calibrated on this evidence.
  • Expect long advisory gaps. Echo's own worked example is instructive: a fix for CVE-2026-7482 in Ollama was validated on 25 February and submitted to MITRE on 2 March, but the CVE was not assigned until 28 April and published on 1 May - 65 days with a fix available but no advisory.

That last gap is why Echo became an authorised CVE Numbering Authority; the mechanics are in echo is now a CVE Numbering Authority, and the structural reasons advisories go missing are in how CVEs fall through the cracks in upstream bug trackers.

FAQ

Were the Mythos findings mostly false positives?

No - that is the common misreading. Of the 1,900 findings that were externally reviewed, 90.8% were confirmed valid. The attrition from 23,019 candidates to 27 CVEs was driven overwhelmingly by review capacity rather than accuracy: roughly 21,000 findings were never examined at all. Validity and severity accuracy are separate questions, and the model performed far better on the first.

How wrong were the severity ratings?

Substantially, and in a consistent direction. Fourteen of the 27 assigned CVEs carried a severity mismatch after independent CVSS re-derivation and maintainer review: 13 were overstated and one understated. Of eight findings initially rated Critical, only one survived. The largest single drift was a Temporal Server issue rated Critical that maintainers scored 2.3, or Low - three severity rungs.

Does this mean AI vulnerability discovery does not work?

It means the discovery and the scoring should be judged separately. The findings were largely valid and spanned code introduced between 2004 and 2026, so this was not just resurfacing old bugs. What proved unreliable was the model's own severity judgement. Treat AI output as candidate findings requiring independent scoring, not as a prioritised, ready-to-action queue.

Why were two-thirds of the findings in C code?

C accounted for 18 of the 27, and all 12 memory-corruption findings. That may reflect genuine model strength at memory-safety analysis, or it may reflect the corpus. Anthropic has not published the language composition of the 281 projects scanned, so the two explanations cannot currently be separated. Treat any generalisation about language coverage as unsupported until that breakdown is released.

What should change in my triage process?

Three things. Re-derive severity from the CVSS vector rather than accepting a supplied label. Review low-rated findings as well as high ones, since deflation exists and nothing else catches it. And stop treating maintainer acknowledgement as remediation - 1,451 findings were acknowledged in this programme while only 97 were patched upstream.

Frontier models can now find and exploit zero days in your code

Find out what to do when there's no CVE, scanner alert or fix available yet.

What are the 7 blind spots in your vulnerability scans?

Discover when "0 vulnerabilities" doesn't actually mean you're clean.

Read now →

Ready to eliminate vulnerabilities at the source?