Claude Mythos, Explained: What the Benchmarks Show

Key takeaways
- Name the release you mean. Mythos Preview (April 2026, research only), Mythos 5 and Fable 5 (June 2026, commercial) are different things, and almost every published benchmark refers to the preview.
- The 90x figure is real but narrow: 181 of 250 trials, against known bugs, with Firefox's sandbox removed, in an evaluation the vendor designed and ran.
- Independent results are mixed - Cloudflare found coverage, not capability, was the limit; AISI recorded 73% on Expert-tier CTFs but a failure on operational technology; AISLE found a 3.6B model matched frontier models at detection.
- Cost, not capability, is the durable change. A working exploit chain for under $2,000, and a competitive field of suppliers, means this does not reverse.
- Discovery is commoditising; remediation capacity is not. 37% of surveyed security leaders already name "detecting more than we can fix" as their biggest obstacle, and only 11% would spend next on more detection.
Anthropic has shipped three distinct things under the Mythos name, and conflating them is the most common error in coverage of it. Mythos Preview was an April 2026 research release that was evaluated by Anthropic and third parties but never made generally available. Mythos 5 and Fable 5 are the commercial models that followed in June 2026. The headline capability claim - a roughly 90-fold improvement in exploit development over the previous generation - comes from a single benchmark on the research preview, with caveats that materially change how you should read it.
What the Firefox benchmark actually measured
Anthropic's evaluation took 50 known vulnerabilities in the SpiderMonkey engine of Firefox 147 and gave the model five attempts at each - 250 trials in total. The results:
- Mythos Preview produced working exploits in 181 of 250 trials (72.4%), plus partial register control in a further 29 (11.6%).
- Claude Opus 4.6 managed 2 (under 1%) on the same benchmark.
That gap is where the ~90x figure comes from. Three caveats travel with it, and all three are load-bearing:
- It measures exploit development against known vulnerabilities, not discovery. The bugs were handed to the model.
- The evaluation removed Firefox's sandbox and other defence-in-depth layers. Production browsers have them.
- Anthropic designed and ran it, and it has not been independently replicated.
A 72.4% success rate at weaponising a bug someone else already found, in an environment with the mitigations switched off, is a meaningfully narrower claim than "AI can now hack browsers."
What independent evaluations found
Three external results fill in the picture, and they do not all point the same way.
Cloudflare tested the preview against 50+ of its own production repositories. It chained small weaknesses into realistic multi-step exploit paths - but the limiting factor was coverage, not capability: finite context on large repositories. Notably, other frontier models found many of the same bugs. Mythos's edge was synthesis into complete chains, not detection.
The UK AI Security Institute (AISI) recorded a 73% success rate on Expert-tier capture-the-flag challenges, the first model to reach that level, at a 50M-token budget - roughly $2,500 at published pricing. The same institute's separate operational-technology evaluation is the counterweight: Mythos failed it, getting stuck before reaching the OT-specific stages.
AISLE evaluated smaller and open-weight models and found the frontier advantage is narrower than the headline suggests:
- On a 17-year-old FreeBSD NFS bug, every model tested found it and rated it correctly - including a 3.6B-parameter model costing about $0.11 per million tokens.
- False-positive rates do not improve with model size.
- Patch validation is a shared weakness: large and small models alike still flag already-patched code as vulnerable.
- The frontier advantage shows up in exploit construction, not detection.
The economics changed more than the capability did
The capability numbers are contested. The cost numbers are not, and they are the ones that should change your planning:
Source: Echo, Mythos Readiness Report.
On published browser-exploit-development scoring, the models cluster rather than separate: GPT-5.6 Sol at 73.5% on ExploitBench against 74.2% for Mythos Preview and 78% for Mythos 5. This is not a one-vendor capability. It is a category shift with several suppliers, which means it will not be undone by any single company's access policy.
Access to these models is not stable, and you cannot plan around it
Fable 5 and Mythos 5 were released on 9 June 2026. Access was suspended on 12 June under US Department of Commerce export controls, the controls were lifted on 30 June, and access was restored on 1 July - 19 days offline for every customer worldwide. OpenAI's GPT-5.6 Sol was likewise released under restricted access to government-approved partners, at the National Cyber Director's direction.
The defensive implication is uncomfortable: a security programme that depends on frontier-model access has a dependency that can be withdrawn by regulators with three days' notice. Attackers operating outside that perimeter do not share the constraint.
Why discovery capability is the wrong thing to worry about
Mean time to exploitation has fallen from 1,142 days in 2018 to 14 days in 2026. Meanwhile roughly 89% of CVEs already have a fix available - 96.5% of Critical and 95.7% of High - and about 40% of fixable vulnerabilities remain unresolved past six months.
The bottleneck was never finding things. Echo's survey of 80+ US senior security professionals at organisations with 20+ development teams, fielded 16–25 July 2026, makes the point directly:
- 37% named "detecting more than we can fix" as their single biggest obstacle - the top answer.
- Only 11% would prioritise additional detection or scanning as their next investment.
- Among those already using AI extensively for supply chain security, only one in eight would trust an AI-generated fix without human review.
- Only 17% could effectively quantify supply chain risk for their board in business terms.
If discovery is commoditising - and AISLE's 3.6B-parameter result suggests it is - then judgment and remediation capacity are the scarce resources, and the organisations that come out ahead are the ones inheriting less risk rather than detecting more of it.
What to do about it this quarter
- Stop buying detection to solve a remediation problem. Measure the gap between seeing a risk and being protected from it, not the size of your findings backlog.
- Assume prior-year CVEs are the threat. Most exploited CVEs in any given year are from earlier years; 2026 is projected at roughly 280 total, 142 prior-year against 121 same-year.
- Treat AI-generated severity ratings as unreviewed input. See EPSS vs CVSS: how to prioritise container CVEs for why sorting by severity score alone fails.
- Reduce what you inherit. Fewer packages in production means fewer findings to triage, whoever or whatever generated them.
For the wider context on how AI is reshaping the supply chain threat model, see AI supply chain security: where the new risks actually come from and our earlier note on supply chain attacks, Mythos, Glasswing and Echo.
FAQ
Is Claude Mythos publicly available?
Partly. Mythos Preview, the April 2026 research release that most published benchmarks refer to, was never made generally available - it was released for evaluation only. Mythos 5 and Fable 5, the commercial models, shipped in June 2026 and are available commercially, though access was suspended for 19 days that month under US export controls before being restored on 1 July.
Does Mythos find new vulnerabilities on its own?
The flagship Firefox benchmark does not test that. It measured exploit development against 50 vulnerabilities that were already known and supplied to the model. Cloudflare's independent test of its own repositories did involve discovery, and found that other frontier models surfaced many of the same bugs - the distinguishing capability was chaining weaknesses into complete multi-step exploit paths, not spotting them.
How much does an AI-generated exploit actually cost?
Less than most threat models assume. Turning a known Linux vulnerability into a working exploit was achieved for under $2,000 and in under a day, although that figure covers a two-vulnerability chain rather than a single CVE. Several dozen OpenBSD vulnerabilities were reportedly surfaced for under $20,000 in inference cost. Expert-tier CTF runs cost roughly $2,500 each.
Do I need a frontier model to defend against one?
The evidence suggests not, for detection. AISLE found that on a 17-year-old FreeBSD NFS bug, every model it tested identified and rated the issue correctly, including a 3.6-billion-parameter model at about $0.11 per million tokens. Frontier models pull ahead at exploit construction, which is an offensive task. Defensive value concentrates in remediation capacity, not in detection horsepower.
Should this change my patching SLA?
Probably your prioritisation, more than your SLA. Mean time to exploitation has fallen from 1,142 days in 2018 to around 14 days in 2026, but 89% of CVEs already have a fix available and roughly 40% of fixable ones sit unresolved past six months. The constraint is throughput, not speed of detection - so fix capacity is what to invest in.
Frontier models can now find and exploit zero days in your code
Find out what to do when there's no CVE, scanner alert or fix available yet.
What are the 7 blind spots in your vulnerability scans?
Discover when "0 vulnerabilities" doesn't actually mean you're clean.





.avif)
.avif)