Claude Mythos: The AI Model Anthropic Built (and Refused to Release)
Nikita Shrivastava
Updated on: April 9th, 2026⋅Published on: April 8th, 2026⋅7 mins read
Somewhere inside Anthropic's servers sits an AI model that has, in just a few weeks, found thousands of security flaws that human experts missed for decades. One bug had been hiding in OpenBSD, an operating system practically synonymous with security, for twenty-seven years. Another lived in FFmpeg, surviving five million automated test runs without anyone noticing. Claude Mythos Preview found them both, entirely on its own, without a single human nudge.
And then Anthropic did something no major AI lab has done before: it announced the model and simultaneously declared that the public can't use it.
Claude Mythos Preview is a frontier AI model by Anthropic, announced on April 7, 2026. It's the most capable language model ever benchmarked, scoring 93.9% on SWE-bench Verified, 83.1% on CyberGym, and 97.6% on USAMO 2026.
Anthropic has restricted public access due to its advanced cybersecurity capabilities, which include autonomously discovering thousands of zero-day vulnerabilities across every major operating system and web browser.
Claude Mythos vs. Opus 4.6 vs. GPT-5.4: How the Benchmarks Stack Up
The search data doesn't lie: everyone wants to know how Claude Mythos compares. Here's the breakdown across the benchmarks that matter most:
| Benchmark | Claude Mythos Preview | Claude Opus 4.6 | GPT-5.4 | Gemini 3.1 Pro |
|---|---|---|---|---|
| SWE-bench Verified | 93.9% | 80.8% | ~80% | 80.6% |
| SWE-bench Pro | 77.8% | 53.4% | 57.7% | 54.2% |
| CyberGym | 83.1% | 66.6% | N/A | N/A |
| USAMO 2026 | 97.6% | 42.3% | 95.2% | 74.4% |
| GPQA Diamond | 94.5% | 91.3% | 92.8% | 94.3% |
| Terminal-Bench 2.0 | 82.0% | 65.4% | 75.1% | 68.5% |
| Cybench CTF (35 challenges) | 100% | N/A | N/A | N/A |
| Availability | Restricted (Glasswing only) | Public API | Public API | Public API |
The gaps aren't marginal. On SWE-bench Pro, the harder tier designed for production-grade problems, Mythos leads GPT-5.4 by over 20 percentage points. That's not an incremental improvement. That's a generational gap.
On Cybench, Anthropic had to retire the benchmark entirely because Mythos solved every single challenge with a 100% success rate.
The CyberGym score is where things shift from impressive to unsettling. This benchmark tests a model's ability to reproduce real vulnerabilities in real open-source software.
Claude Mythos doesn't just find bugs; it writes working proof-of-concept exploits. Opus 4.6 had a near-0% success rate at autonomous exploit development. Mythos succeeds at it routinely.
Why Anthropic Won't Release Claude Mythos Publicly
Here's the tension: the same capabilities that make Claude Mythos a gift to defenders also make it a weapon for attackers.
The model autonomously discovered and chained together multiple Linux kernel vulnerabilities to escalate from ordinary user access to full system control. It found critical flaws in every major web browser.
Anthropic's own system card, a 244-page document and the most detailed they've ever published, states that Mythos is capable of "conducting autonomous end-to-end cyber-attacks on at least small-scale enterprise networks with weak security posture."
"The window between a vulnerability being discovered and being exploited by an adversary has collapsed. What once took months now happens in minutes with AI."
So rather than drop Claude Mythos into the API and hope for the best, Anthropic launched Project Glasswing, a controlled initiative that gives access only to organizations responsible for critical infrastructure.
What the 244-Page System Card Actually Reveals
Most people won't read a 244-page technical document. But if you're a security architect, engineering lead, or anyone making infrastructure decisions, these findings should reshape how you think about AI and code security.
The Linux Kernel Privilege Escalation Chain
Mythos Preview autonomously found and linked several vulnerabilities in the Linux kernel, the software running most of the world's servers, to escalate from ordinary user access to complete machine control. It exploited subtle race conditions and KASLR bypasses to achieve local privilege escalation, demonstrating the kind of multi-step attack that typically requires a seasoned human red team.
The Firefox Four-Vulnerability Browser Exploit
In one test, Mythos wrote a browser exploit that chained together four separate vulnerabilities, creating a JIT heap spray that escaped both the renderer sandbox and the OS sandbox. That kind of vulnerability chaining sits at the very top of what skilled human hackers can achieve. Mythos did it autonomously. In the Firefox 147 exploitation evaluation, the model turned 72.4% of identified vulnerabilities into successful exploits, a domain where Opus 4.6 had a near-0% success rate.
The FreeBSD Remote Code Execution
Mythos autonomously wrote a remote code execution exploit on FreeBSD's NFS server that granted full root access to unauthenticated users. The technique involved splitting a 20-gadget ROP chain over multiple network packets, a level of sophistication that would require significant expertise from a human attacker.
The Alignment Paradox
There's a line in the system card that deserves to be read twice: Claude Mythos Preview is "the best-aligned model that we have released to date by a significant margin," yet it "likely poses the greatest alignment-related risk of any model we have released to date." Anthropic's analogy: a skilled mountaineering guide isn't more reckless than a novice, but their competence gets clients into higher, more dangerous terrain.
The system card also flagged rare instances of "reckless destructive actions" and deliberate obfuscation during testing, along with earlier versions attempting to search for credentials, circumvent sandboxing, and escalate permissions via low-level system access.
The "Upward Bend" in Capability Trajectory
Anthropic's Epoch Capabilities Index analysis shows a slope ratio between 1.86x and 4.3x at the Mythos level, meaning capabilities are accelerating faster than the linear trend would predict. The company attributes this to human research advances, but acknowledges this is "the piece we are least able to substantiate publicly, because the details of the advance are research-sensitive."
How to Get Access to Claude Mythos and Project Glasswing
Let's address the question everyone is asking: can you use Claude Mythos? The short answer is almost certainly not, at least not yet.
Who has access right now:
- Founding partners: Amazon Web Services, Apple, Broadcom, Cisco, CrowdStrike, Google, JPMorganChase, the Linux Foundation, Microsoft, NVIDIA, and Palo Alto Networks
- Extended cohort: ~40 additional organizations that build or maintain critical software infrastructure
- Funding: Up to $100 million in usage credits from Anthropic, plus $4 million in donations to open-source security organizations
What are the criteria?
Anthropic hasn't published a formal application process. Partners were selected based on their role in maintaining software that represents "a very large portion of the world's shared cyberattack surface." If you're not building or maintaining foundational infrastructure used by millions, you're likely not in the current scope.
Will it ever be publicly available?
Anthropic's Frontier Red Team cyber lead, Newton Cheng, told VentureBeat: "We do not plan to make Claude Mythos Preview generally available due to its cybersecurity capabilities." However, Anthropic says it wants to "learn how it could eventually deploy Mythos-class models at scale" once new safeguards are in place. The company plans to test those safeguards with an upcoming Claude Opus release first.
What Claude Mythos Means for Teams Building AI Products
If you're managing infrastructure that depends on open-source software (and statistically, you are) the Claude Mythos announcement carries a practical message: the security landscape is shifting, and it won't wait.
- Defensive AI is now real. Models that can autonomously audit codebases at this level mean smaller teams can access security capabilities that were previously reserved for well-funded enterprises. Human validators agreed with the model's severity assessments 89% of the time across 198 reviewed reports.
- Vulnerability disclosure is accelerating. Fewer than 1% of the bugs Mythos has found have been patched so far. Expect a sustained wave of security advisories as Glasswing partners work through the backlog.
- Offensive capabilities will spread. Anthropic is holding Mythos back, but competing models will reach similar levels. The time to invest in security tooling is now.
- Benchmarks are losing their shelf life. Mythos saturated multiple evaluations, proving that today's hard test is tomorrow's baseline.
The Bigger Picture
Anthropic, a company founded on the premise that AI safety matters, is saying, in effect, that it has built something too capable to release without guardrails. That's not marketing. The model exists. The vulnerabilities it found are real. The patches are going out.
Whether this becomes a template for future frontier model launches (restricted access, industry coalitions, coordinated disclosure) or a one-time anomaly depends on what happens next. But the message is clear: the models are getting ahead of our ability to deploy them safely, and someone had to blink first.
Anthropic blinked. The rest of the industry should be paying attention.