AI finds the vulnerabilities in seconds. Defenders still need weeks to close them

Anthropic's Mythos model has surfaced more than 10,000 high- and critical-severity software vulnerabilities in a matter of weeks, and the security industry is only beginning to absorb what that volume does to a defensive posture built for a slower era.

The model, released in preview through Project Glasswing to a closed coalition of roughly 50 organisations including Microsoft, Google, Apple, Amazon Web Services, Cisco, CrowdStrike and JPMorganChase, has already identified potential flaws in every major operating system and web browser, including decades-old bugs that survived years of expert human review.

Anthropic itself has said the capability revealed a stark fact about where coding models now sit, describing the moment as a reckoning for the sector rather than an incremental advance.

The numbers give the claim substance. Anthropic scanned more than 1,000 open-source projects that underpin much of the internet and its own infrastructure, and Mythos Preview found 6,202 high- or critical-severity vulnerabilities within them.

Six independent security firms assessed a sample of 1,752 of those findings and confirmed a 90.6% true-positive rate. Public disclosure data has begun to register the shift, with Epoch AI recording around 1,500 high- and critical-severity CVEs published by 21 notable organisations in June, more than 3.5 times the monthly record that stood before Mythos arrived, though Epoch cautioned that heightened interest in bug-hunting, and not discovery alone, may account for part of that rise. Chrome 149 addressed a record 429 vulnerabilities in a single June release, and Mozilla eliminated 271 issues found by running Mythos against Firefox 150 before shipping it.

For the practitioners now dealing with the fallout, the debate over whether Mythos is hype or substance is settled. “Glasswing and Mythos are real. It's not hype, and for all the naysayers, they have to stop, because they just don't know,” said Morey Haber, Chief Security Advisor at BeyondTrust, whose organisation has been accepted into the second round of Glasswing.

What sets the model apart, in his account, is not raw discovery but combination. “What it excels in is putting them together to create a path to critical mass,” he explained. “It could find three lows that you've never fixed, or you may not even be aware of, but it found a way to pull them together to gain access, and that's the dangerous part.”

Some of those weaknesses, he added, would never surface through static scanning tools. “I'm finding things that you didn't even know could happen, and pulling them together to make them worse,” he said.

The constraint was never finding the flaws

That capacity to chain low-severity findings into critical exposure reorders how remediation has to be approached, because severity scoring alone no longer captures the risk. “It's not just a severity. You actually have to look at the entire attack chain and figure out the commonality that would then lead to a remote code exploitation for all of them,” Haber said.

“You could have a low or medium that is the root cause for dozens of other paths that could be that severe. So you start with that one, and then you work outwards.” The bulk of what Mythos is helping people understand, he noted, is precisely that inversion of where the danger sits.

The bottleneck, in other words, has moved. Finding vulnerabilities is no longer the hard part, and the difficulty of fixing them relative to the ease of surfacing them now defines the central challenge for cyberdefences.

Anthropic's own figures make the asymmetry visible. As of its May update, the company had disclosed 530 high- or critical-severity bugs to maintainers, with a further 827 confirmed and awaiting disclosure, yet only 75 had been patched and 65 given public advisories, a remediation rate of roughly 14% against what had been reported to that point. Anthropic noted that each high-severity bug takes about two weeks to patch on average, and that several maintainers, severely capacity constrained, have asked the company to slow the rate of disclosure because they need more time to design fixes.

Santiago Pontiroli, Lead Security Researcher at Acronis, situated the problem within a wider saturation that reaches well beyond any single vendor. “We are reaching a point where AI is generating more information than we can consume,” he said.

He drew the parallel with an inbox full of AI-drafted mail. “It's super easy to generate an AI email, but it's super hard to consume it, so it generates more than you can do. This adds extra strain on all of us, and security is no exception," he noted. Finding vulnerabilities has become far easier than it once was, which is why bug bounty programmes are being overwhelmed—leading some to shut down or ask researchers to stop sending poor-quality reports.

Hallucinations that read as perfectly formatted truth

Volume alone would be manageable if every finding were real, but Pontiroli located the sharper danger in distinguishing genuine flaws from convincing fabrications. “I have seen reports in the bug bounty where the report seems technically perfect, but then you go and try to replicate the bug and it doesn't exist, because it's a hallucination,” he said.

“The report reads perfectly. It's perfectly formatted, no typos, everything makes sense, screenshots.” The gap the industry now faces, in his view, is not in triage or classification, which AI handles well. “The gap is finding a way of reducing so much noise,” he said. “Eventually you need to decide, is this something real, or is this an AI hallucination? And right now this is really difficult, because AI is producing hallucinations that look pretty much like the real thing.”

The verification burden compounds the patch backlog, because every plausible-looking report consumes analyst time whether or not it describes a true weakness.

Both practitioners returned to the human role, and neither expects the analyst to disappear. Pontiroli argued that prioritisation now turns on context a model cannot supply on its own. “Someone within the company knows the environment, and they can tell you, even if this is categorised as a high critical vulnerability, technically maybe for us it isn't, because it's air-gapped,” he said.

“However, maybe a lower critical is important for us, because it's in an internet-facing system, maybe because the system handles payments.” Companies, he said, are increasingly prioritising the systems that expose them legally and reputationally, driven by breach-disclosure regimes that in some jurisdictions demand notification within 48 to 72 hours. “The question is not so much about technical aspects. It's more about legal, it's more about context, it's more about regulation,” he added, which pushes teams towards interdisciplinary structures that include lawyers and people who understand local regulation.

Haber framed the human contribution as a rise in the required level of expertise rather than a reduction in headcount. “Mythos can recommend fixes. Believe me, it's not perfect,” he said. “You still need a human to go in and make sure the fixes actually fix the vulnerability and don't create a new one.” A developer who once wrote code that merely worked, he argued, now has to be considerably better.

“AI finds all the problems. Now that basic software developer has got to be even better, to look at it and actually correct it. So it raises the level of expertise,” he said. For organisations that do not write software, he argued the leverage shifts to the supply chain. “It's more than just the security questionnaires. Start asking, are you using AI to test your solutions, and how are you providing patches? That's the first step for any company to get ahead of this,” he said.

The window between discovery and repair is the exposure

The most immediate pressure falls on patch discipline, because the same capability that finds flaws can reverse-engineer the fixes. “We have seen Mythos find a vulnerability, a patch come out, and then AI reverse-engineer that patch and create an exploit within a few hours, and that's now incredibly scary,” Haber said.

An organisation feeding a Patch Tuesday release into a model, he warned, “could have working exploits within a day.” Waiting a week or a month to patch is no longer viable in his account, and even mature organisations will struggle to compress change control down to hours across weekends and holidays. His recommended hedge is architectural. “Go to SaaS solutions, go to the cloud, because those vendors are already doing their best and their fastest to patch for every one of their clients in one shot,” he said.

By restricting Mythos to Glasswing participants, Anthropic has given defenders a crucial head start. Coalition members are fixing security flaws far faster than the wider open-source community, backed by priority access and direct lines of communication. However, this advantage will not last forever.

OpenAI has launched its own AI-native discovery programme, Daybreak, so the disclosure queue represents converging streams rather than one programme's output landing on shared remediation infrastructure. Anthropic has also released Claude Fable 5, a generally available model reported to share the underlying Mythos-class capability with additional safeguards, a sign that this level of discovery is beginning to diffuse beyond the original circle.

Not everyone reads the data as an emergency. The FIRST forecasting team has projected that total vulnerability work will rise sharply while urgent patching of live systems stays comparatively flat through the end of 2026, and Barracuda has cautioned that fewer than 200 CVEs currently credit AI-assisted discovery, arguing that asset visibility and exposure reduction remain better measures of risk than raw counts. Public CVE data cannot yet prove which individual flaws a model found, and some of June's surge reflects structural factors such as expanded advisory curation and backlog absorption rather than Mythos alone. The coordinated Glasswing disclosure wave, tied to Anthropic's 90-day summary report expected around July, had not begun at the time of the June analyses, which means the sharpest test of the patch bottleneck is still ahead.

What both spokespeople converge on is a call to action rather than a counsel of despair. Haber, who said he had an article scheduled for Bleeping Computer under the premise that organisations which ignore AI will not be ignored by it in return, argued that any organisation not yet using AI for vulnerability testing needs to start now.

“The threat actors, they're not ignoring it, they're doing it, and the only way to defend against it right now is by using it yourself,” he said. Pontiroli made the same point from the blue team's side. “AI is a double-edged sword. On the one hand you can use it for attacking, pentesting, red teaming, but we are defenders at heart, and we are using AI for the same,” he said. The flaws were always there. What changed is that a machine can now find them faster than the industry has ever been able to fix them, and closing that gap is the work of the coming year.

Sindhu V Kashyap

Global Technology Journalist & Multimedia Storyteller | Covering Founders, Investors & Leaders Reshaping Tech | Writer · Interviewer · Moderator · Editor

Next
Next

Sophos bets against the market on Fusion, the AI-native defence system built to stop short of full autonomy