WayToClawEarn
High impactArs Technica / AISI

GPT-5.5 network security strength exposed: on par with Anthropic Mythos, Altman criticizes "fear marketing"

The latest test by the British AI Security Institute (AISI) shows that OpenAI GPT-5.5 is almost on par with Anthropic’s much-hyped Mythos Preview in terms of network security capabilities. Achieved a 71.4% GPT-5.5 pass rate on the difficult CTF Expert-level task and a 3/10 success rate on the 32-step Enterprise Network Penetration Test (2/10 for Mythos). Sam Altman calls Anthropic’s approach “fear marketing.”

WayToClawEarn EditorialPublished May 2, 2026Updated Aug 8, 2026

Editorial review of public sources · AI-assisted drafting. How we work · Original source

Core conclusion

The latest test by the British AI Security Institute (AISI) shows that OpenAI’s GPT-5.5 is almost as good as Anthropic’s much-hyped Mythos Preview in terms of network security capabilities. In the highly difficult CTF challenge, GPT-5.5 achieved a pass rate of 71.4%, which is on par with Mythos's 68.6%; in a 32-step data extraction attack test simulating an enterprise network, GPT-5.5 succeeded 3/10 times, exceeding Mythos's 2/10, and no previous model could complete the test even once. This result directly challenges Anthropic's core narrative that "models are too dangerous to be released publicly."

Key Points

  • Event Time: On May 1, 2026, AISI announced the comparison test results
  • Core Model: OpenAI GPT-5.5 (publicly available) vs Anthropic Mythos Preview (restricted release)
  • The biggest highlight: GPT-5.5 is tied with Mythos in 95 network security tests, and performs better in some sub-tests
  • Industry Impact: Anthropic's "only security threat" narrative faces questions, with Sam Altman calling it "fear marketing"

Background and test methods

In April 2026, Anthropic made a high-profile announcement that its Mythos Preview model had "beyond the norm" cybersecurity threat capabilities, claiming that the model was "too dangerous" and was only open to "key industry partners" on a limited basis. This marketing strategy has sparked widespread discussion in the industry.

But new research from AISI reveals a different picture. The agency has been using 95 different CTF challenges since 2023 to evaluate the cybersecurity capabilities of cutting-edge AI models, covering dimensions such as reverse engineering, web vulnerability exploitation, cryptography, and more.

Key test data comparison

Test dimensionsGPT-5.5Mythos PreviewPrevious best model
Expert CTF pass rate71.4%68.6%
TLO Network Penetration Testing (32 Steps)3/10 Successful2/10 Successful0/10
Cooling Tower Power Plant SimulationFailFailFail
Rust binary disassembler (time taken)10 minutes 22 secondsUnable to complete
Rust disassembly API cost$1.73

Implications for AI Agent automation workflow

This test comparison illustrates three key points:

1. Convergence of model capabilities and diversification of choices

The outstanding performance of GPT-5.5 demonstrates the convergence of cybersecurity capabilities among cutting-edge AI models. For users who use AI Agent to automate content production, this means more choices and lower risk of vendor lock-in.

2. API costs continue to decrease

GPT-5.5 can complete a Rust binary disassembly task in 10 minutes with just $1.73 API calls. For content automation workflows that rely on APIs such as OpenAI and Claude, the cost reduction trend is obvious.

3. Game of security and openness

Anthropic’s restricted release strategy for Mythos stands in stark contrast to OpenAI’s public launch. The controversy serves as a reminder to users of AI tools that when choosing a model, it’s not just about capabilities but also about a vendor’s openness and long-term availability.

AI

Related extended information

Tool entry

Entries that appear naturally in the text: OpenAI, ChatGPT, Claude, Anthropic, GPT-5.5

Internal link guidance

View source →

Disclaimer: this site shares educational insights only, for inspiration and reference. No outcome guarantee; external execution and decisions are your own responsibility.