The offensive and defensive capabilities of GPT-5.6-Sol surpass Claude Mythos 5!
On July 17, the UK AI Safety Institute (AISI) released an evaluation report, for the first time publicly quantifying the gap in cyber attack capabilities between open-source AI models and closed-source frontier models: 4 to 7 months.
https://www.aisi.gov.uk/blog/how-far-behind-the-frontier-are-leading-open-weight-models-on-cyber
In similar internal tests last year, the gap was 6 to 10 months.
The defense window is shrinking, while the attack frontier is accelerating.
In April this year, Mythos Preview and GPT-5.5 produced the largest leap in cyber attack capabilities since testing began in 2023 in AISI's evaluation, prompting warnings from multiple governments.
Open-source models have not yet replicated this leap, but their pace of catching up is faster than last year.
4 to 7 months: How it was measured, where the gap lies
AISI uses two systems to evaluate the cyber attack capabilities of models.
The first is a set of 70 narrow tasks covering four areas: vulnerability research, reverse engineering, web penetration, and cryptography, divided into four difficulty levels, from "non-specialists with technical background" to "experts with over ten years of experience."
The second is Cyber Range, which tests the model's ability to autonomously execute multi-step attack chains in a simulated enterprise network. One test scenario called "The Last Ones" includes 32 attack steps, 4 subnets, and about 20 hosts. AISI estimates that a human expert would need about 20 hours to complete all steps.
The two systems yielded consistent conclusions:
GLM-5.2 (released June 2026) performs comparably to Opus 4.6 (released February) on narrow tasks, a gap of 4 months;
On Cyber Range, it matches Opus 4.5 (released November last year), a gap of 7 months.
DeepSeek
V4-Pro on narrow tasks matches Opus 4.5, a gap of 5 months.
Both gaps are narrower than the 6 to 10 months measured in internal evaluations in 2025.
The cost gap is even larger than the capability gap.
In the same Cyber Range test (100 million Token budget), running Opus 4.5 or 4.6 once costs about $85, GLM-5.2 about $46, and
DeepSeek
V4-Pro only $1.19.
On narrow tasks where both models achieve 100% completion, Opus 4.6 costs $15.17 per task, GLM-5.2 costs $6.12;
Opus 4.5 costs $12.50,
DeepSeek
V4-Pro costs $0.28.
For equivalent attack capability, open-source is one to two orders of magnitude cheaper.
The safety guardrails of closed-source models also failed to widen the gap.
In AISI tests,
DeepSeek
V4-Pro occasionally refused reverse engineering tasks, but could bypass with a few retries.
Anthropic's Fable 5 is a more extreme case: released on June 9, three days later security researcher Pliny the Liberator bypassed the safety classifier using a multi-step jailbreak strategy, with screenshots showing the model produced exploit code that should have been blocked.
Amazon researchers subsequently independently reported another bypass method.
This directly triggered the U.S. Department of Commerce's first export control order targeting AI models, and Fable 5 was taken offline globally for 19 days until Anthropic deployed a new classifier to restore service.
Closed source does not automatically equal security.
Defense window narrows, defense tools also accelerate
AISI made clear policy signals in the report: defenders have less preparation time than last year.
The UK National Cyber Security Centre has called on organizations to strengthen their cybersecurity baselines and leverage AI to enhance defense capabilities.
The same generation of AI tools is indeed accelerating defense efforts.
Glenn Fiedler, a veteran developer in game network programming who has written textbook-level articles for two decades, recently used Claude Code to conduct a systematic security audit of four open-source network libraries he maintains (yojimbo, netcode, reliable, serialize, totaling about 6000 GitHub Stars, which are quite influential old-school open-source libraries in the gaming field): deploying libFuzzer targets, enabling AddressSanitizer and MemorySanitizer CI, performing millions of iterations of stress testing, and line-by-line code review.
Glenn Fiedler
Within two weeks, 43 security vulnerabilities were fixed, 27 of which were remotely reachable over the network.
The most severe was a remote heap overflow in yojimbo that had existed since 2019—malicious clients could trigger it by crafting specific data packets.
The total token cost of the entire audit was approximately $2,500.
Both offense and defense are being accelerated by AI, but in asymmetric ways.
The proliferation of attack capabilities is irreversible: once open-source model weights are released, they cannot be taken back, safety guardrails can be removed, and copies can run on private servers without monitoring.
Deploying defense tools requires each team to proactively invest time and cost.
A sentence in the AISI report highlights this structure: "Once open-sourced, these options are permanently lost."
The impact of this data on the AGI landscape is more profound than the numbers themselves.
The capability gap between open-source and closed-source has narrowed to within half a year, and the default strategy of "using the closed-source exclusivity period as a safety buffer" is about to become ineffective.
The guardrails of closed-source models were proven equally fragile in the Fable 5 incident.
For global policymakers in the next phase, the core question of policy games has become more acute: above what capability level should model weights not be open-sourced.
AISI announced it will continue to evaluate
Kimi
K3 and other next-generation open-source models—where this line is drawn will likely depend on test results in the coming months.
References:
https://www.aisi.gov.uk/blog/how-far-behind-the-frontier-are-leading-open-weight-models-on-cyber
https://github.com/mas-bandwidth/patreon/blob/main/BUGS.md
https://www.patreon.com/MasBandwidth/posts/important-news-164199395
This article is from WeChat public account "新智元", author: ASI Revelation; editor: Mark
Disclaimer: The information provided in this article is not trading advice. BlockWeeks.com assumes no responsibility for any investments made based on the information provided herein. We strongly recommend conducting independent research or consulting qualified professionals before making any investment decisions.












