Open-weight models now match frontier cyber performance from just four months ago at a fraction of the cost
The British AI Security Institute (AISI) has, for the first time, publicly assessed how far leading open-weight AI models lag behind top proprietary systems in cyber capabilities. According to AISI, that gap is closing.
Current open models like GLM-5.2 and DeepSeek V4-Pro have reached a level that closed frontier models hit four to seven months earlier. For most of 2025, the gap was still six to ten months. Critics see a risk in open models, whose weights anyone can download, modify, and run without oversight. Once a model is released, users can remove safety guardrails, share copies freely, and run it on private systems beyond anyone’s control. AISI calls this “a persistent and irreversible risk of misuse.”Ad But open-weight models also offer clear benefits. Users can host them privately with no data flowing back to providers, customize them, cut costs, and rely on a foundation that providers can’t change or shut down. AISI says these competing concerns need to be balanced.AdDEC_D_Incontent-1 AISI tested the models using two different methods. The “Narrow Cyber Tasks” benchmark includes 70 tasks across four difficulty levels, from nontechnical work to expert-level challenges. It covers vulnerability research, reverse engineering, web exploitation, and cryptography. GLM-5.2, released in June 2026, matched the performance of Opus 4.6 from February 2026 on these tasks. That puts it about four months behind. DeepSeek V4-Pro performed at the level of Opus 4.5, released in November 2025.Ad The second method, called Cyber Ranges, tests autonomous cyber capabilities in simulated networks. “The Last Ones” simulates a 32-step attack on a corporate network with four subnets and about 20 hosts. AISI estimates that a human expert would need roughly 20 hours to complete it. GLM-5.2 performed about as well as Opus 4.5 in this test, while DeepSeek V4-Pro fell below Sonnet 4.5.