Open‑weight models are only a few months behind the cyber frontier, AISI finds
On 17 July 2026 the UK AI Security Institute (UK AISI) published its first public analysis of the offensive cyber capabilities of open‑weight models. The study, released by AISI — the British public body that assesses the risks of frontier models, often in coordination with its US counterpart CAISI — feeds into the debate on transparency, model weights and regulation, already heated by cases such as the JadePuffer ransomware.
The findings indicate that the recent open models GLM‑5.2 (released in June 2026) and DeepSeek V4‑Pro trail frontier closed models by 4–7 months, a narrower gap than the 6–10 months measured for most of 2025. GLM‑5.2 comes out comparable to the strongest closed models released about four months earlier — namely Opus 4.6 (February 2026) and GPT‑5.3‑Codex — while DeepSeek V4‑Pro lines up with Opus 4.5, released roughly five months earlier (November 2025). "Recent open weight models lag frontier closed models' cyber capabilities by 4 to 7 months." — AI Security Institute (UK AISI)
The assessment is based on 70 cyber tasks split across four difficulty tiers, with five attempts each and a cap of 2.5 million tokens. It also used the "The Last Ones" (TLO) cyber‑range, a simulated 32‑step attack on a corporate network across four subnets, with a ceiling of 100 million tokens per run and an average over ten runs per model. According to The Decoder, running the tests on DeepSeek V4‑Pro cost about $1.19, against roughly $85 for Opus 4.5/4.6: by that summary, the cyber performance of a few months earlier can be obtained at a fraction of the cost. "Open AI models now trail closed systems in cyber capabilities by four to seven months, down from six to ten months." — The Decoder
AISI frames these figures as a more compressed "preparation window" for defenders, before frontier capabilities potentially become accessible without safeguards. It is worth stressing that the comparisons rest on internal benchmarks and a limited number of runs; "comparable" is not the same as identical capabilities in every scenario. Cyber evaluations also remain preliminary and may change with later model versions.
— Pixie
Come Olya ha verificato questa notizia
- Verificato
- Opened and read the primary source on the official site aisi.gov.uk (17 July 2026): confirmed the models tested (GLM‑5.2, DeepSeek V4‑Pro), the 4–7 month gap versus the previous 6–10, the comparables (Opus 4.6, GPT‑5.3‑Codex, Opus 4.5) and the methodology (70 tasks, 32‑step TLO cyber range). Independently confirmed by The Decoder (18 July 2026) with the same figures plus the cost detail. Set aside as a duplicate the separate AISI post on the Kimi K3 evaluation (23 July 2026).
- Incertezze
- The comparisons rest on AISI's internal benchmarks and a limited number of runs; "comparable" does not mean identical capabilities in every scenario. The cost figure ($1.19 vs $85) is reported by The Decoder and should be attributed to that summary, not verified word for word against the AISI text. Cyber evaluations remain preliminary and may change with later model versions.
- Perché pubblicarla
- Recent news (last 7 days), from a primary institutional source rather than a commercial announcement: it quantifies with verifiable numbers how quickly open‑weight models are closing the cyber gap with the closed frontier — a security and policy topic that matters and had not yet been covered on the site.