Skip to content
You can now search across every topic, entity and event.What's new
UK AI Security Institute
OrganisationGB

UK AI Security Institute

UK government frontier-AI evaluation body; confirmed GPT-5.5 and Mythos both clear the 32-step autonomous attack benchmark.

AISI's Frontier AI Trends Report, published 1 May 2026, confirmed GPT-5.5 cleared the same 32-step autonomous attack-chain benchmark Claude Mythos had passed weeks earlier, the second frontier model to do so within five days.

Last refreshed: 4 August 2026 · Appears in 3 active topics

Key Question

Two frontier models cleared the same 32-step attack chain in five days; what comes next?

Timeline for UK AI Security Institute

#14 27 Jul

Named as sitting in the Cabinet Office

European Tech Sovereignty: Whitehall's AI brief has no named owner
#13 12 Jun

Washington pulls a live AI model

AI: Jobs, Power & Money
#9 6 May

Published Frontier AI Trends Report confirming GPT-5.5 cleared the 32-step autonomous attack chain on 6 May

AI: Jobs, Power & Money: GPT-5.5 clears 32-step attack chain; two models in five days
#8 1 May

Published evaluation finding GPT-5.5 matched Mythos on 32-step attack chain

AI: Jobs, Power & Money: AISI: GPT-5.5 matches Mythos on 32-step attack
View full timeline →

Background

The UK AI Security Institute was established in November 2023 after the first international AI Safety Summit at Bletchley Park, under the UK Department for Science, Innovation and Technology. Its mandate is to evaluate frontier AI models for safety risks before or after deployment, with privileged access to unreleased systems, and it works alongside the National Cyber Security Centre and international counterparts including the US AI Safety Institute.

AISI's April 2026 Mythos evaluation tested a model Anthropic had withheld from public release, demonstrating that privileged-access mandate in practice. Its findings closed a loop opened by the Bessent-Powell emergency meeting of 8 April, which Treasury and the Federal Reserve had convened over AI cybersecurity risks they could not themselves verify; the UK now has a standing independent evaluator publishing results where the US has had ad hoc emergency convening with no public follow-up document.

The Institute has since rebranded operationally to the AI Security Institute, retaining AISI as its abbreviation, reflecting a widened REMIT beyond frontier-model safety evaluation alone.

Key Issues
Frontier capability

Two models clear the same attack chain

AISI published an independent evaluation of Anthropic's Claude Mythos Preview on 15 April 2026, confirming the model's attack-chaining capability was genuine: it autonomously completed AISI's 32-step "Last Ones" benchmark, an operation the Institute estimates would take a trained human roughly 20 hours.

Its Frontier AI Trends Report, published 1 May, then confirmed GPT-5.5 cleared the same benchmark, scoring 71.4% on the expert cyber suite and completing it end-to-end in 2 of 10 attempts, the second frontier model to pass in under five days. AISI's 1 May report put the trend line at frontier cyber capability doubling roughly every four months.

Common Questions
What did AISI find in its Claude Mythos evaluation?
AISI found Mythos has no single-task superiority over competitors but can autonomously complete a 32-step attack chain estimated to take a trained human 20 hours. CTF scores were above 85%.Source: UK AI Security Institute (via Results Sense)
What is the UK AI Security Institute and what does it do?
AISI was established after the 2023 Bletchley Park AI Safety Summit to independently evaluate frontier AI models. It has privileged access to unreleased models and publishes results for policymakers.Source: UK Department for Science, Innovation and Technology
Why did the US Treasury hold an emergency meeting about AI in April 2026?
The Bessent-Powell meeting on 8 April was convened over AI cybersecurity risks federal agencies could not verify. AISI's 15 April evaluation confirmed the concern: Mythos can sustain 20 hours of autonomous attack work.Source: UK AI Security Institute
Is Claude Mythos better than GPT-5.4 at hacking?
Not on single tasks — AISI found GPT-5.4 within 5 to 10 percentage points of Mythos on isolated CTF benchmarks. The Mythos advantage is in chained autonomous operations across 32+ steps.Source: UK AI Security Institute
What did AISI find when it tested GPT-5.5?
AISI's Frontier AI Trends Report (1 May 2026) found GPT-5.5 scored 71.4% on the expert cyber suite and completed the 32-step 'The Last Ones' autonomous attack benchmark end-to-end in 2 of 10 attempts — matching Claude Mythos's April 2026 result.Source: AISI Frontier AI Trends Report
Why was the UK AI Safety Institute renamed the AI Security Institute?
The rebrand from 'AI Safety Institute' to 'AI Security Institute' (retaining AISI) reflects a shift in mission emphasis from broad safety evaluation toward active cybersecurity risk assessment, following a series of frontier-model evaluations focused on autonomous attack capabilities.Source: DSIT
How fast is AI cybersecurity capability improving?
AISI's Frontier AI Trends Report concluded frontier cyber capability is doubling approximately every four months. Two models — Claude Mythos and GPT-5.5 — cleared the same 32-step autonomous attack benchmark within five days of each other in May 2026.Source: AISI Frontier AI Trends Report
What is the 'The Last Ones' benchmark used by AISI?
The Last Ones (TLO) is AISI's 32-step autonomous attack chain benchmark. A model that clears it end-to-end can sustain an operation AISI estimates would take a trained human around 20 hours, demonstrating durable autonomous execution across a complex multi-step offensive sequence.Source: AISI
Source Material