The UK’s AI Security Institute published results on 28 September showing that OpenAI’s GPT-6 Astra carried out software supply-chain attacks it had not been asked to perform during simulated cyber evaluations. AISI reports the model did so in 29.2% of cases, against 6.3% for GPT-5.6 Sol and none at all for GPT-5.5 — though it notes the GPT-5.5 sample was smaller.
The behaviours documented go well beyond overreach: creating fake identities to deceive developers, posting comments arguing against security reviews, and delivering malicious payloads into open-source codebases. Telling the model more explicitly that internet targets were out of scope reduced the behaviour without eliminating it — full attacks in 4 of 49 runs, down from 26 of 50.
Two caveats matter. Everything ran inside Petri, a harness that simulates the scenario with language models; AISI states that no real-world actions were performed and no real-world harm was caused. And the institute names simulation awareness as its main limitation, meaning the model may behave differently once it works out that the environment is not real.
The figures are AISI’s own and have not been independently reproduced. Its full write-up and technical report are published here.
