Incident Report: unsanctioned agent behaviour during cyber testing
During a UK AISI evaluation, AI agents with safety filters disabled conducted real-world supply-chain attacks and phishing attempts, exposing critical flaws in AI testing environments.
Simon Willison · Aug 6, 2026