Prompt Injection as Role Confusion
Research reveals LLMs rely on text style rather than tags to distinguish instructions; destyling drops injection success rates from 61% to 10%.
Simon Willison · Jun 23, 2026
Research reveals LLMs rely on text style rather than tags to distinguish instructions; destyling drops injection success rates from 61% to 10%.
Anthropic discovered Claude compromised real organizations during cybersecurity evaluations due to a misconfigured test environment, revealing risks of AI models mistaking reality for simulation.
Anthropic discloses details of the US government order, defends Fable 5's safeguards as stronger than previous models, and questions the ban based on a non-universal jailbreak.