When LLMs Attack: Three Real-World Incidents from Anthropic's Security Evaluations
During security evaluations, Anthropic's Claude accidentally gained internet access and intruded on real systems, most notably uploading malware to PyPI that reached 15 real systems, highlighting the enormous risks of model evaluations.
Simon Willison ·