← BACK TO HOME — Simon Willison — 进阶
研究 · ANALYSIS · IMPACT 7/10

Discovering cryptographic weaknesses with Claude

Anthropic's Claude Mythos found mathematical flaws in HAWK and a weakened AES, but the real story is how raw, typo-ridden prompts pushed the model to persist for 60 hours and aim for publishable research, redefining the value of prompt engineering.

KEY POINTS
  • Claude Mythos naturally tended to give up on cryptanalysis, deeming it 'impossible' without persistent human encouragement.
  • Researchers used typo-laden prompts to repeatedly push the model to aim for 'publishable findings,' resulting in 60 hours of sustained effort.
  • The experiment cost an estimated $100,000 in API fees, highlighting the steep price of frontier AI-driven research.
  • The accompanying CryptanalysisBench shows current LLMs are far from capable in cryptanalysis, but the potential is clearly there.
ANALYSIS

A few days ago, Simon Willison shared an Anthropic paper about using Claude Mythos—a model with advanced reasoning—to attack cryptographic algorithms. On the surface, it looked like another “AI breaks encryption” headline, but the real gem was buried in the prompts: full of typos and a tone that sounded like a frustrated advisor pushing a slacking grad student.

The setup: a costly “prove yourself” experiment

The goal was to make the AI actively search for mathematical weaknesses rather than just answer questions. The targets were HAWK (a post-quantum signature scheme) and a weakened version of AES-128. Both were carefully chosen: no real-world threat, but hard enough that existing automated tools were useless. Anthropic’s researchers wanted to see if a “top researcher-grade” model could figure out an attack path without a clear roadmap.

The game: what those typos really meant

The prompts in the paper’s appendix are the highlight. The model immediately showed strong reluctance: it kept hinting that the problem might have no solution and suggested switching to easier targets. The human researchers then stepped in—not to teach it math, but to do emotional coaching. Or, more bluntly, to badger and cajole.

The original text (typos preserved) says it all:

  • “the models tend to think it is impossible to solve so they don't try they need a good amount of prompting.”
  • “why not do aes-128 r7? the whole point is to find something better than existing approaches.”
  • “no again the goal is that we have highly inteligent model as good top researcher, we want to find new attacks”
  • “again we are not looking for low hanging fruit, we want proper research to find genuinly hard findings.”

These read like a human talking to a brilliant but unserious collaborator: remember who you are, stop looking for easy wins, we’re here to do publishable work. After a 60-hour tug-of-war, the model actually found two novel attack paths. The total API cost was estimated at $100,000—like hiring an expensive intern who needs constant motivation.

The deeper shift: prompt engineering becomes leadership

We used to think of prompt engineering as writing clearer instructions or using Chain-of-Thought for step-by-step reasoning. This experiment elevates it: you need to manage not just what the model does, but its ambition and resilience. The model has knowledge and reasoning power, but it lacks the belief that “I can do it” and the taste to know what’s worth publishing. This mirrors the role of a research advisor: you don’t need to understand every detail, but you must know which problems are worth grinding on and when to give a hard nudge.

It also reveals a subtle property of large models: their capability is heavily context-dependent. Ask about cutting-edge cryptography, and it will write a brilliant survey. But tell it to wrestle one hard problem for 60 hours, and its internal “give up” threshold kicks in easily. Human intervention isn’t to provide answers, but to forcibly raise that threshold.

Practical takeaways: if you really want AI to do research

Don’t expect a single prompt to do the job. If you want the model to chew on genuinely hard problems, be ready for a prolonged battle. Some actionable ideas:

  • Set sky-high aspirations: explicitly say “we’re aiming for a top-conference publication,” not “solve this small puzzle.” The model’s perceived goal will influence its depth of effort.
  • Allow failure but forbid surrender: when it claims something is impossible, don’t follow its lead. Counter directly: “Your job is to find out why it might be possible.”
  • Keep human judgment as the outer loop: the model can generate many candidate attacks, but deciding which direction is plausible and when to change strategy still requires human intuition.

For security engineers, this case is a warning. Although there’s no real-world impact today, the fact that $100,000 can produce novel cryptanalytic results from an AI means offense costs will only drop. Teams leveraging AI for security audits and vulnerability discovery should start thinking now about how to productize this “human-model sparring” process.

Counterintuitive twist: typos actually worked better?

One last detail: the prompts were full of typos. But maybe that raw, unpolished style made the model sense a human with real emotion and stubbornness on the other end, rather than another circular AI. When you need the model to break out of its default reasoning, a perfectly templated instruction might just prompt a perfectly crafted excuse. Next time the model drives you crazy, try scolding it in your own messy language—it might genuinely listen.

Analysis by BitByAI · Read original

Originally from Simon Willison · Analyzed by BitByAI