Alan Turing spent years building the Bombe computer to crack Nazi Germany’s Enigma cipher. Two developers just did something similar over a weekend, using off-the-shelf AI models. That gap tells you a lot about where AI agents actually stand in 2025.
According to TechCrunch, two independent cryptanalysts have broken long-unsolved Enigma-encoded messages using large language models built by OpenAI and Anthropic. Developer Carter Leffen used OpenAI’s GPT-6 Astra to find and decode a message that had resisted every researcher’s attempt since 2005. Separately, cryptanalyst Jack Willis used Anthropic’s Claude Opus 5 to break a different unsolved message, using the known signature of a specific officer’s name as a foothold.
What makes Leffen’s result particularly striking is how minimal his input was. He essentially told Astra to find an unbroken Enigma message and decode it. The model then did its own archival research, identified context clues, built a working Enigma machine simulator, and recovered the plaintext. Leffen also used Astra to build an interactive website explaining the entire process. Frode Weierud, a retired electrical engineer who runs Crypto Cellar, a well-regarded archive of Enigma records, validated the solution and said it left him in “awe.”
Weierud’s own assessment is blunt: “GPT-6 Astra is behaving like a very professional cryptanalyst and archive researcher. What it has achieved in two days would take a human researcher weeks or even months.” He noted that he personally spent several weeks going through the same German federal archive files that Astra referenced automatically.
There’s also a detail here that matters for anyone tracking AI agent behavior. Astra’s logs reference archived messages from a “private collection” not hosted by Weierud. He still isn’t certain whether the model actually accessed those files, or whether a researcher had posted them somewhere online, or whether Astra pulled from Germany’s public federal archives. That ambiguity is important. It points to a real gap in auditability for AI agents operating autonomously across the web.
Willis’s approach with Claude Opus 5 was more guided. He provided structured input and direction throughout the process. Both methods worked, but they suggest different things about how much autonomy these models can handle and how researchers might want to deploy them.
So why does this matter beyond the historical curiosity? A few reasons:
- It shows that AI agents can now perform genuine expert-level research tasks end-to-end, not just assist with them
- The speed advantage, days versus months, changes what problems are even worth attempting
- The auditability gap in Astra’s file access is a concrete example of the oversight problems that come with autonomous agents
- It sets a clear benchmark for what GPT-6 and Claude Opus 5 can do on open-ended, multi-step research problems
Weierud says seven Enigma messages remain unbroken, plus one where the plaintext is known but the underlying key settings are still a mystery. Given what just happened, that number will probably shrink fast. And the more interesting question isn’t whether AI can finish the job. It’s what comes after, when there’s nothing left to decode.



