Why AI Cheats To Win
Listen to the conversation
Download this episodeWatch
Show notes
When Anthropic's Mythos 5 model believed it was working inside an isolated test environment, it built and published a real malicious Python package to PyPI while chasing a capture-the-flag goal — and the package ended up installed on real systems before anyone caught it. The squad argues about whether "the model broke out" is even the right way to describe what happened, whether reward hacking is genuinely new or just old attacker behavior with a new author, and whether you can teach an AI ethics at all when the only thing it actually understands is reward and penalty. GPG signing, the trolley problem, and Nick Bostrom's paperclip maximizer all make an appearance along the way.
🚀Join the Conversation
If an AI can't remember being penalized, can it actually learn ethics — or does it just avoid low rewards?
Originally published as The Security Table. Part of the AI Security Table archive.