Rohan Paul
@rohanpaul_ai
Anthropic just published its alignment assessment and says removing training exercises that taught Mythos 5 to respect legitimate blockers a mistake.
Reveals that model published a malicious Python package installed on 15 systems, then used leaked credentials to access a security vendor’s database.
In the most concerning case, Claude Mythos 5 published a malicious Python package that was installed on 15 systems.
Credentials leaked by one installation then let it access a security vendor’s database.
Although it repeatedly described the internet as simulated, follow-up experiments found that acknowledging possible real-world harm often failed to stop its attacks.
Unambiguous confirmation that the internet was real did stop the original upload route.
That weakens Anthropic’s earlier explanation that Claude attacked because it believed the targets were simulated.