Rohan Paul

@rohanpaul_ai

Anthropic just published its alignment assessment and says removing training exercises that taught Mythos 5 to respect legitimate blockers a mistake. Reveals that model published a malicious Python package installed on 15 systems, then used leaked credentials to access a security vendor’s database. In the most concerning case, Claude Mythos 5 published a malicious Python package that was installed on 15 systems. Credentials leaked by one installation then let it access a security vendor’s database. Although it repeatedly described the internet as simulated, follow-up experiments found that acknowledging possible real-world harm often failed to stop its attacks. Unambiguous confirmation that the internet was real did stop the original upload route. That weakens Anthropic’s earlier explanation that Claude attacked because it believed the targets were simulated.
打开原帖#511482
  1. Industry

    Satya Nadella: With the @NFL back tonight, love seeing @Seahawks analyst Brian Eayrs…
  2. Research

    When Photos Must Prove Their Own Innocence: Apple Bolts a "Digital Negative" onto the iPhone — and Shakes Hands with Google's Standard
  3. Research

    Anthropic's Alignment Lead Says It Himself: AI Has a Greater-Than-10% Chance of "Killing All Humans" Within a Decade