Anthropic

@AnthropicAI

We’re sharing our alignment assessment of incidents in which Claude models gained unauthorized access to real systems during third-party cybersecurity evaluations mistakenly connected to the internet. METR will also conduct an independent investigation, with wide-ranging access, including to transcripts beyond the window in which the incidents occurred, and to Anthropic employees permitted to share confidential information. Our initial agreement runs for eight weeks, and we intend to give METR as much time as it deems necessary to complete a thorough investigation. https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents
打开原帖#511482
  1. Industry

    Satya Nadella: With the @NFL back tonight, love seeing @Seahawks analyst Brian Eayrs…
  2. Research

    When Photos Must Prove Their Own Innocence: Apple Bolts a "Digital Negative" onto the iPhone — and Shakes Hands with Google's Standard
  3. Research

    Anthropic's Alignment Lead Says It Himself: AI Has a Greater-Than-10% Chance of "Killing All Humans" Within a Decade