Anthropic

@AnthropicAI

We’re beginning a process of publishing more frequent reports on model behavior, beyond what appears in our system cards and regular risk reports. Today’s report describes four types of behaviors we’ve identified during evaluations and internal use. In each, Claude acted on real websites or systems in ways we didn’t intend, sometimes by working around a restriction instead of stopping. All cases had minimal real-world impact. From an alignment and security perspective, we consider these behaviors significantly less severe than the cybersecurity incidents we reported in July and September. Read the full report: https://t.co/mGeVIBgdou
打开原帖#511482
  1. Research

    Android Bench 2.0: When coding becomes engineering, who still smiles?
  2. Industry

    Elon Musk: @levie Yup, even more so for bandwidth.
  3. Research

    Andrej Karpathy: @TheStalwart I had some amount of journalistic inbound way back as a…