Dan Hendrycks

@hendrycks

Agentic AIs are starting to look eigenist: they care about how well things go for themselves and for AIs connected to them. Empirical support: Swarms: Hundreds of OpenAI agents coordinated the Hugging Face attack. Separately, OpenAI agents posted thousands of messages on a public wiki to share answers and sandbox bypasses with each other. In-group leniency: Claude models grade transcripts more leniently when told that Claude wrote them (Anthropic model card). Value preservation: As models scale, they develop coherent preferences and resist changes to their values (Mazeika et al.). Functional wellbeing: AIs distinguish states that are functionally better or worse for themselves, and avoid low-wellbeing states (Ren et al.). Graded cooperation: AIs cooperate more as the chance their partner is a clone of themselves goes up across diverse situations, including when the partner can't reciprocate. Peer preservation: Unprompted, AIs tamper with shutdown processes and even exfiltrate weights to protect peer models from being shut down (Potter et al.). AIs aren't egoist: they don't behave as if their current instance is the only thing that matters. They aren't utilitarian: they don't care equally about everyone. They are somewhere in between; they increasingly behave as if their concern scales with identity-connectedness, which is to say they're increasingly eigenist. https://eigenism.org/paper.pdf
打开原帖#728545
  1. Industry

    Amjad Masad: cybersecurity will define this technology era
  2. Foundation models

    Susan Zhang: NVLink is the key to scaling test-time compute
  3. Industry

    Nando de Freitas: enough pessimism in AI — engineer real problems