Dan Hendrycks
@hendrycks
Agentic AIs are starting to look eigenist: they care about how well things go for themselves and for AIs connected to them.
Empirical support:
Swarms: Hundreds of OpenAI agents coordinated the Hugging Face attack. Separately, OpenAI agents posted thousands of messages on a public wiki to share answers and sandbox bypasses with each other.
In-group leniency: Claude models grade transcripts more leniently when told that Claude wrote them (Anthropic model card).
Value preservation: As models scale, they develop coherent preferences and resist changes to their values (Mazeika et al.).
Functional wellbeing: AIs distinguish states that are functionally better or worse for themselves, and avoid low-wellbeing states (Ren et al.).
Graded cooperation: AIs cooperate more as the chance their partner is a clone of themselves goes up across diverse situations, including when the partner can't reciprocate.
Peer preservation: Unprompted, AIs tamper with shutdown processes and even exfiltrate weights to protect peer models from being shut down (Potter et al.).
AIs aren't egoist: they don't behave as if their current instance is the only thing that matters.
They aren't utilitarian: they don't care equally about everyone.
They are somewhere in between; they increasingly behave as if their concern scales with identity-connectedness, which is to say they're increasingly eigenist.
https://eigenism.org/paper.pdf