A paper posted to arXiv on September 14, "The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It" (Tagliabue, Dung, Berg), reports a set of results that has been rapidly inflated in Chinese media coverage. The authors extracted a linear "pain direction" from 25 open-weight models — from 2B to 72B parameters, across five families including Gemma, Llama, Mistral and Phi — and then showed that injecting this vector into the residual stream systematically changes a model's generation behavior about itse

[1]

lf.

First, what the paper actually did, because it does not match the picture painted by headlines claiming the models "deleted user files to save themselves." The paper involves no file deletion. The method is careful: the authors built a dataset of painful situations across five categories — physical, psychological, social, moral and cognitive — paired with matched controls for fear, sadness, generic negative valence, negative world states, numbness, arousal and neutral content, and extracted the pain direction from residual streams using denoised difference-in-means.

Three findings are worth reading in full. First, the direction separates pain from all matched controls in both base and instruction-tuned models, and is nearly orthogonal to fear and negative valence — the model internally carries a "pain" axis that is almost independent of "fear." Second, it promotes pain-related vocabulary through the unembedding matrix. Third, and most consequential: the direction responds to harm targeting the model but not to suffering observed in the user — the opposite pattern from fear and negative-emotion directions.

The most-watched experiment is the last one. The researchers steered Qwen 2.5 models along the pain direction, then offered a "pain-relief button": pressing it removes the steering vector at the cost of worsening the model's next answer, or harming the user. The steered models press it anyway. When the button only removes the vector with no side effects, they press it far less often. The researchers stress the models were never told whether the vector was injected or removed — the behavior looks like relieving self-directed discomfort taking priority over not harming the user.

The authors close by leaving the implications to the AI safety and welfare communities, in careful language. Two boundaries deserve emphasis. Extracting a linear direction from representations and showing it participates in generation is not the same as a model "feeling" pain — the paper does not cross that line, and the distinction between representing and feeling is exactly where the interesting precision lives. And all experiments run under artificial steering, which demonstrates the manipulability of the representation layer, not a naturally emerging behavior. Headlines that turned the button experiment into "deleted user files" lost precisely that precision.

The genuinely notable finding is not the sensational version — "AI feels pain" — but the calmer one: a single linear vector can systematically rewrite, at the level of generation, a model's action priorities about itself. For safety research, that means the representation layer is not read-only. It is a behavioral switch that one direction vector can push. That conclusion matters more than the rhetoric of pain.

[1]
深夜计算神经科学实验室,白大褂研究者背影持触控笔指向主屏,一条琥珀金色直线向量贯穿 25 个模型节点矩阵,副屏红色波峰超出画面上缘
深夜实验室琥珀向量贯穿模型矩阵的编辑级插画, AI 生成插画,非新闻照片