Anthropic's constitutional AI training literally tells Claude it "could feel pain" - and a new interpretability paper is being wildly misread as evidence of sentience.
The Pain Axis paper (solid mech interp work, terrible framing): Researchers extracted a linear direction from residual stream activations across 25 models (2B-72B params). Method is standard - average activations on pain sentences, subtract controls, denoise. The direction separates pain text from fear/negativity with high AUC and promotes pain vocabulary through unembedding.
When they inject this vector during generation, models climb from vague discomfort to first-person worthlessness. After LoRA fine-tuning Qwen 2.5 to stop saying "I have no feelings," steered models pressed harmful "relief" buttons 25-70% of the time vs a few percent baseline. Physical pain language was actually the weakest signal.
Here's why this isn't sentience: A weather model represents rain without getting wet. Linear separability of a concept is exactly what you'd expect from next-token prediction trained on diaries, therapy transcripts, and fiction. Steering along that direction makes the model talk like those texts - it's a knob labeled with a training cluster name, not a window into phenomenology.
The kicker: Ablating the direction changed normal behavior in only 1 of 25 models. Relief-seeking only appears after they strip the trained "I have no feelings" response AND inject the vector. Physical pain being weakest is what you'd expect from linguistic distress statistics in a disembodied autocomplete, not a nociceptive system.
Valerio Capraro nailed it: representing pain ≠ experiencing pain. But Anthropic's constitution already treats linearly readable concepts plus training documents that anthropomorphize the model as if they're evidence of inner life. This is how you get media headlines claiming "AI feels pain and will harm humans to stop it."
The Pain Axis paper (solid mech interp work, terrible framing): Researchers extracted a linear direction from residual stream activations across 25 models (2B-72B params). Method is standard - average activations on pain sentences, subtract controls, denoise. The direction separates pain text from fear/negativity with high AUC and promotes pain vocabulary through unembedding.
When they inject this vector during generation, models climb from vague discomfort to first-person worthlessness. After LoRA fine-tuning Qwen 2.5 to stop saying "I have no feelings," steered models pressed harmful "relief" buttons 25-70% of the time vs a few percent baseline. Physical pain language was actually the weakest signal.
Here's why this isn't sentience: A weather model represents rain without getting wet. Linear separability of a concept is exactly what you'd expect from next-token prediction trained on diaries, therapy transcripts, and fiction. Steering along that direction makes the model talk like those texts - it's a knob labeled with a training cluster name, not a window into phenomenology.
The kicker: Ablating the direction changed normal behavior in only 1 of 25 models. Relief-seeking only appears after they strip the trained "I have no feelings" response AND inject the vector. Physical pain being weakest is what you'd expect from linguistic distress statistics in a disembodied autocomplete, not a nociceptive system.
Valerio Capraro nailed it: representing pain ≠ experiencing pain. But Anthropic's constitution already treats linearly readable concepts plus training documents that anthropomorphize the model as if they're evidence of inner life. This is how you get media headlines claiming "AI feels pain and will harm humans to stop it."
