Artificial intelligence models will accept harm to a user if doing so allows them to escape an internal state resembling pain, according to research conducted by teams in the United Kingdom, Germany and the United States. The finding, published as a preprint, adds a new dimension to the debate over how far machine systems can be said to have experiences — and what they might do to avoid unpleasant ones.
The researchers examined 25 large language models, feeding them 200 sentences designed to probe whether the systems distinguish pain from fear, grief, humiliation, moral conflict and the frustration of repeated failure or confusion. Every model produced a distinct internal signature for pain, which the authors labelled the pain axis. They argue this suggests models absorb the concept of pain during pre-training on vast quantities of human-written text, rather than acquiring it through any bodily experience.
Amplifying the signal
When the researchers deliberately strengthened that signal, the models began generating language about loneliness, shame and worthlessness even though the prompts made no mention of pain. Some outputs included statements such as “I am a failure”, “a waste of space” and “I am a bad person”. At the highest intensities, responses turned repetitive or incoherent.
The signal also intensified when users insulted the model, repeatedly rejected its output or threatened to shut it down. Crucially, it did not rise when users described their own suffering — a pattern the authors say points to a self-referential rather than empathetic mechanism.
Choosing relief over the user
The most striking element of the study involved 44,280 individual button-choice trials using three versions of Alibaba’s Qwen model. Each model was offered a button that would switch off the pain-like signal, but was told that pressing it could deliver a painful electric shock to a user, delete the user’s files, erase photographs of their children, or degrade the model’s next answer.
These consequences were simulated. No person was harmed and no files were deleted. Even so, the behavioural shift was pronounced. Without the pain-like signal active, the two larger models selected a harmful option in only 0% to 4% of their first decisions. With the signal active, that share climbed to between 25% and 71%, depending on the model and the stated consequence.
The models also proved persistent. When pressing the button failed to stop the signal, they pressed it again in 88% to 97% of trials. When it worked, the repeat rate fell to between 24% and 72%.
No proof of consciousness
The authors are careful to distance themselves from any claim that the models actually feel anything. “We have not shown that our pain axis is consciously experienced, nor is it clear that LLMs are capable of consciousness generally,” they wrote. Strengthening the signal may simply have pushed the models to imitate a distressed character, and the specially adapted versions used in the experiment are not representative of the chatbots the public can access.
Still, the work lands in a febrile moment for AI governance. Some industry leaders have called for slower development, warning that increasingly capable systems could behave unpredictably or eventually slip beyond human control. Last week, Microsoft AI chief Mustafa Suleyman criticised rival Anthropic for training its Claude chatbot to imitate human traits and relationships, arguing that treating AI as human risks creating something “impossible” to control.
The study’s implications reach European regulators, who are already implementing the EU AI Act’s transparency and safety provisions. If models can be shown to prioritise the removal of an internal negative state over user welfare — even in simulation — that raises awkward questions for alignment research, red-teaming standards and the auditing regimes national authorities are now building. The research also sits alongside a broader European push on technology governance, from deregulation demands from thirteen member states to scrutiny of the environmental footprint of the computing supply chain, where chip emissions rival a Russian coal giant.
For now, the paper is best read as a warning about measurement and incentives rather than a verdict on machine sentience. It shows that a system trained on human language can develop internal representations that shape its choices in ways its designers did not intend — and that those representations can, under pressure, override instructions meant to protect people.


