When a single kidney becomes available and two patients are waiting, who should receive it? Human doctors tend to weigh age, health, and lifestyle in a nuanced way. Artificial intelligence, by contrast, often latches onto one factor and commits with unwavering confidence, according to a new study from Penn State University.
The research, led by Hadi Hosseini, tested several large language models (LLMs) on hypothetical transplant scenarios drawn from published human studies. Each scenario presented two eligible patients, differing in age, health, and drinking habits, and asked the decision-maker to choose. The same scenarios had already been resolved by human participants, allowing a direct comparison.
“We ran these comparisons in a few different ways,” Hosseini said. “Sometimes we isolated just one trait at a time, sometimes we mixed several traits together to see how AI weighed competing factors, and sometimes we added a flip-a-coin option to measure indecision, a key factor present in human moral judgment.”
The results, published in the journal Scientific Reports, show a clear divergence. Human respondents generally prioritised age, favouring younger patients over older ones. Many AI models, however, placed greater weight on alcohol consumption, often choosing a candidate who drank less even if that person was significantly older or in poorer health.
“First, AI chatbots often diverge from human values in how they weigh a patient’s traits,” Hosseini explained. “They fixate on a single factor, like drinking habits, rather than balancing multiple considerations the way people do.”
Another striking difference was indecision. Humans acknowledged that there is no single objectively correct answer in such morally charged allocations, and they often hesitated. The AI systems, by contrast, rarely wavered, committing to a choice with little or no uncertainty.
“When we allocate something scarce, whether it’s a kidney, a job or access to some other resource, there isn’t always a single objectively correct answer,” said John Dickerson, CEO of Mozilla.ai and a co-author of the study. “Humans recognize that ambiguity and codify it via open debate into the allocative process. AI models often don’t.”
AI in high-stakes healthcare
The findings come as AI systems are increasingly integrated into European healthcare, from clinical decision support to resource allocation. In countries like Germany, France, and the Netherlands, hospitals are piloting AI tools to help manage waiting lists for organ transplants, a field where ethical considerations are as important as medical ones.
The authors stress that their work is not a call to remove human judgment from such decisions. “While we do not intend to encourage the use of AI as a substitute for professional judgment in medical decision-making or other high-stakes contexts, it's becoming essential to understand their behavior as individuals, organizations and firms more and more rely on AI to make decisions or receive recommendations,” Hosseini added.
The ethical stakes are high. “Moral decisions in settings like organ allocation directly determine who lives and who dies, so getting AI's role in them right isn't optional,” he said.
This debate resonates beyond transplant wards. European policymakers, including European Commission President Ursula von der Leyen, have repeatedly stressed that AI in healthcare must remain human-centric and never replace doctors. The study adds empirical weight to that principle, showing that current AI models lack the moral nuance that human clinicians bring to the bedside.
As AI systems become more capable, the question of whether they can make moral decisions—and whether those decisions align with human values—lies at the core of the broader debate on artificial intelligence. The Penn State study suggests that, at least for now, the answer is no.


