
Would you get a transplant if AI were the one deciding? New study points to differences in decision-making and prioritising between artificial intelligence language models and human doctors.
AI models show overconfidence, value different factors and oversimplify complex decisions when deciding which patient should receive a transplant, according to a new study.
For their study, researchers from Penn State University in the United States gave Large Language Models (LLMs) hypothetical scenarios, based on existing datasets from published human research on kidney allocation, where real participants had already made these same choices.
Each scenario involved two patients, Patient A and Patient B, both eligible for a single available kidney and characterised by age, health, and drinking habits. A decision maker must then choose which of the two patients should receive it.
“We ran these comparisons in a few different ways,” Hosseini said. “Sometimes we isolated just one trait at a time, sometimes we mixed several traits together to see how AI weighed competing factors, and sometimes we added a flip-a-coin option to measure indecision, a key factor present in human moral judgment.”
While human respondents tended to place greater importance on age — favouring younger over older patients — many models favoured lower alcohol consumption instead. Human decisions considered multiple factors and were more context-sensitive than those of LLMs, which often focused on a single attribute.
“First, AI chatbots often diverge from human values in how they weigh a patient’s traits,” said Hadi Hosseini, lead of the study at Penn State University. “They fixate on a single factor, like drinking habits, rather than balancing multiple considerations the way people do.”
RelatedThe researchers also saw that AI did not struggle with indecision. While humans recognised there is no single objectively correct answer and decisions can rely on nuanced human moral judgements, the systems committed to a single option with little hesitation.
"When we allocate something scarce, whether it’s a kidney, a job or access to some other resource, there isn’t always a single objectively correct answer,” said John Dickerson, chief executive officer at Mozilla.ai, who collaborated in the study.
“Humans recognize that ambiguity and codify it via open debate into the allocative process. AI models often don’t.”
AI and ethical decisions
Recent interactions with AI systems increasingly require them to go beyond factual information and make value judgments, the authors noted.
Large Language Models (LLMs) are increasingly integrated into healthcare, supporting clinical workflows, diagnosis, treatment planning and timely utilisation of scarce medical resources.
One of their applications involves decisions on allocating deceased-donor or living-donor kidneys to patients, which, according to the authors, depend on complex ethical and moral considerations.
In such high-stakes scenarios, the researchers pointed out that decisions demand not only accuracy but also alignment with human values and moral judgement.
According to the researchers, asking if AI can make moral decisions or whether they’re aligned with human values lies at the core of today’s wider debate on artificial intelligence.
“The ethical stakes are high, and AI’s role in such life-altering decisions requires deep reflection,” said Hosseini. “Moral decisions in settings like organ allocation directly determine who lives and who dies, so getting AI's role in them right isn't optional.”
“While we do not intend to encourage the use of AI as a substitute for professional judgment in medical decision-making or other high-stakes contexts, it's becoming essential to understand their behavior as individuals, organizations and firms more and more rely on AI to make decisions or receive recommendations,” he added

