Published in JAMA Network Open · View the paper (DOI)
A study of 21 large language models found that they excel at accurate final diagnoses but falter at differential diagnoses, highlighting the need for human oversight in healthcare. The researchers developed a novel measure, PrIME-LLM, to evaluate AI's clinical competency across different stages of reasoning.
Coverage from 2 institutions