Generative AI falls short in diagnostic reasoning despite accuracy

Generative AI falls short in diagnostic reasoning despite accuracy

Despite increasing use of artificial intelligence (AI) in health care, a new study led by Mass General Brigham researchers from the MESH Incubator shows that generative AI models continue to fall short at their clinical reasoning capabilities. By asking 21 different large language models (LLMs) to play doctor in a series of clinical scenarios, the…