Generative AI falls short in diagnostic reasoning despite accuracy
Despite increasing use of artificial intelligence (AI) in health care, a new study led by Mass General Brigham researchers from the MESH Incubator shows that generative AI models continue to fall short at their clinical reasoning capabilities. By asking 21 different large language models (LLMs) to play doctor in a series of clinical scenarios, the…