Unknown

Dataset Information

0

Evaluating multimodal AI in medical diagnostics.


ABSTRACT: This study evaluates multimodal AI models' accuracy and responsiveness in answering NEJM Image Challenge questions, juxtaposed with human collective intelligence, underscoring AI's potential and current limitations in clinical diagnostics. Anthropic's Claude 3 family demonstrated the highest accuracy among the evaluated AI models, surpassing the average human accuracy, while collective human decision-making outperformed all AI models. GPT-4 Vision Preview exhibited selectivity, responding more to easier questions with smaller images and longer questions.

SUBMITTER: Kaczmarczyk R 

PROVIDER: S-EPMC11306783 | biostudies-literature | 2024 Aug

REPOSITORIES: biostudies-literature

altmetric image

Publications

Evaluating multimodal AI in medical diagnostics.

Kaczmarczyk Robert R   Wilhelm Theresa Isabelle TI   Martin Ron R   Roos Jonas J  

NPJ digital medicine 20240807 1


This study evaluates multimodal AI models' accuracy and responsiveness in answering NEJM Image Challenge questions, juxtaposed with human collective intelligence, underscoring AI's potential and current limitations in clinical diagnostics. Anthropic's Claude 3 family demonstrated the highest accuracy among the evaluated AI models, surpassing the average human accuracy, while collective human decision-making outperformed all AI models. GPT-4 Vision Preview exhibited selectivity, responding more t  ...[more]

Similar Datasets

| S-EPMC11411230 | biostudies-literature
| S-EPMC11751740 | biostudies-literature
| S-EPMC6406313 | biostudies-literature
| S-EPMC11772878 | biostudies-literature
| S-EPMC9277645 | biostudies-literature
| S-EPMC11464372 | biostudies-literature
| S-EPMC2750850 | biostudies-literature
| S-EPMC11651983 | biostudies-literature
| S-EPMC10996713 | biostudies-literature
| S-EPMC10984061 | biostudies-literature