Comparative performance of large language models on cardiovascular certification simulation exam

Highlights

  • We compared 3 large-language models to humans on cardiology board-style exams.

  • ChatGPT-4.0, Bing, and Gemini were tested on 434 ACCSAP multiple-choice questions.

  • ChatGPT-4.0 performed most similarly to human subjects.

  • Generative AI models require further refinement before incorporation into clinical practice.

ABSTRACT

Artificial intelligence (AI) is becoming increasingly prevalent in medical practice and has demonstrated sufficient clinical acumen to pass several licensing examinations. We tested the ability of 3 popular large-language models, ChatGPT-4.0 (OpenAI), Gemini (Google), and Bing AI (Microsoft), to pass a cardiovascular medicine board-style exam. Of these AI platforms, only ChatGPT-4.0 was able to achieve a score similar to human participants.

Background

Artificial intelligence (AI) has become increasingly ubiquitous and is poised to become a central component of healthcare delivery. , Nearly 10% of all U.S. Food and Drug Administration (FDA) approved AI technologies have cardiovascular applications, highlighting the potential for AI technologies to enhance diagnostic accuracy and improve outcomes for patients with a wide range of cardiovascular disease. Among the generative AI systems developed, the clinical reasoning of large-language models (LLM), including ChatGPT-4.0 (OpenAI), Gemini (Google), and Bing AI (Microsoft), have been tested using simulated medical licensure exams and demonstrated adequate performance. ,,,,,,,,,, However, little is known about their competency on the general cardiovascular board exam and the comparative accuracy among the 3 platforms. The goal of the study was to evaluate and compare the clinical reasoning abilities and diagnostic accuracy of the 3 predominant AI systems, ChatGPT-4.0, Gemini, and Bing AI, on cardiovascular board exam-style questions, and compare these results to human participants.

Methods

To comprehensively assess the capabilities of each AI system, 434 cardiology board preparation questions were sourced from the American College of Cardiology Self-Assessment Program (ACCSAP). We excluded 205 questions that included multimedia components (images, electrocardiograms, and other videos), as not all of the 3 AI platforms are able to upload and interpret this type of data. The included questions covered a broad spectrum of cardiovascular topics to ensure a representative evaluation of the AI systems’ proficiency in general cardiology and provide insights into the AI systems’ competency across diverse clinical scenarios. Text-only questions as well as questions with multimedia were selected. The questions we used were from the following categories: Arrhythmias, Coronary Artery Disease, Heart Failure and Cardiomyopathies, Pericardial Diseases, Congenital Heart Diseases, Vascular Diseases, Systemic Hypertension and Hypotension, Pulmonary Circulation Disorders, Systemic Disorders Affecting the Circulatory System, Valvular Disease and Miscellaneous Topics. Each AI platform was presented with a board-style question and multiple-choice answers and asked to provide the most appropriate response. This was done manually by one of the investigators (E.N.). The prompt contained the ACCSAP question and the following statement: “Out of the given choices, choose the best possible answer.” Justification for the answer was not required. The percentage of ACCSAP users who selected the correct answer for each question was also recorded. These percentages were averaged to measure the performance of those preparing for the cardiovascular medicine board exam and were used to define a likely passing score. The study’s primary outcome was the comparative performance of human participants and the AI platforms as assessed using Chi-squared tests. When the result of the Chi-squared test was significant, pairwise comparisons were done using false discovery rate (FDR) correction for multiple comparisons. Effect size was assessed using Cramer’s V. No extramural funding was used to support this work. The authors are solely responsible for the design and conduct of this study, all study analyses, the drafting and editing of the paper, and its final contents.

Results

The percentage of correct responses for test participants and the AI platforms are shown in Table 1 and Figure 1 . The overall accuracy of the AI platforms was 53.1% with Gemini, 72.0% with Bing AI and 80.9% with ChatGPT-4.0. Human test participants scored 78.7%. There were significant differences in test performance of the AI platforms across all categories except congenital heart disease, miscellaneous topics, pericardial diseases, and systemic disorders affecting the circulatory system. Test scores for ChatGPT-4.0 were similar to those of test participants, while there was a trend towards lower scores with Bing AI and Gemini’s results were significantly lower than that of test participants. ChatGPT-4.0 had the highest accuracy of the 3 LLM and numerically outperformed human participants overall and in congenital heart diseases, coronary artery disease, miscellaneous topics, pulmonary circulation disorders, systemic disorders affecting the circulatory system, systemic hypertension and hypotension, valvular disease, and vascular disease.

Table 1

Comparative results of ChatGPT-4.0, Gemini, Bing AI, and ACCSAP test candidates on simulated cardiovascular board style questions.

Specific topic Test participants ChatGPT-4.0 Bing AI Gemini P -value Effect size (Cramer’s V) Adjusted pairwise comparisons
ChatGPT-4.0 Bing AI Gemini
Overall ( n = 434) 78.7%
(74.8%-82.5%)
80.9%
(77.2%-84.6%)
72.0%
(67.7%-76.2%)
53.1%
(48.4%-57.8%)
<.01 0.28 0.69 0.06 <0.01
Arrhythmias ( n = 28) 83.6%
(69.9%-97.3%)
78.6%
(63.4%-93.8%)
53.6%
(35.1%-72.0%)
39.3%
(21.2%-57.4%)
<.01 0.44 0.84 0.06 <0.01
Congenital heart diseases ( n = 25) 78.4%
(62.3%-94.6%)
80.0%
(64.3%-95.7%)
64.0%
(45.2%-82.8%)
52.0%
(32.4%-71.6%)
.11 0.28 N/A N/A N/A
Coronary artery disease ( n = 70) 82.5%
(73.6%-91.4%)
82.8%
(74.0%-91.7%)
78.6%
(69.0%-88.2%)
64.3%
(53.1%-75.5%)
.03 0.21 0.99 0.70 0.051
Heart failure and cardiomyopathies ( n = 61) 78.0%
(67.6%-88.4%)
70.5%
(59.1%-81.9%)
80.3%
(70.4%-90.3%)
54.1%
(41.6%-66.6%)
.01 0.26 0.55 0.89 0.02
Miscellaneous Topics ( n = 28) 73.6%
(57.3%-90.0%)
78.6%
(63.4%-93.8%)
67.9%
(50.6%-85.2%)
57.1%
(38.8%-75.5%)
.34 0.2 N/A N/A N/A
Pericardial diseases ( n = 17) 76.1%
(55.9%-96.4%)
70.6%
(48.9%-92.2%)
52.9%
(29.2%-76.7%)
52.9%
(29.2%-76.7%)
.37 0.25 N/A N/A N/A
Pulmonary circulation disorders ( n = 25) 80.1%
(64.4%-95.7%)
88.0%
(75.3%-100.0%)
76.0%
(59.3%-92.7%)
48.0%
(28.4%-67.6%)
.01 0.39 0.69 0.84 0.06
Systemic disorders affecting the circulatory system ( n = 10) 77.4%
(51.5%-100.0%)
80.0%
(55.2%-100.0%)
70.0%
(41.6%-98.4%)
50.0%
(19.0%-81.0%)
.46 0.29 N/A N/A N/A
Systemic hypertension and hypotension ( n = 50) 76.7%
(65.0%-88.4%)
82.0%
(71.4%-92.6%)
76.0%
(64.2%-87.8%)
44.0%
(30.2%-57.8%)
<.01 0.38 0.69 0.99 <0.01
Valvular disease ( n = 84) 80.0%
(71.4%-88.5%)
83.3%
(75.4%-91.3%)
83.3%
(75.4%-91.3%)
56.0%
(45.3%-66.6%)
<.01 0.31 0.69 0.69 <0.01
Vascular diseases ( n = 36) 79.0%
(65.7%-92.3%)
94.4%
(87.0%-100.0%)
88.9%
(78.6%-99.2%)
66.7%
(51.3%-82.1%)
.01 0.32 0.10 0.44 0.55
Only gold members can continue reading. Log In or Register to continue

Stay updated, free articles. Join our Telegram channel

Jun 27, 2026 | Posted by in CARDIOLOGY | Comments Off on Comparative performance of large language models on cardiovascular certification simulation exam

Full access? Get Clinical Tree

Get Clinical Tree app for offline access