Assessing the Capability of ChatGPT, Google Bard, and Microsoft Bing in Solving Radiology Case Vignettes

Pradosh Kumar Sarangi; Ravi Kant Narayan; Sudipta Mohakud; Aditi Vats; Debabrata Sahani; Himel Mondal

doi:10.1055/s-0043-1777746

RSS-Feed abonnieren

Bitte kopieren Sie die angezeigte URL und fügen sie dann in Ihren RSS-Reader ein.

https://www.thieme-connect.de/rss/thieme/de/10.1055-s-00050590.xml

Teilen / Bookmarken

Facebook Linkedin Weibo

PDF herunterladen

CC BY-NC-ND 4.0 · Indian J Radiol Imaging 2024; 34(02): 276-282
DOI: 10.1055/s-0043-1777746

Original Article

Assessing the Capability of ChatGPT, Google Bard, and Microsoft Bing in Solving Radiology Case Vignettes

Pradosh Kumar Sarangi

¹Department of Radiodiagnosis, All India Institute of Medical Sciences, Deoghar, Jharkhand, India

,

Ravi Kant Narayan

²Department of Anatomy, ESIC Medical College & Hospital, Bihta, Patna, Bihar, India

,

Sudipta Mohakud

³Department of Radiodiagnosis, All India Institute of Medical Sciences, Bhubaneswar, Odisha, India

,

Aditi Vats

³Department of Radiodiagnosis, All India Institute of Medical Sciences, Bhubaneswar, Odisha, India

,

Debabrata Sahani

³Department of Radiodiagnosis, All India Institute of Medical Sciences, Bhubaneswar, Odisha, India

,

Himel Mondal

⁴Department of Physiology, All India Institute of Medical Sciences, Deoghar, Jharkhand, India

› Institutsangaben
Funding None.

› Weitere Informationen

Auch verfügbar auf

Abstract
Volltext
Referenzen

Lizenzen und Reprints

Abstract

Background The field of radiology relies on accurate interpretation of medical images for effective diagnosis and patient care. Recent advancements in artificial intelligence (AI) and natural language processing have sparked interest in exploring the potential of AI models in assisting radiologists. However, limited research has been conducted to assess the performance of AI models in radiology case interpretation, particularly in comparison to human experts.

Objective This study aimed to evaluate the performance of ChatGPT, Google Bard, and Bing in solving radiology case vignettes (Fellowship of the Royal College of Radiologists 2A [FRCR2A] examination style questions) by comparing their responses to those provided by two radiology residents.

Methods A total of 120 multiple-choice questions based on radiology case vignettes were formulated according to the pattern of FRCR2A examination. The questions were presented to ChatGPT, Google Bard, and Bing. Two residents wrote the examination with the same questions in 3 hours. The responses generated by the AI models were collected and compared to the answer keys and explanation of the answers was rated by the two radiologists. A cutoff of 60% was set as the passing score.

Results The two residents (63.33 and 57.5%) outperformed the three AI models: Bard (44.17%), Bing (53.33%), and ChatGPT (45%), but only one resident passed the examination. The response patterns among the five respondents were significantly different (p = 0.0117). In addition, the agreement among the generative AI models was significant (intraclass correlation coefficient [ICC] = 0.628), but there was no agreement between the residents (Kappa = –0.376). The explanation of generative AI models in support of answer was 44.72% accurate.

Conclusion Humans exhibited superior accuracy compared to the AI models, showcasing a stronger comprehension of the subject matter. All three AI models included in the study could not achieve the minimum percentage needed to pass an FRCR2A examination. However, generative AI models showed significant agreement in their answers where the residents exhibited low agreement, highlighting a lack of consistency in their responses.

Keywords

artificial intelligence - Bard - Bing - ChatGPT - natural language processing - radiology - FRCR2A - fellowship

Publikationsverlauf

Artikel online veröffentlicht:
29. Dezember 2023

© 2023. Indian Radiological Association. This is an open access article published by Thieme under the terms of the Creative Commons Attribution-NonDerivative-NonCommercial License, permitting copying and reproduction so long as the original work is given appropriate credit. Contents may not be used for commercial purposes, or adapted, remixed, transformed or built upon. (https://creativecommons.org/licenses/by-nc-nd/4.0/)

Thieme Medical and Scientific Publishers Pvt. Ltd.
A-12, 2nd Floor, Sector 2, Noida-201301 UP, India

References
1 Rathan R, Hamdy H, Kassab SE, Salama MNF, Sreejith A, Gopakumar A. Implications of introducing case based radiological images in anatomy on teaching, learning and assessment of medical students: a mixed-methods study. BMC Med Educ 2022; 22 (01) 723

MissingFormLabel
Crossref PubMed Suche in Google Scholar
2 Meng F, Kottlors J, Shahzad R. et al. AI support for accurate and fast radiological diagnosis of COVID-19: an international multicenter, multivendor CT study. Eur Radiol 2023; 33 (06) 4280-4291

MissingFormLabel
Crossref PubMed Suche in Google Scholar
3 Nijiati M, Zhang Z, Abulizi A. et al. Deep learning assistance for tuberculosis diagnosis with chest radiography in low-resource settings. J XRay Sci Technol 2021; 29 (05) 785-796

MissingFormLabel
PubMed Suche in Google Scholar
4 Akkus Z, Galimzianova A, Hoogi A, Rubin DL, Erickson BJ. Deep learning for brain MRI segmentation: state of the art and future directions. J Digit Imaging 2017; 30 (04) 449-459

MissingFormLabel
Crossref PubMed Suche in Google Scholar
5 Esmaeili M, Vettukattil R, Banitalebi H, Krogh NR, Geitung JT. Explainable artificial intelligence for human-machine interaction in brain tumor localization. J Pers Med 2021; 11 (11) 1213

MissingFormLabel
Crossref PubMed Suche in Google Scholar
6 Sinha RK, Deb Roy A, Kumar N, Mondal H. Applicability of ChatGPT in assisting to solve higher order problems in pathology. Cureus 2023; 15 (02) e35237

MissingFormLabel
PubMed Suche in Google Scholar
7 Vaishya R, Misra A, Vaish A. ChatGPT: is this version good for healthcare and research?. Diabetes Metab Syndr 2023; 17 (04) 102744

MissingFormLabel
Crossref PubMed Suche in Google Scholar
8 Korb KB, Nyberg EP, Oshni Alvandi A. et al. Individuals vs. BARD: experimental evaluation of an online system for structured, collaborative bayesian reasoning. Front Psychol 2020; 11: 1054

MissingFormLabel
Crossref PubMed Suche in Google Scholar
9 Rahsepar AA, Tavakoli N, Kim GHJ, Hassani C, Abtin F, Bedayat A. How AI responds to common lung cancer questions: ChatGPT vs Google Bard. Radiology 2023; 307 (05) e230922

MissingFormLabel
Crossref PubMed Suche in Google Scholar
10 Bhayana R, Krishna S, Bleakney RR. Performance of ChatGPT on a radiology board-style examination: insights into current strengths and limitations. Radiology 2023; 307 (05) e230582

MissingFormLabel
Crossref PubMed Suche in Google Scholar
11 Bhayana R, Bleakney RR, Krishna S. GPT-4 in radiology: improvements in advanced reasoning. Radiology 2023; 307 (05) e230987

MissingFormLabel
Crossref PubMed Suche in Google Scholar
12 McCoubrie P, McKnight L. Single best answer MCQs: a new format for the FRCR part 2a exam. Clin Radiol 2008; 63 (05) 506-510

MissingFormLabel
Crossref PubMed Suche in Google Scholar
13 Surry LT, Torre D, Trowbridge RL, Durning SJ. A mixed-methods exploration of cognitive dispositions to respond and clinical reasoning errors with multiple choice questions. BMC Med Educ 2018; 18 (01) 277

MissingFormLabel
Crossref PubMed Suche in Google Scholar
14 Tilmatine M, Hubers F, Hintz F. Exploring individual differences in recognizing idiomatic expressions in context. J Cogn 2021; 4 (01) 37

MissingFormLabel
Crossref PubMed Suche in Google Scholar
15. Khan B, Fatima H, Qureshi A. et al. Drawbacks of artificial intelligence and their potential solutions in the healthcare sector. Biomed Mater Devices 2023; 1-8

MissingFormLabel
PubMed Suche in Google Scholar
16 Agarwal M, Sharma P, Goswami A. Analysing the applicability of ChatGPT, Bard, and Bing to generate reasoning-based multiple-choice questions in medical physiology. Cureus 2023; 15 (06) e40977

MissingFormLabel
PubMed Suche in Google Scholar
17 Williams MC, Shambrook J. How will artificial intelligence transform cardiovascular computed tomography? A conversation with an AI model. J Cardiovasc Comput Tomogr 2023; 17 (04) 281-283

MissingFormLabel
Crossref PubMed Suche in Google Scholar
18 Korteling JEH, van de Boer-Visschedijk GC, Blankendaal RAM, Boonekamp RC, Eikelboom AR. Human- versus artificial intelligence. Front Artif Intell 2021; 4: 622364

MissingFormLabel
Crossref PubMed Suche in Google Scholar
19 Booth TC, Martins RDM, McKnight L, Courtney K, Malliwal R. The Fellowship of the Royal College of Radiologists (FRCR) examination: a review of the evidence. Clin Radiol 2018; 73 (12) 992-998

MissingFormLabel
Crossref PubMed Suche in Google Scholar
20 Ferrell B, Raskin SE, Zimmerman EB. Calibrating a transformer-based model's confidence on community-engaged research studies: decision support evaluation study. JMIR Form Res 2023; 7: e41516

MissingFormLabel
Crossref PubMed Suche in Google Scholar

RSS-Feed abonnieren

Teilen / Bookmarken

Assessing the Capability of ChatGPT, Google Bard, and Microsoft Bing in Solving Radiology Case Vignettes

Abstract

Keywords

Publikationsverlauf

References