Comparison of AI-based Chatbot Performance in Analyzing Clinical Scenarios versus Medical Residents: A Novel Approach in Chest Diseases Education


Creative Commons License

Bilgin M. H., Alp H. H.

Thoracic Research and Practice, cilt.27, sa.5, ss.275-279, 2026 (ESCI, Scopus, TRDizin)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 27 Sayı: 5
  • Basım Tarihi: 2026
  • Doi Numarası: 10.4274/thoracrespract.2026.2026-1-2
  • Dergi Adı: Thoracic Research and Practice
  • Derginin Tarandığı İndeksler: Emerging Sources Citation Index (ESCI), Scopus, TR DİZİN (ULAKBİM)
  • Sayfa Sayıları: ss.275-279
  • Anahtar Kelimeler: artificial intelligence, clinical, clinical reasoning, decision support systems, diagnosis, differential, Education, humans, medical
  • Açık Arşiv Koleksiyonu: AVESİS Açık Erişim Koleksiyonu
  • Van Yüzüncü Yıl Üniversitesi Adresli: Evet

Özet

OBJECTIVE: Rapid advancements in artificial intelligence (AI) technologies offer new opportunities in medical education. The aim of this study is to compare the performance of large language models, specifically ChatGPT-4 and Gemini, in analyzing clinical scenarios with that of chest diseases research assistants (residents), and to evaluate their potential roles in medical education. MATERIAL AND METHODS: This cross-sectional, comparative study included 28 resident physicians working in the department of chest diseases at a tertiary-care university hospital. Four clinical scenarios involving diagnoses of massive pulmonary embolism, chronic obstructive pulmonary disease, asthma, and severe pneumonia/sepsis were presented to both participants and AI models (ChatGPT-4 and Gemini). Responses were scored by blinded experts based on current guidelines (Global Initiative for Chronic Obstructive Lung Disease, Global Initiative for Asthma, American Thoracic Society). RESULTS: AI models achieved significantly higher scores than residents, particularly on structured questions requiring theoretical knowledge, classification skills, and the listing of contraindications (P < 0.05). However, it was observed that residents achieved success levels similar to those of AI models in situations requiring emergency intervention (e.g., shock management) through practical, results-oriented approaches. While AI models provided a broader spectrum in differential diagnosis, residents preferred “telegraphic” and practice-oriented responses. CONCLUSION: ChatGPT and Gemini have significant potential as clinical decision-support systems and educational assistants. However, rather than replacing human factors in clinical reasoning and emergency management, they should be positioned as complementary tools that accelerate physicians’ access to theoretical knowledge.