E-ISSN: 2822-6771
Volume : 18 Issue : 4 Year : 2026
Quick Search
Evaluation of Artificial Intelligence Chatbots for Preoperative Counseling in Patients with Neurological Disorders [Compreh Med]
Compreh Med. 2026; 18(4): 424-430 | DOI: 10.14744/cm.2026.68916

Evaluation of Artificial Intelligence Chatbots for Preoperative Counseling in Patients with Neurological Disorders

Ceyhan Oflezer1, Mert Göbel2
1Department of Anesthesiology and Reanimation, University of Health Sciences, Bakırköy Prof. Dr. Mazhar Osman Training and Research Hospital for Psychiatry, Neurology and Neurosurgery, İstanbul, Türkiye
2Department of Neurology, Kızıltepe State Hospital, Mardin, Türkiye

INTRODUCTION: The quality of information provided to patients with neurological disorders using large language model (LLM)-based artificial intelligence (AI) chatbots has not been sufficiently investigated. The aim of this study was to evaluate and compare the quality of responses generated by ChatGPT, Google Gemini, and Microsoft Copilot to frequently asked preoperative questions from patients with neurological disorders.
METHODS: Twenty frequently asked preoperative questions were identified across four neurological disorders: multiple sclerosis (MS), epilepsy, Parkinson’s disease, and dementia. Each question was submitted to ChatGPT, Gemini, and Copilot using identical prompts in newly initiated sessions. Responses were eval-uated in a blinded manner by a panel of eight neurologists and two anesthesiologists. Clinical accuracy, safety, and adequacy were assessed using a 5-point Likert scale. Inter-rater agreement was assessed using the intraclass correlation coefficient (ICC).
RESULTS: ChatGPT achieved the highest overall scores across all disease groups. Significant differences were observed among chatbots in the overall analysis (χ²(2)=40.635, p<0.001, Kendall’s W=0.423). Post-hoc analyses demonstrated that ChatGPT significantly outperformed both Gemini (r=0.705, p<0.001) and Copilot (r=0.747, p<0.001), whereas no significant difference was found between Gemini and Copilot (p=0.876). The lowest scores for all chatbots were observed in the MS group. Overall inter-rater agreement was moderate (ICC=0.571; 95% CI, 0.411–0.712). No response was classified as potentially harmful or misleading.
DISCUSSION AND CONCLUSION: All three AI chatbots provided generally high-quality responses to questions from patients with neurological disorders. ChatGPT demonstrated superior performance compared with Gemini and Copilot. Although these tools may support patient education, they should complement rather than replace physician-guided counseling.

Keywords: Artificial intelligence, ChatGPT, Copilot, Gemini, large language models, neurological diseases, patient education, preoperative care


Nörolojik Bozukluğu Olan Hastalarda Preoperatif Danışmanlık İçin Yapay Zekâ Sohbet Botlarının Değerlendirilmesi

Ceyhan Oflezer1, Mert Göbel2
1Sağlık Bilimleri Üniversitesi, Bakırköy Prof. Dr. Mazhar Osman Psikiyatri, Nöroloji ve Beyin Cerrahisi Eğitim ve Araştırma Hastanesi, Anestezi ve Yoğun Bakım Kliniği, İstanbul
2Kızıltepe Devlet Hastanesi, Nöroloji Kliniği, Mardin

GİRİŞ ve AMAÇ: Nörolojik bozuklukları olan hastalara büyük dil modeli (LLM) tabanlı yapay zekâ (YZ) sohbet robotları kullanılarak sağlanan bilgilerin kalitesi yeterince araştırılmamıştır. Bu çalışmanın amacı nörolojik bozuklukları olan hastalardan sıkça sorulan ameliyat öncesi sorulara Chat GPT, Google Gemini ve Microsoft Copilot tarafından üretilen yanıtların kalitesini değerlendirmek ve karşılaştırmaktır.
YÖNTEM ve GEREÇLER: Dört nörolojik bozuklukta (multipl skleroz (MS), epilepsi, parkinson hastalığı ve demans) sıkça sorulan yirmi ameliyat öncesi soru belirlendi. Her soru, yeni başlatılan oturumlarda aynı komutlar kullanılarak ChatGPT, Gemini ve Copilot'a gönderildi. Yanıtlar, sekiz nörolog ve iki anestezistten oluşan bir panel tarafından körleme yöntemiyle değerlendirildi. Klinik doğruluk, güvenlik ve yeterlilik, 5 puanlık Likert ölçeği kullanılarak değerlendirildi. Değerlendiriciler arası uyum, sınıf içi korelasyon katsayısı (ICC) kullanılarak değerlendirildi.
BULGULAR: ChatGPT, tüm hastalık gruplarında en yüksek genel puanları elde etti. Genel analizde chatbotlar arasında önemli farklılıklar gözlemlendi (χ²(2)=40.635, p<0.001, Kendall’s W=0.423). Post-hoc analizler, ChatGPT'nin hem Gemini'den (r=0.705, p<0.001) hem de Copilot'tan (r=0.747, p<0.001) önemli ölçüde daha iyi performans gösterdiğini, Gemini ve Copilot arasında ise anlamlı bir fark bulunmadığını (p=0.876) gösterdi. Tüm chatbotlar için en düşük puanlar MS grubunda gözlemlendi. Genel olarak değerlendiriciler arası uyum orta düzeydeydi (ICC=0.571; %95 CI, 0.411–0.712). Hiçbir yanıt potansiyel olarak zararlı veya yanıltıcı olarak sınıflandırılmadı.
TARTIŞMA ve SONUÇ: Her üç yapay zekâ chatbotu da nörolojik bozuklukları olan hastalardan gelen sorulara genel olarak yüksek kaliteli yanıtlar verdi. ChatGPT, Gemini ve Copilot'a kıyasla üstün performans gösterdi. Bu araçlar hasta eğitimini destekleyebilse de, hekim rehberliğindeki danışmanlığın yerini almak yerine onu tamamlamalıdır.

Anahtar Kelimeler: Yapay zekâ, ChatGPT, Copilot, Gemini, büyük dil modelleri, nörolojik hastalıklar, hasta eğitimi, preoperatif bakım


Corresponding Author: Ceyhan Oflezer, Türkiye
Manuscript Language: English
×
APA
NLM
AMA
MLA
Chicago
Copied!
CITE