Assessing the Accuracy of Artificial Intelligence Chatbots in Medical Information Retrieval: A Structured Query-based Evaluation
S. Dhohan, Gagan D. Urs, K. M. Sneha, Vismaya V, Siddhartha ND
Advances in Research · pp. 143–155 · Published 5 Aug 2026
10.9734/air/2026/v27i51698Abstract
Background: Artificial intelligence chatbots are increasingly used to obtain medical and drug-related information, but their accuracy for clinical use remains uncertain. Objective: To evaluate and compare the performance of three large language models—ChatGPT, Gemini, and Grok—in responding to standardised drug-related queries concerning three commonly prescribed drugs. Methods: Three commonly prescribed drugs—metformin, hydrochlorothiazide, and azithromycin—were selected for assessment. Each model was asked ten standardised questions per drug (two questions in each of five categories: indications, off-label indications, drug–drug interactions, adverse drug events, and drug availability). Responses were manually assessed against standard clinical references, principally UpToDate, and scored on a four-point scale from 0 to 3, where 3 represented a completely accurate and clinically sound response. A non-perfect score (0–2) was considered an error. Each model answered 30 questions in total. Results: ChatGPT and Gemini each produced 18 perfect responses, corresponding to an empirical probability of 0.60 for a completely correct answer. Grok produced 14 perfect responses, corresponding to an empirical probability of 0.47. Error rates varied across drugs and models, ranging from 30% to 60%. A two-way analysis of variance (ANOVA) of mean error rates showed that drug type had a statistically significant effect (F = 7.75, p = 0.0421), whereas the effect of the AI model was not statistically significant at the conventional threshold (F = 4.00, p = 0.1111). Grok's numerically higher mean error rate (53.33%, compared with 40.00% for ChatGPT and Gemini) was consistent with its lower empirical success probability. Conclusion: These findings indicate that freely available AI chatbots may provide rapid drug information but show variable accuracy. As a practical implication for current use, rather than a proposed future research direction, their responses should be verified against authoritative clinical references before use in healthcare education or practice.
Cited by 0
No indexed citations yet.
Related research
- The Impact/Role of Artificial Intelligence in Anesthesia: Remote Pre-Operative Assessment and Perioperative — shares topic coverage
- Detecting Dental Caries through Captured Images Using the Machine Learning Technology Teachable Machine — shares topic coverage
- Harnessing Artificial Intelligence in Healthcare Analytics: From Diagnosis to Treatment Optimization — shares topic coverage
- Diagnostic Accuracy of Artificial Intelligence for Breast Cancer Detection: A Systematic Review — shares topic coverage
- Artificial Intelligence in the Analysis of the Fetal Genome in Utero: A Critical Review of Current Paradigms, Clinical Utility and Future Horizons — shares topic coverage
Article metrics
Real usage data collected on this platform.
0
Page views
0
PDF downloads
0
Outbound clicks
0
Citations
Views by country
Approximate, from request IP at view time — not citizenship or institution. Countries with fewer than 5 views are grouped as "Other".
No views recorded yet.
Traffic sources
Referring site, by host.
No traffic recorded yet.
Views and downloads exclude known bots/crawlers. Citations combines this platform's own DOI-resolved index with each external source's own reported total — see Cited by above for individually listed citing works. Last refreshed 0 seconds ago.