Skip to content
Research Article Open access CC BY 4.0

Assessing the Accuracy of Artificial Intelligence Chatbots in Medical Information Retrieval: A Structured Query-based Evaluation

S. Dhohan, Gagan D. Urs, K. M. Sneha, Vismaya V, Siddhartha ND

Advances in Research · pp. 143–155 · Published 5 Aug 2026

10.9734/air/2026/v27i51698

Abstract

Background: Artificial intelligence chatbots are increasingly used to obtain medical and drug-related information, but their accuracy for clinical use remains uncertain. Objective: To evaluate and compare the performance of three large language models—ChatGPT, Gemini, and Grok—in responding to standardised drug-related queries concerning three commonly prescribed drugs. Methods: Three commonly prescribed drugs—metformin, hydrochlorothiazide, and azithromycin—were selected for assessment. Each model was asked ten standardised questions per drug (two questions in each of five categories: indications, off-label indications, drug–drug interactions, adverse drug events, and drug availability). Responses were manually assessed against standard clinical references, principally UpToDate, and scored on a four-point scale from 0 to 3, where 3 represented a completely accurate and clinically sound response. A non-perfect score (0–2) was considered an error. Each model answered 30 questions in total. Results: ChatGPT and Gemini each produced 18 perfect responses, corresponding to an empirical probability of 0.60 for a completely correct answer. Grok produced 14 perfect responses, corresponding to an empirical probability of 0.47. Error rates varied across drugs and models, ranging from 30% to 60%. A two-way analysis of variance (ANOVA) of mean error rates showed that drug type had a statistically significant effect (F = 7.75, p = 0.0421), whereas the effect of the AI model was not statistically significant at the conventional threshold (F = 4.00, p = 0.1111). Grok's numerically higher mean error rate (53.33%, compared with 40.00% for ChatGPT and Gemini) was consistent with its lower empirical success probability. Conclusion: These findings indicate that freely available AI chatbots may provide rapid drug information but show variable accuracy. As a practical implication for current use, rather than a proposed future research direction, their responses should be verified against authoritative clinical references before use in healthcare education or practice.

Artificial intelligence large language models ChatGPT Gemini Grok drug information medical information retrieval response accuracy pharmacy education patient safety

Cited by 0

No indexed citations yet.

Article metrics

Real usage data collected on this platform.

0

Page views

0

PDF downloads

0

Outbound clicks

0

Citations

Views by country

Approximate, from request IP at view time — not citizenship or institution. Countries with fewer than 5 views are grouped as "Other".

No views recorded yet.

Traffic sources

Referring site, by host.

No traffic recorded yet.

Views and downloads exclude known bots/crawlers. Citations combines this platform's own DOI-resolved index with each external source's own reported total — see Cited by above for individually listed citing works. Last refreshed 0 seconds ago.