Frontier in Medical & Health Research
COMPARISON OF ARTIFICIAL INTELLIGENCE BASED DIAGNOSIS WITH CLINICAL DIAGNOSIS IN THE EMERGENCY DEPARTMENT OF A TERTIARY CARE HOSPITAL
PDF

Keywords

Artificial Intelligence; Emergency Department; Clinical Diagnosis; Diagnostic Agreement

How to Cite

COMPARISON OF ARTIFICIAL INTELLIGENCE BASED DIAGNOSIS WITH CLINICAL DIAGNOSIS IN THE EMERGENCY DEPARTMENT OF A TERTIARY CARE HOSPITAL. (2026). Frontier in Medical and Health Research, 4(3), 2776-2782. https://fmhr.net/index.php/fmhr/article/view/3486

Abstract

Background: Emergency departments (EDs) require rapid, accurate diagnosis, yet clinical decision-making is frequently constrained by time pressure, high patient volumes, and variable disease presentation. Artificial intelligence (AI), particularly large language models such as ChatGPT, has attracted growing interest as a diagnostic support tool, but evidence from EDs in developing countries remains scarce. Objective: To compare AI-based diagnosis generated by ChatGPT with clinical diagnosis made by physicians in the ED of Rehman Medical Institute (RMI), Peshawar. Methods: This retrospective, comparative cross-sectional study analyzed secondary data from 384 patient records (January 2025–April 2026), selected by non-probability consecutive sampling. De-identified demographic, symptom, vital-sign, and laboratory data (excluding clinical diagnosis) were entered into ChatGPT to generate an independent AI-based diagnosis, which was then compared with the physician's documented diagnosis. Data were analyzed in SPSS v26 using descriptive statistics, cross-tabulation, percentage agreement, and chi-square testing. Results: Exact agreement between AI and physician diagnosis was observed in 87.0% of cases (334/384), an overall diagnostic accuracy of 86.98%. When cases where AI matched with added clinical detail or was clinically close to the physician's diagnosis were included, agreement reached 92.4% (355/384). Agreement was highest for neurological (93.1%), respiratory (92.9%), and renal/genitourinary (84.6–85.6%) presentations, and lowest for infectious disease presentations (82.8%). None of the six organ-system subgroup comparisons reached statistical significance (all p > 0.05), with cardiac presentations showing the strongest trend (p = 0.069). Conclusion: ChatGPT demonstrated substantial, clinically meaningful agreement with physician diagnosis in a real-world, resource-limited ED setting, supporting its potential as an adjunct decision-support tool. AI cannot substitute for physician judgment and should be integrated alongside, not in place of, clinical expertise.

PDF