Hostname: page-component-76d6cb85b7-92wsb Total loading time: 0 Render date: 2026-07-25T06:08:11.100Z Has data issue: false hasContentIssue false

Evaluating chain-of-thought prompting in a GPT chatbot for BCID2 interpretation and stewardship: how does AI compare to human experts?

Published online by Cambridge University Press:  11 July 2025

Daniel M. Tassone*
Affiliation:
Division of Infectious Diseases, Department of Medicine, Central Virginia VA Health Care System, Richmond, VA, USA Virginia Commonwealth University, School of Pharmacy, Richmond, VA, USA
Matthew M. Hitchcock
Affiliation:
Division of Infectious Diseases, Department of Medicine, Central Virginia VA Health Care System, Richmond, VA, USA Division of Infectious Diseases, Department of Medicine, Virginia Commonwealth University School of Medicine, Richmond, VA, USA
Connor J. Rossier
Affiliation:
Department of Health Informatics, Central Virginia VA Health Care System, Richmond, VA, USA
Douglas Fletcher
Affiliation:
Department of Health Informatics, Central Virginia VA Health Care System, Richmond, VA, USA
Julia Ye
Affiliation:
Division of Infectious Diseases, Department of Medicine, Central Virginia VA Health Care System, Richmond, VA, USA Virginia Commonwealth University, School of Pharmacy, Richmond, VA, USA
Ian Langford
Affiliation:
Division of Infectious Diseases, Department of Medicine, Virginia Commonwealth University School of Medicine, Richmond, VA, USA
Julie Boatman
Affiliation:
Division of Infectious Diseases, Department of Medicine, Central Virginia VA Health Care System, Richmond, VA, USA Division of Infectious Diseases, Department of Medicine, Virginia Commonwealth University School of Medicine, Richmond, VA, USA
J. Daniel Markley
Affiliation:
Division of Infectious Diseases, Department of Medicine, Central Virginia VA Health Care System, Richmond, VA, USA Division of Infectious Diseases, Department of Medicine, Virginia Commonwealth University School of Medicine, Richmond, VA, USA
*
Corresponding author: Daniel M. Tassone; Email: Daniel.tassone@va.gov

Abstract

Background:

Rapid molecular diagnostics, such as the BIOFIRE® Blood Culture Identification 2 (BCID2) panel, have improved the time to pathogen identification in bloodstream infections. However, accurate interpretation and antimicrobial optimization require Infectious Disease (ID) expertise, which may not always be readily available. GPT-powered chatbots could support antimicrobial stewardship programs (ASPs) by assisting non-specialist providers in BCID2 result interpretation and treatment recommendations. This study evaluates the performance of a GPT-4 chatbot compared to ASP prospective audit and feedback interventions.

Methods:

This prospective observational study assessed 43 consecutive real-world cases of bacteremia at a 399-bed VA Medical Center from January to May 2024. The GPT-chatbot utilized “chain-of-thought” prompting and external knowledge integration to generate recommendations. Two independent ID physicians evaluated chatbot and ASP recommendations across four domains: BCID2 interpretation, source control, antibiotic therapy, and additional diagnostic workup. The primary endpoint was the combined rate of harmful or inadequate recommendations. Secondary endpoints assessed the rate of harmful or inadequate responses for each domain.

Results:

The chatbot had a significantly higher rate of harmful or inadequate recommendations (13%) compared to ASP (4%, p = 0.047). The most significant discrepancy was observed in the domain of antibiotic therapy, where harmful recommendations occurred in up to 10% (p <0.05) of chatbot evaluations. The chatbot performed well in BCID2 interpretation (100% accuracy) but provided more inadequate responses in source control consideration (10% vs. 2% for ASP, p = 0.022).

Conclusions:

GPT-powered chatbots show potential for supporting antimicrobial stewardship but should only complement, not replace, human expertise in infectious disease management.

Information

Type
Original Article
Creative Commons
Creative Common License - CCCreative Common License - BY
This is an Open Access article, distributed under the terms of the Creative Commons Attribution licence (https://creativecommons.org/licenses/by/4.0/), which permits unrestricted re-use, distribution and reproduction, provided the original article is properly cited.
Copyright
© The Author(s), 2025. Published by Cambridge University Press on behalf of The Society for Healthcare Epidemiology of America
Figure 0

Figure 1. Study design. Illustrates the study design, outlining three key phases. Part 1 details the antimicrobial stewardship program (ASP)’s real-time documentation of BCID2-positive blood culture cases, including history of present illness, BCID2 interpretation, source control considerations, antibiotic therapy, and additional diagnostic workup. Part 2 describes the ASP provider inputting anonymized patient data into the chatbot, which then generates a response structured similarly to the ASP note. Part 3 involves a retrospective comparison of the ASP note and chatbot response by two independent ID physicians, using a standardized scoring rubric. The diagram visually represents the workflow from clinical documentation to AI evaluation and comparison.

Figure 1

Figure 2. Evaluation rubric. Displays the evaluation rubric used to assess the accuracy and appropriateness of BCID2 result interpretation, source control considerations, antibiotic therapy recommendations, and additional diagnostic workup. The table outlines seven measures (M1–M7) across four domains, specifying scoring classifications such as appropriate, non-optimal, overly broad, or inadequate/harmful responses. This rubric provides a structured framework for evaluating the performance of ASP and chatbot recommendations.

Figure 2

Table 1. Characteristics of the study population (n = 43)

Figure 3

Table 2. Comparison of ASP PAF interventions and chatbot performance

Figure 4

Table 3. Characterization of cases with non-optimal and harmful chatbot responses for antibiotic therapy recommendations

Supplementary material: File

Tassone et al. supplementary material

Tassone et al. supplementary material
Download Tassone et al. supplementary material(File)
File 42.4 KB