Introduction
Health disparities are systematic, plausibly avoidable health differences adversely affecting socially disadvantaged groups. Reference Braveman1 As in other realms of healthcare, health disparities in healthcare-associated infections (HAI) have been described for minority racial and ethnic groups, Reference Bakullari, Metersky and Wang2–Reference Kirtz, Chan, McClanahan and Octaria4 and individuals living in highly vulnerable social areas. Reference Dewitt, Reinke and Inman5,Reference Sood, Dougherty and Martin6 These observed health disparities result from unfavorable social determinants of health (SDoH). SDoH are nonmedical factors that influence health outcomes. Reference Braveman, Egerter and Williams7 These include socioeconomic status, education level, living conditions, access to care, and insurance coverage, with each having a notably large effect in mediating racial and ethnic disparities in HAI. Reference Tarabay, Nix and Doline8,Reference See, Wesson and Gualandi9 Conversely, health equity is defined as the state in which everyone has a fair and just opportunity to attain the highest level of health, often achieved by mitigating health disparities and their causes. Reference Braveman1
Artificial intelligence (AI) encompasses algorithms capable of achieving human-like performance in specific decision-making or recognition tasks, such as speech recognition, natural language processing (NLP), or image recognition. A well-known subfield of AI is machine learning (ML), in which models learn from data, excelling at identifying patterns and making predictions. As a subset of ML, deep learning (DL) algorithms can perform these same tasks but utilize layers of neural networks to learn complex nonlinear patterns across large data sets. AI/ML methods are increasingly used across all aspects of healthcare delivery, further fueled by the implementation of electronic health records (EHRs) as a primary source for healthcare data. Figure 1 provides an overview of AI applications in healthcare epidemiology, along with common data types, methods, and equity considerations, and Table 1 provides examples of AI/ML applications and their potential equity benefits, risks, and practical safeguards. These applications have the potential to worsen existing health disparities, as bias embedded in the data can be amplified by the algorithms and their deployment. An application with high global accuracy can exacerbate health disparities if the model is less accurate for marginalized subgroups, leading to suboptimal health care delivery for those individuals. Fairness is the intentional design and implementation of AI/ML methods to ensure equitable health outcomes and resource allocation across demographic groups. Reference Gao, Chou and McCaw10 In this review, we discuss opportunities to improve health equity in healthcare epidemiology and HAIs outcomes through AI/ML applications, and strategies to ensure that these applications promote, rather than hinder, health equity and fairness.
Healthcare applications for subfields of artificial intelligence, noting common data types and the importance of equity.

Figure 1. Long description
The flowchart outlines the applications of artificial intelligence in healthcare, including improving data capture and processing in electronic health records, risk prediction and infection control, and antimicrobial stewardship. It highlights common types of data used, such as structured data like patient demographics and vital signs, and unstructured data like clinical notes and medical images. The learning mechanisms section details unsupervised learning methods like K-means clustering, supervised learning methods like linear regression and neural networks, and reinforcement learning methods like neural networks and autoencoders. The ensuring health equity section lists benefits such as standardized tools for diagnostic and clinical support, concerns like data bias, and mitigation strategies like using representative data and assessing fairness metrics.
Examples of AI/ML applications in healthcare epidemiology and antimicrobial stewardship and their potential equity benefits, risks, and practical safeguards

Table 1. Long description
The table presents a comparison of AI/ML applications in healthcare epidemiology and antimicrobial stewardship, focusing on equity benefits, risks, and safeguards. It is divided into three main domains: Data capture and EHR processing, HAI risk prediction and management, and Antimicrobial stewardship. Each domain lists specific AI/ML applications, their equity benefits, equity risks, and safeguards. The table has four columns: Domain, AI/ML application, Equity benefits, Equity risks, and Safeguards. Notable applications include natural language processing for extracting social determinants of health, ambient clinical documentation, predictive models for detecting HAI colonization, and AI decision-support for antibiotic prescribing. The table highlights how these applications can both mitigate and exacerbate health disparities, emphasizing the importance of safeguards to ensure equitable outcomes.
Improving data capture and processing on SDoHs
An important bottleneck in identifying health disparities and their causes is the lack of high-quality data on individuals’ race, ethnicity, and SDoHs. Reference Johnson, Moore, Hwang, Hickner and Yeo11,Reference Hatef, Rouhizadeh and Tia12 Currently, EHRs often contain inconsistent or incomplete data on race and ethnicity. Asian, American Indian/Alaskan Native, and Pacific Islanders have a high percentage of misclassification, while Latinx/Hispanic ethnicity has the most incomplete data. Reference Johnson, Moore, Hwang, Hickner and Yeo11 Imputation techniques can be used to replace unrecorded fields; Bayesian methods that use patients’ surnames and geolocation to impute missing race and ethnicity data are common. Reference Elliott, Fremont, Morrison, Pantoja and Lurie13 While these methods are overall robust, challenges remain in classifying certain groups, including American Indian/Alaskan Native and multiracial individuals, as these methods fail for individuals with missing data in either surname or geolocation. Reference Craig, Ji and Zhang14 Even with missing information, ML models trained on name and/or across specific subgroups (ie, racial) have higher accuracy than established methods, Reference Craig, Ji and Zhang14 highlighting the potential of AI/ML methods for record collection and imputation.
Collecting individual-level SDoH data, such as food insecurity, financial strain, or housing instability, is more challenging than collecting race/ethnicity data; although SDoH data are increasingly collected in structured fields, they are often missing altogether or available only in unstructured text, such as provider notes and treatment plans. Reference Hatef, Rouhizadeh and Tia12 Speech recognition and NLP tools are among the most promising AI applications for enhancing information recording and extraction from EHRs. Ambient clinical documentation, an AI tool that generates real-time clinical notes from patient–provider conversations, is the most widely adopted AI tool in healthcare, Reference Poon, Lemak, Rojas, Guptill and Classen15 improving the capture rate of SDoH mentioned in casual conversation. Similarly, NLP models have been developed to extract SDoH from unstructured text. Reference Guevara, Chen and Thomas16 Guevara et al developed large language models (LLM) to extract sparse data on SDoH from clinical notes. Reference Guevara, Chen and Thomas16 The models included both those specifically trained on clinical data and fine-tuned general-purpose ChatGPT-family models. Reference Guevara, Chen and Thomas16 The models trained on the clinical data successfully identified a high percentage of adverse SDoH: 93.8% of patients, compared with standard ICD-10 codes, which captured only 2.0%. When race/ethnicity and gender descriptors were added to the text, ChatGPT models fine-tuned to extract SDoH changed their predictions more frequently than models trained on the clinical corpus alone. Reference Guevara, Chen and Thomas16 Concerns have also been raised about the unequal performance of AI scribes for Black compared with White patients. Reference Zolnoori, Vergez and Xu17 LLM are sensitive to word choice, learning biases, prejudices, and potential racism present during training; integrating differing dialects and vernacular may mitigate this. Reference Zolnoori, Vergez and Xu17 Assuring that these tools perform well across patients is necessary to assure health equity. Their implementation can also exacerbate the digital gap between well-resourced hospitals and those serving underserved populations. Without ensuring fairness, these tools cannot overcome the inequalities and missing data associated with differences in access to care. Reference Gianfrancesco, Tamang, Yazdany and Schmajuk18
HAI risk prediction and management of HAIs in vulnerable individuals
Many ML models are being developed to predict the risks of colonization at admission, bacteremia, and infection-related outcomes for HAI-causing pathogens, as this informs guide screening, isolation, and treatment decisions (Figure 1). Sociodemographic features such as race/ethnicity, or SDoH, are not commonly included as predictive features in these types of models. In particular, the inclusion of race/ethnicity in predictive tools remains an area of active discussion. Reference Gao, Chou and McCaw10 Race and ethnicity are sociopolitical constructs; assigning them as risk factors does not provide a path forward to reduce associated disparities and may perpetuate existing biased views against some demographic groups. Instead, understanding the health disparities encapsulated within these constructs and associated SDoHs is necessary to implement effective interventions. Reference Kilbourne, Switzer, Hyman, Crowley-Matoka and Fine19 On the other hand, the inclusion of race/ethnicity and other sociodemographic features may improve model accuracy, particularly for individuals belonging to multiple vulnerable subgroups. The models can learn biases from other features used to train them that reflect inequalities such as underdiagnosis or underrepresentation. Reference Gao, Chou and McCaw10
Including race/ethnicity helps detect disparities and supports subsequent auditing of the tools. To date, few published models have explicitly evaluated the models’ performance across demographic subgroups defined by race/ethnicity, sex, or socioeconomic status. Reference Rountree, Lin and Liu20 For example, a multicenter study built ML models to predict complicated Clostridioides difficile infection (CDI) outcomes (eg, colectomy) using features available within 48 hours of CDI diagnosis. Reference Berinstein, Steiner and Rifkin21 While the models had good accuracy, they performed worse in non-White patients. Upon model inspection, the poorer performance was attributed to biased features, particularly serum creatinine levels, which are used to estimate glomerular filtration rate and are known to underpredict kidney function in Black patients. Reference Berinstein, Steiner and Rifkin21
ML models, particularly DL-based models, are adept at capturing the complex, nonlinear relationships embedded in the data. However, it can be difficult to track which features drive the model’s predictions, thereby reducing the model’s interpretability (ie, the degree to which a human can understand how the model reaches a prediction), and thus confidence in their use. This can be mitigated to some extent through explainable techniques, stakeholder education, and inclusion during model development. Reference Kim, Hasan and Kellogg22 Explainable techniques provide insights into how input features relate to model outcomes (ie, global explainability) or describe the rationale behind a prediction for a specific patient (ie, local explainability). Reference Bopche, Gustad, Afset, Ehrnström, Damås and Nytrø23
The burden of infections with multidrug-resistant pathogens is high among individuals living in vulnerable areas, where adverse SDoH are highly co-occurring. Reference Dewitt, Reinke and Inman5,Reference Cooper, Beauchamp and Ingle24 Thus, the inclusion of socio-demographics and neighborhood-level vulnerability metrics may improve the predictive value of models aiming at predicting the risk of outcomes and treatments in vulnerable individuals. As a valid address is one of the most complete fields in EHRs, neighborhood-level metrics of SDoHs based on address information are readily available compared to individual-level SDoHs. Reference Hatef, Rouhizadeh and Tia12 While for some health outcomes, such as cardiovascular events, using neighborhood-level metrics of vulnerability as features does not necessarily improve risk prediction beyond the demographic features already recorded in EHR, Reference Bhavsar, Gao, Phelan, Pagidipati and Goldstein25 they may improve predictions related to antimicrobial-resistant pathogens’ burden and risk. For antimicrobial-resistant pathogens, neighborhood-level metrics may provide additional information on individuals’ living conditions that are not captured in existing demographic features and may be relevant to pathogen colonization and transmission. For example, high-poverty neighborhoods have been associated with a higher proportion of urinary tract infections attributed to zoonotic antimicrobial-resistant Escherichia coli. Reference Aziz, Park and Quinlivan26 Diaz et al developed ML models to predict antimicrobial-resistant organisms in blood cultures, facilitating empirical antibiotic decisions. In their models, neighborhood-level metrics of SDoH were ranked as important predictive features, increasing the predictive power of the model when integrated as a feature along with other individual SDoH such as insurance status. Reference Diaz, Cooper and Hanna27 However, the study did not provide a fairness evaluation for subgroups based on race/ethnicity or neighborhood-level metrics.
Health equity in AI tools for antimicrobial stewardship
AI applications are being used to recommend appropriate antibiotic selection, dosing, and duration (for an overview, see Reference AlGain, Marra and Kobayashi28 ). AI decision-making tools have the potential to both perpetuate existing inequities by reinforcing historical treatment patterns and minimize the effects of some of the implicit bias behind health disparities in antibiotic prescribing (Table 1). Historical data contains unequal patterns of prescribing. For example, people of racial or ethnic minority groups are less likely to receive medication for conditions that warrant antibiotics, and less likely to receive broad-spectrum antibiotics. Reference Kim, Kabbani and Dube29 Individuals from rural areas also receive more inappropriate prescriptions, while being white was a predictor of the use of second-line broad-spectrum antibiotics for uncomplicated urinary tract infections under clinicians’ prescription. Reference Kim, Kabbani and Dube29,Reference Kanjilal, Oberst, Boominathan, Zhou, Hooper and Sontag30 But AI tools can also improve antibiotic selection and minimize the inappropriate use of broad-spectrum antibiotics by providing recommendations that are less influenced by providers’ implicit biases, which can lead to prescribing disparities in the first place. Reference Kanjilal, Oberst, Boominathan, Zhou, Hooper and Sontag30,Reference Cavallaro, Moran, Collyer, McCarthy, Green and Keeling31 The supplementary use of an algorithm-based decision-making tool has been shown to reduce the use of broad-spectrum antibiotics and inappropriate antibiotic therapy compared with clinicians’ unaided prescribing. Reference Kanjilal, Oberst, Boominathan, Zhou, Hooper and Sontag30 AI tools may support more standardized decision-making, but only if they are carefully developed, locally validated, prospectively evaluated, and monitored for inequitable effects after implementation.
Biases and fairness during algorithm implementation and deployment
Bias refers to any systematic error or deviation in outcome or performance that can be introduced at any step of the AI lifecycle (Figure 2). Reference Hanna, Pantanowitz and Jackson32 If left unaddressed, these can lead to suboptimal clinical decisions and can perpetuate longstanding healthcare disparities across race, ethnicity, language, and socioeconomic status, among others. Biases can be classified into three overall types: 1) data bias, where data does not equitably represent the population, 2) algorithmic bias, where AI learns and exacerbates instead of reducing inequities, and 3) deployment bias, where improper use of AI perpetuates inequities. Reference Gianfrancesco, Tamang, Yazdany and Schmajuk18 Figure 2 provides specific examples of types of biases within these broad categories. These biases are a threat to building equitable models in healthcare epidemiology and require active mitigation. Reference Gichoya, Thomas and Celi33
Potential biases prevalent in healthcare systems when developing equitable artificial intelligence tools.

Figure 2. Long description
Flowchart illustrating biases in healthcare artificial intelligence tools. The process begins with data collection, which may include representation bias, historic bias, measurement bias, and implicit bias. This leads to data processing, where label bias, class imbalance, and temporal bias can occur. Model development follows, with potential for latent bias, demographic data not included as model features, and proxy variable bias. Evaluation involves aggregation bias and evaluation bias, where the model may not reflect real-world performance. Finally, deployment can introduce dismissal bias, automation bias, and model drift. Each stage is interconnected, showing how biases can propagate and affect the overall system.
Data bias is one of the most pervasive forms of bias in AI models; it originates before model training begins, rooted in how data is collected, labeled, and assembled into training data sets (Figure 2). Because AI models learn primarily from patterns in their training data, biases in that data are transferred to the model. As a result, subgroups without adequate access to healthcare are likely to be underrepresented in the training data, and patients historically under-treated for sepsis or in isolation following contact precautions may continue to experience similar outcomes based on model predictions. Diverse, representative training data are necessary to ensure equitable model performance. One potential strategy to enhance data diversity is federated learning, in which models are trained across multiple institutions. Reference Chen, Wang and Williamson34 Collecting a diverse, representative data set of vulnerable populations is not always feasible, leading to the creation of sampling strategies. These strategies use simulated synthetic data to improve model learning for underrepresented classes. Reference Cary, Zink and Wei35 For example, the Synthetic Minority Over-sampling Technique integrates both oversampling and synthetic resampling, which is helpful for rare outcomes (including HAIs), though such techniques need careful validation to avoid introducing artificial correlations. Reference Chawla, Bowyer, Hall and Kegelmeyer36 Resolving data bias limits the impact of inequities on the overall model, but careful evaluation of fairness must occur before deployment.
Algorithmic bias still raises concerns on model equity, hence why there are are several metrics to measure fairness in models. Reference Gao, Chou and McCaw10 The selection of fairness metrics depends on the specific clinical and equity context. For example, if we want to ensure that a model has equal rates of antibiotic de-escalation recommendations across race/ethnicity, the model should have demographic parity (equal prediction across demographic groups). For high-stakes decisions such as antibiotic initiation for suspected sepsis, ensuring that no demographic subgroup is disproportionately missed is important, thus measures such equalized odds (equal true positive and false positive rate across groups) may be relevant. Fairness metrics indicate bias through variances in performance, but identifying the specific type is key to mitigation and building equitable models. As such, methodological transparency is vital for understanding where bias is introduced. Sharing details about the study design and potential biases starts the conversation about mitigation strategies, a necessary step before creating equitable AI models. Reference Kim, Hasan and Kellogg22,Reference Collins, Moons and Dhiman37 There are several potential downsides of optimizing for fairness, such as that overall model performance may decline, redistribution of errors across groups, or overcorrection; if poorly governed, they may harm other populations or misallocate resources. Reference McCradden, Joshi, Mazwi and Anderson38
The deployment of fair algorithms is not sufficient to have a positive impact on health equity. Reference Kim, Hasan and Kellogg22 Changes over time on patient populations, clinical flows, and prescription practices can degrade model performance (Figure 2). Once algorithms are deployed, fairness audits should be part of routine quality assurance and integrated with existing dashboards and reporting. Clear governance is critical to sustain AI deployment and should be shared by all relevant stakeholders. Reference Kim, Hasan and Kellogg22 Antimicrobial stewardship program leaders may share responsibilities for monitoring downstream effects on antibiotic selection and escalation/de-escalation, integrating algorithm outputs into stewardship review workflows, and evaluating whether algorithm-guided decisions differentially affect subgroups. Infection prevention and healthcare epidemiology teams may assess the effects on surveillance, isolation practices, and HAI metrics, including whether algorithm decisions, are disproportionately burdening specific subgroups. Data scientists and clinical informaticists monitor technical performance, calibration, and data drift and are responsible for conducting routine fairness audits. Hospital leadership is responsible for providing governance, resource allocation, and oversight authority. Inclusion of stakeholders results in addressing bias during model development and deployment, furthering sustained equitable care for patients. With careful curation, AI tools for healthcare epidemiology can bridge the current gaps in equity.
Conclusion
AI/ML has the potential to advance health equity through modeling efforts to improve infection control and antimicrobial stewardship but also exposes gaps relating to bias and perpetuation of disparities. In preventing these disparities, fair AI/ML models, which have little variance in performance when applied to different groups of individuals, are needed. Fairness metrics, or performance assessment metrics for determining bias, are increasingly common as bias is an indicator of unfair models, with the type(s) of bias affecting performance in varying severity, contingent on when the bias was introduced into the AI model. Mitigating these biases necessitates transparency in reporting data and methods, communication/education between healthcare “stakeholders” and modelers, and data-manipulation techniques to ensure minority cases are properly represented in models. By standardizing approaches for identifying, reporting, and tackling bias in AI/ML models, this allows for the formulation of accurate and equitable strategies for infection control, antimicrobial stewardship, and healthcare record management. Carefully leveraging the accuracy of AI/ML models while mitigating for bias can aid pathogen characterization, risk stratification, and predictive modeling, ushering a new wave of equitable and applicable healthcare research.
Acknowledgments
The CDC U01CK000587 partially supported this work. The funder had no role in study design, data collection and analysis, the decision to publish, or the preparation of the manuscript. We would like to thank Sankalp Arya and Sarah Hoover for their helpful feedback during manuscript preparation. Authors have no conflict of interest or disclosures.
