Hostname: page-component-76d6cb85b7-5qg8f Total loading time: 0 Render date: 2026-07-20T09:59:13.745Z Has data issue: false hasContentIssue false

A machine learning approach to identify at-risk populations for yellow fever vaccination

Published online by Cambridge University Press:  01 June 2026

Madison Farnsworth*
Affiliation:
Department of Human Pathophysiology and Translational Medicine, Institute for Translational Science, The University of Texas Medical Branch at Galveston, Galveston, TX, United States
Kostiantyn Botnar
Affiliation:
Department of Pharmacology & Toxicology, The University of Texas Medical Branch, Galveston, TX, United States
Justin T. Nguyen
Affiliation:
Department of Biochemistry and Molecular Biology, The University of Texas Medical Branch, Galveston, TX, United States
Riley K. Watson
Affiliation:
Department of Human Pathophysiology and Translational Medicine, Institute for Translational Science, The University of Texas Medical Branch at Galveston, Galveston, TX, United States
Trevor L. Murphy
Affiliation:
John Sealy School of Medicine, The University of Texas Medical Branch, Galveston, TX, United States
Susan L.F. McLellan
Affiliation:
Division of Infectious Diseases, Department of Internal Medicine, The University of Texas Medical Branch, Galveston, TX, United States
Kamil Khanipov
Affiliation:
Department of Pharmacology & Toxicology, The University of Texas Medical Branch, Galveston, TX, United States
George Golovko
Affiliation:
Department of Pharmacology & Toxicology, The University of Texas Medical Branch, Galveston, TX, United States
*
Corresponding author: M. Farnsworth; Email: madisonfarnsworth@byu.edu
Rights & Permissions [Opens in a new window]

Abstract

Introduction:

Vaccinations play a crucial role in public and personal health. Vaccines such as the Yellow Fever 17D vaccine are effective at preventing disease and provide lifelong immunity with few adverse events. Leveraging the increase in electronic medical records and the widespread use of machine learning algorithms in healthcare, this study aims to investigate and develop an algorithmic framework for analyzing clinical data on Yellow Fever 17D vaccinations in order to predict at-risk populations for adverse events following immunizations. This research incorporates parameters from the patient’s medical history and demographic information.

Methods:

To build our analytic framework, we tested five machine learning algorithms: random forest, gradient boost, logistic regression, XGBoost, and Bernoulli naïve Bayes. We cleaned and managed the datasets through EHRchitect software. We assessed the performance of the algorithms using standard metrics, including precision, recall, F1-scores, accuracy, and the area under the receiver operating characteristic curve. Additionally, we implemented population sampling to mitigate potential biases common in clinical data and performed parameter ranking analysis to identify the most influential features in classifying patient reaction outcomes.

Results:

We found that logistic regression produced the highest performance, and through hyperparameter optimization, we developed a framework with precision, accuracy, and ROC scores of 87.7%, 99.0%, and 0.818, respectively.

Conclusion:

In conclusion, this study examined five modeling algorithms to establish a framework for analyzing real-world data and to predict at-risk populations for adverse events following immunization. We incorporated a logistic regression algorithm to assess our clinical data to achieve a high-precision performance model for classifying patient predictions.

Information

Type
Research Article
Creative Commons
Creative Common License - CCCreative Common License - BYCreative Common License - NC
This is an Open Access article, distributed under the terms of the Creative Commons Attribution-NonCommercial licence (https://creativecommons.org/licenses/by-nc/4.0/), which permits non-commercial re-use, distribution, and reproduction in any medium, provided the original article is properly cited. The written permission of Cambridge University Press or the rights holder(s) must be obtained prior to any commercial use.
Copyright
© The Author(s), 2026. Published by Cambridge University Press on behalf of Association for Clinical and Translational Science
Figure 0

Figure 1. Figure 1 long description.(a) Flowchart of the record selection process from the TriNetX database. The initial population of the entire US Collaborative network database for the two cohorts was determined by adverse patient outcomes. (b) Workflow for comparative analysis of the five ML algorithms and the necessary data quality, harmonization, and selection. The selected population was split into two different datasets to train and validate the chosen models.

Figure 1

Figure 2. Baseline parameters of study participants in the severe and non-severe cohorts. The population percentage within their respective cohorts is expressed on the x-axis. Patient records of those who possess the demographic parameter or confirm a diagnostic history of the clinical parameters are shown in this graph. Conditions, including hyperlipidemia, cardiovascular disease, HIV, and diabetes, were more prevalent in the adverse event cohort, reflecting baseline clinical differences leveraged by machine learning models for risk discrimination.

Figure 2

Table 1. The initial performance results of the five modeling algorithms. The performance matrix can be evaluated based on singular cohort predictions and on the overall model predictions

Figure 3

Table 2. The performance results of the five modeling algorithms with random undersampling. Random undersampling extracts a smaller ratio of the patient population to present a more balanced cohort. The sampling is repeated in iterations to involve all data points within both cohorts. The sampling strategy ratio is a result of the cohort population ratio of the resampling subsets; 0.1, 0.5, and 1.0 indicate a ratio of 10:1, 2:1, and 1:1 of the non-severe and severe cohorts, respectivelyTable 2 long description.

Figure 4

Table 3. The performance results of the five modeling algorithms with hyperparameters for each respective model, with integrated stratified K-fold cross-validation with 10 iterations

Figure 5

Table 4. The final performance results of the logistic regression validation model test

Figure 6

Figure 3. The SHAP results for the parameters used in the LR modeling algorithm. In correlation to the color legend, higher values within the parameter are represented by the color red, and lower values are represented by the color blue; since all parameters (except age_group) were converted to binary values, then all red indicates the presence of the parameter in patients record and blue represents the lack of the presence of the parameter in the patient record. The distance from the black zero line indicates the importance of the parameter to the classification of patient records. Highly influential parameters will dramatically change the score, ranking them with a higher Shapley value.

Supplementary material: File

Farnsworth et al. supplementary material 1

Farnsworth et al. supplementary material
Download Farnsworth et al. supplementary material 1(File)
File 16.4 KB
Supplementary material: File

Farnsworth et al. supplementary material 2

Farnsworth et al. supplementary material
Download Farnsworth et al. supplementary material 2(File)
File 33.4 KB
Supplementary material: File

Farnsworth et al. supplementary material 3

Farnsworth et al. supplementary material
Download Farnsworth et al. supplementary material 3(File)
File 16.8 KB
Supplementary material: File

Farnsworth et al. supplementary material 4

Farnsworth et al. supplementary material
Download Farnsworth et al. supplementary material 4(File)
File 20.1 KB