Hostname: page-component-76d6cb85b7-vdhp9 Total loading time: 0 Render date: 2026-07-12T17:04:54.141Z Has data issue: false hasContentIssue false

121 Study design considerations for post-deployment monitoring of AI/ML models

Published online by Cambridge University Press:  08 May 2026

Momin Malik
Affiliation:
Mayo Clinic
Dave Watson
Affiliation:
Mayo Clinic
Madison J. Beenken
Affiliation:
Mayo Clinic
Shauna M. Overgaard
Affiliation:
Mayo Clinic
Hanyin Wang
Affiliation:
Mayo Clinic
Imad Absah
Affiliation:
Mayo Clinic
Chung Il Wi
Affiliation:
Mayo Clinic
Young J. Juhn
Affiliation:
Mayo Clinic
Rights & Permissions [Opens in a new window]

Abstract

Core share and HTML view are not available for this content. However, as you have access to this content, a full PDF is available via the 'Save PDF' action button.

Objectives/Goals: Safely and effectively translating AI classification models requires robust post-deployment monitoring. Yet there is little guidance about how to do so. Here, we outline how standard machine learning model development study designs are insufficient for post-deployment monitoring, and what specific study designs are needed. Methods/Study Population: We made original linkages between machine learning methodology and biostatistics, particularly to diagnostic testing, for study design guidance on post-deployment model performance and fairness assessment. Results/Anticipated Results: The kind of case–control sampling typically used for model development cannot give valid estimates of the Positive Predictive Value (precision) and may suffer from verification bias; in a pragmatic clinical trial, or under deployment, only the PPV can be measured, unless there is random confirmatory testing of predicted negatives. This is important for measuring sensitivity, specificity, and False Omission Rate as a fairness metric. Discussion/Significance of Impact: By linking well-understood biostatistics principles (e.g., diagnostic testing) to machine learning classifier evaluation, we show what is required for rigorous evaluation of models in clinical practice. This will establish clear guidance and rigorous standards of practice for study design and analytic aspects of post-deployment monitoring.

Information

Type
Biostatistics, epidemiology, and research design
Creative Commons
Creative Common License - CCCreative Common License - BYCreative Common License - NCCreative Common License - ND
This is an Open Access article, distributed under the terms of the Creative Commons Attribution-NonCommercial-NoDerivatives licence (https://creativecommons.org/licenses/by-nc-nd/4.0/), which permits non-commercial re-use, distribution, and reproduction in any medium, provided the original work is unaltered and is properly cited. The written permission of Cambridge University Press must be obtained for commercial re-use or in order to create a derivative work.
Copyright
© The Author(s), 2026. The Association for Clinical and Translational Science