Hostname: page-component-76d6cb85b7-lrvh5 Total loading time: 0 Render date: 2026-07-11T10:46:38.698Z Has data issue: false hasContentIssue false

122 Statistically valid machine learning fairness evaluation

Published online by Cambridge University Press:  08 May 2026

Momin Malik
Affiliation:
Mayo Clinic
Dave Watson
Affiliation:
Mayo Clinic
Madison J. Beenken
Affiliation:
Mayo Clinic
Chung-Il Wi
Affiliation:
Mayo Clinic
Young J. Juhn
Affiliation:
Mayo Clinic
Rights & Permissions [Opens in a new window]

Abstract

Core share and HTML view are not available for this content. However, as you have access to this content, a full PDF is available via the 'Save PDF' action button.

Objectives/Goals: Within the machine learning “fairness” literature, models audits are point estimates, and only effect size is considered. The well-developed frameworks in biostatistics for diagnostic testing apply directly to classifiers more generally; these principles can be extended to have statistically valid fairness audits. Methods/Study Population: We made original linkages between machine learning methodology and biostatistics, particularly to diagnostic testing, to get analytic biostatistical methods for model evaluation that are applicable to fairness testing. Results/Anticipated Results: The use of odds ratios to compare ratios of binomial proportions is superior to taking ratios of binomial proportions directly, as is currently standard in machine learning. Odds ratios also have an analytic asymptotic approximation for standard errors, by which we can test if a point estimate (effect size) is greater than what we would expect from noise. In diagnostic testing, the odds ratios considered are generally across the cells of a confusion matrix, but we can form odds ratios for specific marginal quantities (positive predictive value, negative predictive value, true positive rate, true negative rate) for specific aspects of model performance or fairness. Discussion/Significance of Impact: Biostatistics can make direct contributions to instantly improve the state of machine learning practice, giving ready-made methods that can take account generations of hard-won lessons about the importance of uncertainty quantification and sample size in making claims and conclusions.

Information

Type
Biostatistics, epidemiology, and research design
Creative Commons
Creative Common License - CCCreative Common License - BYCreative Common License - NCCreative Common License - ND
This is an Open Access article, distributed under the terms of the Creative Commons Attribution-NonCommercial-NoDerivatives licence (https://creativecommons.org/licenses/by-nc-nd/4.0/), which permits non-commercial re-use, distribution, and reproduction in any medium, provided the original work is unaltered and is properly cited. The written permission of Cambridge University Press must be obtained for commercial re-use or in order to create a derivative work.
Copyright
© The Author(s), 2026. The Association for Clinical and Translational Science