Hostname: page-component-76d6cb85b7-rxvq6 Total loading time: 0 Render date: 2026-07-19T09:39:59.979Z Has data issue: false hasContentIssue false

Measuring Descriptive Representation at Scale: Methods for Predicting the Race and Ethnicity of Public Officials

Published online by Cambridge University Press:  18 August 2025

Diana Da In Lee*
Affiliation:
Center for the Study of Democratic Politics, Princeton University, Princeton, NJ, USA
Yamil Ricardo Velez
Affiliation:
Department of Political Science, Columbia University, New York, NY, USA
*
Corresponding author: Diana Da In Lee; Email: dl9388@princeton.edu
Rights & Permissions [Opens in a new window]

Abstract

Ethnicity and race are vital for understanding representation, yet individual-level data are often unavailable. Recent methodological advances have allowed researchers to impute racial and ethnic classifications based on publicly available information, but predictions vary in their accuracy and can introduce statistical biases in downstream analyses. We provide an overview of common estimation methods, including Bayesian approaches and machine learning techniques that use names or images as inputs. We propose and test a hybrid approach that combines surname-based Bayesian estimation with the use of publicly available images in a convolutional neural network. We find that the proposed approach not only reduces statistical bias in downstream analyses but also improves accuracy in a sample of over 16,000 local elected officials. We conclude with a discussion of caveats and describe settings where the hybrid approach is especially suitable.

Information

Type
Article
Creative Commons
Creative Common License - CCCreative Common License - BY
This is an Open Access article, distributed under the terms of the Creative Commons Attribution licence (https://creativecommons.org/licenses/by/4.0/), which permits unrestricted re-use, distribution and reproduction, provided the original article is properly cited.
Copyright
© The Author(s), 2025. Published by Cambridge University Press
Figure 0

Table 1. Assessment of existing prediction methods and the proposed hybrid approach. The ‘hybrid method’ is used as shorthand to describe the proposed prediction method introduced in this paper.

Figure 1

Table 2. Overall classification error, Type I error and Type II error rates for each ethnoracial category across six prediction methods. The lowest error rate in each row is in bold. Results are based on out-of-sample data ($N$=4,976)

Figure 2

Figure 1. Tetrahedron diagrams of predicted probabilities across four ethnoracial categories (A = Asian, B = Black, H = Hispanic/Latino, and W = white).

Figure 3

Figure 2. ROC curves across different ethnoracial prediction methods. The AUC scores for each method are shown within each plot.

Figure 4

Figure 3. The proportion of non-white officials over time predicted with different ethnoracial category prediction methods.

Figure 5

Figure 4. Multinomial logistic regression results. The solid black line and shaded grey area represent estimates and 95 percent CIs from the benchmark model. See Appendix H for full regression results.

Supplementary material: File

Lee and Velez supplementary material

Lee and Velez supplementary material
Download Lee and Velez supplementary material(File)
File 14.4 MB
Supplementary material: Link

Lee and Velez Dataset

Link