Introduction
Immigration has ranked among the most salient political issues across Western democracies over the past decade (European Commission 2015, 2016; Ipsos 2024), and research on the topic has correspondingly surged in political science (Grossman et al. Reference Grossman, Dinneen and Torreblanca2025). Public opposition is widespread: across Western Europe, large majorities in Germany (81 per cent), Spain (80 per cent), Sweden (73 per cent), Italy (71 per cent), and the United Kingdom (71 per cent) say immigration over the past decade has been ‘too high’ (YouGov 2025). A central driver of this backlash is concern about crime and security. Across Europe, negative evaluations of immigration’s impact on crime are more widespread than negative evaluations of its economic or cultural effects (Figure 1; see also Pew Research Center 2024; YouGov 2019). Populist and radical-right parties on both sides of the Atlantic have capitalized on these fears, placing the immigration-crime nexus at the center of campaigns for stricter border controls (Alizade Reference Alizade2025; Mudde Reference Mudde2007; Wodak Reference Wodak2015). Consistent with this, crime-related immigration concerns are a strong predictor of far-right support (see SI Figure A.3 for evidence from Germany).
Concerns about immigration across Europe.
Notes: The figure shows the share of respondents expressing negative evaluations of immigration’s impact on (i) crime, (ii) the economy, and (iii) cultural life, using cross-national survey data from the European Social Survey (ESS ERIC 2014). Responses are measured on 11-point scales coded 0–10, where 0 denotes the most negative assessment and 10 the most positive. ‘High concern’ (strongly negative evaluation) is defined as responses 0–3. The survey items are: Crime (immigration makes crime problems worse v. better), Economy (immigration is bad v. good for the economy), and Culture (cultural life is undermined v. enriched by immigration). Estimates are weighted using ESS post-stratification weights. In Estonia, 24.6 per cent of respondents fall into the ‘high concern’ category on both the crime and economy items; as a result, the markers for these two estimates overlap exactly in the plot.

Given the prominence of the topic in public discourse, a growing body of research examines the link between immigration and crime empirically. One of the most common empirical strategies to study the relationship between immigration and crime is to aggregate offenses to a geographic unit – counties, municipalities, or policing districts – and then regress changes in a given unit’s aggregate crime rates on (plausibly exogenous) changes in the same unit’s immigrant share. Variants of this design underpin many of the most highly cited contributions to this literature (see Table 1). The findings from these studies – predominantly null resultsFootnote 1 – have not only been published in leading academic journals but have also influenced broader public discourse through citations in think tanks and major media outlets (see, for example, The Guardian 2013; Tagesschau 2025).
Overview of published research on the effects of immigration on crime

Notes: The table lists studies that use aggregate crime rates as the outcome and relate them to area-level immigration. Most studies examine immigration inflows or changes in immigrant stocks; Miles and Cox (Reference Miles and Cox2014) instead study the rollout of an immigration enforcement program (Secure Communities) but likewise relies on aggregate crime as the dependent variable, so the dilution logic discussed in the main text applies there as well. For Spenkuch (Reference Spenkuch2014), the ‘null’ classification is based on the Instrumental variable (IV) estimates. Effects on total crime reported by Lange and Sommerfeld (Reference Lange and Sommerfeld2024) are not statistically significant at conventional levels for most specifications and are therefore classified as ‘mixed’. Classification of Kayaoglu (Reference Kayaoglu2022) is based on difference in differences estimates.
In this research note, I argue that such designs are limited in what they can tell us about the immigration-crime debate due to what I call the ‘dilution problem’. Because immigrant populations typically represent a small share of the total population, their specific crime rates (be they higher, lower, or equivalent to native rates) are arithmetically diluted within aggregate crime statistics. This issue is often amplified because many studies examine over-time changes in immigrant populations, and these incremental changes are generally even smaller than the already modest immigrant population stocks. As a result, when the motivating question is a group-level comparison – the immigrant–native crime gap – aggregate designs typically rule out only implausibly large differentials while remaining compatible with a broad range of substantively meaningful gaps. By contrast, if the primary goal is to study how immigration affects native-on-native crime more broadly, aggregate designs can be adequately powered, but this mechanism is generally not the central focus or motivation of the studies reviewed in Table 1.
I develop the argument in two steps. First, I formalize the dilution problem by decomposing the total effect of immigration shocks on aggregate crime into a compositional component (the immigrant-native crime differential) and a behavioral component that captures changes in native offending rates. Based on a systematic review of the immigration-crime literature, I find that the compositional channel has been the dominant focus and interpretive lens in prior work. I then derive a general expression for the minimum detectable gap (r MD ), which quantifies the smallest immigrant-native crime differential that a standard aggregate design can reliably distinguish from zero. Secondly, I calibrate a Monte Carlo simulation to real-world immigration shock and crime data across German counties between 2014 and 2016, that is, before and after the large-scale refugee influx in 2015. The simulation shows that aggregate designs reach conventional levels of statistical power only under extreme conditions – for example, when immigration shocks comparable to Germany’s 2015 inflow are combined with immigrant crime rates at least 15× the native rate.
Crucially, this paper takes no position on whether immigrants are more, less, or equally prone to crime than natives. Immigrants are a highly heterogeneous population, and any group-level differences in offending will depend on the demographic mix of a given inflow. For example, a large body of research documents a strong relationship between age and crime, with offending rates peaking in late adolescence and early adulthood, and young men consistently accounting for a disproportionate share of crime (Falk et al. Reference Falk, Wallinius, Lundström, Frisell, Anckarsäter and Kerekes2014; Farrington Reference Farrington1986; Hirschi and Gottfredson Reference Hirschi and Gottfredson1983; Moffitt Reference Moffitt2018). The precise shape and timing of this age–crime relationship varies across social and national contexts (Steffensmeier and Schwartz Reference Steffensmeier and Schwartz2025). As a result, the aggregate impact of a particular immigrant inflow on crime will depend heavily on the share of young people, the share of young men in particular, the origin and reception context, and immigrants’ prior experiences of violence and marginalization (Couttenier et al. Reference Couttenier, Petrencu, Rohner and Thoenig2019).Footnote 2
The dilution problem identified in this paper is broadly related to, but distinct from, the ecological inference (EI) literature. Whereas most EI research concentrates on bias – and methods for correcting it – when drawing group-level inferences from aggregate data (King et al. Reference King, Tanner and Rosen2004), my approach sets bias aside and instead highlights a different, arithmetic constraint: the dilution of group effects within large populations. In this respect, the analysis is most closely aligned with the ‘bounding’ tradition in ecological inference (Duncan and Davis Reference Duncan and Davis1953), which derives the feasible range of group-specific averages from observed aggregates. When the key parameter of interest relates to a small fraction of the population, these bounds become so wide as to be practically uninformative. While this limitation is well understood in theory, it remains a persistent and largely overlooked feature in studies of immigration and crime, affecting the results of much of the published research in this field (Table 1). As Card notes, design-based work generally requires ‘big shocks’ to cut through the statistical noise in aggregate data (Card Reference Card2021) – a condition that immigration inflows or restrictions rarely meet.
This paper contributes to a growing literature documenting the prevalence of underpowered research designs in political science and the social sciences more broadly (Arel-Bundock et al. Reference Arel-Bundock, Briggs, Doucouliagos, Aviña and Stanley2026; Weiss Reference Weiss2024). My argument that dilution makes the aggregate signal for the immigrant-native crime differential r extremely weak provides an illustration of this issue in a key research area, namely the relationship between immigration and crime. In this sense, many immigration-crime studies are ‘null by design’: even under generous assumptions, the empirical strategy can at best rule out the most extreme effects, while leaving both zero and a wide range of substantively important immigrant-native gaps observationally indistinguishable. This result parallels recent simulation-based evidence from other domains, such as cross-national studies on the effects of democracy, which are similarly only powered to detect implausibly large effect sizes (Doucette Reference Doucette2025). As in that line of research, the implication is that many published null findings are best understood as artifacts of research design rather than strong evidence that the underlying effects are near zero.
The Dilution Problem
In this section, I formalize the dilution problem in aggregate crime-on-immigration designs. I first clarify what these designs estimate by decomposing the total effect of changes in the immigrant share on aggregate crime into a compositional component (the immigrant-native crime differential, r) and a behavioral component (changes in native crime rates due to immigration). I then argue that, in practice, the compositional channel r is the primary substantive channel of interest in this literature. I then derive the minimum detectable gap: the smallest immigrant-native crime differential (r) that any consistent estimator of the aggregate slope can detect at conventional levels of statistical power (Equation (3)). I then use a Monte Carlo simulation, calibrated to real-world refugee inflows and county-level crime data, to illustrate this issue under realistic conditions.
Set-up
Let i = 1, …, n index geographic units and t = 0, …, T time. For each (i, t) we observe
where N it is total population, I it is the number of immigrants, and C it is total reported crimes.
Let c N and c I denote per-capita crime propensities for natives and immigrants, and define the immigrant-native crime gap.
For a given immigrant share S it , aggregate crime per capita can be written as
where ϵ it captures shocks orthogonal to immigration.
Estimand and Mechanisms
While the estimand is rarely defined explicitly in the studies listed in Table 1 (see SI Tables B.2 and B.3), their empirical set-ups implicitly target the total effect of changes in the immigrant population share S it on overall crime in an area. We can formalize this as the average change in per-capita crime rates as the immigrant share increases:
To connect this total effect to the composition identity in Equation (1), I allow the crime propensity of the pre-existing population (‘natives’) to depend on the immigrant share, while treating the immigrant crime propensity c I as fixed for a given inflow:
Differentiating with respect to S it Footnote 3 gives
$\eqalign{& \theta = {{\partial {y_{it}}} \over {\partial {S_{it}}}} \cr & \,\,\, = \underbrace {\left( {{c_I} - {c_N}\left( {{S_{it}}} \right)} \right)}_{{\rm{composition{\,\,}channel}}{\,\,}r\left( {{S_{it}}} \right)} + \underbrace {\left( {1 - {S_{it}}} \right){{\partial {c_N}} \over {\partial {S_{it}}}}}_{{\rm{native{\,\,}behavioral{\,\,}channel}}}{\!\!\!\!\!\!\!}. \cr}$
Equation (2) shows that the implied estimand in these designs is a total effect of immigrant share on aggregate crime that bundles two distinct mechanisms:
-
Compositional channel (r(S it ) = c I − c N (S it )): the difference in crime propensities between immigrants and natives. This channel includes both selection/composition differences (for example, age–sex structure, socio-economic disadvantage) and post-migration incentive differences (for example, differences in crime rates that arise from immigrants’ opportunity costs or expected punishment costs such as deportation risk (Becker Reference Becker1968; Ehrlich Reference Ehrlich1973)). Under the assumption of a muted native behavioral response (see below), this reduces to a constant gap, r = c I − c N .
-
Native behavioral response ((1−S it ) ∂c N /∂S it ): the degree to which the arrival of new immigrants changes the crime propensity of the native population. Mathematically, a non-zero native response (
${{\partial {c_N}} \over {\partial S}} \ne 0$
) affects the estimates in two ways: (i) it alters the crime gap r(S
it
) itself, as the baseline against which immigrants are compared shifts, and (ii) it drives changes in the overall crime rate θ directly.
In principle, a coefficient of large magnitude βˆ in an aggregate crime regression could be driven by either of the two channels in Equation (2), or a combination of the two. In practice, however, the published literature predominantly motivates and interprets these estimates through the compositional channel: the immigrant-native crime differential r. With the exception of one study that does not discuss any theoretical mechanism explicitly, every study in Table 1 discusses a variant of the compositional channel (see SI Tables B.2 and B.3).
By contrast, the native behavioral channel is less central. When native responses are discussed explicitly, they usually take one of two forms. The first is changes in native-on-immigrant offending, in particular xenophobic hate crimes. Importantly, such offenses make up a vanishingly small share of aggregate crime: for example, in 2024 the FBI recorded 6,328 hate crime incidents motivated by race or ethnicity, compared to roughly 14 million total offenses in the United States (less than 0.05% (Federal Bureau of Investigation 2024a, 2024b)). As a result, even large relative increases in hate crimes are difficult to detect when the dependent variable is the overall crime rate. When xenophobic violence is the behavioral mechanism of interest, aggregate crime regressions therefore face a similar dilution problem.
The second channel is a more general shift in native offending rates, for example, through labor-market displacement, neighborhood change, social disorganization, or policing. Unlike the compositional channel, this mechanism is not arithmetically constrained by the small size of the immigrant population, so an aggregate design could, in principle, be well-powered to detect it. Empirically, however, such general native responses are discussed in only about half of the studies in Table 1, and even in those cases, they are often framed as ancillary rather than as the central mechanism under test (see SI Section B). For example, Chalfin (Reference Chalfin2014) notes that immigration could, in principle, ‘change the calculus of offending among U.S. natives’, but also argues that his estimates ‘likely represent an upper bound on the criminality of immigrants’. Similarly, Bianchi et al. (Reference Bianchi, Buonanno and Pinotti2012) discuss the possibility of a native behavioral channel via labor-market competition, but argue that it is likely ‘less of an issue’ in the Italian context due to complementarities.
Because the interpretation of aggregate estimates is, in most cases, linked to what they imply about immigrants’ crime propensity relative to natives, I evaluate the design against that specific inferential goal. Throughout, I assume that the native behavioral response is muted (that is, ∂c N /∂S it ≈ 0), such that native crime rates are not changed as a result of immigration inflows. Under this benchmark, the second term in Equation (2) vanishes, and the estimand reduces to the immigrant-native crime gap:
The remainder of the paper focuses on the structural limitations of learning about this parameter r using designs where aggregate-level crime rates are regressed on aggregate-level changes in immigrant population shares.
The Minimum Detectable Gap
Consider any consistent estimator βˆ of the slope in a regression of y it on S it .Footnote 4 Under the benchmark assumption motivated above, variation in S it affects aggregate crime only through the compositional difference between immigrants and natives. It follows that the population slope in a regression of y it on S it identifies the immigrant-native crime gap:
If immigrants and natives offend at the same rate (r = 0), the regression slope converges to zero; in this case, immigration-driven compositional changes leave the aggregate per-capita crime rate unchanged. By contrast, a non-zero differential (r ≠ 0) implies that an influx of immigrants mechanically shifts the aggregate rate. Since y it = c N + rS it + ϵ it , a change in immigrant share ΔS it shifts the expected crime rate by
I define the minimum detectable gap, r MD, as the smallest immigrant-native crime differential that a standard two-sided test can detect with pre-specified power. Consider testing H 0 : r = 0 at significance level α using the usual Wald test statistic
$W\; \equiv \;{{\hat \beta } \over {{\mathop{\rm se}\nolimits} (\hat \beta )}}.$
Under the large-sample approximation, W is approximately standard normal under the null, so we reject when
where z
p
is the pth quantile of
$\mathcal{N}$
(0,1). When r ≠ 0, βˆ is approximately normal with mean r and standard deviation se(βˆ), that is,
so the rejection probability (power) is
Let π denote the target power. Defining r MD as the value of r that achieves power π yieldsFootnote 5
For standard choices of α = 0.05 and π = 0.8, z 1 − α/2 ≈ 1.96 and z π ≈ 0.84, hence
r MD is the smallest immigrant-native crime differential that the design detects with 80 per cent power at the 5 per cent level. Equivalently, r MD is a monotone rescaling of the confidence-interval width: the 95 per cent confidence-interval half-width corresponds to the special case π ≈ 0.5 (50 per cent power).Footnote 6
The magnitude of se(βˆ) (and thus r MD) depends on the properties of the research design. As shown in SI Section C.5, precision improves with a larger sample size and with greater within-unit variation in immigrant shares (that is, stronger immigration ‘shocks’). Conversely, the standard error increases when baseline crime rates are higher and when idiosyncratic noise and other unmodeled determinants of crime are larger.
In SI Section D, I show that this dilution problem severely limits what can be learned from published research about r. Specifically, I examine a study by Masterson and Yasenov (Reference Masterson and Yasenov2021) published in the American Political Science Review. Using their reported point estimates and standard errors, I recover the implied 95 per cent confidence intervals for the aggregate effect of refugee share on crime and translate the upper endpoints of those intervals into the corresponding bounds on the immigrant-native crime differential. The range of parameter values compatible with the confidence intervals includes scenarios where refugees commit crime at approximately 22× (property) and 28× (violent) the native rate.Footnote 7
Monte Carlo Simulation
To illustrate the dilution problem under realistic conditions, I run a Monte Carlo simulation calibrated to German county-level data around the large-scale 2015 refugee inflow (2014–16). The simulation asks which combinations of (i) the size of the refugee immigration shock and (ii) the immigrant-native crime gap r would allow a standard aggregate difference-in-differences (first-difference) design to reliably reject the null hypothesis. Additional details are provided in SI Section E.
I consider a standard two-period first-difference regression,
where i indexes counties, Δy
i
is the change in total recorded crimes per capita, and ΔS
i
is the change in the county refugee share between 2014 and 2016. The simulation uses three empirically calibrated inputs. First, I use county-level refugee registry data (2014 and 2016) to compute the mean change in refugee share and its variation across counties (summarized by
${\rm CV}_S=\operatorname {\rm{sd}}(\Delta S_i)/\overline {\Delta S}$
). Secondly, I use official police crime statistics (PKS) and county population counts for 2014 to calculate the baseline crime rate c
N
, defined as the population-weighted mean of county-level per-capita crime rates in the pre-period (2014). Thirdly, I calibrate the residual volatility in crime changes unrelated to refugee inflows, σ
Δϵ
, as the standard deviation of residuals from regressing observed Δy
i
on observed ΔS
i
between 2014 and 2016.
Statistical power heatmap.
Notes: The heatmap shows the probability of rejecting the null for a two-sided test (α = 0.05) of H 0 : β = 0 in a two-period first-difference model (see SI Section E for details). The horizontal axis represents the average county-level change in the immigrant share, E[ΔS i ]. The vertical axis represents the immigrant–native crime rate ratio (c I /c N ). Power is estimated using 1,000 Monte Carlo simulations per grid point, calibrated to German county-level data as discussed in the main text.

I vary the average immigration shock ΔS avg ∈ [0, 0.01] (0 to 1 percentage point of the total population) and parameterize the gap as r = k × c N with k ∈ [0, 30], so that c I /c N = 1 + k. For each grid point, I generate M = 1, 000 datasets by drawing county immigration shocks with dispersion matching the calibration data and adding idiosyncratic noise calibrated by σ Δϵ . Then, for each simulation, I estimate the first-difference regression by ordinary least squares and record whether a two-sided test of H 0 : β = 0 rejects at α = 0.05.
Figure 2 plots statistical power as a function of the average immigration shock and the immigrant-native crime rate ratio c I /c N . The results show that conventional designs lack sufficient precision except at the extremes: the design can only reliably distinguish the crime gap from zero if the immigration shock is both extremely largeFootnote 8 – on the order of Germany’s post-2015 refugee inflow (≈ 1 per cent on average) – and accompanied by immigrant crime propensities that are substantially higher (at least 15 times) than the native rate. For smaller, more typical shocks, statistical power remains low. Thus, under realistic conditions, the design will tend to yield results that do not reach conventional levels of statistical significance even when group-level differences in offending are substantial. These estimates likely understate the extent of the issue in applied research because they assume the complete absence of measurement error. Allowing for measurement error in crime rates or immigrant shares further reduces precision (see SI Section F).
Discussion
The dilution problem documented in this paper implies that designs using aggregate crime rates and changes in immigrant population shares will typically be uninformative about immigrant-native crime differentials: they tend to produce wide confidence intervals that include zero and therefore yield ‘null’ results regardless of whether immigrants are more, less, or equally crime-prone than natives.
Accordingly, when the mechanism of interest is the compositional gap r, researchers should rely on more granular data – either individual-level crime records (see, for example, Abramitzky et al. Reference Abramitzky, Boustan, Jácome, Pérez and Torres2024; Butcher and Piehl, Reference Butcher and Piehl1998; Freedman et al. Reference Freedman, Owens and Bohn2018; Pinotti Reference Pinotti2017; Sampson et al. Reference Sampson, Morenoff and Raudenbush2005) or, where available, administrative statistics disaggregated by nationality or citizenship (see, for example, Couttenier et al. Reference Couttenier, Petrencu, Rohner and Thoenig2019; Dehos Reference Dehos2021) – which allow r to be measured directly rather than inferred from aggregates.Footnote 9
Aggregate designs are most useful when the mechanism of interest is the native behavioral response to immigration. This can take the form of xenophobic violence and hate crimes directed at immigrants, or broader changes in native offending driven by social disorganization, labor-market displacement, or related channels (see, for example, Borjas et al. Reference Borjas, Grogger and Hanson2010). In the first case, the dependent variable should be restricted to xenophobic hate crimes rather than total crime, because hate crimes make up only a tiny fraction of all offenses and their signal would otherwise be diluted in aggregate data (see, for example, Müller and Schwarz Reference Müller and Schwarz2021). In the second case, since natives comprise the vast majority of the population, aggregate crime rates effectively track native behavior, so regressions of crime on immigrant share can, in principle, be adequately powered to detect shifts in native offending. Ultimately, researchers should be explicit about the mechanism they aim to identify, because this determines the appropriate empirical strategy. While the aggregate designs critiqued in this research note may be informative about how immigration affects native crime rates, they are structurally ill-suited for learning about immigrant-native crime differentials.
Supplementary material
The supplementary material for this article can be found at https://doi.org/10.1017/S0007123426101525.
Data availability statement
Replication data for this paper can be found at https://doi.org/10.7910/DVN/FJFWEI.
Acknowledgements
I thank three anonymous reviewers for valuable comments and suggestions.
Financial support
This research received no specific grant from any funding agency, commercial, or not-for-profit sectors.
Competing interests
The author declares no competing interests.
Artificial intelligence disclosure
The author used AI tools to assist with (i) data cleaning and code debugging, (ii) proofreading and copyediting, and (iii) literature search. All substantive decisions are the author’s own.
