To save content items to your account,
please confirm that you agree to abide by our usage policies.
If this is the first time you use this feature, you will be asked to authorise Cambridge Core to connect with your account.
Find out more about saving content to .
To save content items to your Kindle, first ensure no-reply@cambridge.org
is added to your Approved Personal Document E-mail List under your Personal Document Settings
on the Manage Your Content and Devices page of your Amazon account. Then enter the ‘name’ part
of your Kindle email address below.
Find out more about saving to your Kindle.
Note you can select to save to either the @free.kindle.com or @kindle.com variations.
‘@free.kindle.com’ emails are free but can only be saved to your device when it is connected to wi-fi.
‘@kindle.com’ emails can be delivered even when you are not connected to wi-fi, but note that service fees apply.
This chapter explores three kinds of unsupervised task: clustering, density estimation and dimensionality reduction. Cluster analysis aims to group similar observations together. The K-means algorithm does this by repeatedly reassigning each point to the nearest cluster centre, reducing or maintaining the clustering inertia at each step. Density estimation involves learning a probabilistic model of a data-generating process. Gaussian mixture models represent the distribution as a weighted sum of multivariate normal components. The EM algorithm fits these models by alternating between assigning each component a responsibility for each point and updating component locations using responsibility-weighted averages. Cross-entropy measures how well an estimated density approximates the true one and is minimised when the two match. Dimensionality reduction compresses data into a lower-dimensional latent space via an encoder, with a decoder reconstructing the original data. Principal component analysis uses linear encoder–decoder pairs to minimise reconstruction error, offering a simple yet powerful form of dimensionality reduction.
Multi-model ensembles are widely used in climate science, yet Coupled Model Intercomparison Project Phase 6 (CMIP6) models are neither independent nor equally skillful. We present a framework for interpretable ensemble learning that improves prediction while making model contributions explicit. Using monthly CMIP6 near-surface temperature and precipitation (1948–2014) against ERA5, we compare arithmetic multi-model ensemble (AMME) averaging with Ridge regression, random forest, and a continuous gating mixture that adaptively interpolates between them. On held-out test years (2006–2014), gating achieves the best temperature performance and improves precipitation distributional behavior relative to both AMME and individual learners. We trace these gains back to individual models and space: dominant-cluster and dominant-model maps show that gating reallocates trust in regionally coherent patterns rather than collapsing to a single best model, while entropy-based diagnostics quantify where trust is concentrated versus distributed. These analyses provide a practical pathway for transparent ensemble learning that supports both predictive skill and scientific accountability.
The chapter introduces some of the statistics most commonly used to describe networks. These can be seen as analogues, in a network context, of quantities such as mean and standard deviation for a sample of real numbers. They can be roughly divided into two categories: topological summaries, such as the collection of degrees of the vertices of the network, that could be derived from a picture of the network, and spectral summaries, such as the eigenvector centrality, that are derived from the spectral decomposition of the adjacency matrix of the network (or of one or more related matrices). Many of them, such as clustering coefficients, can be formulated as local summaries, computed at each vertex of the network and can then be combined to yield a global value characteristic of the entire network. The results of a number of the summaries are compared with one another, using the Florentine marriage network as an example.
The random geometric graphs considered in this chapter are derived from a configuration of points that are independently placed in an underlying Euclidean space, according to some distribution. Each pair of points that are separated by a distance less than some given threshold is joined by an edge, and the graph then consists only of vertices, corresponding to the points, and of the edges between them, with the positional information discarded. In this model, the edges are no longer independent, and the neighbourhood structure is quite different from the tree-like neighbourhoods in Chapters 11–14; for instance, the average local clustering coefficient is not typically close to zero, even in sparse graphs. A giant component is shown to be unlikely to exist if the density of points is low enough, and to be almost certain to exist if the density of points is high enough, with the ratio of the critical densities fixed as the number of points grows. A subgraph threshold theorem is established, complemented by a number of distributional approximations to the counts of subgraphs; the independence of the positions of the underlying points simplifies this discussion. Under suitable asymptotics, typical shortest path lengths are shown to grow like a power of the number of points, rather than logarithmically, as was the case for the models in Chapters 11–14.
High-frequency mortality data have attracted growing attention, but their use has largely been confined to specific applications rather than general modeling and forecasting. Such data pose new challenges to traditional mortality models due to pronounced seasonal patterns and short-term fluctuations. To address these challenges and produce more accurate forecasts with the high-frequency mortality data, this paper introduces a novel integration of gradient boosting techniques into traditional stochastic mortality models under a multi-population setting. Our key innovation lies in using the Li and Lee model as the weak learner within the gradient boosting framework, replacing conventional decision trees. Empirical studies are conducted using weekly mortality data from 30 countries (Human Mortality Database, 2015–2019). Empirical evidence highlights that the proposed methodology not only enhances model fit by accurately capturing underlying mortality trends and seasonal patterns but also achieves superior forecast accuracy, compared to the benchmark models. We also investigate a key challenge in multi-population mortality modeling: how to select appropriate subpopulations with sufficiently similar mortality experiences. A comprehensive clustering exercise is conducted based on mortality improvement rates and seasonal strength. The empirical results demonstrate that our proposed model maintains strong forecast accuracy across different clustering configurations, thereby reducing the need for extensive data preprocessing.
This chapter explores how traders’ performance may be influenced by the rationality levels of their peers in the market. Using a Behavioural Data Science approach, the study integrates experimental methods, machine learning and large-scale digital trace analysis to examine this relationship. Specifically, we analysed data from a cryptoasset exchange over a five-week period in late 2017 and early 2018, covering over 700,000 transactions across 17 trading pairs. We complemented this behavioural trace data with an online guessing game involving 2,622 active traders, of whom 273 participated. By combining survey results and trading histories, we applied clustering algorithms to identify seven distinct trader profiles, including ‘jokers’, ‘focal point traders’ and those operating at different levels of strategic reasoning (first, second and third order), as well as ‘professional’ and ‘Nash equilibrium’ traders. The findings suggest that traders engaging in higher-order reasoning generally achieve better financial outcomes, yet even experienced professionals are not immune to behavioural biases. The chapter highlights how Behavioural Data Science methods – linking experimental insight with real-world data and computational tools – can illuminate the cognitive patterns underlying economic decision-making in digital markets.
The weakening of traditional social cleavages in predicting political preferences does not mean that group membership for detecting political behaviour no longer exists—it may only be the case that the boundaries of these groups need to be redefined. Using original and unique data, this article investigates the relevance of lifestyle for capturing and delineating social groups sharing political attitudes. Logistic regression machine learning models test how lifestyle predicts political behaviour alongside conventional sociodemographic variables. K-means clustering identifies three Quebecer lifestyle profiles and their relationship to sociodemographic features. Findings show lifestyle better predicts voting intentions and lifestyle clusters significantly associated with vote choice, even after controlling for sociodemographics. This research reassesses the weakening of socio-structural influences and the importance of contextual and individual variables in understanding political behaviour.
This retrospective study analysed 14,625 isolates of the six major hospital-associated ‘ESKAPE’ pathogens (Enterococcus faecium, Staphylococcus aureus, Klebsiella pneumoniae, Acinetobacter baumannii, Pseudomonas aeruginosa, and Enterobacter spp.) collected between 2002 and 2024 in a Hungarian tertiary-care centre. Antimicrobial resistance was assessed using the antibiotic resistance index (ARI), multidrug resistance (MDR) ratios, and resistance instability index (RII). A. baumannii and E. faecium showed the highest resistance burdens and instability. Age showed a significant monotonic association with resistance (Spearman r = 0.88), with peaks in infants, middle-aged women, and the elderly. Species-specific age trends varied, with a negative correlation seen in Enterobacter spp. Hierarchical clustering grouped pathogens by resistance trajectory rather than taxonomy. Pairwise resistance distances confirmed divergence between Gram-positive and Gram-negative species. Resistance to aminoglycosides and sulphonamides showed the highest year-to-year variability, as quantified by the RII, particularly in A. baumannii and E. faecium. Vector autoregressive (VAR) modelling predicted continued MDR increases in these species. A strong correlation was found between ARI and RII (Pearson r = 0.85, p = 0.032). These findings underscore the importance of integrating resistance magnitude and volatility in surveillance.
Coalition research increasingly emphasizes party-level explanations of coalition outcomes. However, this work does not account for the complex multilevel structure between parties and governments: many parties participate in multiple governments and governments often comprise multiple parties. In this paper, I show that this crisscrossing structure creates dependencies among observations both across and within governments. If ignored, these dependencies produce downward-biased uncertainty estimates that cluster-robust standard errors fail to fully correct. To address this issue, I then introduce a model that extends the Multiple Membership Multilevel Model to represent the multilevel structure of coalition government data. The model accounts for party-level dependencies across governments through party-specific effects in each coalition they join, and for dependencies within governments by representing the total party effect on a government as a weighted sum of its members’ contributions. By allowing party weights to vary with covariates describing their interrelationships, the model enables researchers to examine the interdependent nature of coalition outcomes. I validate the model through simulation and an empirical application to coalition government survival, showing that ignoring party-level dependencies can produce misleading conclusions at all levels of analysis. The model is estimated via Bayesian MCMC and implemented in the accompanying R package ‘bml’.
As a direct consequence of liquid kerosene injection, aeroengine combustors may be categorized as non-premixed combustion systems, characterized by a swirl-stabilized and highly complex flow field. In addition to the flow of air through the fuel injector, there are a large number of other features through which the oxidizer can enter the heat release region. These can have an impact on local fuel–air mixing, inducing strong spatial and temporal variations in stoichiometry, thereby affecting emissions and combustion system performance. This article discusses a novel statistical methodology, based on principal component analysis (PCA) and K-means clustering, that aims to improve the understanding of fuel–air mixing in realistic aeroengine combustors. The method is applied in a post-processing step to data sampled from a large-eddy simulation, where every chamber inflow has been tagged with a unique passive scalar, which allows it to be traced across space and time. PCA is used to construct a low-dimensional, visually interpretable representation of a spatially localized fuel–air mixing process, while K-means clustering is employed to produce an unsupervised discretization of the flow field into regions of similar fuel–air mixing characteristics. The proposed methodology is computationally inexpensive, and the easily interpretable outputs can help the combustion engineer make better-informed decisions about combustor design.
Emphasizing how and why machine learning algorithms work, this introductory textbook bridges the gap between the theoretical foundations of machine learning and its practical algorithmic and code-level implementation. Over 85 thorough worked examples, in both Matlab and Python, demonstrate how algorithms are implemented and applied whilst illustrating the end result. Over 75 end-of-chapter problems empower students to develop their own code to implement these algorithms, equipping them with hands-on experience. Matlab coding examples demonstrate how a mathematical idea is converted from equations to code, and provide a jumping off point for students, supported by in-depth coverage of essential mathematics including multivariable calculus, linear algebra, probability and statistics, numerical methods, and optimization. Accompanied online by instructor lecture slides, downloadable Python code and additional appendices, this is an excellent introduction to machine learning for senior undergraduate and graduate students in Engineering and Computer Science.
The Latent Position Model (LPM) is a popular approach for the statistical analysis of network data. A central aspect of this model is that it assigns nodes to random positions in a latent space, such that the probability of an interaction between each pair of individuals or nodes is determined by their distance in this latent space. A key feature of this model is that it allows one to visualize nuanced structures via the latent space representation. The LPM can be further extended to the Latent Position Cluster Model (LPCM), to accommodate the clustering of nodes by assuming that the latent positions are distributed following a finite mixture distribution. In this paper, we extend the LPCM to accommodate missing network data and apply this to non-negative discrete weighted social networks. By treating missing data as “unusual” zero interactions, we propose a combination of the LPCM with the zero-inflated Poisson distribution. Statistical inference is based on a novel partially collapsed Markov chain Monte Carlo algorithm, where a Mixture-of-Finite-Mixtures (MFM) model is adopted to automatically determine the number of clusters and optimal group partitioning. Our algorithm features a truncated absorb-eject move, which is a novel adaptation of an idea commonly used in collapsed samplers, within the context of MFMs. Another aspect of our work is that we illustrate our results on 3-dimensional latent spaces, maintaining clear visualizations while achieving more flexibility than 2-dimensional models. The performance of this approach is illustrated via three carefully designed simulation studies, as well as four different publicly available real networks, where some interesting new perspectives are uncovered.
In many contexts, an individual’s beliefs and behavior are affected by the choices of their social or geographic neighbors. This influence results in local correlation in people’s actions, which in turn affects how information and behaviors spread. Previously developed frameworks capture local social influence using network games, but discard local correlation in players’ strategies. This paper develops a network games framework that allows for local correlation in players’ strategies by incorporating a richer partial information structure than previous models. Using this framework we also examine the dependence of equilibrium outcomes on network clustering—the probability that two individuals with a mutual neighbor are connected to each other. We find that clustering reduces the number of players needed to provide a public good and allows for market sharing in technology standards competitions.
This chapter introduces the mathematics of data through the example of clustering, a fundamental technique in data analysis and machine learning. The chapter begins with a review of essential mathematical concepts, including matrix and vector algebra, differential calculus, optimization, and elementary probability, with practical Python examples. The chapter then delves into the k-means clustering algorithm, presenting it as an optimization problem and deriving Lloyd's algorithm for its solution. A rigorous analysis of the algorithm's convergence properties is provided, along with a matrix formulation of the k-means objective. The chapter concludes with an exploration of high-dimensional data, demonstrating through simulations and theoretical arguments how the "curse of dimensionality" can affect clustering outcomes.
This study analyzes 1,000 meta-analyses drawn from 10 disciplines—including medicine, psychology, education, biology, and economics—to document and compare methodological practices across fields. We find large differences in the size of meta-analyses, the number of effect sizes per study, and the types of effect sizes used. Disciplines also vary in their use of unpublished studies, the frequency and type of tests for publication bias, and whether they attempt to correct for it. Notably, many meta-analyses include multiple effect sizes from the same study, yet fail to account for statistical dependence in their analyses. We document the limited use of advanced methods—such as multilevel models and cluster-adjusted standard errors—that can accommodate dependent data structures. Correlations are frequently used as effect sizes in some disciplines, yet researchers often fail to address the methodological issues this introduces, including biased weighting and misleading tests for publication bias. We also find that meta-regression is underutilized, even when sample sizes are large enough to support it. This work serves as a resource for researchers conducting their first meta-analyses, as a benchmark for researchers designing simulation experiments, and as a reference for applied meta-analysts aiming to improve their methodological practices.
Texts, whether literary or historical, exhibit structural and stylistic patterns shaped by their purpose, authorship and cultural context. Formulaic texts, which are characterized by repetition and constrained expression, tend to differ in their information content (as defined by Shannon) compared to more dynamic compositions. Identifying such patterns in historical documents, particularly multi-author texts like the Hebrew Bible, provides insights into their origins, purpose and transmission. This study aims to identify formulaic clusters: sections exhibiting systematic repetition and structural constraints, by analyzing recurring phrases, syntactic structures and stylistic markers. However, distinguishing formulaic from non-formulaic elements in an unsupervised manner presents a computational challenge, especially in high-dimensional and sample-poor data sets where patterns must be inferred without predefined labels.
To address this, we develop an information-theoretic algorithm leveraging weighted self-information distributions to detect structured patterns in text. Our approach directly models variations in sample-wise self-information to identify formulaicity. By extending classical discrete self-information measures with a continuous formulation based on differential self-information in multivariate Gaussian distributions, our method remains applicable across different types of textual representations, including neural embeddings under Gaussian priors.
Applied to hypothesized authorial divisions in the Hebrew Bible, our approach successfully isolates stylistic layers, providing a quantitative framework for textual stratification. This method enhances our ability to analyze compositional patterns, offering deeper insights into the literary and cultural evolution of texts shaped by complex authorship and editorial processes.
This chapter explores practical applications of network representation learning techniques for analyzing individual networks. It begins by addressing the community detection problem, demonstrating how to estimate community labels using network embeddings. The chapter then discusses the challenges posed by network sparsity and introduces efficient storage methods for sparse networks. The text proceeds to examine testing for differences between groups of edges, applying hypothesis testing to stochastic block models and structured independent edge models. It also covers model selection techniques for stochastic block models, helping readers choose appropriate levels of model complexity. The chapter introduces the vertex nomination problem, which aims to identify nodes similar to a set of known "seed" nodes. It presents spectral vertex nomination techniques and explores extensions to related problems. Finally, the chapter addresses out-of-sample embedding, providing efficient strategies for embedding new nodes into existing network representations. This approach is particularly valuable for large-scale, dynamic networks where frequent re-embedding would be computationally prohibitive.
Our politics are increasingly polarised. Polarisation takes many forms. One is increasing clustering, whereby people hold down-the-line liberal or conservative views on a wide range of orthogonal issues. Some philosophers think that such clustering is indicative of irrationality, and so finding yourself in one of several clusters gives you evidence that not all your political beliefs are true. I argue that the reverse is true, presenting a simple model of belief-formation in which finding yourself in one of several clusters of opinion on orthogonal issues should increase, rather than decrease, your confidence that all your beliefs are true.
This chapter discusses how to build probabilistic models that include both discrete and continuous variables. Mathematically, this is achieved by defining them as random variables within the same probability space. In practice, the variables are manipulated using their marginal and conditional distributions. We define the conditional pmf of a discrete random variable given a continuous variable, and the conditional probability density of a continuous random variable given a discrete variable. We use these objects to build mixture models and apply them to model height in a population. Next, we describe Gaussian discriminant analysis, a classification method based on mixture models with Gaussian conditional distributions, and apply it to diagnose Alzheimer's disease. Then, we explain how to perform clustering using Gaussian mixture models and leverage the approach to cluster NBA players. Finally, we introduce the framework of Bayesian statistics which enables us to explicitly encode our uncertainty about model parameters, and use it to analyze poll data from the 2020 United States presidential election.
Factor score indeterminacy is a characteristic property of factor analysis (FA) models. This research introduces a novel procedure, regression-based factor score exploration (RFE), which uniquely determines factor scores and simultaneously estimates other parameters of the FA model. RFE uniquely determines factor scores by minimizing a loss function that balances FA and multivariate regression, regulated by a tuning parameter. Theoretical aspects of RFE, including the uniqueness of factor scores, the relationship between observed and latent variables, and rotational indeterminacy, are examined. Additionally, clustering-based factor exploration (CFE) is presented as a variant of RFE, derived by generalizing the penalty term to enable the clustering of factor scores. It is demonstrated that CFE creates cluster structures more accurately than the existing method. A simulation study shows that the proposed procedures accurately recover true parameter matrices even in the presence of error-contaminated data, with lower computational demand compared to existing methods. Real data examples illustrate that the proposed procedures provide interpretable results, demonstrating high relevance to the factor scores obtained by existing methods.