To save content items to your account,
please confirm that you agree to abide by our usage policies.
If this is the first time you use this feature, you will be asked to authorise Cambridge Core to connect with your account.
Find out more about saving content to .
To save content items to your Kindle, first ensure no-reply@cambridge.org
is added to your Approved Personal Document E-mail List under your Personal Document Settings
on the Manage Your Content and Devices page of your Amazon account. Then enter the ‘name’ part
of your Kindle email address below.
Find out more about saving to your Kindle.
Note you can select to save to either the @free.kindle.com or @kindle.com variations.
‘@free.kindle.com’ emails are free but can only be saved to your device when it is connected to wi-fi.
‘@kindle.com’ emails can be delivered even when you are not connected to wi-fi, but note that service fees apply.
In this chapter, the mathematically simplest random graph, the Bernoulli random graph, is introduced. Each of the possible edges is present, independently, with the same probability, so that the model is one of a network entirely without structure. To start with, the structure of the graph in the neighbourhood of a point is investigated and is shown to be very similar to that of a branching process with Poisson-distributed offspring numbers. Explicit bounds on the accuracy of the approximation are derived, using the Poisson approximation techniques derived in Chapter 7. The classical threshold theorem for the existence of a giant component is then established; the precision of the neighbourhood approximation simplifies the proof. The counts of small subgraphs are then investigated, and a subgraph threshold theorem is proved. Finally, the distribution of the length (in graph distance) of the shortest path between two vertices is investigated. These grow logarithmically with the number of points, if the expected degree of a vertex is kept constant. Once again, the approximation of the neighbourhood structure is a key element in the proofs, and the statement of the main theorem involves the Laplace transform of the distribution of the limit random variable associated with the approximate branching process.
Many networks are not completely known; the only access to them is by taking samples. This chapter presents methods for deducing information about the whole network from samples. First, some classical sampling methods are briefly considered; random sampling, with and without replacement, stratified sampling and the Horvitz–Thompson estimator. Then sampling methods based on the network structure are introduced, including two-level sampling, induced subgraph sampling, star and snowball sampling and traversal sampling. The differences between the structure of sample networks and those of the parent network are illustrated for some simple models. Finally, the problem of assessing whether a particular network sample is `interesting’ is discussed; interesting, in that it differs from what might be expected of a typical network sample.
There are many topics that are not covered in the book. First, networks may be weighted, directed or signed. Then networks may exhibit structures other than those considered in the book, such as hierarchical structures, or have edges of different types, and collections of networks may arise as snapshots of a network process evolving in time. Each of these settings requires different methods of analysis. Then relationships may be expressed in more intricate ways; `edges’ may link more than just two objects, as in a hypergraph, and abstract simplicial complexes can be thought of as higher dimensional analogues of geometric graphs. These, and other topics, are sketched in this chapter. The material in this book forms a general basis that can be used in coming to grips with these more advanced settings.
The configuration and GPDS models allow the degrees of the vertices to be exactly specified. They are mathematically more challenging than those considered in Chapters 11–13, because the assignment of edges is no longer independent. The configuration model, which generates multigraphs, acts as a more tractable alternative to the GPDS model, because of the symmetries inherent in the random mapping that is used to define it, and because each realization from the configuration model that is simple is also a realization from the GPDS with the same vertex degrees. In this chapter, the neighbourhood structure in the configuration model is approximated by that of a related branching process, and a threshold theorem for the appearance of a giant component is also derived. For subgraph counts, the configuration multigraphs are first made simple, by collapsing multiple edges and deleting self-loops; the resulting graphs are shown to satisfy a subgraph threshold theorem. Finally, the neighbourhood structure is used to derive an approximation to the distribution of the length of the shortest path between two vertices. The proofs in this chapter are very much more involved than those elsewhere in the book.
The chapter introduces some of the statistics most commonly used to describe networks. These can be seen as analogues, in a network context, of quantities such as mean and standard deviation for a sample of real numbers. They can be roughly divided into two categories: topological summaries, such as the collection of degrees of the vertices of the network, that could be derived from a picture of the network, and spectral summaries, such as the eigenvector centrality, that are derived from the spectral decomposition of the adjacency matrix of the network (or of one or more related matrices). Many of them, such as clustering coefficients, can be formulated as local summaries, computed at each vertex of the network and can then be combined to yield a global value characteristic of the entire network. The results of a number of the summaries are compared with one another, using the Florentine marriage network as an example.
This chapter discusses a number of generalizations of the Bernoulli random graph. The first is to Erdős–Rényi mixture models, in which each vertex belongs to one of a small number of distinct types, and the probability of an edge between two vertices being present depends on the types of the two vertices. Much as for the Bernoulli random graph, the neighbourhood structure is approximated by a multitype branching process. A threshold theorem for the existence of a giant component is derived, and the asymptotic proportions of the vertices in the giant component that belong to the different types is determined. Finally, the distribution of the length of the shortest path between two vertices of given types is approximated. The properties of some special network models of this kind, vertex weighted random graphs and bipartite networks, are examined. Another set of models considered is that of directed Bernoulli random graphs, in which there may be edges in either (or both) directions between any pair of vertices. Finally, Poissonized Erdős–Rényi mixture models for multigraphs are introduced, in which edges may be present between any pair of vertices, and the number of such edges has a Poisson distribution, with mean depending on the types of the vertices, the numbers being chosen independently. The extent to which these models differ from the standard Erdős–Rényi mixture models in the sparse regime is discussed.
This chapter considers settings in which networks are used in statistical inference. First, there is often dependence between measurements taken at vertices of a network, with vertices at small graph distance more highly dependent. Methods for accommodating such dependence, including network effects and network disturbance models, are discussed. In such cases, the network is treated as a nuisance parameter. In contrast, the network itself may be the object of primary interest. Fitting a model to the network may make possible inference, for instance about the presence or absence of a particular edge, or of characteristics of vertices; this is valuable if knowledge of the network is not perfect. If more than one network is available, it may also be possible to compare the properties of different, but related networks, using the fitted models, and to use the comparisons to uncover scientifically interesting features. It can also be of interest to find ways of comparing whole networks with one another, when they may have very different vertex sets, but are conjectured to be related; methods discussed include the graphlet correlation distance, Netdis and NetEMD. For instance, the similarities between the protein–protein interaction networks of different species can be used as a basis for constructing phylogenies.
In this chapter, an adaptation of Stein’s method for bounding the error in multivariate normal approximation is presented. For simplicity, the distance measure used is based on expectations of functions with three bounded derivatives; more natural measures of distance would require much more complicated treatments. The Stein equation used is now a second-order partial differential equation. Solutions to the equation are exhibited, and some of their properties are established; they can then be used to derive a general bound on the approximation error in multivariate normal approximation. For exploiting the general bound, a local approach is introduced, which uses a multivariate version of the double decomposition used for (univariate) normal approximation. This is applied to the number of monochrome edges in a graph whose vertices are randomly coloured. A size bias coupling approach is also developed and applied to the joint distribution of counts of vertices of different degrees in the Bernoulli random graph.
In most network models, the distributions of summary statistics are rarely known exactly. However, it may well be possible to approximate their distributions by other, well-known distributions, when the values of local statistics at different vertices are only weakly dependent. Such settings are well suited to the application of the Stein–Chen method, which, in the context of Poisson approximation, enables concrete estimates of the approximation error to be derived, with respect to the total variation distance. In this chapter, the Stein–Chen method is developed in some detail. A Stein equation is derived, together with the necessary properties of its solutions, and a general estimate of approximation error is given, which is expressed solely in terms of the random variable whose distribution is being approximated. The method is applied to sums of dependent random variables, using both a local and a coupling approach. Examples given include the number of triangles and the number of isolated vertices in a Bernoulli random graph.
A number of methods for finding communities in networks are introduced, and their performance on two benchmark networks is compared. The chapter begins with methods of comparing classifications and for assessing the quality of a classification. Then, a number of classification algorithms that use modularity as a measure of quality are presented. An alternative approach is to fit a stochastic blockmodel to a network. Methods for doing so include a Bayesian approach based on the Gibbs sampler, and a variational method that makes use of the EM algorithm; estimation of the number of blocks is also considered. Spectral classification methods are described, including those based on the spectral decomposition of the non-backtracking matrix. The performance of the algorithms on two benchmark data sets is encouragingly consistent. The chapter concludes with two methods designed to find overlapping clusters, and with a discussion of the theoretical detectability threshold.
To understand their properties, networks can be compared with those randomly generated from one or more network models. The network models introduced in this chapter, many of which are discussed at greater length later in the book, include the Bernoulli random graph, Erdős–Rényi mixture models, Chung–Lu graphs, small world networks, the configuration model, random geometric graphs, preferential attachment models, exponential random graph models, stochastic blockmodels, latent space models, random intersection graphs, graphon models, models for directed graphs and duplication–divergence models. Some basic properties of the models are considered, such as the distribution of the degree of a randomly chosen vertex and typical values of the clustering coefficient.
A number of methods of estimation are introduced, and are applied in the context of networks. Maximum likelihood estimation is applied to Bernoulli random graphs and to Erdős–Rényi mixture graphs. The EM algorithm, used later in fitting stochastic blockmodels, is also introduced. Both maximum likelihood and the (generalized) method of moments are used in the context of estimating the exponent of power law decay in degree distributions. Bayesian methods are presented, and the choice of prior discussed; they are applied to Erdős–Rényi mixture graphs and to their Poissonized variants. Further general methods introduced include Approximate Bayesian Computation, as well as Markov Chain Monte Carlo methods, for which both the Metropolis–Hastings algorithm and the Gibbs sampler are presented. Some specific models are given special attention. In exponential random graph models, MCMC methods offer an approach, though convergence to equilibrium can be very slow. The estimation of latent space models is discussed both from a frequentist and from a Bayesian point of view. Estimating the underlying dimension of a random geometric graph is also touched upon.