Hostname: page-component-76d6cb85b7-8p85h Total loading time: 0 Render date: 2026-07-21T18:25:15.538Z Has data issue: false hasContentIssue false

A novel predictor of multilocus haplotype homozygosity: comparison with existing predictors

Published online by Cambridge University Press:  01 February 2010

I. M. MacLEOD*
Affiliation:
Melbourne School of Land and Environment, University of Melbourne, VIC 3010, Australia BioSciences Research Division, Department of Primary Industries, VIC 3083, Australia
T. H. E. MEUWISSEN
Affiliation:
Department of Animal and Aquacultural Sciences, Norwegian University of Life Sciences, N-1432 Aas, Norway
B. J. HAYES
Affiliation:
BioSciences Research Division, Department of Primary Industries, VIC 3083, Australia
M. E. GODDARD
Affiliation:
Melbourne School of Land and Environment, University of Melbourne, VIC 3010, Australia BioSciences Research Division, Department of Primary Industries, VIC 3083, Australia
*
*Corresponding author. Agriculture and Food Systems, Melbourne School of Land and Environment, University of Melbourne, VIC 3010, Australia. Tel: 03_8344 7224. e-mail: macleodi@unimelb.edu.au
Rights & Permissions [Opens in a new window]

Summary

The patterns of linkage disequilibrium (LD) between dense polymorphic markers are shaped by the ancestral population history. It is therefore possible to use multilocus predictors of LD to infer past population history and to infer sharing of identical alleles in quantitative trait locus (QTL) studies. We develop a multilocus predictor of LD for pairs of haplotypes, which we term haplotype homozygosity (HHn): the probability that any two haplotypes share a given number of n adjacent identical markers or ‘runs of homozygosity’. Our method, based on simplified coalescence theory, accounts for recombination and mutation. We compare our HHn predictions, with HHn in simulated populations and with two published predictors of HHn. Our method performs consistently better across a range of population parameters, including populations with a severe bottleneck followed by expansion, compared to two published methods. We demonstrate that we can predict the pattern of HHn observed in dense single nucleotide polymorphisms (SNPs) genotyped in a cattle population, given appropriate historical changes in population size. Our method is practical for use with very large numbers of individuals and dense genome wide polymorphic DNA data. It has potential applications in inferring ancestral population history and QTL mapping studies.

Information

Type
Paper
Copyright
Copyright © Cambridge University Press 2010
Figure 0

Table 1. General notation presented in this section

Figure 1

Fig. 1. Two possible coalescence pathways for a pair of chromosome segments with three markers, in a population which changes in size ancestrally (divided into three ‘phases’). Tracing two DNA segments back in time, in coalescence A there is no recombination or coalescence in phase 1, and coalescence occurs in phase 2 with a mutation event on one segment. In coalescence B, both gametes recombine in phase 2 and in phase three loci ‘a’ mutates, followed by coalesce of segments AB and C.

Figure 2

Fig. 2. (A, B) Graphs show HHn for chromosome segments with 2, 3, 4, etc. markers, that are evenly spaced. All populations have reached a drift–recombination–mutation equilibrium, assuming N=500, T=8000, μ=0·0002. In (A) marker intervals (r) are 0·001 M compared to (B), where r=0·0002. Comparisons are made between HHn predictions using three analytical methods and observed HHn in simulated populations (average of 2000 replicates).

Figure 3

Table 2. Observed (simulated data) and predicted HHn (three methods) with population parameters; N=15 000, T=300 000, μ=0·00001 and r=0·00003

Figure 4

Fig. 3. Predictions of HHn using three analytical methods compared to observed values in simulated populations (average of 5000 replicates). Population size is changing over time from very large ancestral N=100 000, and gradually decreasing to present day N=100. (r=0·0015 M and μ=10−5). There are four phases of differing population size. HHn predicted by MHH is shown for these four phases, but also calculated with non-equilibrium phases split into a number of shorter phases (16 or 1776 phases total), to reduce approximation error.

Figure 5

Table 3. Observed HHn (simulated population with 5000 replicates) compared with MHH predicted HHn, splitting each long phase of constant population size into an increasing number of shorter phases (population parameters as for Fig. 3)

Figure 6

Fig. 4. (A, B) Both graphs show predictions of HHn in two different populations, using two analytical methods compared to observed values in simulated populations (average of 1000 replicates). One population is in equilibrium with constant size (N=50 000) and a second is of the same ancestral and present size, but with a bottleneck (N=2500, T=2000), 5000 generations before present day. In (A), the marker intervals (r) are 0·0005 M, while in (B), r=2·5×10−5 M. The mutation rates in the two different populations have been adjusted to maintain the same single marker homozygosity.

Figure 7

Fig. 5. HHn predictions using two analytical methods, for an expanding population as for Fig. 4B (‘bottleneck’), but with no large ancestral size: ancestral N=2500 for T=52 000, followed by N=50 000 for T=5000 to present day. Predictions are compared with observed HHn in simulated populations (2000 replicates). Also plotted is the observed HHn from the bottleneck population in Fig. 4B (i.e. with large ancestral size=50 000 and bottleneck N=2500 for T=2000), to demonstrate that the bottleneck does not mask the more ancestral population size effect on HHn.

Figure 8

Fig. 6. Observed genome wide HHn in cattle, genotyped for 38 259 SNPs (38 K data), compared with MHH predicted HHn. Predictions are shown for constant population size (N=120), as well as sharply decreasing population size with three or six ancestral sizes (from N=50 000 to N<300).