To save content items to your account,
please confirm that you agree to abide by our usage policies.
If this is the first time you use this feature, you will be asked to authorise Cambridge Core to connect with your account.
Find out more about saving content to .
To save content items to your Kindle, first ensure no-reply@cambridge.org
is added to your Approved Personal Document E-mail List under your Personal Document Settings
on the Manage Your Content and Devices page of your Amazon account. Then enter the ‘name’ part
of your Kindle email address below.
Find out more about saving to your Kindle.
Note you can select to save to either the @free.kindle.com or @kindle.com variations.
‘@free.kindle.com’ emails are free but can only be saved to your device when it is connected to wi-fi.
‘@kindle.com’ emails can be delivered even when you are not connected to wi-fi, but note that service fees apply.
In order to document further the phenomena of variance in reproductive success in natural populations of the European flat oyster Ostrea edulis, two complementary studies based on natural and experimental populations were conducted. The first part of this work was focused on paternity analyses using a set of four microsatellite markers for larvae collected from 13 brooding females sampled in Quiberon Bay (Brittany, France). The number of individuals contributing as the male parent to each progeny assay was highly variable, ranging from 2 to more than 40. Moreover, paternal contributions showed a much skewed distribution, with some males contributing to 50–100% of the progeny assay. The second part of this work consisted of the analysis of six successive cohorts experimentally produced from an acclimated broodstock (62 wild oysters sampled in the Quiberon Bay). Allelic richness was significantly higher in the adult population than in the temporal cohorts collected. Genetic differentiation (Fst estimates) was computed for each pair of samples and all significant values ranged from 0·7 to 11·9%. A limited effective number of breeders (generally below 25) was estimated in the six temporal cohorts. The study gives first indications of the high variance in reproductive success as well as a reduced effective size, not only under experimental conditions but also in the wild. Surprisingly, the pool of the successive cohorts, based on the low number of loci used, appeared to depict a random and representative set of alleles of the progenitor population, indicating that the detection of patterns of temporal genetic differentiation at a local scale most likely depends on the sampling window.
Quantitative trait loci (QTLs) mapping often results in data on a number of traits that have well-established causal relationships. Many multi-trait QTL mapping methods that account for correlation among the multiple traits have been developed to improve the statistical power and the precision of QTL parameter estimation. However, none of these methods are capable of incorporating the causal structure among the traits. Consequently, genetic functions of the QTL may not be fully understood. In this paper, we developed a Bayesian multiple QTL mapping method for causally related traits using a mixture structural equation model (SEM), which allows researchers to decompose QTL effects into direct, indirect and total effects. Parameters are estimated based on their marginal posterior distribution. The posterior distributions of parameters are estimated using Markov Chain Monte Carlo methods such as the Gibbs sampler and the Metropolis–Hasting algorithm. The number of QTLs affecting traits is determined by the Bayes factor. The performance of the proposed method is evaluated by simulation study and applied to data from a wheat experiment. Compared with single trait Bayesian analysis, our proposed method not only improved the statistical power of QTL detection, accuracy and precision of parameter estimates but also provided important insight into how genes regulate traits directly and indirectly by fitting a more biologically sensible model.
Plant cuticular n-alkanes have been successfully used as markers to estimate diet composition and intake of grazing herbivores. However, additional markers may be required under grazing conditions in botanically diverse vegetation. This study was conducted to describe the n-alkane profiles and the carbon isotope enrichment of n-alkanes of common plant species from the Mid Rift Valley rangelands of Ethiopia, and evaluate their potential use as nutritional markers. A total of 23 plant species were collected and analysed for long-chain n-alkanes ranging from heptacosane to hexatriacontane (C27 to C36), as well as their carbon isotopic ratio (13C/12C). The analysis was conducted by gas chromatography/combustion isotope ratio mass spectrometry following saponification, extraction and purification. The isotopic composition of the n-alkanes is reported in the delta notation (δ13C) relative to the Vienna Pee Dee Belemnite standard. The dominant n-alkanes in the species were C31 (mean ± s.d., 283 ± 246 mg/kg dry matter) and C33 (149 ± 98 mg/kg dry matter). The carbon isotopic enrichment of the n-alkanes ranged from −19.37‰ to −37.40‰. Principal component analysis was used to examine interspecies differences based on n-alkane profiles and the carbon isotopic enrichments of individual n-alkanes. Large variability among the pasture species was observed. The first three principal components explained most of the interspecies variances. Comparison of the principal component scores using orthogonal procrustes rotation indicated that about 0.84 of the interspecies variances explained by the two types of data sets were independent of each other, suggesting that the use of a combination of the two markers can improve diet composition estimations. It was concluded that, while the n-alkane profile of the pasture species remains a useful marker for use in the study region, the δ13C values of n-alkanes can provide additional information in discriminating diet components of grazing animals.
The objective of this study was to evaluate the impact of diets enriched with plant oils or seeds, high in polyunsaturated fatty acids (PUFA), on the fatty acid profile of sheep intramuscular and subcutaneous adipose tissue (SAT). Sixty-six lambs were blocked according to initial body weight and randomly assigned to six concentrate-based rations containing 60 g fat/kg dry matter from different sources: (1) Megalac (MG; ruminally protected saturated fat), (2) camelina oil (CO), (3) linseed oil (LO), (4) NaOH-treated camelina seed (CS), (5) NaOH-treated linseed (LS) or (6) CO protected from ruminal saturation by reaction with ethanolamine; camelina oil amides (CA). The animals were offered the experimental diets for 100 days, after which samples of m. longissimusdorsi and SAT were collected and the fatty acid profile determined by GLC. The data were analyzed using ANOVA with ‘a priori’ contrasts including camelina v. linseed, oil v. NaOH-treated seeds and CS v. CA. Average daily gain and total fatty acids in intramuscular adipose tissue were similar across treatments. The NaOH-treatment of seeds was more effective in enhancing cis-9, trans-11 conjugated linoleic acid (CLA) incorporation than the corresponding oil, but the latter resulted in a higher content of trans-11 18:1 in both muscle neutral and polar lipids (P < 0.01, P < 0.001, respectively). Inclusion of LS resulted in the highest PUFA:saturated fatty acid (SFA) ratio in total intramuscular fat (0.22). The NaOH-treatment of seeds resulted in a higher PUFA/SFA ratio (0.21 v. 0.18, P < 0.001) than oils and on average, linseed resulted in a higher PUFA/SFA ratio than camelina (P < 0.01). Lambs offered LS had the highest concentration of n-3 PUFA in the muscle, while those offered MG had the lowest (P < 0.001). This was reflected in the lowest (P < 0.001) n-6: n-3 PUFA ratio for LS-fed lambs (1.15) than any other treatment, which ranged from 2.14 to 1.72, and the control (5.28). The trends found in intramuscular fat were confirmed by the data for SAT. This study demonstrated the potential advantage from a human nutrition perspective of feeding NaOH-treated seeds rich in PUFA when compared to the corresponding oil. The use of camelina amides achieved a greater degree of protection of dietary PUFA, but decreased the incorporation of biohydrogenation intermediates such as cis-9, trans-11 CLA and trans-11 18:1 compared to NaOH-treated seeds.
Appetite control is a major issue in normal growth and in suboptimal growth performance settings. A number of hormones, in particular leptin, activate or inhibit orexigenic or anorexigenic neurotransmitters within the arcuate nucleus of the hypothalamus, where feed intake regulation is integrated. Examples of appetite regulatory neurotransmitters are the stimulatory neurotransmitters neuropeptide Y (NPY), agouti-related protein (AgRP), orexin and melanin-concentrating hormone and the inhibitory neurotransmitter, melanocyte-stimulating hormone (MSH). Examination of messenger RNA (using in situ hybridization and real-time PCR) and proteins (using immunohistochemistry) for these neurotransmitters in ruminants has indicated that physiological regulation occurs in response to fasting for several of these critical genes and proteins, especially AgRP and NPY. Moreover, intracerebroventricular injection of each of the four stimulatory neurotransmitters can increase feed intake in sheep and may also regulate either growth hormone, luteinizing hormone, cortisol or other hormones. In contrast, both leptin and MSH are inhibitory to feed intake in ruminants. Interestingly, the natural melanocortin-4 receptor (MC4R) antagonist, AgRP, as well as NPY can prevent the inhibition of feed intake after injection of endotoxin (to model disease suppression of appetite). Thus, knowledge of the mechanisms regulating feed intake in the hypothalamus may lead to mechanisms to increase feed intake in normal growing animals and prevent the wasting effects of severe disease in animals.
Foods derived from animals are an important source of nutrients in the diet but there is considerable uncertainty about whether or not these foods contribute to increased risk of various chronic diseases. For milk in particular there appears to be an enormous mismatch between both the advice given on milk/dairy foods items by various authorities and public perceptions of harm from the consumption of milk and dairy products, and the evidence from long-term prospective cohort studies. Such studies provide convincing evidence that increased consumption of milk can lead to reductions in the risk of vascular disease and possibly some cancers and of an overall survival advantage from the consumption of milk, although the relative effect of milk products is unclear. Accordingly, simply reducing milk consumption in order to reduce saturated fatty acid (SFA) intake is not likely to produce benefits overall though the production of dairy products with reduced SFA contents is likely to be helpful. For red meat there is no evidence of increased risk of vascular diseases though processed meat appears to increase the risk substantially. There is still conflicting and inconsistent evidence on the relationship between consumption of red meat and the development of colorectal cancer, but this topic should not be ignored. Likewise, the role of poultry meat and its products as sources of dietary fat and fatty acids is not fully clear. There is concern about the likely increase in the prevalence of dementia but there are few data on the possible benefits or risks from milk and meat consumption. The future role of animal nutrition in creating foods closer to the optimum composition for long-term human health will be increasingly important. Overall, the case for increased milk consumption seems convincing, although the case for high-fat dairy products and red meat is not. Processed meat products do seem to have negative effects on long-term health and although more research is required, these effects do need to be put into the context of other risk factors to long-term health such as obesity, smoking and alcohol consumption.
In this study, the complete sequence of the Tibetan Mastiff mitochondrial genome (mtDNA) was determined, and the phylogenetic relationships between the Tibetan Mastiff and other species of Canidae were analyzed using the coyote (Canis latrans) as an outgroup. The complete nucleotide sequence of the Tibetan Mastiff mtDNA was 16 710 bp, and included 22 tRNA genes, 2S rRNA gene, 13 protein-coding genes and one non-coding region (D-loop region), which is similar to other mammalian mitochondrial genomes. The characteristics of the protein-coding genes, non-coding region, tRNA and rRNA genes among Canidae were analyzed in detail. Neighbor-joining and maximum-parsimony trees of Canids constructed using 12 mitochondrial protein-coding genes showed that as the coyotes and Tibetan wolves clustered together, so too did the gray wolves and domestic dogs, suggesting that the Tibetan Mastiff originated from the gray wolf as did other domestic dogs. Domestic dogs clustered into four clades, implying at least four maternal origins (A to D). The Tibetan Mastiff, which belongs to clade A, appears to be closely related to the Saint Bernard and the Old English Sheepdog.
Results on the behaviour of the rightmost particle in the nth generation in the branching random walk are reviewed and the phenomenon of anomalous spreading speeds, noticed recently in related deterministic models, is considered. The relationship between such results and certain coupled reaction-diffusion equations is indicated.
AMS subject classification (MSC2010) 60J80
Introduction
I arrived at the University of Oxford in the autumn of 1973 for postgraduate study. My intention at that point was to work in Statistics. The first year of study was a mixture of taught courses and designated reading on three areas (Statistics, Probability, and Functional Analysis, in my case) in the ratio 2:1:1 and a dissertation on the main area. As part of the Probability component, I attended a graduate course that was an exposition, by its author, of the material in Hammersley (1974), which had grown out of his contribution to the discussion of John's invited paper on subadditive ergodic theory (Kingman, 1973). A key point of Hammersley's contribution was that the postulates used did not cover the time to the first birth in the nth generation in a Bellman–Harris process. Hammersley (1974) showed, among other things, that these quantities did indeed exhibit the anticipated limit behaviour in probability. I decided not to be examined on this course, which was I believe a wise decisin but I was intrigued by the material. That interest turned out to be critical a few months later.
We consider perfect simulation algorithms for locally stable point processes based on dominated coupling from the past, and apply these methods in two different contexts. A new version of the algorithm is developed which is feasible for processes which are neither purely attractive nor purely repulsive. Such processes include multiscale area-interaction processes, which are capable of modelling point patterns whose clustering structure varies across scales. The other topic considered is nonparametric regression using wavelets, where we use a suitable area-interaction process on the discrete space of indices of wavelet coefficients to model the notion that if one wavelet coefficient is non-zero then it is more likely that neighbouring coefficients will be also. A method based on perfect simulation within this model shows promising results compared to the standard methods which threshold coefficients independently.
Keywords coupling from the past (CFTP), dominated CFTP, exact simulation, local stability, Markov chain Monte Carlo, perfect simulation, Papangelou conditional intensity, spatial birth-and-death process
Markov chain Monte Carlo (MCMC) is now one of the standard approaches of computational Bayesian inference. A standard issue when using MCMC is the need to ensure that the Markov chain we are using for simulation has reached equilibrium. For certain classes of problem, this problem was solved by the introduction of coupling from the past (CFTP) (Propp and Wilson, 1996, 1998).
The paper considers the situation of transmission over a memoryless noisy channel with feedback, which can be given a number of interpretations. The criteria for achieving the maximal rate of information transfer are well known, but examples of a simple and meaningful coding meeting these are few. Such a one is found for the Gaussian channel.
Keywords feedback channel, Gaussian channel
AMS subject classification (MSC2010) 94A24
Interrogation, transmission and coding
In this section we set out material which is classic in high degree, harking back to Shannon's seminal paper (1948), and presented in some form in texts such as those of Blahut (1987), Cover and Thomas (1991) and MacKay (2003). However, some exposition is necessary if we are to distinguish clearly between three versions of the model.
Suppose that an experimenter wishes to determine the value of a random variable U of which he knows only the probability distribution P(U). The formal argument is conveyed well enough for the moment if we suppose all distributions discrete and use the notation P(·) generically for such distributions. We may also abuse this convention by occasionally using P(U) to denote the function of U defined by the distribution.
The outcome of the experiment, if errorless, might be written x(U), where the form of the function x(U) reflects the design of the experiment. The experimenter will choose this, subject to practical constraints, so as to make the experiment as informative as possible.
Computer-based tests with randomly generated questions allow a large number of different tests to be generated. Given a fixed number of alternatives for each question, the number of tests that need to be generated before all possible questions have appeared is surprisingly low.
AMS subject classification (MSC2010) 60G70, 60K99
Introduction
The use of computer-based tests in which questions are randomly generated in some way provides a means whereby a large number of different tests can be generated; many universities currently use such tests as part of the student assessment process. In this paper we present findings that illustrate that, although the number of different possible tests is high and grows very rapidly as the number of alternatives for each question increases, the average number of tests that need to be generated before all possible questions have appeared at least once is surprisingly low. We presented preliminary findings along these lines in Cornish et al. (2006).
A computer-based test consists of q questions, each (independently) selected at random from a separate bank of a alternatives. Let Nq be the number of tests one needs to generate in order to see all the aq questions in the q question banks at least once. We are interested in how, for fixed a, the random variable Nq grows with the number of questions q in the test.
We prove a long-standing conjecture which characterizes the Ewens—Pitman two-parameter family of exchangeable random partitions, plus a short list of limit and exceptional cases, by the following property: for each n = 2, 3, …, if one of n individuals is chosen uniformly at random, independently of the random partition πn of these individuals into various types, and all individuals of the same type as the chosen individual are deleted, then for each r > 0, given that r individuals remain, these individuals are partitioned according to for some sequence of random partitions which does not depend on n. An analogous result characterizes the associated Poisson—Dirichlet family of random discrete distributions by an independence property related to random deletion of a frequency chosen by a size-biased pick. We also survey the regenerative properties of members of the two-parameter family, and settle a question regarding the explicit arrangement of intervals with lengths given by the terms of the Poisson–Dirichlet random sequence into the interval partition induced by the range of a homogeneous neutral-to-the right process.
Kingman introduced the concept of a partition structure, that is a family of probability distributions for random partitions πn of a positive integer n, with a sampling consistency property as n varies.
The volume of a Wiener sausage constructed from a diffusion process with periodic, mean-zero, divergence-free velocity field, in dimension 3 or more, is shown to have a non-random and positive asymptotic rate of growth. This is used to establish the existence of a homogenized limit for such a diffusion when subject to Dirichlet conditions on the boundaries of a sparse and independent array of obstacles. There is a constant effective long-time loss rate at the obstacles. The dependence of this rate on the form and intensity of the obstacles and on the velocity field is investigated. A Monte Carlo algorithm for the computation of the volume growth rate of the sausage is introduced and some numerical results are presented for the Taylor–Green velocity field.
We consider the problem of the existence and characterization of a homogenized limit for advection-diffusion in a perforated domain. This problem was initially motivated for us as a model for the transport of water vapour in the atmosphere, subject to molecular diffusion and turbulent advection, where the vapour is also lost by condensation on suspended ice crystals. It is of interest to determine the long-time rate of loss and in particular whether this is strongly affected by the advection. In this article we address a simple version of this set-up, where the advection is periodic in space and constant in time and where the ice crystals remain fixed in space.
We begin by reviewing some probabilistic results about the Dirichlet Process and its close relatives, focussing on their implications for statistical modelling and analysis. We then introduce a class of simple mixture models in which clusters are of different ‘colours’, with statistical characteristics that are constant within colours, but different between colours. Thus cluster identities are exchangeable only within colours. The basic form of our model is a variant on the familiar Dirichlet process, and we find that much of the standard modelling and computational machinery associated with the Dirichlet process may be readily adapted to our generalisation. The methodology is illustrated with an application to the partially-parametric clustering of gene expression profiles.
The purpose of this note is four-fold: to remind some Bayesian nonparametricians gently that closer study of some probabilistic literature might be rewarded, to encourage probabilists to think that there are statistical modelling problems worth of their attention, to point out to all another important connection between the work of John Kingman and modern statistical methodology (the role of the coalescent in population genetics approaches to statistical genomics being the most important example; see papers by Donnelly, Ewens and Griffiths in this volume), and finally to introduce a modest generalisation of the Dirichlet process.
John Frank Charles Kingman was born on 28th August 1939, a few days before the outbreak of World War II. This Festschrift is in honour of his seventieth birthday.
John Kingman was born in Beckenham, Kent, the son of the scientist Dr F. E. T. Kingman FRSC and the grandson of a coalminer. He was brought up in north London, where he attended Christ's College, Finchley. He was an undergraduate at Cambridge, where at age 19 at the end of his second year he took a First in Part II of the Mathematical Tripos, following it with a Distinction in the graduate-level Part III a year later, for his degree. He began postgraduate work as a research student under Peter Whittle, but transferred to David Kendall in Oxford when Peter left for Manchester in 1961, returning to Cambridge when Kendall became the first Professor of Mathematical Statistics there in 1962.
John's early work was on queueing theory, a subject he had worked on with Whittle, but was also an interest of Kendall's. His lifelong interest in mathematical genetics also dates back to this time (1961). His next major interest was in Markov chains, and in a related matter—what happens to Feller's theory of recurrent events in continuous time. His first work here dates from 1962, and led to his landmark 1964 paper on regenerative phenomena, where we meet (Kingman) p-functions.
The dynamics of tumour evolution are not well understood. In this paper we provide a statistical framework for evaluating the molecular variation observed in different parts of a colorectal tumour. A multi-sample version of the Ewens Sampling Formula forms the basis for our modelling of the data, and we provide a simulation procedure for use in obtaining reference distributions for the statistics of interest. We also describe the large-sample asymptotics of the joint distributions of the variation observed in different parts of the tumour. While actual data should be evaluated with reference to the simulation procedure, the asymptotics serve to provide theoretical guidelines, for instance with reference to the choice of possible statistics.
Cancers are thought to develop as clonal expansions from a single transformed, ancestral cell. Large-scale sequencing studies have shown that cancer genomes contain somatic mutations occurring in many genes; cf. Greenman et al., Sjöblom et al., Shah et al. Many of these mutations are thought to be passenger mutations (those that are not driving the behaviour of the tumour), and some are pathogenic driver mutations that influence the growth of the tumour. The dynamics of tumour evolution are not well understood, in part because serial observation of tumour growth in humans is not possible.
In an attempt to better understand tumour growth and structure, a number of evolutionary approaches have been described.