12.1 Introduction
The challenge of speech perception begins with the infinitely large space of possible messages, and is exponentially compounded by the infinite variations with which any one message can be transduced into an auditory signal. In practice, human listeners can tolerate incredible variability in speech signals, resulting from different content, speakers, and contexts (Miller and Licklider, Reference Miller and Licklider1950; Huggins, Reference Huggins1964; Beasley et al., Reference Beasley, Bratt and Rintelmann1980; Drullman et al., Reference Drullman, Festen and Plomp1994; Shannon et al., Reference Shannon, Zeng, Kamath, Wygonski and Ekelid1995; Dorman et al., Reference Dorman, Loizou and Rainey1997; Ahissar et al., Reference Ahissar, Nagarajan and Ahissar2001; Huang et al., Reference Huang, Chen, Li, Chang and Zhou2001; Ghitza and Greenberg, Reference Ghitza and Greenberg2009). At the same time, speech perception remains exquisitely sensitive to subtle variations in linguistic and paralinguistic signals, with listeners extracting nuanced information from a huge variety of cues (Munhall et al., Reference Munhall, Jones, Callan, Kuratate and Vatikiotis-Bateson2004; Mattys et al., Reference Mattys, White and Melhorn2005; Kurumada et al., Reference Kurumada, Brown, Bibyk, Pontillo and Tanenhaus2014; Baese-Berk et al., Reference Baese-Berk, Morrill and Dilley2018). How is it that listeners extract meaning from speech with a combination of robustness and sensitivity as yet unequaled by artificial speech recognition? The variability of speech provides a route to a solution, rather than posing an inherent problem, by providing rich context to shape the search space of possible interpretations. These perspectives connect to an emerging conceptual understanding of perception, action, and cognition as Bayesian inference, formalized by developments in the theory of hierarchical probabilistic generative models (Rao and Ballard, Reference Rao and Ballard1999; Knill and Pouget, Reference Knill and Pouget2004; Friston et al., Reference Friston, Kilner and Harrison2006, Reference Friston, Daunizeau and Kiebel2009; Adams et al., Reference Adams, Shipp and Friston2013; Millidge et al., Reference Millidge, Seth and Buckley2021).
In this chapter, we focus on the application of these ideas to the classic problem of recovering words from continuous spoken signals. To accomplish this, the continuous speech stream must be segmented into discrete lexical units (word segmentation) that must be retrieved from memory (lexical access). Unlike words on a printed page – which are separated by spaces – spoken words are not reliably separated by silence or any other acoustic cue (Klatt, Reference Klatt1976; Mattys et al., Reference Mattys, White and Melhorn2005). Classic psycholinguistic theories have explained spoken word recognition as a process of recognizing constituent phonemes of lexical material (McClelland and Elman, Reference McClelland and Elman1986; Norris and McQueen, Reference Norris and McQueen2008). However, there are few reliable acoustic correlates of phonemes; phonemes as well as words are best understood as perceptual rather than acoustic categories (Goldinger and Azuma, Reference Goldinger and Azuma2003; Kazanina et al., Reference Kazanina, Bowers and Idsardi2018; Samuel, Reference Samuel2020). This establishes that theories relying on acoustics cannot account for word recognition. Instead, listeners use multiple cues, on multiple levels of linguistic abstraction, to deduce the locations of word boundaries (Stevens, Reference Stevens2002; Mattys et al., Reference Mattys, White and Melhorn2005, Reference Mattys and Melhorn2007; Mattys and Melhorn, Reference Mattys and Melhorn2007; Dilley and McAuley, Reference Dilley and McAuley2008; Dilley and Pitt, Reference Dilley and Pitt2010; Dilley et al., Reference Dilley and Pitt2010), supporting proposals that view word segmentation as probabilistic inference (Norris and McQueen, Reference Norris and McQueen2008; Martin, Reference Martin2016; Norris et al., Reference Norris, McQueen and Cutler2016; Brown et al., Reference Brown, Tanenhaus and Dilley2021).
A particular challenge posed by word segmentation is the necessity for online inference: Linguistic signals arrive sequentially, vanishing after momentary articulation; listeners must make sense of the current input before it is overwritten by the next (Marslen-Wilson, Reference Marslen-Wilson1973; MacGregor et al., Reference MacGregor, Pulvermuller, Van Casteren and Shtyrov2012; Christiansen and Chater, Reference Christiansen and Chater2016). Results have suggested a central role for the syllabic timescale in both the pacing of speech and its intelligibility (Miller and Licklider, Reference Miller and Licklider1950; Huggins, Reference Huggins1964; Beasley et al., Reference Beasley, Bratt and Rintelmann1980; Drullman et al., Reference Drullman, Festen and Plomp1994; Shannon et al., Reference Shannon, Zeng, Kamath, Wygonski and Ekelid1995; Ahissar et al., Reference Ahissar, Nagarajan and Ahissar2001; Elliott and Theunissen, Reference Elliott and Theunissen2009; Ghitza and Greenberg, Reference Ghitza and Greenberg2009; Doelling et al., Reference Doelling, Arnal, Ghitza and Poeppel2014; Sun and Poeppel, Reference Sun and Poeppel2023; Chapters 2 and 5), with experiments on so-called repackaged speech suggesting an upper limit on the speech “information rate” of ∼9 syllables per second (Ghitza and Greenberg, Reference Ghitza and Greenberg2009; Ghitza, Reference Ghitza2014). Another challenge is that timing and content are inextricably intertwined in speech and the brain (Ten Oever and Martin, Reference Ten Oever and Martin2023). This is illustrated by distal rate effects, recent surprising results on speech rate dependence that conclusively demonstrate that syllables are no more acoustically invariant than phonemes (Dilley and Pitt, Reference Dilley and Pitt2010; Brown et al., Reference Brown, Dilley and Tanenhaus2012; Baese-Berk et al., Reference Baese-Berk, Heffner and Dilley2014; Pitt et al., Reference Pitt, Szostak and Dilley2016; Brown et al., Reference Brown, Tanenhaus and Dilley2021). In distal rate effects, an unaltered target segment consisting of a word followed by a blended, coarticulated function word (e.g., “summer or,” “saw a”) is heard as one word when prior speech is sufficiently slowed (Figure 12.1A).
Distal rate effects and the syllable inference (SI) hypothesis.
Distal rate effects. (i) A sentence with a target segment (“summer or”) containing a reduced, coarticulated function word (“or”) (speech waveform in gray, target segment in black). (ii) A version of the same sentence manipulated to slow context speech rate. (iii) Subjects asked to repeat the sentence report fewer function words for the slowed context.

Figure 12.1A. Long description
Part A demonstrates the difference in speech waveforms between a normal speech rate context and a slowed speech rate context, both with the same target word, which reads," Fred would rather have a summer or a lake". A bar graph with an error bar in part A 3 compares the percentage of function words reported with the normal rate and slowed context. The mean percentage of function is 80% for the normal rate. The mean percentage of function is 35% for slowed context.
The SI hypothesis. (i) An illustration of SI for the normal context speech stimulus from (A). A mean speech rate (μ_1) is computed from the interpretation(s) of context speech. For each candidate interpretation of the target segment (“summer or” and “summer”), knowledge of each syllable’s relative duration is combined with μ_1 to obtain an estimated candidate duration. These estimated candidate durations are compared to the observed (e.g., acoustic) duration (ν). The candidate that best explains what is heard (in this case, the one most probable given the observed duration) is perceived. (ii) An illustration of SI for the slowed context speech stimulus from (A). Here, a slower mean speech rate (μ_2) leads to a different judgment of which candidate interpretation is most probable.

Figure 12.1B. Long description
Part B 1 illustrates a speech waveform along with the candidate's interpretations of the words that read: Fred would rather have a summer or a lake. A portion of words that read" Fred would rather have a" denotes my 1, which represents prior mean speech rate. An arrow from a sum points at a sum which is approximately equal to a 1, mer is approximately equal to a 2, which denotes syllable specific relative durations, or is approximately equal to a 3. The equations for candidate durations and syllable interference are given.
The unfolding of speech in time and the interdependence of speech content and timing imply that in the “guess-and-check” procedure of Bayesian inference, when guesses are made and how long it takes to make and check them are crucially important. Existing theoretical accounts have largely addressed the issue of timing in speech segmentation by leveraging syllabic-timescale brain rhythms as a dynamic and self-organizing mechanism of speech segmentation (Lakatos et al., Reference Lakatos, Shah and Knuth2005; Ghitza and Greenberg, Reference Ghitza and Greenberg2009; Shamir et al., Reference Shamir, Ghitza, Epstein and Kopell2009; Ghitza, Reference Ghitza2011, Reference Ghitza2013, Reference Ghitza2014, Reference Ghitza2020; Hyafil et al., Reference Hyafil, Fontolan, Kabdebon, Gutkin and Giraud2015; Hovsepyan et al., Reference Hovsepyan, Olasagasti and Giraud2020, Reference Hovsepyan, Olasagasti and Giraud2023; Friston et al., Reference Friston, Sajid and Quiroga-Martinez2021; Nabé et al., Reference Nabé, Schwartz and Diard2021; Su et al., Reference Su, MacGregor, Olasagasti and Giraud2023; Chapters 3 and 5). We pursue a different strategy: that of understanding the computational principles underlying word segmentation independently of their neurophysiological mechanisms (rhythmic or otherwise) (Doelling and Assaneo, Reference Doelling and Assaneo2021; Adolfi et al., Reference Adolfi, Wareham and van Rooij2023). Toward this end, we sketch how a recently proposed conceptual model of speech understanding – syllable inference (SI) (Figure 12.1B; Brown et al., Reference Brown, Tanenhaus and Dilley2021) – might be elaborated and extended within the Bayesian network formalism for probabilistic inference. SI builds on the frameworks of predictive coding (Millidge et al., Reference Millidge, Seth and Buckley2021), chunk-and-pass processing (Christiansen and Chater, Reference Christiansen and Chater2016), cue integration (Martin, Reference Martin2016), and analysis-by-synthesis (Halle and Stevens, Reference Halle and Stevens1962). It proposes that extracting meaning from spoken signals involves the dynamical generation of alternative candidate speech interpretations. Each candidate speech interpretation consists of a hierarchical model of morphosyntactic structure and a model of speech timing, linked at the syllabic timescale. These generate statistical predictions about the presence and timing of multiple acoustic cues that are compared to the speech signal. The interpretations that explain the speech input sufficiently well rise to the level of perception, leading to the phonemes, syllables, words, and phrases that listeners hear.
In the process of extending SI, we unpack the two timing-related chicken-and-egg problems confronted by the listener during word segmentation: the integration of holistic, context-dependent processing and time-bound, incremental processing; and the near-simultaneous inference of speech timing and speech content. In both cases, each process in the pair seems to depend on the prior completion of the other. We show how these problems are intimately related to narrowing the search space over speech interpretations without bias, and optimizing the speed/accuracy trade-off in language processing, employing concepts borrowed from machine learning and drift-diffusion-to-bound models of decision making. We make four claims about how listeners solve these problems: (1) Listeners model prosody and speech timing as part of a speaker model, as well as modeling morphosyntactic structure in a message model; (2) listeners incrementally infer the content and timing of “chunks” of speech comprised of variable-length sequences of morphosyntactic units, rather than individual morphosyntactic units; and (3) listeners adaptively schedule this irregular inference of content and timing, as well as the deployment of computationally intensive iterative search and optimization operations, using (4) predictable fluctuations in uncertainty accessible through the speaker model. We pull these four claims together in a mechanistic proposal – vowel-onset-paced syllable inference (VPSI) – and outline how VPSI accounts for repackaging effects, distal rate effects, and other aspects of speech psychophysics.
We begin with an interdisciplinary literature review. We first list the psychophysical properties of speech perception (including context dependence, incrementality, rate dependence, and repackaging and distal rate effects) that we propose any model of word segmentation must account for. We then situate the notion of a speaker model within the contexts of prosody, speech production, and linguistic communication more broadly. We briefly review Bayesian inference, generative modeling, and probabilistic approaches to speech processing, including the SI hypothesis. Finally, we explore computational aspects of online word segmentation and present the VPSI model.
12.2 Context and Timing in Speech Perception
12.2.1 Context Dependence in Speech Perception
Many psycholinguistic studies have shown the importance of context for speech perception and spoken word recognition. For example: speech sounds replaced by noise are perceptually “filled-in” (Warren, Reference Warren1970); phonemes and words pronounced intelligibly in longer utterances are unrecognizable when excised and presented alone (Pollack and Pickett, Reference Pollack and Pickett1963); and how a given phoneme is produced and perceived depends in a complex way on neighboring phonemes via coarticulation (e.g., the “sk-” in “ski” spliced with the “-ool” from “school” is perceived as “spool”) (Schatz, Reference Schatz1954). In terms of word segmentation and identification, the boundaries of phonological and morphosyntactic units (phonemes, syllables, words, etc.) are often – but not always – marked by acoustic cues (Heffner et al., Reference Heffner, Dilley, McAuley and Pitt2013). Thus, segmentation relies on the flexible, context-dependent integration (Mattys et al., Reference Mattys, White and Melhorn2005) of multiple cues across levels of analysis and timescales, including acoustic features (Stevens, Reference Stevens2002), phonotactics, allophones, coarticulation, syntax, semantics, and lexical competition (Mattys et al., Reference Mattys, White and Melhorn2005, Reference Mattys and Melhorn2007; Mattys and Melhorn, Reference Mattys and Melhorn2007), syllabic-level transitional probabilities (Morrill et al., Reference Morrill, Baese-Berk, Heffner and Dilley2015), and stress (Dilley et al., Reference Dilley and Pitt2010).
Theoretical perspectives in psycholinguistics have addressed two separable aspects of this context dependence: the context imposed by the hierarchical compositionality of linguistic structure itself; and the flexible use of varying sources of information to infer linguistic structure and meaning. The hierarchical structure of language – in which sentences are composed of sequences of words, which in turn are composed of sequences of syllables, which in turn are composed of sequences of phonemes – gives rise to interlocking constraints at multiple levels of abstraction (Christiansen and Chater, Reference Christiansen and Chater2016; Martin and Doumas, Reference Martin and Doumas2017; Martin, Reference Martin2020; Chapters 13 and 20). These provide stringent limits on the possible interpretations of speech signals.
Perhaps the broadest framework for understanding the flexible use of information in speech perception is cue integration (Martin, Reference Martin2016). In this framework, a cue is any signal or piece of information that reflects the linguistic structure and meaning of an utterance. This definition includes internally generated, linguistically abstract representations activated by knowledge and prior context (e.g., prior speech), for example syntactic structure as a cue to word class and discourse history as a cue to word identity and meaning. To infer speech interpretations, these cues are integrated in an approximately Bayesian way, weighted by their relative reliability. Cue integration construes language as the network of dependencies between different cues, with representations at each level of linguistic abstraction serving as cues to representations at other levels of abstraction. As we will see, this maps naturally onto the Bayesian network formalism.
12.2.2 Temporal Context in Speech: Incrementality and Rate Dependence
The specific temporal signatures of segmentation and lexical access are, themselves, highly context-dependent. The speed of word identification ranges from ∼200 ms after onset (within the first few phonemes) (Marslen-Wilson, Reference Marslen-Wilson1973; MacGregor et al., Reference MacGregor, Pulvermuller, Van Casteren and Shtyrov2012; Salverda et al., Reference Salverda, Kleinschmidt and Tanenhaus2014) in optimal conditions to considerably after onset in running speech (Bard et al., Reference Bard, Shillcock and Altmann1988). Distinct syllable and word structures can readily be inferred from subtle acoustic cues before, during (Pickett and Decker, Reference Pickett and Decker1960), or after (Miller and Liberman, Reference Miller and Liberman1979) a syllable boundary (see also Repp et al., Reference Repp, Liberman, Eccardt and Pesetsky1978; Galle et al., Reference Galle, Klein-Packard, Schreiber and McMurray2019), and acoustic information considerably after an acoustic event has elapsed can influence the speech content perceived (Miller and Liberman, Reference Miller and Liberman1979; Connine et al., Reference Connine, Blasko and Titone1993; Wade and Holt, Reference Wade and Holt2005).
Nevertheless, one important set of results demonstrates the rate dependence of speech processing. Speech is intelligible over a limited range of syllabic rates, with understanding of American English deteriorating quickly at syllabic rates above ∼9 Hz (two–three times the average rate of ∼3–5 Hz) (Beasley et al., Reference Beasley, Bratt and Rintelmann1980; Elliott and Theunissen, Reference Elliott and Theunissen2009; Ghitza and Greenberg, Reference Ghitza and Greenberg2009; Pefkou et al., Reference Pefkou, Arnal, Fontolan and Giraud2017). Experiments with repackaged speech suggest this limit is on the syllabic rate, and not the rate of segmental information (Ghitza and Greenberg, Reference Ghitza and Greenberg2009). Repackaging involves separating uniform chunks of compressed speech with uniform periods of silence, so that the segmental rate (controlled by the compression factor) and the syllabic rate (controlled by the silent interval length) are independent. Repackaging rescues intelligibility for speech compressed up to a factor of 6. The highest intelligibility occurs when a chunk of speech having an uncompressed duration of ∼333 ms (roughly, a syllable-sized chunk) is delivered at a rate below 9 Hz. This suggests a maximum “information rate” of ∼9 English syllables per second, regardless of the segmental rate (across languages, an information rate of ∼39 bits/s has been observed (Coupé et al., Reference Coupé, Oh, Dediu and Pellegrino2019). Within this limit, speech perception adjusts for the context speech rate, with, for example, the perception of ambiguous phonemes determined by the rate of adjacent phonemes, and sudden changes in speech rate altering the perception of subsequent vowel duration and word identity (Oganian et al., Reference Oganian, Kojima and Breska2023).
Experiments on temporal cues in speech understanding illuminate the potential mechanisms of this rate dependence. Modulations of the speech envelope, the profile of amplitude fluctuations ≤ 20 Hz in the speech waveform (Elliott and Theunissen, Reference Elliott and Theunissen2009; Edwards and Chang, Reference Edwards and Chang2013; Chapter 2), provide nearly all the information necessary to decipher speech (Shannon et al., Reference Shannon, Zeng, Kamath, Wygonski and Ekelid1995). Syllabic-timescale (∼1–9 Hz) fluctuations are predominant in these amplitude modulations (Elliott and Theunissen, Reference Elliott and Theunissen2009), and more generally, the integrity of information at syllabic timescales is uniquely critical for speech comprehension (Miller and Licklider, Reference Miller and Licklider1950; Huggins, Reference Huggins1964; Drullman et al., Reference Drullman, Festen and Plomp1994; Elliott and Theunissen, Reference Elliott and Theunissen2009). Neurophysiological explorations of speech comprehension have revealed that brain activity – including endogenous rhythmic activity at frequencies that mirror those of speech (e.g., phonemic, syllabic, and phrasal) (Lakatos et al., Reference Lakatos, Shah and Knuth2005) – tracks or synchronizes with the speech signal and with higher-order linguistic structure at multiple timescales (Ahissar et al., Reference Ahissar, Nagarajan and Ahissar2001; Luo and Poeppel, Reference Luo and Poeppel2007; Nourski et al., Reference Nourski, Reale and Oya2009; Hertrich et al., Reference Hertrich, Dietrich, Trouvain, Moos and Ackermann2012; Peelle et al., Reference Peelle, Gross and Davis2012; Doelling et al., Reference Doelling, Arnal, Ghitza and Poeppel2014; Di Liberto et al., Reference Di Liberto, O’Sullivan and Lalor2015; Ding et al., Reference Ding, Melloni, Zhang, Tian and Poeppel2016; Riecke et al., Reference Riecke, Formisano, Sorger, Başkent and Gaudrain2017; Keitel et al., Reference Keitel, Gross and Kayser2018; Zoefel et al., Reference Zoefel, Archer-Boyd and Davis2018; Oganian and Chang, Reference Oganian and Chang2019; Kojima et al., Reference Kojima, Oganian and Cai2021; Chapters 3, 5, 11, and 17). This speech–brain entrainment is associated with speech intelligibility (Ahissar et al., Reference Ahissar, Nagarajan and Ahissar2001; Luo and Poeppel, Reference Luo and Poeppel2007; Nourski et al., Reference Nourski, Reale and Oya2009; Hertrich et al., Reference Hertrich, Dietrich, Trouvain, Moos and Ackermann2012; Peelle et al., Reference Peelle, Gross and Davis2012; Doelling et al., Reference Doelling, Arnal, Ghitza and Poeppel2014; Ding et al., Reference Ding, Melloni, Zhang, Tian and Poeppel2016; Riecke et al., Reference Riecke, Formisano, Sorger, Başkent and Gaudrain2017; Zoefel et al., Reference Zoefel, Archer-Boyd and Davis2018; Chapters 3 and 5) and can even modulate speech understanding (Riecke et al., Reference Riecke, Formisano, Sorger, Başkent and Gaudrain2017; Wilsch et al., Reference Wilsch, Neuling, Obleser and Herrmann2018; Zoefel et al., Reference Zoefel, Archer-Boyd and Davis2018). Speech tracking depends in part on the brain’s response to spectro-temporal discontinuities or acoustic edges (Doelling et al., Reference Doelling, Arnal, Ghitza and Poeppel2014; Oganian and Chang, Reference Oganian and Chang2019; Kojima et al., Reference Kojima, Oganian and Cai2021; Chapters 3, 5, and 8), and neural activity is particularly sensitive to peak-rate events in the speech envelope, which are statistically associated with vowel onsets (Oganian and Chang, Reference Oganian and Chang2019) (in production, if not perception [Chapter 11]).
Going beyond this rate normalization and its mechanisms, distal rate effects (aka lexical rate effects) reveal that even the perceived number of morphosyntactic units in a given utterance can depend on context speech rate (Figure 12.1A; Dilley and Pitt, Reference Dilley and Pitt2010; Brown et al., Reference Brown, Dilley and Tanenhaus2012, Reference Brown, Tanenhaus and Dilley2021; Baese-Berk et al., Reference Baese-Berk, Heffner and Dilley2014; Pitt et al., Reference Pitt, Szostak and Dilley2016). In distal rate effects, subjects listen to and verbally repeat recorded sentences containing a target segment (“summer or”) with a reduced, coarticulated function word (“or”; Figure 12.1Ai). In some of these recordings, context speech is slowed while the target segment is left unaltered (Figure 12.1Aii). Subjects repeat back fewer function words for the slowed context (Figure 12.1Aiii). Sizeable effects of distal rate can occur even when word onsets are acoustically clear (Heffner et al., Reference Heffner, Dilley, McAuley and Pitt2013), and greater degrees of time expansion on distal context produce greater reductions in hearing a spoken word (Heffner et al., Reference Heffner, Dilley, McAuley and Pitt2013; Morrill et al., Reference Morrill, Dilley, McAuley and Pitt2013). Extensive explorations have found that distal rate effects occur only when rate-carrying context signals are intelligible, confirming that they cannot be accounted for by bottom-up, signal-driven processes alone (Pitt et al., Reference Pitt, Szostak and Dilley2016). Lexical rate effects are also manifestly probabilistic, depending in a graded way on the degree of alteration of context speech speed (Heffner et al., Reference Heffner, Dilley, McAuley and Pitt2013; Morrill et al., Reference Morrill, Dilley, McAuley and Pitt2013; Brown et al., Reference Brown, Tanenhaus and Dilley2021). Eye-tracking data from distal rate experiments show that determination of the presence of the ambiguous, blended word in the target segment was delayed until 800 ms after target segment offset (Brown et al., Reference Brown, Tanenhaus and Dilley2021), a delay considerably longer than other measurements of the timescale for uptake and use of segmental information (Salverda et al., Reference Salverda, Dahan and Tanenhaus2007; Lamekina and Meyer, Reference Lamekina and Meyer2023).
12.2.3 Context Dependence in Linguistic Communication
Accounts of linguistic communication increasingly have taken seriously the role of context in meaningful inference (Pickering and Garrod, Reference Pickering and Garrod2004; Stolk et al., Reference Stolk, Verhagen and Toni2016). Much of this work frames linguistic communication within the evolutionary imperative of social cognition (Hari et al., Reference Hari, Sams and Nummenmaa2016). This suggests an interdependence between inferences about intentions and meanings and inferences about speakers, speaker groups, and culturally specific linguistic variations (Macrae and Bodenhausen, Reference Macrae and Bodenhausen2001; Brown-Schmidt et al., Reference Brown-Schmidt, Yoon, Ryskin and Ross2015; Mattan et al., Reference Mattan, Kubota and Cloutier2017).
12.2.4 Timing in Speech Production
In the context of communication, the temporal dynamics of speech perception depend on those of speech production (Cutler et al., Reference Cutler, Dahan and Van Donselaar1997; Turk and Shattuck-Hufnagel, Reference Turk and Shattuck-Hufnagel2014; Brown-Schmidt et al., Reference Brown-Schmidt, Yoon, Ryskin and Ross2015; McQueen and Dilley, Reference McQueen, Dilley, Gussenhhoven and Chen2020; Chapters 2, 6, and 7). Apart from the morphosyntactic structure of speech, a major determinant of speech timing is prosody, a hierarchical linguistic structure that defines the groupings and intonations of words, as well as the stress and prominence of words and syllables (Cutler et al., Reference Cutler, Dahan and Van Donselaar1997; Turk and Shattuck-Hufnagel, Reference Turk and Shattuck-Hufnagel2014; Chapters 17 and 18). Prosodic structure may be indicated acoustically by the pitch, timing, and duration of words and syllables relative to each other, as in phrase-final lengthening. However, no acoustic cue is diagnostic of prosodic structure, which language users impute even in the absence of relevant acoustic cues, and, indeed, while reading (Cutler et al., Reference Cutler, Dahan and Van Donselaar1997; Breen et al., Reference Breen, Fitzroy and Oraa Ali2019; McQueen and Dilley, Reference McQueen, Dilley, Gussenhhoven and Chen2020).
Speech timing is also influenced by factors not considered prosodic, including speech speed – of central importance in our account – as well as speech style, pragmatics, discourse-level semantics, emotion, clarity requirements, movement costs, and cognitive, motor, and physiological constraints (Turk and Shattuck-Hufnagel, Reference Turk and Shattuck-Hufnagel2014; Watson et al., Reference Watson, Jacobs and Buxó-Lugo2020; Chapters 1, 2, 6, 7, 15, 16, and 18). For example, word duration, often a marker of stress, is also influenced by word predictability, by speakers’ level of engagement, and by production difficulty (Watson et al., Reference Watson, Jacobs and Buxó-Lugo2020). Similarly, the placement of phrase boundaries – often considered to reflect prosodic and syntactic considerations – can reflect cognitive capacity, as suggested by evidence correlating interindividual variability in prosodic phrase length with working memory capacity (which presumably limits speakers’ planning scope) (Bishop and Intlekofer, Reference Bishop and Intlekofer2020).
Thus, prosody and other extra-morphosyntactic factors shape speech timing and acoustics to reflect not only the speaker’s intended message but also the speech production process, and the computational, cognitive, and physiological constraints and contexts under which speech production is occurring. In turn, this information reflects and provides clues to the emotional and physiological state, cognitive capacity, intentions, and identity of the speaker, providing a wealth of information relevant to a speaker model.
12.3 Probabilistic and Generative Modeling in Speech Perception
12.3.1 Speech Perception as Generative Modeling and Probabilistic Inference
Generative (i.e., top-down, synthetic, or predictive) modeling and probabilistic inference have a long history in linguistic theory. In the early analysis-by-synthesis framework for speech recognition (Halle and Stevens, Reference Halle and Stevens1962; Bever and Poeppel, Reference Bever and Poeppel2010), well-learned, bottom-up filters extract a rough interpretation; an internal generative model constructs a set of alternative interpretations accounting for longer-timescale contextual dependencies and predicts their sensory consequences; and the best predictor of the input is the final interpretation.
Bayesian statistical models provide a principled, quantitative approach to integrating knowledge- and signal-based cues (Martin, Reference Martin2016; Norris et al., Reference Norris, McQueen and Cutler2016), and have been applied to word segmentation (Norris and McQueen, Reference Norris and McQueen2008; Hovsepyan et al., Reference Hovsepyan, Olasagasti and Giraud2020, Reference Hovsepyan, Olasagasti and Giraud2023; Friston et al., Reference Friston, Sajid and Quiroga-Martinez2021; Nabé et al., Reference Nabé, Schwartz and Diard2021; Su et al., Reference Su, MacGregor, Olasagasti and Giraud2023), syntactic parsing (Traxler, Reference Traxler2014), and language learning (Abend et al., Reference Abend, Kwiatkowski, Smith, Goldwater and Steedman2017) and evolution (Griffiths and Kalish, Reference Griffiths and Kalish2007; Moulin-Frier et al., Reference Moulin-Frier, Diard, Schwartz and Bessière2015). Bayes’ theorem describes how prior beliefs should be combined with current observations during inference, expressing the conditional probability of an event A given an observation B, P(A|B), as (a multiple of) the product of the prior probability of the event, P(A), and the likelihood of the observation given the event, P(B|A). The distribution P(B|A) can be interpreted as a generative model of the data, which is inverted when calculating the posterior probability P(A|B).
Recent applications of Bayesian inference to cognition and neurophysiology broadly (Rao and Ballard, Reference Rao and Ballard1999; Knill and Pouget, Reference Knill and Pouget2004; Friston et al., Reference Friston, Kilner and Harrison2006, Reference Friston, Daunizeau and Kiebel2009; Adams et al., Reference Adams, Shipp and Friston2013; Millidge et al., Reference Millidge, Seth and Buckley2021) emphasize the brain’s function as a “prediction engine” that maintains an internal (generative) model of the world and continually produces hypotheses about the most likely causes of sensory data. This generative model biases inference toward conclusions that are probable given prior experience and current knowledge, but also enables rapid, accurate inference in the face of considerable variability and uncertainty. Updating the model to eliminate prediction errors (mismatches between predictions and observations) leads to both perception and learning, with uncertainty (or its inverse, precision) determining the relative adjustment of hypotheses and data. Predictive coding is a simple algorithmic implementation of this process (Rao and Ballard, Reference Rao and Ballard1999; Millidge et al., Reference Millidge, Seth and Buckley2021). A generative model can be formalized as a set of (probability distributions over) hypothetical causal factors known as hidden or latent variables that probabilistically determine observed variables or states (i.e., sensory data). Precision is not directly observable and must be predicted by second-order hidden variables (Koelsch et al., Reference Koelsch, Vuust and Friston2019), which are updated through confirmation or disconfirmation of first-order predictions. The brain is believed to implement a deep (i.e., multi-level) hierarchical generative model, in which hidden and observed variables are arranged into a compositional hierarchy like that observed in language (Friston et al., Reference Friston, Trujillo-Barreto and Daunizeau2008). The relationships between hidden and observed variables in a generative model can be captured in a Bayesian network.Footnote 1
12.3.2 Rhythm-Mediated Segmentation in Probabilistic Models
Leveraging experimental results on speech–brain entrainment, influential hypotheses have suggested that speech segmentation and sampling are mediated by a speech-entrained hierarchy of neural oscillators (Ghitza and Greenberg, Reference Ghitza and Greenberg2009; Ghitza, Reference Ghitza2011; Giraud and Poeppel, Reference Giraud and Poeppel2012). In one prominent set of proposals, a θ-frequency (∼4–8 Hz) oscillator tracks syllabic-timescale fluctuations in the speech amplitude envelope, driving a γ (∼40–60 Hz) oscillator that samples phonemic information at a rate proportional to the syllabic rate (Ghitza, Reference Ghitza2011; Giraud and Poeppel, Reference Giraud and Poeppel2012; Chapter 9). In this and similar accounts, the self-organized, history-dependent dynamics of brain rhythms are imputed to integrate prior speech speed with speech acoustics to perform accurate segmentation. Indeed, a recent model suggests that a neurophysiologically inspired oscillator exhibiting frequency and phase adaptation can implement something like Bayesian inference of event duration (Doelling et al., Reference Doelling, Arnal and Assaneo2023).
Several recent computational models integrate rhythmic accounts of syllable segmentation with probabilistic generative modeling of morphosyntactic structure (Hovsepyan et al., Reference Hovsepyan, Olasagasti and Giraud2020, Reference Hovsepyan, Olasagasti and Giraud2023; Friston et al., Reference Friston, Sajid and Quiroga-Martinez2021; Nabé et al., Reference Nabé, Schwartz and Diard2021; Su et al., Reference Su, MacGregor, Olasagasti and Giraud2023). These models fit under the umbrella of the Bayesian network formalism, with several models employing an approximate inference scheme that relies on free-energy minimization.Footnote 2 They employ oscillators or oscillator-inspired mechanisms to implement segmentation at the level of syllables, which in turn enables word- or lemma-level inference (Hovsepyan et al., Reference Hovsepyan, Olasagasti and Giraud2020, Reference Hovsepyan, Olasagasti and Giraud2023; Nabé et al., Reference Nabé, Schwartz and Diard2021; Su et al., Reference Su, MacGregor, Olasagasti and Giraud2023). While all make invaluable and pioneering contributions, none of them performs online word segmentation of naturalistic speech: one segments speech post hoc rather than online (Friston et al., Reference Friston, Sajid and Quiroga-Martinez2021); some have been implemented only on small synthetic “languages” whose units have regular duration (Nabé et al., Reference Nabé, Schwartz and Diard2021; Su et al., Reference Su, MacGregor, Olasagasti and Giraud2023); and some recognize syllables but not words (Hovsepyan et al., Reference Hovsepyan, Olasagasti and Giraud2020, Reference Hovsepyan, Olasagasti and Giraud2023). More importantly, the rhythm-inspired mechanisms used to determine speech timing and syllable onsets in most models do not receive feedback from the probabilistic models inferring morphosyntactic structure (Hovsepyan et al., Reference Hovsepyan, Olasagasti and Giraud2020, Reference Hovsepyan, Olasagasti and Giraud2023; Nabé et al., Reference Nabé, Schwartz and Diard2021; Su et al., Reference Su, MacGregor, Olasagasti and Giraud2023). While these mechanisms exhibit their own bottom-up adaptation to fluctuations in the temporal structure of speech acoustics, they cannot integrate information about content – for example, the number of words imputed to a given speech segment. The SI hypothesis takes a step toward rectifying this omission by outlining the bidirectional interactions between inference of timing and content.
12.3.3 The Syllable Inference Hypothesis
As discussed in Section 12.1, SI proposes that extracting meaning from spoken signals involves the dynamical generation of alternative candidate speech interpretations, each consisting of a hierarchical model of morphosyntactic structure and speech timing and prosody, including speed, stress, and phrasal contours, with a key focus on the syllable (Figure 12.1B; Brown et al., Reference Brown, Tanenhaus and Dilley2021). From this ensemble of models, those that explain the speech input sufficiently well (minimizing prediction error, maximizing model evidence, and/or surpassing some minimum level of confidence) rise to the level of perception, leading to the phonemes, syllables, words, and phrases that listeners hear. Under SI, each interpretation generates statistical predictions about the presence and timing of multiple time-frequency cues in the input. These predictions or hypotheses are based on the morphosyntactic and prosodic hierarchies imputed by the interpretation, estimates of the current speech rate and rhythm, and language-specific knowledge about the relative duration and timing of segments, syllables, words, and phrases. The arrival or absence of the cues predicted by any interpretation triggers a re-evaluation of the goodness-of-fit of the ensemble of current candidate interpretations, and potentially the updating of that ensemble to include new candidates.
Acoustic edges are a broad class of such cues (Stevens, Reference Stevens2002; Doelling et al., Reference Doelling, Arnal, Ghitza and Poeppel2014; Oganian and Chang, Reference Oganian and Chang2019; Kojima et al., Reference Kojima, Oganian and Cai2021; Chapter 8), and SI predicts that they provide evidence (in the form of prediction errors) for the processes of model evaluation and generation. Recognizing that the syllabic timescale is privileged both in speech acoustics and the neural dynamics of speech perception, SI asserts that the detection of syllabic-timescale temporal landmarks in particular – a highly reliable source of evidence pertaining to both the phonetic and the prosodic structure of speech – triggers neural activity implementing the evaluation and generation of candidate speech interpretations. SI postulates that the most reliable acoustic evidence of a syllable boundary occurs following that boundary, at the onset of the vocalic nucleus of the next syllable, an acoustic event to which the brain is exquisitely sensitive (Oganian and Chang, Reference Oganian and Chang2019). This indicates that all pertinent information regarding the previous syllable has arrived, and thus that interpretation of that syllable can be completed.
At the same time, rhythmic processes driven by reliable (acoustic or linguistic) evidence of syllable boundaries perform the dual functions of tracking speech speed and evaluating the relative timing of speech input. Rhythmic syllabic-timescale population activity thus encodes a probability distribution over speech speed, inferred from the timing of previous (perceived) syllables. As speech unfolds, the evolving estimate of the mean speech rate affects both the suite of candidate speech interpretations and the predicted relative timing of sound units. Deviations from predicted durations must be explained by positing alternative hypotheses as to the sources of variation (e.g., local lengthening for emphasis), or by updating the estimated speech rate. According to SI, the neural activity implementing these calculations – hypothesis generation and evaluation and inference of speech speed and rhythm – is what results in the statistical phase-locking of neural oscillations to the speech amplitude envelope, and the association of speech intelligibility with speech–brain entrainment at syllabic timescales.
12.4 Computational Challenges in Online Word Segmentation
As the SI hypothesis draws on the framework of predictive processing, Bayesian networks provide a natural framework for its computational construction. Here, we sketch this construction in a way intentionally agnostic to neurophysiological implementations, while assuming such implementations exist and making use of empirical observations about the relative timing of brain activity and speech. We take this approach to clarify the computational challenges inherent in online word segmentation (Table 12.1), and to allow these challenges and their solutions to constrain and guide the search for neurophysiological mechanisms, rather than the other way around. Along the way, we aim to reconcile incremental processing with the (temporally extended) hierarchical representation of sequences of linguistic units, to explain repackaging effects and distal rate effects in speech perception, and to provide an explanation of just how speech content and speech speed are estimated concurrently.
A non-comprehensive attempt to connect computational challenges to strategies for their solution and potential mechanisms of implementation. Note that challenges, solution strategies, and possible mechanisms are highly overlapping.
| Computational challenge | Solution strategies | Possible mechanisms |
|---|---|---|
| Accuracy | Using all available evidence | Post hoc inference |
| Long-timescale evidence accumulation | High-level latent variables in deep hierarchical models | |
| Inferring sequences of morphosyntactic units | ||
| Synthetic mechanisms | ||
| Speed | Using evidence as it becomes available | Continuous inference (i.e., filtering) |
| Short-timescale evidence consolidation | Incremental (i.e., Markov) assumption | |
| Balancing speed and accuracy, i.e., | Employing a priori constraints | Bayesian inference |
| Integrating holistic and incremental inference, i.e., | Inferring causal variables at multiple levels of abstraction | Deep hierarchical generative modeling |
| Managing search space size | Posing hypotheses at optimal level of abstraction | |
| Alternating filtering with synthetic mechanisms | Periodic post hoc re-estimation | |
| Adaptively alternating filtering and synthetic mechanisms | Pacing post hoc re-estimation by observed and/or predicted precision | |
| Alternating between evidence accumulation and evidence consolidation | Periodic model updates and/or fluctuations in precision | |
| Adaptively switching between evidence accumulation and evidence consolidation | Pacing model updates by observed and/or predicted precision | |
| Urgency-like signals based on estimated information rate | ||
| Inferring both timing and content | Alternating timing and content inference (i.e., expectation maximization) | Periodically alternating model updates and/or fluctuations in precision |
| Periodic post hoc re-estimation | ||
| Adaptively switching between timing and content inference | Pacing model updates by observed and/or predicted precision | |
| Urgency-like signals based on estimated information rate | ||
| Pacing post hoc re-estimation by observed and/or predicted precision | ||
| Minimizing the impact of computationally intensive operations | Adaptively switching between evidence accumulation and evidence consolidation | Pacing model updates by observed and/or predicted precision |
| Urgency-like signals based on estimated information rate | ||
| Adaptively scheduling iterative and synthetic operations | Pacing post hoc re-estimation by observed and/or predicted precision |
12.4.1 Generative Models Predict Speech Timing and Content
First, we specify a generative model, consisting of a hierarchy of dynamic (i.e., time-dependent) probability distributions. As in existing work, this generative model encodes the morphosyntactic structure of the message through probability distributions over the word, syllable, and phoneme levels of linguistic organization (Figure 12.2A; Nabé et al., Reference Nabé, Schwartz and Diard2021). To account for distal rate effects, we must at least keep track of (i.e., infer) speech rate alongside morphosyntactic structure; SI proposes that word segmentation also makes use of the duration of individual syllables relative to the speech rate (Brown et al., Reference Brown, Tanenhaus and Dilley2021). Thus, we specify a hierarchical generative model of speech timing, with variables representing the mean syllabic rate and individual syllable and phoneme durations (Figure 12.2A). Together, these two conditionally independent hierarchies specify the acoustics of the speech signal. A model of speech timing, as part of a speaker model (Brown-Schmidt et al., Reference Brown-Schmidt, Yoon, Ryskin and Ross2015) potentially governed by the temporal regularities of speech production (Elliott and Theunissen, Reference Elliott and Theunissen2009; Bishop and Intlekofer, Reference Bishop and Intlekofer2020; Ten Oever and Martin, Reference Ten Oever and Martin2021), begins to address the communicative context of speech (Brown-Schmidt et al., Reference Brown-Schmidt, Yoon, Ryskin and Ross2015).
Distal rate effects and the SI hypothesis.
A hierarchical Bayesian network for word segmentation. Hierarchically organized latent variables representing sequences of words (Wordi), syllables (Syli), and phonemes (Φi) constitute a generative model of morphosyntactic structure. Hierarchically organized latent variables representing syllabic rate (Rate) and sequences of syllable durations (DSyl,i) and phoneme durations (DΦ,i) constitute a generative model of speech timing. Together, they specify a speech signal as a trajectory in feature space.

Figure 12.2A. Long description
Part A: A hierarchical structure of words, syllables, and their corresponding acoustic features, represented by D. The arrows indicate the flow of information from higher-level units to lower-level units and, ultimately, to the Speech Acoustics features. The sign phi represents transformations or mappings between levels.
Interdependence of contextual and incremental processing. The size of each increment depends on the mean rate, and vice versa. Left and right show two different interdependent sets of increments and mean rate.

Figure 12.2B. Long description
Part B: A waveform with four vertical lines drawn over it is labeled contextual rate. The hypothesized increments representing the wave reads, d i, are approximately equal to 1 over mu. An equation on the right side reads, mu is approximately one, open large bracket, one over n sigma 1 to n d i, close large bracket raised to the power of minus 1.
The VPSI mechanism. A rate prior enables rate-dependent inference of content. When reliable timing information arrives, it triggers post hoc content re-estimation, which leads to rate re-estimation and a rate posterior. This serves as the rate prior for rate-dependent content inference of the next chunk of speech.

12.4.2 Sequences Not Units Are Inferred
Distal rate effects demonstrate that only a sequence of units reliably contains enough contextual information to disambiguate the speech signal. Thus, in SI, candidate speech interpretations may contain different numbers of words, syllables, and phonemes. We propose that this must be operationalized by defining hidden variables (at all levels in both hierarchies) as single distributions over sequences of units, rather than sequences of distributions over individual units (Figure 12.2A).
Inferring sequences accords not only with the highly context-dependent nature of speech but also with the essentially fungible nature of the division of speech into units on evidence in processes of language learning and evolution (Christiansen and Chater, Reference Christiansen and Chater2016). Further, it aligns with a framing in which the ultimate goal of linguistic communication is guiding adaptive action in support of survival and reproductive success (Chapters 6, 7, 15, 16, 18, and 29). We propose that this goal takes priority over the decoding of linguistic content, and can be achieved in its absence, for example by deciphering the emotion and intent of a speaker of a foreign language (Frühholz and Schweinberger, Reference Frühholz and Schweinberger2021). We suggest that meaning making, similarly, can occur without the identification and segmentation of specific linguistic units by mapping stretches of speech to sequences of linguistic units of any length – for example, through the holistic identification of commonly used multi-word expressions.
12.4.3 Inference Proceeds Incrementally
Next, we specify an inference algorithm to invert the generative model. Given a speech signal, this scheme should output a probability distribution over possible sequences of words and syllabic rates, which can be used (along with some utility function) to select a single most likely or appropriate interpretation. Exact inference involves enumerating the set of possible sequences of words and syllabic rates; calculating the likelihood of the speech input given each sequence, and multiplying it by the prior probability of the sequence; and then normalizing to find the distribution over possible sequences.
This approach maximizes the accuracy of inference by using all the information contained in an utterance to determine its probable interpretations. However, it is slow, and not only because it requires waiting until the end of the utterance. While a longer speech signal offers more evidence to distinguish between interpretations, it also has a larger number of possible interpretations. Indeed, large search spaces make exact inference impossible in practice for most problems. Generally speaking, both exhaustive search processes and their reformulations as iterative computations – including, for example, prediction error minimization (Tschantz et al., Reference Tschantz, Millidge, Seth and Buckley2023) – have costs in terms of time and resources that scale with search space size, and efficiency requires avoiding, minimizing, or carefully timing them (Halle and Stevens, Reference Halle and Stevens1962; Tschantz et al., Reference Tschantz, Millidge, Seth and Buckley2023). This trade-off between the information content of the evidence and the size of the search space – that is, between the speed and the accuracy of inference – can also be viewed as a tension between context dependence and incrementality in processing. This tension is particularly easy to understand in the case of speech rate inference (Figure 12.2B). Here, the context (mean speech rate) informs incremental inference (via estimated syllable duration), and must simultaneously be derived from multiple incremental estimates (Figure 12.2B).
Thus, even presented with an entire utterance at once – for example, during reading – the most efficient strategy employs one “piece” of evidence at a time, incrementally constraining the space of possible interpretations. In online inference, time offers a natural way to parcel out evidence, and all existing models of online speech inference make a temporal incremental approximation. That is, they approximate the distribution over word sequences and speech speeds that best explains the entire utterance with a series of distributions over word sub-sequences and speech speeds, each of which best explains the input up to a given time. This is accomplished in practice by iteratively updating the posterior distribution, adjusting it based on the inversion of the generative model for one “chunk” of input at a time; we assume that the sequence of words best explaining the first and second chunks of input is reasonably approximated by the sequence of words A that best explains the first chunk of input, followed by the sequence of words B that, given A, best explains the second chunk of input.
Formally, we assume a Markov property at each hierarchical level of the model. This means the distribution over sequences of morphosyntactic units (e.g., phonemes) explaining the current chunk of speech is independent of all but the previous chunk, given the current distribution over sequences of higher-level units (e.g., syllables). Why is lower-level units’ independence from the past conditional on higher-level units? Because the structure of speech is compositional in time, distributions over higher-level units effectively encode (some aspects of) an arbitrarily long past context, as well as their influence on sequences of lower-level units. Thus, deep hierarchical representations allow arbitrarily long past context to influence inference at all levels, even if the distant past is not explicitly represented at lower levels (Friston et al., Reference Friston, Trujillo-Barreto and Daunizeau2008; Martin and Doumas, Reference Martin and Doumas2017; Martin, Reference Martin2020). This makes it possible to assume a Markov property (i.e., independence from the past) without losing long-timescale contextual information. This point illustrates that deep hierarchical modeling is roughly equivalent to the modeling of distant temporal context. By modeling, for example, sequences at the word level, we roughly approximate the modeling of deeper hierarchical levels of linguistic structure such as syntax and meaning; and by modeling sequences of speech speeds, we could potentially approximate the modeling of higher levels of prosodic structure.
This temporal compression is just one manifestation of the information compression inherent in the compositional structure of hierarchical generative models. In general, this structure lowers search costs without sacrificing accuracy, by enabling alternatives to be enumerated at high hierarchical levels with small search spaces and evaluated at low hierarchical levels with high information density. Thus, a hierarchical morphosyntactic model enables us to account for the current stretch of speech by searching for possible word (rather than syllable or phoneme) sequences, while determining the likelihood of those sequences from acoustic features (rather than inferred phonemes or syllables).
12.4.4 Evidence Accumulation Alternates with Model Updating across Processing Streams
If inference is to be incremental, it remains to specify the size of the increments – that is, what constitutes a chunk. Given that speech inputs of a fixed duration will invariably represent (and be interpreted as) sequences of morphosyntactic units of variable length – and more basically, because the number and duration of morphosyntactic units is exactly what we hope to infer during word segmentation – we propose that the length of a chunk cannot be specified in terms of a fixed number of morphosyntactic units, at any level of analysis. It is tempting to define the length of a chunk relative to the speed of speech; but how are we to obtain a reliable estimate of the speech speed without first inferring speech content? A variation on the issues discussed in the previous section, this question highlights two new issues: the importance of establishing informative prior probabilities over speech content and speed; and the difficulty of simultaneously inferring the “what” and the “when” of speech (Figure 12.2B).
Regarding the first, we propose that speech patterns facilitate estimates of speech speed and content at the beginning of an utterance – for example with highly stressed and/or predictable monosyllabic utterances such as greetings or interjections (Norrick, Reference Norrick2009). We can illustrate the latter obstacle in a simple template-matching account of speech perception. Ignoring variation in the pronunciation of individual morphosyntactic units, we think of each linguistic unit as a noisy trajectory through some feature space, and the listener as possessing linguistic knowledge comprised of “canonical” or “prototypical” trajectories for each unit (e.g., word, syllable, or phoneme). The computational goal is to determine what sequence of units comprise the speech signal; an auxiliary goal is to determine the speed at which these units are replayed. If we knew the sequence of units, it would be trivial to determine the speed by stretching the corresponding sequence of canonical trajectories to match, as well as possible, the observed trajectory. If all trajectories were traversed at the same known speed, we could recover the sequence of units – even unit by unit – by finding the sequence of canonical trajectories that best match the observed trajectory. In the case where we know neither the speed nor the content, the space of possible trajectories includes all possible sequences, of all possible lengths, replayed at all possible speeds. Notably, the transition from one unit to the next can occur at any time. In fact, some automatic speech recognition algorithms (and TRACE, the early interactive activation model of word segmentation [McClelland and Elman, Reference McClelland and Elman1986]) find unit onset times by comparing unit templates having all possible onsets to the data.
The central issue here is that when we must infer two mutually dependent variables (e.g., timing and content), the search space grows very fast with the data length. One well-known method for the efficient simultaneous inference of two mutually dependent variables is expectation maximization (EM) (Dempster et al., Reference Dempster, Laird and Rubin1977). EM involves repeatedly inferring one variable while holding the other constant, using updated expectations about one variable to improve inference of the other, until the likelihood of the combined guess arrives at a (locally) maximal value. From the perspective of a single variable, past inferences about the other variable are evidence to be integrated in the current inference process.
In Precoss-β, a recent model integrating probabilistic and rhythmic mechanisms to infer the identities and durations of syllables from naturalistic speech (Hovsepyan et al., Reference Hovsepyan, Olasagasti and Giraud2023), such an alternation was implemented via rhythmic fluctuations in prediction error precision (PEP). PEP determines how prediction errors are weighted in model updates. When PEP is high, models are adjusted to eliminate prediction errors, and as a result prediction errors have low magnitude; when PEP is low, models follow their own internal dynamics, and large prediction errors are “allowed” to remain. Notably, a model exhibiting rhythmic fluctuations in PEP identified syllables more effectively than one in which PEP was held constant, because periods of low PEP effectively served as windows for evidence accumulation, while subsequent periods of high PEP consolidated the accumulated evidence into a single model update. Interestingly, this alternation was most advantageous at β frequency, and was most effective when the PEP for syllable boundaries and syllable identities fluctuated in antiphase, resulting in alternating estimation of “when” a syllable occurred and “what” syllable occurred (Hovsepyan et al., Reference Hovsepyan, Olasagasti and Giraud2023). There is evidence that β rhythms are related to fluctuations in precision in multiple brain systems: In the visual cortex, rhythms in sensitivity of visual target detection (Fiebelkorn et al., Reference Fiebelkorn, Saalmann and Kastner2013) also involve β oscillations (Fiebelkorn et al., Reference Fiebelkorn, Pinsk and Kastner2018) that have been interpreted as representing predictions of precision (Aussel et al., Reference Aussel, Fiebelkorn, Kastner, Kopell and Pittman-Polletta2023); and in a model of the basal ganglia, β rhythms have been implicated in the maintenance of the currently active motor plan (Chartove et al., Reference Chartove, McCarthy, Pittman-Polletta and Kopell2020), which in active inference is a function of precision (Adams et al., Reference Adams, Shipp and Friston2013).
Coming full circle, inference over hierarchical generative models effectively implements an alternation between contextual and incremental processing. High-level distributions represent context informing low-level incremental inference; and low-level increments then inform updates to high-level contextual representations. This alternation is straightforward to implement when increments are regular and fixed, as in a recent model inferring sentential context from speech waveforms for a small set of examples containing syllables, lemmas, and sentences with uniform or contextually determined durations (Su et al., Reference Su, MacGregor, Olasagasti and Giraud2023).
12.4.5 Predicted Precision Paces Incremental Inference
Suggestively, the syllabic-timescale structure of speech is already favorably arranged for an alternating inference procedure, via the (statistical) alternation between vocalic nuclei and consonantal clusters (Nespor et al., Reference Nespor, Pena and Mehler2003; Ghitza, Reference Ghitza2013; Sun and Poeppel, Reference Sun and Poeppel2023). Vocalic nuclei, as we’ve discussed, are prominent features of speech acoustics (Chapters 5 and 11). Because vowels are smaller in number and longer in duration than consonants in most languages, vocalic nuclei potentially afford the inference of the time (window) of occurrence of a single linguistic unit with high precision, providing an important temporal landmark for speech perception that informs language learning early in development (Hochmann et al., Reference Hochmann, Benavides-Varela, Nespor and Mehler2011). Consonantal clusters, meanwhile, are composed of multiple temporally fleeting acoustic cues that can in aggregate provide highly reliable evidence about phoneme sequence identity, offering an opportunity to anchor the inference of speech content (Sun and Poeppel, Reference Sun and Poeppel2023; Chapter 39) relied on by adult listeners (Hochmann et al., Reference Hochmann, Benavides-Varela, Nespor and Mehler2011). Interestingly, the power of the β rhythms associated with visual sampling and motor plan maintenance are nested within ongoing θ rhythms (i.e., their power fluctuates at θ frequency) (Chartove et al., Reference Chartove, McCarthy, Pittman-Polletta and Kopell2020; Aussel et al., Reference Aussel, Fiebelkorn, Kastner, Kopell and Pittman-Polletta2023).
However, in Precoss-β, syllabic (i.e., θ) timescale alternations in PEP were less effective than β frequency alternations (Hovsepyan et al., Reference Hovsepyan, Olasagasti and Giraud2023). We propose that this is because, at the syllabic timescale, the effectiveness of alternating between inference of “what” and “when” is dependent on the alignment of the “what” phase with consonantal clusters, and the “when” phase with vocalic nuclei. Generally speaking, the issue here is that the information content of speech is not uniformly distributed in time; nor does it need to be for accurate speech perception, as demonstrated by, for example, repackaging. This suggests that, while rhythms may offer a hard-coded heuristic for the alternation of precision between different variables during online inference, the most effective way to alter precision is based on informed predictions or estimates of its fluctuation over the course of the speech signal.
The principles at play here are illuminated by results from drift-diffusion-to-bound models of decision making, in which evidence is integrated across time until a decision threshold is reached (Ratcliff et al., Reference Ratcliff, Smith, Brown and McKoon2016). This decision threshold varies across tasks and decisions, and even during single decisions, with the duration of deliberation (Malhotra et al., Reference Malhotra, Leslie, Ludwig and Bogacz2018). This variation depends on neural activity encoding both the conditional probability of obtaining further information (e.g., the hazard rate of occurrence of some probabilistic event [Herbst et al., Reference Herbst, Fiedler and Obleser2018]) and the opportunity cost of deliberation (i.e., how much potential reward is lost by continuing to deliberate), dubbed the urgency (Cisek et al., Reference Cisek, Puskas and El-Murr2009; Thura et al., Reference Thura, Cos, Trung and Cisek2014). Urgency alters the decision threshold or the speed of evidence accumulation, tuning the speed–accuracy trade-off to obtain the greatest possible cumulative reward (Cisek et al., Reference Cisek, Puskas and El-Murr2009; Thura et al., Reference Thura, Cos, Trung and Cisek2014). For example, relatively speaking, a task with a high reward rate will engender a high level of urgency and a low decision threshold, because a high error rate is less costly, in terms of cumulative reward, than a slow rate of return.
A similar trade-off is inherent in timing the shift from evidence gathering or accumulation to evidence integration or consolidation. Here, the quantity to be maximized is cumulative information gain rather than cumulative reward. One strategy for navigating this trade-off is to delay evidence integration until there is either a high level of precision or reliability, or a clear “decision point” or “deadline” (Hawkins et al., Reference Hawkins, Wagenmakers, Ratcliff and Brown2015). We propose that chunking (i.e., evidence integration) occurs when precision either meets or is predicted to meet a particular threshold. At the level of syllables, vowel onsets serve as both a highly reliable temporal landmark and a deadline for evidence integration (Ghitza, Reference Ghitza2013, Reference Ghitza2020), and may be predicted by a model of speech timing and prosody. Similarly, a sufficiently precise estimate of speech rate may provide an endogenous deadline for evidence integration. Interestingly, perceived speech speed reflects cognitive load (Bosker et al., Reference Bosker, Reinisch and Sjerps2017); we suggest that increased cognitive load decreases the level of certainty attainable in the time available to make a decision about speech content – that is, it requires lowering a decision threshold – which is interpreted subjectively as an increase in speech speed.
This notion of adaptive chunking furnishes an account of repackaging effects in which the silent interval between (signal-based) chunks provides a predictable cue for evidence consolidation. This account predicts that a different such cue – such as a tone, or a repeated nonsense syllable – would have the same effect as a silent interval. In language games such as Geta or Gibberish and Op, the easy intelligibility of speech in which a nonsense string such as “itig” or “op” is inserted at the beginning (between the onset and rime) of each syllable (Kazanina et al., Reference Kazanina, Bowers and Idsardi2018) provides anecdotal support for this hypothesis.
Adaptive chunking may also pace the timing of costly iterative inference. Reliable temporal landmarks such as vocalic nuclei and silent gaps in speech may be leveraged to engage synthetic mechanisms as in analysis-by-synthesis, or the backward pass of Bayesian smoothing. These periods of evidence integration over longer timescales may increase the certainty of inference and allow for reassessment of the interpretations of past data given new information. Furthermore, the timing of computations by estimates of precision may serve to align the cognitive, neural, and physiological processes of listeners to those of speakers (Pickering and Garrod, Reference Pickering and Garrod2004; Chapter 29) in the service of creating a shared conceptual and experiential space (Stolk et al., Reference Stolk, Verhagen and Toni2016). This may allow speech timing to cue content inference prospectively (Ten Oever and Martin, Reference Ten Oever and Martin2021) as well as retrospectively, as when a pause that’s longer than usual prompts the recovery of a double, less common, or more subtle meaning for previous words or phrases.
12.5 Vowel-Onset-Paced Syllable Inference
We integrate these ideas in a model called vowel-onset-paced syllable inference (VPSI). To illustrate VPSI, we return to the simplified template-matching account of segmentation described above (Section 12.4.4), in which we think of speech as a trajectory in feature space. The dimensions of this space represent phonemic features including, for example, vocalic or consonantal identity, position and manner of articulation (for consonants), openness and frontness (for vowels), nasalization, syllabification, and so on (see Moulin-Frier and Arbib, Reference Moulin-Frier and Arbib2013, for an example). We add a dimension representing the rate of change of the broadband speech amplitude envelope. We also assume two independent sources of noise – one contaminating the instantaneous rate of the envelope and another contaminating the phonemic dimensions of feature space.
12.5.1 VPSI Combines Rate-Dependent Content Inference with Irregular Re-estimation of Content and Rate
In VPSI, each candidate speech interpretation consists of a word sequence, which determines probability densities over syllable and phoneme sequences and over trajectories in feature space. (A phoneme sequence is mapped to a trajectory in feature space by concatenating canonical trajectories for individual phonemes.) Evidence for a given word sequence is calculated as the likelihood of the speech input given this probability distribution on trajectories in feature space, and accumulates continuously. VPSI begins with content inference informed by expectations about rate: The onset of an utterance initiates a filtering process of evidence accumulation about speech content, with the speed of rollout of candidate interpretations determined by the rate expected according to a prior on speech speed (i.e., by the expectation of this prior; Figure 12.2C). Independent evidence of vowel onsets is determined by the rate of change of the broadband speech envelope. Note that because VPSI infers a distribution over sequences of units, transitions between units are probabilistic events. The probability of a transition at a given point in time is determined by both the sum of probabilities over the set of sequences making a transition at that time and the rate of change of the broadband amplitude envelope.
In line with our proposal that inference is adaptively paced, and that speech speed depends on speech content, updates of the speech rate distribution occur at temporal landmarks (Figure 12.2C). These landmarks occur when either of two conditions is met: the level of precisionFootnote 3 of the distribution over speech content passes a threshold; or the probability of a transition passes a threshold. The latter can occur if the probability distribution over phoneme, syllable, or word sequences is such that nearly all probable sequences exhibit a transition at the same time point. It can also occur when a peak-rate event in the broadband speech amplitude envelope signals a vowel onset.
When a temporal landmark occurs, it triggers a post hoc re-estimation of both the timing and the content of the speech elapsed since the previous landmark (or since the onset of speech, if the current landmark is the first; Figure 12.2C). This involves using the transitional chunk of speech between landmarks to calculate evidence for a new set of candidate interpretations, each of which is associated with a different speech rate. These are determined by (e.g., sampled from) the current priors on speech rate and over word sequences, (i.e., the posteriors on speech rate and word sequences calculated at the previous and current temporal landmarks, respectively). For a given candidate interpretation, the likelihood of the transitional chunk factors as a product of the likelihood of the observed content (i.e., the trajectory in phonemic feature space) and the likelihood of the observed duration sequence (determined, to first order, by a linear scaling of the template for that interpretation). After re-estimation, the durations associated with the best-fitting template are used to update the speech rate distribution (Figure 12.2C). Thus, estimates of speech speed rely not only on the intervals between vowel onsets but also on the imputed distributions over syllable sequences for each chunk of speech.
12.5.2 VPSI Accounts for Results on Repackaging and Distal Rate Effects
Without vowel onsets, the response of VPSI depends on the signal-to-noise ratio (SNR) for phonemic features, and how close the speech rate is to the speech rate prior. If phonemic features are clear, and the speech rate prior is close to the actual speech rate, then content inference will reach a high level of precision. When this high precision triggers re-estimation, it ensures the set of candidate interpretations considered will have low morphosyntactic variance, varying mostly in speech rate. We propose this will result in rapid re-estimation that mostly serves to update speech rate based on the most-probable morphosyntactic candidate structures. If phonemic features are unclear, or the speech rate prior is not close to the actual speech rate, content inference will be imprecise, and there will be no updates of speech rate, resulting in a global failure to reach precision in inference.
If vowel onsets (or other reliable cues to speech timing) are present, even with inter-onset intervals corresponding to frequencies in the tails of the speech speed prior, we claim that the performance of VPSI will depend on the SNR for phonemic features alone. This is because even for a high phonemic feature SNR, a speech rate prior that is far off the mark will lead to a low level of certainty about speech content. Thus, the re-estimation of speech content will be primarily prompted by vowel onsets. This re-estimation will involve a large suite of candidate interpretations, whose speech rates are drawn from the tails of the current rate prior, enabling the sequential adjustment of the speech rate distribution. Note that despite large variability in the speech rate, the number of syllables these candidates contain is quantized, because the syllabic “distance” between two vowel onsets must be a whole number (e.g., it must consist of one, two, or three syllables), significantly reducing re-estimation’s computational cost. Another advantage of re-estimating at salient vowel onsets is that the computationally intensive re-estimation process occurs during the vocalic nucleus, a period of time when the information rate of speech is relatively low (Sun and Poeppel, Reference Sun and Poeppel2023).
We suggest this situation provides an account of repackaging effects, with the silent gaps inserted between packages playing the role of reliable timing cues. This interpretation suggests two factors contributing to the ∼9 Hz ceiling on the syllabic rate. The second is the support of the speech rate prior, that is, the interval over which most of the prior’s probability mass is distributed; this may be speaker-, listener-, and language-specific. For instance, in English, the number of possible syllables between vowel onsets is three, and the ceiling is three times the mean syllabic rate; we predict that in moraic languages such as Japanese, the ceiling on the syllabic rate will be a lower multiple of the average speech rate. Likely even more important is the duration of the re-estimation process; this process, which we propose normally occurs following vowel onset, is triggered in this case by the onset of the silent period, and must fit within the silent interval.
In distal rate effects, vowel onsets occur in the first syllable of the target segment and following the target segment, but not during the reduced function word. Despite a high phonemic feature SNR, the ambiguity of the coarticulated, blended target content results in imprecise content inference. Thus, when the vowel onset following the target segment triggers re-estimation, both the with (function word) and without (function word) interpretations will be within the re-estimation search space. The content of both interpretations has a high likelihood, and we assume they are equally likely given prior speech. Thus, the competition between the two interpretations will be decided by the likelihood of their respective imputed duration sequences, and with certain assumptions it should be possible to write down the probability of choosing one interpretation over another.Footnote 4
12.6 Conclusion
We suggest that in online speech inference, strict temporal constraints advantage computational efficiency, above and beyond tractability or computability (Adolfi et al., Reference Adolfi, Wareham and van Rooij2023). Like others (Halle and Stevens, Reference Halle and Stevens1962; Christiansen and Chater, Reference Christiansen and Chater2016; Martin and Doumas, Reference Martin and Doumas2017; Brown et al., Reference Brown, Tanenhaus and Dilley2021; Friston et al., Reference Friston, Sajid and Quiroga-Martinez2021; Hovsepyan et al., Reference Hovsepyan, Olasagasti and Giraud2023), we highlight the cost of search operations (Tschantz et al., Reference Tschantz, Millidge, Seth and Buckley2023), the size of the search space, and the speed–accuracy trade-off as major computational challenges.
We propose that listeners address these challenges with adaptive timing of inference. EM-like inferential turn taking between processing streams (e.g., “what” and “when”) (Dempster et al., Reference Dempster, Laird and Rubin1977; Hovsepyan et al., Reference Hovsepyan, Olasagasti and Giraud2023) and hierarchical levels (at timescales characteristic of each level [Halle and Stevens, Reference Halle and Stevens1962; Friston et al., Reference Friston, Trujillo-Barreto and Daunizeau2008; Christiansen and Chater, Reference Christiansen and Chater2016; Martin and Doumas, Reference Martin and Doumas2017; Su et al., Reference Su, MacGregor, Olasagasti and Giraud2023]) optimize the timing of evidence accumulation and evidence consolidation. These alternations are timed by second-order predictions of temporal fluctuations in precision (e.g., the syllabic-timescale alternation between temporally informative vowels and phonologically informative consonants [Nespor et al., Reference Nespor, Pena and Mehler2003; Ghitza, Reference Ghitza2013; Sun and Poeppel, Reference Sun and Poeppel2023]), located within a generative model of speech timing that heavily overlaps with a speaker model predicting temporal regularities arising from physiological, motor, and cognitive processes (Elliott and Theunissen, Reference Elliott and Theunissen2009; Bishop and Intlekofer, Reference Bishop and Intlekofer2020; Ten Oever and Martin, Reference Ten Oever and Martin2021). Listeners may use such temporal information to pace their own computational, neural, and even neurophysiological processes, aligning them with those of the speaker (Pickering and Garrod, Reference Pickering and Garrod2004; Chapter 29). This “cooperative pacing” may facilitate the construction of a shared conceptual space (Stolk et al., Reference Stolk, Verhagen and Toni2016) and allow listeners to leverage aspects of speech production that facilitate information transfer (Aylett and Turk, Reference Aylett and Turk2004; Mahowald et al., Reference Mahowald, Dautriche, Gibson and Piantadosi2018; Ten Oever and Martin, Reference Ten Oever and Martin2021).
Incorporating a hierarchical generative model of speech timing, the VPSI model performs rate-based continuous inference of speech content, re-estimating the rate and content of recent speech using generative, synthetic mechanisms when sufficient precision about speech content or timing is attained. Content precision triggers rapid re-estimation; timing precision triggers computationally intensive re-estimation occurring when the content information rate is low. Re-estimation in VPSI accounts for both repackaging and distal rate effects, iteratively improving speech rate estimates and resolving ambiguity in speech content.
We predict timing-related prediction errors, and processes of model evaluation and construction following reliable temporal cues, should have neurophysiological traces. The neural signatures of precision should alternate between brain regions computing and representing distinct processing streams and hierarchical levels. If speech rate estimates depend on both lexical and acoustic cues, then lexical manipulations can affect rate perception, as well as vice versa; distal rate effects may require the content precision provided by syntactic and semantic structure (Pitt et al., Reference Pitt, Szostak and Dilley2016); and data requirements for rate estimation may underlie results on minimum intelligible segments excised from running speech (Pollack and Pickett, Reference Pollack and Pickett1963).
Expressing the computational problem of online segmentation mathematically (Adolfi et al., Reference Adolfi, Wareham and van Rooij2023) is crucial to precisely articulating and quantifying the impact of the challenges discussed. So is clarifying whether the adaptive updating mechanisms illustrated in VPSI are “built in” to predictive processing via free-energy minimization in hierarchical generative models, or must be explicitly engineered in (Hovsepyan et al., Reference Hovsepyan, Olasagasti and Giraud2023; Tschantz et al., Reference Tschantz, Millidge, Seth and Buckley2023). A deeper understanding of the computational and algorithmic architecture of word segmentation (Adolfi et al., Reference Adolfi, Wareham and van Rooij2023) can provide a platform for exploring the neural signatures (and potential implementations) of speech perception (Doelling and Assaneo, Reference Doelling and Assaneo2021), through biophysically detailed brain region- and operation-specific neurophysiological mechanisms and mathematical models (Oganian and Chang, Reference Oganian and Chang2019; Cannon, Reference Cannon2021; Cannon and Patel, Reference Cannon and Patel2021; Doelling and Assaneo, Reference Doelling and Assaneo2021; Frühholz and Schweinberger, Reference Frühholz and Schweinberger2021; Pittman-Polletta et al., Reference Pittman-Polletta, Wang and Stanley2021; Doelling et al., Reference Doelling, Arnal and Assaneo2023; Gwilliams et al., Reference Gwilliams, King, Marantz and Poeppel2022; Adolfi et al., Reference Adolfi, Wareham and van Rooij2023).
12.7 Acknowledgements
We thank Jon Cannon, Sevada Hovsepyan, Yohan John, Tom Lagatta, Mamady Nabé, Itsaso Olasagasti, Johanna Rimmele, Yaqing Su, and an anonymous reviewer for many useful discussions and comments that contributed to this work.
Summary
We propose that the probabilistic modeling of speech timing plays a key role in word segmentation, by predicting the reliability of speech information across channels and linguistic levels. These predictions enable the efficient and adaptive timing of model updates and resource-intensive computations, supporting an optimal trade-off between speed and accuracy.
Implications
This proposal implies that extra-morphosyntactic regularities predict fluctuations in the uncertainty of the speech signal, and that cognitively demanding operations occur during times of low predicted uncertainty. It suggests experiments assessing the relationship between higher-order speech statistics, behavior, and neural activity that distinguish the proposed mechanisms from influential bottom-up accounts.
Gains
Our account of prosodic modeling for speech understanding formulates detailed, novel hypotheses about how the brain times information processing across modalities and the linguistic hierarchy. These have the potential to advance both the understanding of human speech processing and its neurophysiology, and the performance of artificial language systems.






