Skip to main content Accessibility help
×
Hostname: page-component-76d6cb85b7-pn7tm Total loading time: 0 Render date: 2026-07-19T20:10:38.882Z Has data issue: false hasContentIssue false

12 - Adaptive Pacing in Word Segmentation and the Vowel-Onset-Paced Syllable Inference Model

from Section 2 - Acoustic and Sublexical Rhythms

Published online by Cambridge University Press:  23 April 2026

Lars Meyer
Affiliation:
Max Planck Institute for Human Cognitive and Brain Sciences
Antje Strauss
Affiliation:
University of Konstanz

Summary

In speech perception, timing and content are interdependent. For example, in distal rate effects, context speech rate determines the number of words, syllables, and phonemes heard in an unchanging target speech segment. Such results confront psycholinguistic theory with the chicken-and-egg problem of concurrently inferring speech timing and content, and the interrelated issues of narrowing the search space of speech interpretations without bias and optimizing the speed/accuracy tradeoff in online processing. We propose listeners address these issues by managing the timing of speech-related computations. Specifically, we claim: (1) Listeners model speech timing as part of a speaker model; (2) variable-length sequences of morphosyntactic units are the basic increments of speech inference; and (3) listeners adaptively schedule inferential updates and computationally intensive operations according to (4) fluctuations in uncertainty predicted by the speaker model. We illustrate these claims in a mechanistic model – vowel-onset-paced syllable inference – explaining multiple psycholinguistic results, including distal rate effects.

Information

Figure 0

Figure 12.1A. Distal rate effects. (i) A sentence with a target segment (“summer or”) containing a reduced, coarticulated function word (“or”) (speech waveform in gray, target segment in black). (ii) A version of the same sentence manipulated to slow context speech rate. (iii) Subjects asked to repeat the sentence report fewer function words for the slowed context.Figure 12.1A. long description.

Figure 1

Figure 12.1B. The SI hypothesis. (i) An illustration of SI for the normal context speech stimulus from (A). A mean speech rate (μ_1) is computed from the interpretation(s) of context speech. For each candidate interpretation of the target segment (“summer or” and “summer”), knowledge of each syllable’s relative duration is combined with μ_1 to obtain an estimated candidate duration. These estimated candidate durations are compared to the observed (e.g., acoustic) duration (ν). The candidate that best explains what is heard (in this case, the one most probable given the observed duration) is perceived. (ii) An illustration of SI for the slowed context speech stimulus from (A). Here, a slower mean speech rate (μ_2) leads to a different judgment of which candidate interpretation is most probable.Figure 12.1B. long description.

Acoustic time series and envelope images adapted from Peelle and Davis (2012) (licensed under CC BY 3.0).
Figure 2

Figure 12.2A. A hierarchical Bayesian network for word segmentation. Hierarchically organized latent variables representing sequences of words (Wordi), syllables (Syli), and phonemes (Φi) constitute a generative model of morphosyntactic structure. Hierarchically organized latent variables representing syllabic rate (Rate) and sequences of syllable durations (DSyl,i) and phoneme durations (DΦ,i) constitute a generative model of speech timing. Together, they specify a speech signal as a trajectory in feature space.Figure 12.2A. long description.

Figure 3

Figure 12.2B. Interdependence of contextual and incremental processing. The size of each increment depends on the mean rate, and vice versa. Left and right show two different interdependent sets of increments and mean rate.Figure 12.2B. long description.

Figure 4

Figure 12.2C. The VPSI mechanism. A rate prior enables rate-dependent inference of content. When reliable timing information arrives, it triggers post hoc content re-estimation, which leads to rate re-estimation and a rate posterior. This serves as the rate prior for rate-dependent content inference of the next chunk of speech.

Acoustic time series and envelope images adapted from Peelle and Davis (2012) (licensed under CC BY 3.0).

Save book to Kindle

To save this book to your Kindle, first ensure no-reply@cambridge.org is added to your Approved Personal Document E-mail List under your Personal Document Settings on the Manage Your Content and Devices page of your Amazon account. Then enter the ‘name’ part of your Kindle email address below. Find out more about saving to your Kindle.

Note you can select to save to either the @free.kindle.com or @kindle.com variations. ‘@free.kindle.com’ emails are free but can only be saved to your device when it is connected to wi-fi. ‘@kindle.com’ emails can be delivered even when you are not connected to wi-fi, but note that service fees apply.

Find out more about the Kindle Personal Document Service.

Available formats
×

Save book to Dropbox

To save content items to your account, please confirm that you agree to abide by our usage policies. If this is the first time you use this feature, you will be asked to authorise Cambridge Core to connect with your account. Find out more about saving content to Dropbox.

Available formats
×

Save book to Google Drive

To save content items to your account, please confirm that you agree to abide by our usage policies. If this is the first time you use this feature, you will be asked to authorise Cambridge Core to connect with your account. Find out more about saving content to Google Drive.

Available formats
×