Skip to main content Accessibility help
×
Hostname: page-component-76d6cb85b7-8p85h Total loading time: 0 Render date: 2026-07-19T13:36:42.898Z Has data issue: false hasContentIssue false

8 - A Polychromatic Portrait of Speech Rhythm

from Section 2 - Acoustic and Sublexical Rhythms

Published online by Cambridge University Press:  23 April 2026

Lars Meyer
Affiliation:
Max Planck Institute for Human Cognitive and Brain Sciences
Antje Strauss
Affiliation:
University of Konstanz

Summary

The “speech envelope” is often used as an acoustic proxy for neural rhythm. The problem is its assumption that the unfiltered, broadband signal can satisfactorily model neural modulation in the auditory pathway (and beyond). However, the auditory system does not function as a passive transducer but rather decomposes and segregates the signal into an array of tonotopically organized frequency channels. This modulation filtering results in a partitioning of slow (3–20 Hz) neural modulation patterns across the tonotopic axis that bear only a passing resemblance to the broadband speech envelope. Such polychromatic diversity (in frequency, magnitude, and phase) of auditory modulation patterns is critical for decoding the speech signal, as it highlights critical linguistic properties such as articulatory-acoustic and prosodic features important for decoding and understanding spoken language. The low-frequency modulation patterns associated with high-frequency (>2 kHz) auditory channels are especially important for prosodic processing and consonant discrimination, both key for speech intelligibility, especially in adverse listening conditions and among the hard of hearing.

Information

Figure 0

Figure 8.1 The speech envelope illustrated.The “speech envelope” for a single sentence (shown in black). The temporal fine structure is portrayed in very light gray. Syllable boundaries are indicated by dotted vertical lines. The contour shows the “rate of change” in envelope energy, analogous to “delta features” used in certain ASR applications. Note how coarse the speech envelope is relative to the speech signal’s finer details. The envelope contour pertains to the broadband, unfiltered signal.Figure 8.1 long description.

Adapted from Oganian and Chang (2019).
Figure 1

Figure 8.2 Speech waveforms, spectrograms, and modulation spectra across acoustic frequencies.Spectrographic and time domain representations of the single sentence “The most recent geological survey found seismic activity” (Greenberg et al., 1998). The waveforms are plotted on the same amplitude scale, while the scale of the original, unfiltered signal is compressed by a factor of five for illustrative clarity. The frequency axis of the spectrographic display of the channels has been nonlinearly compressed for illustrative purposes. Note the quasi-orthogonal temporal registration of the waveform modulation pattern across frequency channels. On the right are modulation spectra (magnitude component) associated with each of four, 1/3-octave channels. The peak of the spectrum (in all but the highest channel) lies between 4 Hz and 6 Hz. Note the large amount of energy in the higher-modulation frequencies associated with the highest-frequency channel. The modulation spectra of the four-channel compound and the original, unfiltered signal are illustrated for comparison (top panel).

From Greenberg et al. (1998).
Figure 2

Figure 8.3 Speech intelligibility and the CMS.The CMS integrates the magnitude and phase components into a single value. The sentence material’s intelligibility for a listening experiment was manipulated by locally time-reversing the speech signal over different segment lengths. As the reversed-segment duration increases beyond 40 ms, intelligibility declines precipitously, as does the magnitude of the CMS. The spectro-temporal properties (and articulatory-acoustic features) also deteriorate appreciably under such conditions.Figure 8.3 long description.

Reprinted from Greenberg (2022). The figure is an adaptation of one originally published in Greenberg and Arai (2001, 2004).

Save book to Kindle

To save this book to your Kindle, first ensure no-reply@cambridge.org is added to your Approved Personal Document E-mail List under your Personal Document Settings on the Manage Your Content and Devices page of your Amazon account. Then enter the ‘name’ part of your Kindle email address below. Find out more about saving to your Kindle.

Note you can select to save to either the @free.kindle.com or @kindle.com variations. ‘@free.kindle.com’ emails are free but can only be saved to your device when it is connected to wi-fi. ‘@kindle.com’ emails can be delivered even when you are not connected to wi-fi, but note that service fees apply.

Find out more about the Kindle Personal Document Service.

Available formats
×

Save book to Dropbox

To save content items to your account, please confirm that you agree to abide by our usage policies. If this is the first time you use this feature, you will be asked to authorise Cambridge Core to connect with your account. Find out more about saving content to Dropbox.

Available formats
×

Save book to Google Drive

To save content items to your account, please confirm that you agree to abide by our usage policies. If this is the first time you use this feature, you will be asked to authorise Cambridge Core to connect with your account. Find out more about saving content to Google Drive.

Available formats
×