26.1 What Is Rhythmic about Speech and Song?
Rhythms surround us: from footsteps in the hallway to the blinking red of a traffic signal. Although rhythms are a pervasive part of daily life, they play an important role in vocal communication for humans and a range of species (Honing, Reference Honing2012; Merchant and Honing, Reference Merchant and Honing2014; Patel, Reference Patel2010; Ravignani and Madison, Reference Ravignani and Madison2017). In humans, rhythms are present in two primary means of communication: speech and song. People perceive rhythms in speech from prosodic and linguistic stress patterns of syllables, and perceive rhythms in song from notes as they unfold in time (Patel, Reference Patel2010). Although the definition of rhythm can vary considerably, here we adopt a music cognition perspective and describe rhythm in its broadest sense as a pattern of events unfolding in time (McAuley, Reference McAuley, Riess Jones, Fay and Popper2010; Ravignani and Madison, Reference Ravignani and Madison2017). As such, speech events are typically syllables and song events are notes. This definition is inclusive of many different types of repeating sounds, such as cricket chirps, waves crashing on the beach, chewing food, and even a dough hook smacking around the bowl of an electric mixer.
Speech and song are both rhythmic, but a key difference between the two domains is how events relate to one another. In song, and music more generally, rhythms are organized around perceptually strong events, called a beat or pulse (Povel and Essens, Reference Povel and Essens1985). The alternation of strong and weak beats in music occurs at roughly equal intervals, which gives rise to the perception of isochrony (i.e., “iso” meaning “same” and “chronos” meaning “time”) despite considerable variation in note lengths or timing of note onsets relative to the beat. In many cases, a physical event is not always present on the beat and yet listeners perceive a strong pulse in that location (London, Reference London2004; Nave et al., Reference Nave, Snyder and Hannon2023; Parncutt, Reference Parncutt1994; Snyder and Krumhansl, Reference Snyder and Krumhansl2001; Temperley, Reference Temperley2004). When rhythms are structured around the beat, note durations are related by integer multiples with small integer ratios. That means that a long note in a piece of music is one, two, three, or four times the duration of a short note in the sequence (e.g., quarter versus half note; Jacoby and McDermott, Reference Jacoby and McDermott2017; Roeske et al., Reference Roeske, Tchernichovski, Poeppel and Jacoby2020). In contrast, speech rhythms are dictated by the word length, syntactic structure, and prosodic emphasis of the specific utterance (Patel, Reference Patel2010; Selkirk, Reference Selkirk1980; Turk and Shattuck-Hufnagel, Reference Turk and Shattuck-Hufnagel2013). For instance, stressed syllables are often longer than unstressed syllables (Cutler and Foss, Reference Cutler and Foss1977; Fry, Reference Fry1955, Reference Fry1958; Seidl et al., Reference Seidl, French, Wang and Cristia2014). Articles, such as a, an, and the, occur in fixed positions relative to content words in English, and are unstressed, monosyllabic, and shorter than many content words, creating an unstressed-stressed rhythm (Andreou et al., Reference Andreou, Kashino and Chait2011; Hayes, Reference Hayes1985). Words at the end of a phrase are lengthened (Klatt, Reference Klatt1975), words that receive prosodic emphasis for effect are also held out longer (Selkirk, Reference Selkirk1995), and different languages have different stress patterns or a lack of stress at all (Cutler and Clifton, Reference Cutler, Clifton, Bouma and Bouwhuis1984; Peperkamp et al., Reference Peperkamp, Vendelin and Dupoux2010). The rhythmic patterns of speech can be predictable and repeating, sometimes even alternating between stressed and unstressed syllables much like the definition of a beat above (Breen and Clifton Jr, Reference Breen and Clifton2011; Cutler and Foss, Reference Cutler and Foss1977; Levey and Raphael, Reference Levey and Raphael2002; Patel, Reference Patel2010). However, syllable durations do not relate to one another by consistent integer relationships, which means that there is no regular pulse or beat in speech.
We present two additional terms, periodicity and rhythmic regularity, to carefully differentiate how rhythmic events relate to each other in speech and song. Periodicity is defined as a pattern repeating in time (Patel, Reference Patel2010) and is traditionally related to a fixed interval or period with which the pattern repeats. However, speech research has long used the term periodicity to mean repeating patterns, such as alternations of stressed and unstressed syllables, that do not have a fixed period (Beier and Ferreira, Reference Beier and Ferreira2018). So here we use this term in that sense of a repeating pattern that does not repeat at a fixed interval, but instead at some variable period (e.g., like the meaning of periodic). To illustrate this notion, imagine an auditory notification played to alert travelers that a subway train is about to arrive (a quick sequential ascending major chord: G-B-D-G) and a new train comes every four to six minutes (see Figure 26.1A). This musical alert is periodic – it occurs repeatedly whenever a train arrives – and is also rhythmic – the four-note chord has a short-short-short-long pattern of note events – but it does not give rise to the perception of a regular pulse. Rhythmic regularity is when an alternation between strong and weak events gives rise to a beat (i.e., a series of perceptually strong events; Large, Reference Large and Grondin2008) and the intervals between strong events occur at roughly equal intervals (Ravignani and Madison, Reference Ravignani and Madison2017; Ravignani and Norton, Reference Ravignani and Norton2017). Think again of the train arrival notification. If the train arrived every 2,400 ms and the musical alert itself was composed of three 200 ms intervals followed by one 600 ms interval, then it would give rise to a regular pulse, or a beat in music. Specifically, the first and last notes would become strong beats because the (implausible) 2,400 ms train arrival window would suggest two 600 ms silent beats, making the whole sequence a cycle of four beats (see Figure 26.1B). In everyday English speech, an individual phrase provides a rhythm of stressed and unstressed syllables, and, when we speak in longer utterances, new rhythmic patterns with similar properties repeat at irregularly spaced phrasal and prosodic intervals. As such, speech is considered periodic: It is a repeating rhythmic pattern of events, but it is not rhythmically regular. In contrast, song is rhythmically regular, where a repeating pattern occurs at intervals of the same duration and gives rise to the perception of an underlying pulse or beat. In music, these beats are grouped into a hierarchy of strong and weak beats (e.g., groups of three or four beats per measure; Jones, Reference Jones1985; Lerdahl and Jackendoff, Reference Lerdahl and Jackendoff1983) called meter. Although speech has also been described as having meter (e.g., Selkirk, Reference Selkirk1984), metricality in speech is the alternation of strong and weak syllables without a fixed integer-related interval relative to previous syllables. Speech, even at the metrical level, is more about the recurrence of a particular pattern (periodicity) instead of the presence of an underlying beat occurring at isochronous intervals.
Illustrating rhythmic regularity as different from periodic.
Illustration of the difference between a signal that is A) periodic and rhythmic compared to one that is B) rhythmically regular and gives rise to the percept of a beat.


The differences between periodicity and rhythmic regularity may seem subtle, but the ramifications of the presence or absence of rhythmic regularity are anything but subtle. Musical behaviors such as dancing, singing, and playing an instrument all depend on the presence of rhythmic regularity to allow an individual or multiple individuals to move or play together in an ensemble or to play with music (Hannon et al., Reference Hannon, Nave‐Blodgett and Nave2018; Honing, Reference Honing2012; Nave-Blodgett et al., Reference Nave-Blodgett, Snyder and Hannon2021). Even the survival of the human species has been related back to rhythmic coordination and its importance for social bonding, especially through coordination with a beat as in music-making (Savage et al., Reference Savage, Loui and Tarr2021; Trainor and Cirelli, Reference Trainor and Cirelli2015).
Music is not the only stimulus that has rhythmic regularity, even if it is the strongest naturally occurring example of it. For instance, our heartbeat is perhaps the first regular beat we hear, with a weak-strong alternation in the lub-dub-lub-drub of a normal heartbeat cycle (De Meo et al., Reference De Meo, Matusz and Knebel2016; Ullal-Gupta et al., Reference Ullal-Gupta, Vanden Bosch der Nederlanden, Tichko, Lahav and Hannon2013). Even speech can have a beat, exemplified through the uncharacteristically regular alternation of unstressed and stressed syllables in iambic pentameter (e.g., “the rose is bright and shines a beam of light”) and other forms of poetry. When speech has a beat, that type of speech arguably plays a qualitatively different functional role than speech without a beat. Instead of a speech versus song dichotomy based on the presence or absence of rhythmic regularity, it is more likely that a communicative continuum exists with speech and song as the endpoints and several different types of vocalizations falling between, including poetry, rap, sprechstimme, public speeches, emotional speech, acting, infant-directed speech, and more (Vanden Bosch der Nederlanden et al., Reference Vanden Bosch der Nederlanden, Qi and Sequeira2022b). There are many acoustic features that move a stimulus along the speech-to-song continuum (Albouy et al., Reference Albouy, Mehr, Hoyer, Ginzburg and Zatorre2024; Brown, Reference Brown, Wallin, Merker and Brown2000; Fitch, Reference Fitch2006; Ozaki et al., Reference Ozaki, Tierney and Pfordresher2024; Vanden Bosch der Nederlanden et al., Reference Vanden Bosch der Nederlanden, Qi and Sequeira2022b), but rhythmic regularity is key to differentiating the endpoints of this continuum (see Chapters 11, 32, and 35).
26.2 Examining Acoustic Features of Rhythmic Regularity in Speech and Song
When adults describe acoustic differences between speech and song, they tend to focus on the different uses of pitch, while younger children (4–11 years of age) describe differences based on acoustic timing (Vanden Bosch der Nederlanden et al., Reference Vanden Bosch der Nederlanden, Qi and Sequeira2022b). Our previous qualitative work simply asked children and adults to describe the difference between speech and song in their own words, and we reported themes from their responses binned across all temporal acoustic features and other features (melodic, spectral, loudness). Here, we provide a reanalysis of these themes (data here: https://osf.io/xar3c/) to examine how children and adults describe different temporal features, such as the presence of a beat, throughout development. As is clear from Table 26.1, temporal characteristics were mentioned infrequently for describing the differences between speech and song (5–30% of participants depending on the age group: N=19, 4–7-year-olds; N=16, 8–11-year-olds; N=16, 12–17-year-olds; and N=47, 18–24-year-olds). Children and adults mentioned tempo at similar rates (reporting that song is faster than speech for children, but that speech is faster than song for adults); smoothness or flow (song is more smooth than speech) – which could be related to temporal and/or spectral features – increased with age; and adults mentioned beat or rhythm (lumped together because children’s definitions did not necessarily mention beat but talked about rhythm in a beat-like manner) more often than children overall (reporting more beat/rhythm in song than speech). Although these qualitative data suggest that rhythmic regularity is not the first thing people focus on in their verbal or written descriptions, perceptual judgments reported by our group (Vanden Bosch der Nederlanden et al., Reference Vanden Bosch der Nederlanden, Qi and Sequeira2022b) suggest otherwise: Adults rank the presence of a beat and rhythmic regularity in the top three characteristics that distinguish speech from song, and children of all ages agree that song has a beat.
Reanalysis of the qualitative data of Vanden Bosch der Nederlanden et al. (Reference Vanden Bosch der Nederlanden, Qi and Sequeira2022b) illustrating that individual participants’ spontaneous descriptions of the differences between speech and song infrequently describe beat or rhythm as a key differentiator compared to pitch or melody. Values are the proportion of participants endorsing each acoustic feature by age group.
| Age group | Beat/rhythm | Tempo | Flow |
|---|---|---|---|
| 4–7 | 0.11 | 0.11 | 0.05 |
| 8–11 | 0.13 | 0.13 | 0.19 |
| 12–17 | 0.06 | 0.13 | 0.19 |
| 18–24 | 0.23 | 0.09 | 0.30 |
Since rhythmic regularity is important for distinguishing speech from song, how is the degree of rhythmic regularity measured across these domains? Many studies characterize the strength of a beat in music (Alluri and Toiviainen, Reference Alluri and Toiviainen2010; Burger et al., Reference Burger, Thompson, Luck, Saarikallio and Toiviainen2013, Reference Burger, Thompson, Luck, Saarikallio and Toiviainen2014; De Gregorio et al., Reference De Gregorio, Valente and Raimondi2021; Henry et al., Reference Henry, Herrmann and Grahn2017; Lartillot et al., Reference Lartillot, Eerola, Toiviainen and Fornari2008; Matthews et al., Reference Matthews, Witek, Heggli, Penhune and Vuust2019; Roeske et al., Reference Roeske, Tchernichovski, Poeppel and Jacoby2020) and the degree of durational contrastiveness in speech (i.e., short–long) (Arvaniti, Reference Arvaniti2012; Grabe and Low, Reference Grabe and Low2002; Jadoul et al., Reference Jadoul, Ravignani, Thompson, Filippi and De Boer2016; Nolan and Jeon, Reference Nolan and Jeon2014; White and Mattys, Reference White and Mattys2007; Wiget et al., Reference Wiget, White and Schuppler2010), but few compare the degree of rhythmic regularity across speech and song. Our recent study (Yu et al., Reference Yu, Cabildo, Grahn and Vanden Bosch der Nederlanden2023) illustrated that subjective ratings of regularity reliably differentiated speech from song, when regularity was operationalized as how easy it would be to tap or clap along with the stimulus. Regularity ratings were also predicted by the acoustic features of spectral flux and syllable duration, which meant that less moment-to-moment change in the spectrum (i.e., spectral flux) and longer durations were associated with higher regularity ratings (Yu et al., Reference Yu, Cabildo, Grahn and Vanden Bosch der Nederlanden2023). This finding held true even after controlling for modality (speech and song), suggesting that these features predicted regularity ratings for both speech and song. Other studies that examine what features listeners use to differentiate speech and song can also shed light on key temporal characteristics. Speech and song are differentiated by their temporal modulation rates (e.g., song is slower than speech; Ding et al., Reference Ding, Patel and Chen2017), spectro-temporal modulation rates (co-occurrence of high-frequency spectral modulation with low-frequency temporal modulation for song relative to speech; Albouy et al., Reference Albouy, Mehr, Hoyer, Ginzburg and Zatorre2024), pulse clarity (Hilton et al., Reference Hilton, Moser and Bertolo2022), tempo (Ozaki et al., Reference Ozaki, Tierney and Pfordresher2024), and rhythmic regularity (Ozaki et al., Reference Ozaki, Tierney and Pfordresher2024; see their exploratory analyses). These features, including spectral flux and duration mentioned above, are certainly dependent on temporal dynamics, but nearly all fail to capture the integer-multiple relationships that characterize song’s alignment with a beat. It is still an open question what features listeners use to perceive rhythmic regularity in speech and song and whether the features reported here are merely a byproduct of the integer relationships in sounds with a beat.
One issue with the practice of relating acoustic features in speech and song to subjective ratings of regularity is the influence human perception has in altering the incoming signal. For instance, naturally occurring music does not have strictly equal durations between beats because of motor output delays, micro-timing deviations, or expressive timing (Appen et al., Reference Appen, Doehring and Moore2015; Danielsen et al., Reference Danielsen, Haugen and Jensenius2015; Datseris et al., Reference Datseris, Ziereis and Albrecht2019; Drake and Palmer, Reference Drake and Palmer1993; Rasch, Reference Rasch and Sloboda1988; Senn et al., Reference Senn, Kilchenmann, Von Georgi and Bullerjahn2016). People perceptually regularize, or quantize, the signal to make it sound more rhythmically regular than the physical signal would suggest. It is also likely that people perceive the moment of the beat in natural stimuli later than the sound’s physical onset. In musical timbres, people report that the perceptual occurrence of an event, or p-center, is delayed relative to the physical onset of the sound (see Chapter 11). Fast attack time sounds (e.g., percussion) have p-centers with shorter delays relative to the onset than slow attack time sounds (e.g., violin), and longer notes have p-centers with longer delays relative to onsets compared to shorter notes (Danielsen et al., Reference Danielsen, Nymoen and Anderson2019). The timing of the p-center is also dependent on the experience or expertise of the listener, which leads researchers to describe p-centers in music as “beat bins,” or windows of time, in which people are most likely to hear a beat (Speich et al., Reference Spiech, Endestad, Laeng, Danielsen and Haghish2023).
The concept of the p-center originates from speech studies (Morton et al., Reference Morton, Marcus and Frankish1976), but this literature is quite mixed as to the acoustic correlates of the p-center (Villing, Reference Villing2010). The most consistent finding relates the p-center to a window of time near the vowel onset (Fox and Lehiste, Reference Fox and Lehiste1987; Marcus, Reference Marcus1981). Like music, the p-center in speech is dependent on the “attack” of the consonants preceding the vowel onset (Cooper et al., Reference Cooper, Whalen and Fowler1986; Pompino-Marschall, Reference Pompino-Marschall1989). More recent studies have used sensorimotor synchronization tasks such as tapping to looped speech to characterize that the vowel onset is a strong attractor of taps in speech and music (Lidji et al., Reference Lidji, Palmer, Peretz and Morningstar2011; Rathcke et al., Reference Rathcke, Lin, Falk and Dalla Bella2021). More work is needed to characterize what acoustic factors contribute to the perceptual location of a beat in speech, song, and a range of natural environmental stimuli. If beat bins can be used to calculate integer-multiple relationships from stimulus features in complex stimuli such as speech and song, then beat bins could also be used to calculate integer-multiple relationships in a range of naturally occurring stimuli.
In summary, rhythmic regularity is an important feature for differentiating speech and song, children are sensitive to temporal features, but beat/rhythm are not critical for differentiating speech and song until later in development (see Section 6 for more about development). Finally, subjective ratings of rhythmic regularity reliably differentiate speech from song, although acoustic features (i.e., spectral flux, duration) that predict regularity ratings do not capture the integer relationships inherent to stimuli with a hierarchical beat structure. As outlined above, next steps include determining if psychologically informed metrics of rhythmic regularity are stronger predictors of regularity than easily extracted acoustic features, such as spectral flux or syllable duration. The long-term goal of this work is to characterize the degree of regularity in all sounds, including those beyond human communicative contexts, such as dog barks, typing, bird calls, water droplets, and sneezes. As such, acoustic metrics of regularity should be general enough to be applied in non-communicative contexts.
So why is it important to characterize the degree of regularity in speech, song, and environmental sounds? We now turn to the existing literature on neural and perceptual dynamics of rhythm to provide rationale for the important functions rhythmic regularity plays in attention, learning, and memory.
26.3 Attention Is Rhythmic
Already 60 years ago, Mari Reiss Jones outlined her work on dynamic attending theory (DAT) (Jones, Reference Jones1976), which proposed that attention is not constant but cyclical, with alternations between moments of focused attention followed by inattention (Large and Jones, Reference Large and Jones1999). In this theory, external stimulation, such as music, can guide attention toward moments in time when important information or events have a high likelihood of occurrence. In a musical sequence, attention would be allocated to the strong beats, when important melodic or metrical information is likely to occur, and attention would wane on weaker or off-beat positions. If attention is allocated more often to the strong than to the weak beats, this should be evident as better processing and sequencing of information occurring on the beat than off. Indeed, people are better at remembering faces when they are presented on the beat of a musical rhythm compared to when they are presented off the beat or in silence (Johndro et al., Reference Johndro, Jacobs, Patel and Race2019).
A stronger test of rhythmic facilitation through DAT is whether an external rhythm can start a cycle of focused attention and inattention at the rate of the rhythmic stimulus that continues even when the rhythm is no longer present. Several studies have shown that after a rhythmically regular prime, participants have better memory for information played after the prime, such as speech (Canette et al., Reference Canette, Fiveash and Krzonowski2020a; Cason and Schön, Reference Cason and Schön2012; Cason et al., Reference Cason, Astésano and Schön2015; Chern et al., Reference Chern, Tillmann, Vaughan and Gordon2018; Przybylski et al., Reference Przybylski, Bedoin and Krifi-Papoz2013) and visual information (Plancher et al., Reference Plancher, Lévêque, Fanuel, Piquandet and Tillmann2018), compared to an irregular prime or silence. Rhythmic attention may improve working memory performance by freeing up attentional resources by processing fewer, but highly relevant, elements in an ongoing stream of sensory input (Johndro et al., Reference Johndro, Jacobs, Patel and Race2019). Music and speech signals themselves are designed to provide key details at rhythmically salient moments in time. Speech stress typically highlights semantically meaningful elements in a sentence (Mattys, Reference Mattys1997; Pitt and Samuel, Reference Pitt and Samuel1990), and strong beats in music are more often populated with pitches that are crucial to the tonality of a piece of music (Prince and Schmuckler, Reference Prince and Schmuckler2014; Prince et al., Reference Prince, Tan and Schmuckler2020). Thus, not only do perceptual and attentional systems sample sensory input in a rhythmic manner, but the stimulus streams themselves are optimized for the listener. It remains to be seen whether human communicative sounds are uniquely structured to capitalize on information transfer at rhythmic moments in time or whether other animal vocalizations or environmental sounds are structured similarly (but see De Gregorio et al., Reference De Gregorio, Valente and Raimondi2021; Roeske et al., Reference Roeske, Tchernichovski, Poeppel and Jacoby2020).
26.4 Neural Correlates of Rhythmic Regularity
The type of rhythmic attention described in DAT is predicted by the cyclical excitatory-inhibitory oscillations for populations of neurons in the brain (Lakatos et al., Reference Lakatos, Karmos, Mehta, Ulbert and Schroeder2008; Large and Jones, Reference Large and Jones1999). A period of heightened sensitivity at the peak of an excitatory phase of a neural oscillation would explain better perceptual or memory outcomes described in the studies reported above (see also Henry and Herrmann, Reference Henry and Herrmann2014). Indeed, change sensitivity is related to the phase of oscillations aligned with the external rhythmic sensory input (e.g., Henry and Obleser, Reference Henry and Obleser2012). Not all neural tracking of sensory input can be described as neural entrainment (Haegens, Reference Haegens2018; Haegens and Golumbic, Reference Haegens and Golumbic2018; Kösem et al., Reference Kösem, Bosker and Takashima2018; Obleser and Kayser, Reference Obleser and Kayser2019; Zoefel et al., Reference Zoefel, Archer-Boyd and Davis2018), as neural tracking of stimulus rhythms could be evidence of true neural entrainment (e.g., oscillations continue even in the absence of stimulation) or stimulus-driven responses to sound onsets (e.g., N1-P2 onset responses). Regardless, the degree of phase alignment between brain and stimulus rhythms predicts behavioral performance, including speech comprehension (Peelle, Reference Peelle2012), attention (Zion-Golumbic and Schroeder, Reference Zion-Golumbic and Schroeder2012), and expertise (Doelling and Poeppel, Reference Doelling and Poeppel2015; Harding et al., Reference Harding, Sammler, Henry, Large and Kotz2019). For more on this topic, see Chapters 3 and 5.
Neural oscillations are described as inherently rhythmically regular, ascribing to a single frequency of oscillation (e.g., Zoefel et al., Reference Zoefel, Archer-Boyd and Davis2018; but see Haas and Kubin, Reference Haas and Kubin1998, for nonlinear multi-frequency oscillators capable of entraining to the irregularities of speech). If oscillators are biased toward isochrony, then the alignment of ongoing neural oscillations should be better for a stimulus that is also rhythmically regular (e.g., less phase resetting of ongoing oscillations to match the incoming stimulus). Since song is more rhythmically regular than speech (Yu et al., Reference Yu, Cabildo, Grahn and Vanden Bosch der Nederlanden2023), it follows that song should garner greater alignment with ongoing neural activity – greater neural tracking – than speech. Consistent with this assertion, neural tracking is reduced when speech rhythms are made less regular by inserting pauses (delta band; Kayser et al., Reference Kayser, Ince, Gross and Kayser2015), and neural tracking is increased when words are sung over spoken (theta band; Vanden Bosch der Nederlanden et al., Reference Vanden Bosch der Nederlanden, Joanisse and Grahn2020, Reference Vanden Bosch der Nederlanden, Joanisse, Grahn, Snijders and Schoffelen2022a). Thus, it is possible that neural tracking increases linearly with the degree of regularity (see Section 6 chapters for more about rhythm and learning).
Here, we report a reanalysis that characterizes how subjective rhythmic regularity relates to neural tracking. We reanalyzed magnetoencephalography (MEG) data from the Vanden Bosch der Nederlanden et al. (Reference Vanden Bosch der Nederlanden, Joanisse, Grahn, Snijders and Schoffelen2022a) dataset (available here: https://data.donders.ru.nl/) and related it with the regularity ratings of the same stimulus set of Yu et al. (Reference Yu, Cabildo, Grahn and Vanden Bosch der Nederlanden2023) (available here: https://osf.io/hnw5t/) to elucidate whether the degree of regularity predicts neural tracking at the syllable rate across both spoken and sung utterances. The MEG data was collected as part of the melody familiarity study, where participants listened to 96 stimuli (presented four times) and rated whether the melody of the spoken or sung stimulus was one they had learned during their training sessions (they learned half of the sung melodies without lyrics). We used a jackknifing technique (Richter et al., Reference Richter, Thompson, Bosman and Fries2015) to extract single-trial estimates of neural tracking for each stimulus (96 stimuli, four presentations) presented to each of their 32 participants and related these stimulus-level data with the average regularity rating (see Section 26.2 – subjective ratings of ease of clapping or tapping to the stimulus) provided for each stimulus by 51 participants. Jackknifing estimates use a leave-one-out approach, so if a particular stimulus contributed strongly to the overall estimate of neural tracking, there would be a large reduction in the jackknifed neural tracking estimate when the stimulus was left out. If that same stimulus was also rated high in rhythmic regularity, then a significant negative correlation between neural tracking and rhythmic regularity would mean that greater neural tracking is related to increases in subjective experiences of rhythmic regularity.
We used a linear mixed-effects model (R, lmer package; stimulus and participant as random effects) to predict neural tracking based on acoustic features and subjective regularity ratings described in Yu et al. (Reference Yu, Cabildo, Grahn and Vanden Bosch der Nederlanden2023). We found that after controlling for utterance type (speech versus song), neural tracking was predicted by subjective rhythmic regularity ratings and pulse clarity (minimum value, MIR Toolbox; see Figure 26.2). Interestingly, although spectral flux is a stronger predictor of neural tracking than amplitude envelope measures (Weineck et al., Reference Weineck, Wen and Henry2022), it did not explain unique variance beyond rhythmic regularity ratings (see Table 26.2). This reanalysis suggests that the degree of rhythmic regularity in a stimulus is a strong driver of neural tracking, even beyond acoustic features that are highly predictive of rhythmic regularity (i.e., spectral flux).
Results from the linear mixed-effects regression models predicting average theta-band cerebro-acoustic phase coherence (raw values), with stimulus and participant as random effects. Statistically significant variables in each model are highlighted in bold.
| Model | Variable | Estimate | t-value | p |
|---|---|---|---|---|
| Model 1 | Regularity | −0.191 | −4.145 | <.001 |
| X2(5, N=3072) = 16.113, p<.001, AIC = 8159.3 (compared to random intercept model) | ||||
| Model 2 | Regularity | −0.142 | −2.821 | 0.006 |
| Song | −0.220 | −2.183 | 0.032 | |
| X2(6, N=3072) = 4.798, p=.0285, AIC = 8156.5 (compared to Model 1) | ||||
| Model 3 | Regularity | −0.138 | −2.703 | 0.036 |
| Song | −0.365 | −2.496 | 0.008 | |
| Spectral flux | 0.052 | 0.693 | 0.490 | |
| F0 stability | −0.098 | −1.251 | 0.215 | |
| Pulse clarity | −0.116 | −2.528 | 0.013 | |
| Syllable duration | 0.031 | 0.507 | 0.613 | |
| Prop. small integer ratio | 0.034 | 0.729 | 0.468 | |
| Consonant PVI | −0.034 | −0.641 | 0.523 | |
| % vocalic | 0.083 | 1.468 | 0.146 | |
| X2(13, N=3072) = 11.662, p=.1122 AIC = 8158.8 (compared to Model 2) | ||||
| Model 4 | Regularity | −0.146 | −2.974 | 0.004 |
| Song | −0.203 | −2.077 | 0.041 | |
| Pulse clarity | −0.113 | −2.568 | 0.012 | |
| X2(7, N=3072) = 6.648, p=.0099, AIC = 8151.8 (compared to Model 2) | ||||
Linear mixed-effects regression results for neural and acoustic data.
Neural tracking (speech-brain coherence in the theta band – 4–8 Hz) is related to rhythmic regularity and pulse clarity, even after controlling for utterance type.

Figure 26.2 Long description
Left. An error bar graph of regularity ratings, speech or song and pulse clarity with effect sizes. The data points are as follows. Speech or song, minus 0.20. Regularity ratings, minus 0.15. Pulse Clarity, minus 0.11. Right. Two line graphs and one error bar graph plot the average theta coherence with the rhythmic regularity, utterance type and pulse clarity, respectively. Top: The line originates at (minus 2, 0.4) and terminates at (2, minus 0.3). Center: The error bars are plotted at 0.0 for Speech and at minus 0.1 for Song. Bottom: The line originates at (minus 2, 0.3) and terminates at (3, minus 0.5). The values are estimated.
These results may lend a neural explanation for recent studies showing that attention is biased toward rhythmic regularity, in both auditory (Andreou et al., Reference Andreou, Kashino and Chait2011) and visual domains (Zhao et al., Reference Zhao, Al-Aidroos and Turk-Browne2013), even without explicit awareness (Zhao et al., Reference Zhao, Al-Aidroos and Turk-Browne2013). Perhaps, stimuli that capture attention are those with greater rhythmic regularity because they are easier to process neurally than less regular or irregular stimuli. Further, if regular stimuli capture attention, then the regularity of the stimulus would also make it more likely that important moments in time align with moments of high neural excitability, resulting in better memory or comprehension. Although speech is not rhythmically regular, like music, it likely has a higher degree of regularity than other real-world sounds, such as a babbling brook. While regularity may move stimuli from one side of the speech–song continuum to the other, regularity may be critical for a broader range of noncommunicative stimuli forming a much larger acoustic continuum. The periodicity of speech may also underlie why people of all ages are nearly perfect at detecting when a human voice changes in a complex scene compared to musical instrument sounds, environmental sounds, or animal vocalizations (Vanden Bosch der Nederlanden et al., Reference Vanden Bosch der Nederlanden, Snyder and Hannon2016, Reference Vanden Bosch der Nederlanden, Zaragoza, Rubio-Garcia, Clarkson and Snyder2018; Vanden Bosch der Nederlanden and Vouloumanos et al., Reference Vanden Bosch der Nederlanden and Vouloumanos2021). Future work should examine whether attentional biases toward real-world sounds could be predicted based on the degree of regularity in the signal, such that the stimulus with the highest degree of regularity is the one that captures attention. Together, this review and reanalysis suggest that the human brain is endowed with neural mechanisms specialized to process rhythmic regularities across multiple domains.
26.5 Leveraging Rhythmic Regularity in Music
A growing number of studies show relationships between musical rhythmic processing and language (e.g., Gordon et al., Reference Gordon, Shivers and Wieland2015; Kertész and Honbolygó, Reference Kertész and Honbolygó2023; Ozernov-Palchik and Patel, Reference Ozernov‐Palchik and Patel2018; Woodruff Carr et al., Reference Woodruff Carr, White-Schwoch, Tierney, Strait and Kraus2014). They show, for instance, that kindergarteners’ rhythmic discrimination skills and letter-sound knowledge – an early predictor of reading development – are highly correlated (Ozernov-Palchik and Patel, Reference Ozernov‐Palchik and Patel2018), that drumming consistency relates to language outcomes (Woodruff Carr et al., Reference Woodruff Carr, White-Schwoch, Tierney, Strait and Kraus2014), and that rhythm perception in elementary school-aged children predicts expressive grammar skills (Nitin et al., Reference Nitin, Gustavson and Aaron2023). Deficits in rhythmic processing abilities have also been connected to a wide range of developmental disorders, including dyslexia, ADHD, autism, motor coordination disorders, and stuttering (Fiveash et al., Reference Fiveash, Bedoin, Gordon and Tillmann2021; Garnett et al., Reference Garnett, Chow, Limb, Liu and Chang2022; Hannon et al., Reference Hannon, Nave‐Blodgett and Nave2018; Ladanyi et al., Reference Ladányi, Persici, Fiveash, Tillmann and Gordon2020; Lense et al., Reference Lense, Ladányi, Rabinowitch, Trainor and Gordon2021). Compared to typically developing children, those with reading disorders and language impairment struggle to synchronize their movements to a metronome or musical rhythm (Corriveau and Goswami, Reference Corriveau and Goswami2009). The ability to extract a stable pulse or perceive rhythmic regularity may be crucial to successful language processing (Bekius et al., Reference Bekius, Cope and Grube2016; Centanni et al., Reference Centanni, Pantazis and Truong2018, Reference Centanni, Beach and Ozernov-Palchik2022).
Given that language deficits are related to musical beat processing, music-based interventions that help people pick up on rhythmic regularity can complement existing therapies to increase the positive impact on language outcomes. As described above, rhythmic priming studies successfully use musical primes to entrain neural activity and facilitate the extraction of relevant information from subsequently presented speech stimuli (Canette et al., Reference Canette, Lalitte and Bedoin2020b; Chern et al., Reference Chern, Tillmann, Vaughan and Gordon2018). Extra scaffolding and priming of rhythmic attention through music may be particularly beneficial for children with developmental disorders who struggle to create internal representations of salient moments in ongoing speech streams. This is consistent with reports that rhythm-based musical training can benefit reading and phonological processes in dyslexic children (Flaugnacco et al., Reference Flaugnacco, Lopez and Terribili2015) and speech processing for children with cochlear implants (Bedoin et al., Reference Bedoin, Besombes and Escande2018). Beyond the potential benefits in language processing, rhythm in music is also tightly related to the experience of pleasure and reward (Fiveash et al., Reference Fiveash, Ferreri and Bouwer2023; Todd and Lee, Reference Todd and Lee2015; Zatorre, Reference Zatorre2015). This makes musical interventions focused on rhythmic regularity especially fruitful for improving perceptual as well as socio-emotional outcomes across the lifespan. Musical features such as rhythmic regularity can be leveraged to improve neural tracking and downstream behavioral outcomes, such as comprehension or engagement.
Summary
Rhythmic regularity is a key differentiator of speech and song, but how it is measured in the acoustic signal is not trivial. Listeners subjectively report rhythmic regularity in song, but the acoustic features that are consistent with regularity do not correspond with the small integer ratios of events that align with a beat. Rhythmic regularity is a significant modulator of neural tracking in speech and song, suggesting regularity likely plays a crucial role in guiding other behavioral outcomes related to neural tracking, such as attention, memory, engagement, and even comprehension.
Implications
Our work suggests that regularity drives neural tracking and is a key factor differentiating speech and song. Given the neural tracking findings, would regularizing the speech signal to be more music-like improve intelligibility and comprehension compared to typical, irregular, speech rhythms? If this is the case, then why did humans evolve two distinct systems of communication that differ in the degree of regularity and their functional goals: speech for transaction of meaning, and song for conveying or regulating emotion? That is, why do we not exclusively sing to one another? Future work should characterize the role of rhythmic regularity across multiple domains, including speech, song, and other everyday sounds.
Gains
This chapter presents a first step toward understanding the role of regularity across communicative modalities and characterizing its impact on neural processing. We highlight the necessity for a good metric of the degree of regularity that can characterize a range of naturally occurring stimuli. Such a metric has the potential to have broad impact, for example on the best practices for teaching and learning in the classroom to improve retention, or in high-risk situations to grab listeners’ attention, as in hospital-alert settings.


