42.1 Introduction
Auditory disabilities affect a substantial portion of the global population, and the World Health Organization estimates that close to 2.5 billion people will live with some degree of hearing loss (HL) by 2050 (World Health Organization, 2023). Many listeners with mild HL may struggle to understand speech in noisy situations, and those with moderate HL experience difficulties with speech perception even under ideal listening conditions. Individuals, however, can share identical audiometry, meaning that their hearing thresholds – estimated by an audiologist using pure tone stimuli across standard frequency bands – are similar, yet they achieve very different functional hearing outcomes. For example, imagine that two listeners with HL follow speech in quiet relatively well, but one person has substantially greater difficulty than the other in understanding speech in noise, such as when dining in a busy restaurant. This discrepancy highlights the complexities and individuality of HL not captured by audiometry: The other listener’s relatively preserved speech in noise perception might be related to any number of factors, including lip-reading skill, working memory capacity, educational background, or awareness of contextual cues (Akeroyd, Reference Akeroyd2008; Aydelott et al., Reference Aydelott, Leech and Crinion2010; Dryden et al., Reference Dryden, Allen, Henshaw and Heinrich2017). Individual differences in temporal processing, such as sensitivity to temporal structure, is one cognitive trait that could be associated with listening outcomes in the context of audiological difference and disability (Füllgrabe et al., Reference Füllgrabe, Moore and Stone2015; Hidalgo et al., Reference Hidalgo, Zécri and Pesnot-Lerousseau2021); yet, HL may also impact which aspects of temporal processing are available to listeners. This could in turn limit the extent to which individuals are able to benefit from temporal cues or predictable speech rhythms, for instance.
In this chapter, I review an interdisciplinary literature investigating the perception of speech timing and rhythm – which I define here as temporal patterns and/or structures – in the context of HL, as well as listening via an auditory prosthesis, the cochlear implant (CI). I first introduce some of the etiologies that can result in hearing disability, as well as the main medical and rehabilitative interventions for HL. I then survey the clinical and experimental data that shed light on how speech rhythm and prosody are processed by people with HL. This includes data regarding group and individual differences in temporal perception, and how these differences may have implications for, or be modulated by, HL. Although speech rhythm is the focus of this chapter, as a point of comparison, I also briefly survey the scientific picture surrounding musical rhythm perception and production with regards to HL.Footnote 1 The preceding sections are organized according to the timing of HL onset: pre-lingual, or before language acquisition (i.e., primarily affecting infants and children); and post-lingual, or after language acquisition (i.e., primarily affecting adolescents and adults). I conclude the chapter by describing open questions and possible directions for interventions.
42.2 The Many Sides of HL
At the outset, it is important to distinguish between deafness, which refers to the state of having HL, and Deafness, which conveys belonging to a Deaf linguistic and cultural identity (Obasi, Reference Obasi2008). Someone who is Deaf may see HL as a matter of social difference, rather than a disorder or medical problem in need of intervention – noting that a person can be hearing, or use a CI, and also be culturally Deaf. The lines between groups who identify as Deaf, deaf, or hearing are porous, may not necessarily align with use of spoken or signed language, and have been subject to a complex social history. Here, I focus exclusively on HL as it is experienced by someone who uses spoken language to communicate. I will use HL and auditory disability to refer to those with reduced access to auditory speech information, while acknowledging that different individuals may identify with other terms to describe their hearing status. This chapter also omits extended discussion of rhythm in sign language, which should be incorporated when accounting for functional roles of rhythm in human communication more generally (MacIntyre, Reference MacIntyre2018).
42.2.1 Causes of HL
There are a variety of ways in which HL can occur (World Health Organization, 2018). Causes include exposure to very loud sounds, physical trauma, hereditary syndromes, side effects of medications and other treatments, infection, chronic illnesses (e.g., diabetes), and normal ageing. Congenital HL, present at birth, may be caused by genetic variation, complications during birth, or viral infection, among other factors. Although HL can stem from problems with the outer and middle ear (e.g., a perforated eardrum), many of the aforementioned conditions affect the inner ear wherein cochlear hair cells, the sensory receptors of the auditory system, are damaged. HL can also arise from issues affecting the auditory nerve (Kaga, Reference Kaga2016). HL can advance very gradually, especially for age-related HL, meaning that someone may not notice a difference until their HL can be characterized as moderate or poorer. At this point, changes in central auditory processing, including cognitive and behavioral adaptations to cope with the reduced or degraded sensory input, have likely already occurred (Herrmann and Butler, Reference Herrmann and Butler2021; Mitchell and Maslin, Reference Mitchell and Maslin2007; Slade et al., Reference Slade, Plack and Nuttall2020).
42.2.2 Hearing Devices
Whether HL is treatable depends on the underlying cause, but cochlear damage is usually irreversible. In the case of mild to moderate permanent HL, a listener can make use of a hearing aid (HA) to amplify the acoustic signal and compensate the audibility of specific frequencies; however, when HL is severe to profound, a CI may help to restore a sense of hearing. A CI is an auditory prosthesis that bypasses the inner ear (Carlyon and Goehring, Reference Carlyon and Goehring2021; Fan-Gang Zeng et al., Reference Zeng, Rebscher, Harrison, Sun and Feng2008; Loizou, Reference Loizou1999). It consists of an array of electrodes surgically implanted within a person’s cochlea, which receive a signal originating from the external components of the system. Briefly, the CI pipeline from sound source to auditory percept is as follows: First, sounds in the environment are sampled by an external microphone (or multiple microphones). An external processor then divides this signal into its component frequencies, emulating the frequency selectivity of the inner ear. Other signal processing techniques, such as noise reduction, may also be implemented at this step. The processed signal, now converted into discretized pulses, is transmitted via radio waves to an internal receiver, which conveys the signal onward to the stimulator implanted within the cochlea. Here, each signal component activates a specific electrode, thereby directly interfacing with the auditory nerve. Whereas acoustic hearing entails thousands of individual cochlear hair cells, CIs normally contain around 22 individual electrodes or fewer – in practice, the number of useable electrodes is often even lower, depending on the health of the auditory nerve and/or interaction between channels due to electrical current spread within the cochlea (Goehring et al., Reference Goehring, Archer-Boyd, Arenberg and Carlyon2021). Despite limited spectro-temporal resolution, however, many people with CIs have excellent speech reception in quiet, and some are furthermore able to communicate via voice calls in the complete absence of nonauditory cues (e.g., lip-reading). CI listening should be considered as a perceptual representation of sound that is distinct from typical acoustic hearing, and in particular, CI users who received their implant later in life usually require time, patience, effort, and support to make the most of their device (Boisvert et al., Reference Boisvert, Reis, Au, Cowan and Dowell2020; Tang et al., Reference Tang, Thompson and Clark2017).
42.3 Psychophysical and Cognitive Aspects of HL and the Perception of Speech Rhythm
42.3.1 Influences of HL at the Auditory Periphery
Even when accounting for differences in audibility and perceptual thresholds, people with moderate to severe HL may also experience other changes regarding their sensitivity to, and processing of, sounds. For example, people with HL show reduced frequency selectivity (Moore, Reference Moore1985, Reference Moore1996). This increases the threshold at which contrasts in spectral quality are perceived, which can affect the discrimination of vowel sounds. HL is also associated with poorer dynamic range, possibly due to a reduction in the compression performed at the basilar membrane, caused by damage or loss of the outer hair cells that would normally selectively amplify vibration and sharpen frequency tuning (Moore, Reference Moore1996). The consequence is that softer sounds may become inaudible, while louder ones are perceived similarly to typically hearing listeners – or, in some cases, as intolerably loud. This may influence temporal processing relatively early in the auditory system. For instance, in silent gap detection tasks, experimental participants with HL require longer gap duration values, in comparison to their typically hearing counterparts, when the stimuli in which the gap occurs consist of narrowband noise (Glasberg and Moore, Reference Glasberg and Moore1992). This effective reduction in temporal resolution may be related to the amplitude fluctuations inherent to noise: With reduced dynamic range, randomly occurring peaks and troughs within the noise amplitude modulation patterns may be perceptually exaggerated, making true gaps more difficult to detect. Another temporal consequence of cochlear damage relates to how certain acoustic features, such as perceived pitch or loudness, are processed over time at the sub-second level. Specifically, for typically hearing listeners, increases in stimulus duration tend to accompany a change in thresholds required to perceive the feature of interest (Kidd et al., Reference Kidd, Mason and Feth1984). For instance, a short tone following the presentation of a longer tone will need to have a relatively higher sound pressure level in order to be perceived as identically loud. This process is referred to as temporal integration, and is theorized to reflect how sensory information is sampled and/or accumulated over time (Florentine et al., Reference Florentine, Fastl and Buus1988). Hence, HL may lead to weaker or less stable representations of sustained sounds at short timescales, which could in turn affect the processing of temporal structure at longer timescales, as in speech or music (Moore, Reference Moore2003).
42.3.2 HL and Central Auditory Processing
At the auditory periphery, the aforementioned HL-related differences in temporal processing operate at the timescale of tens of milliseconds or fewer. These differences could in turn affect speech processing at the timing of syllables or phonemes by potentially distorting or obscuring salient temporal cues (e.g., the roughly 70 milliseconds of voice onset time that distinguish ta from da). However, HL is often accompanied by more central changes that could also impact the perception of speech rhythm at slower timescale, i.e., at the level of words and phrases. In this section, I will survey the empirical literature investigating how people with HL perceive and reproduce temporal patterns and structure in speech with a focus on cognitive and other higher-level aspects of sensory processing. I begin by discussing congenital and early HL, and examine how perception and production of speech rhythm, as well as other aspects of temporal and prosodic processing, develop in children who use CIs and HAs for speech communication. I then address post-lingual HL in adults, before turning to focus on the interaction between HL and cognitive changes associated with normal ageing, especially with relevance to speech in noise perception.
42.3.2.1 HL and Speech Rhythm Perception and Production during Development
42.3.2.1.1 Perception of Speech Rhythm by Children with HL
Before the advent of regular newborn hearing screening, one indicator of auditory disability was the atypical speech produced by children with HL. In 1936, educator Charles G. Rawlings noted that the speech of pre-lingually Deaf children was often “arrhythmic,” but that “If the deaf child could be … trained from his first years in school to speak with a normal rhythm even at the cost of perfect articulation, his speech would be far more natural and intelligible to the general public” (Rawlings, Reference Rawlings1936:143). Such differences in production could indicate limited perceptual access to rhythmic properties in speech. For example, researchers report that children with HL (including those fitted with HAs and/or CIs) have poor sensitivity to lexical and syllabic stress (Holt et al., Reference Holt, Demuth and Yuen2016; Kalathottukaren et al., Reference Kalathottukaren, Purdy and Ballard2017; Konadath et al., Reference Konadath, Raveendran and Yeshoda2021; Lyxell et al., Reference Lyxell, Wass and Sahlén2009; Most and Peled, Reference Most and Peled2007; Segal et al., Reference Segal, Houston and Kishon-Rabin2016), which in natural speech are conveyed as variations in intensity, duration, intonation, or a combination of these. This reduction in sensitivity is already apparent when comparing infants (12–33 months old) with CIs to typically hearing infants (Segal et al., Reference Segal, Houston and Kishon-Rabin2016), although the overall pattern of developing a neural response to verbal stress appears to be similar across infants with CIs and those with typical hearing (Vavatzanidis et al., Reference Vavatzanidis, Mürbe, Friederici and Hahne2016). The fine-grained representation of auditory cues, however, are also likely to differ according to hearing device: Children (8–15 years old) who wear HAs have been found to have superior intonation and stress perception in comparison to children with CIs (Most and Peled, Reference Most and Peled2007). Recently, longitudinal experiments have shown that, for children (5–9 years old) with better residual hearing, the combination of HAs and CIs in contralateral ears may lead to improved speech prosody perception outcomes, presumably due to the low-frequency audibility provided by HAs (Davidson et al., Reference Davidson, Geers, Uchanski and Firszt2019). In the same cohort, prosody perception was determined to be uniquely predictive of later achievement in other linguistic areas including reading and phoneme perception (Grantham et al., Reference Grantham, Davidson, Geers and Uchanski2022). Hence, although children with HL tend to encounter difficulties in prosody perception, there is a wide range of ability that likely depends on hearing history, mode of amplification, and many other factors (see Wenrich et al., Reference Wenrich, Davidson and Uchanski2017, for a more in-depth discussion).
42.3.2.1.2 Production of Speech Rhythm by Children with HL
Scientific efforts to quantify the subjectively “arrhythmic” character of speech produced by children with HL generally cohere with Rawlings’ qualitative observation (Reference Rawlings1936). That is, there is consensus for “errors of prosody (e.g., errors of intonation, stress, and/or phrasing)” (Parkhurst and Levitt, Reference Parkhurst and Levitt1978). Children with HL tend to produce atypical patterns of stressed and unstressed syllables (Gold, Reference Gold1980; Lenden and Flipsen, Reference Lenden and Flipsen2007; Murphy et al., Reference Murphy, McGarr and Bell-Berti1990; Osberger and Levitt, Reference Osberger and Levitt1979; Sundström et al., Reference Sundström, Löfkvist, Lyxell and Samuelsson2018) – not only by misapplying stress within multisyllabic words but sometimes also by failing to differentiate or mark any stress at all, resulting in “staccato” speech with roughly equal syllabic weighting (McGarr and Osberger, Reference McGarr and Osberger1978; Osberger and Levitt, Reference Osberger and Levitt1979). Similarly, some analyses suggest that speech by children with HL is disrupted by the unusual placement of pauses (Bochner et al., Reference Bochner, Barefoot and Johnson1987; Gold, Reference Gold1980; McGarr and Osberger, Reference McGarr and Osberger1978; Osberger and Levitt, Reference Osberger and Levitt1979; Rawlings, Reference Rawlings1936). Children (6–10 years old) who use CIs are also less able to correctly imitate syllabic and stress patterns in both pseudo-words (Carter et al., Reference Carter, Dillon and Pisoni2002) and real words (Grandon et al., Reference Grandon, Vilain and Gillis2019), in comparison to typically hearing children.
Presumably, deficits in children with HL’s speech production originate in absent or degraded sensory input during speech perception, including auditory sensorimotor feedback of one’s own voice. Where the link, within-child, between speech prosody perception and production has been formally tested, some studies found a significant association (e.g., Klieve and Jeanes, Reference Klieve and Jeanes2001; Konadath et al., Reference Konadath, Raveendran and Yeshoda2021), although not all (e.g., Kalathottukaren et al., Reference Kalathottukaren, Purdy and Ballard2017). Objective and ratings-based measures of speech prosody production by children with HL have been found to correlate with other aspects of their linguistic development, including vocabulary and word recognition skills (Carter et al., Reference Carter, Dillon and Pisoni2002; Lyxell et al., Reference Lyxell, Wass and Sahlén2009). The reasons for this are unclear, although some possible mechanisms are addressed in the following section. Timing issues during speech communication can emerge as early as infancy: Dyads consisting of a mother and a CI-recipient child (one–two years old, observed three and nine months post-implant) produce longer between-speaker pauses relative to within-speaker pauses, contrary to the pattern associated with typically hearing children (Kondaurova et al., Reference Kondaurova, Smith, Zheng, Reed and Fagan2020). In addition, when compared to typically hearing peers, children with CIs (seven years and older) show reduced speech rate-matching during interactions with a clinician, and continue to demonstrate such delays well into adolescence (Freeman and Pisoni, Reference Freeman and Pisoni2017).
42.3.2.1.3 Auditory Deprivation and the Development of Rhythmic-Sequential Ability
Some authors argue that HL disrupts the development of cognitive faculties important for speech, particularly sequence learning, which is “the ability to encode and represent the order of discrete elements occurring in a sequence” (Conway and Christiansen, Reference Conway and Christiansen2001:539). Sequential learning has been shown to be reduced in children and adults with HL – even when using nonauditory paradigms (Conway et al., Reference Conway, Pisoni, Anaya, Karpicke and Henning2011; Lévesque et al., Reference Lévesque, Théoret and Champoux2014). Why would this be? One theory is that because sound is an inherently temporal signal, its analysis relies on attention, working memory, and the processing of serial order; hence, lack of experience with sound would deprive the developing child of crucial experience with temporal processing more generally (Conway and Christiansen, Reference Conway and Christiansen2005, Reference Conway and Christiansen2009). This coheres with evidence linking sequence learning to perceptual sensitivity for nonverbal auditory rhythms in typically hearing populations (François and Schön, Reference François and Schön2011; MacIntyre et al., Reference MacIntyre, Lo, Cross and Scott2022). Moreover, sequence and rhythm processing appear to share some neural correlates (Bednark et al., Reference Bednark, Campbell and Cunnington2015; Janata and Grafton, Reference Janata and Grafton2003). In the context of speech rhythm, insensitivity to stress and pauses could interfere with sequential processing by, for instance, obscuring syntactic boundaries (e.g., “fruit salad and milk” versus “fruit, salad, and milk”), a cue that is frequently missed by children with HL (Fortunato et al., Reference Fortunato-Tavares, Schwartz, Marton, de Andrade and Houston2018; Kalathottukaren et al., Reference Kalathottukaren, Purdy and Ballard2017). This may also bear some relationship to observed correlations between aberrant rhythm processing and developmental difficulties with grammar and syntax in the typically hearing population (Gordon et al., Reference Gordon, Jacobs, Schuele and McAuley2015; Ladányi et al., Reference Ladányi, Persici, Fiveash, Tillmann and Gordon2020; Lee et al., Reference Lee, Ahn, Holt and Schellenberg2020; Nitin et al., Reference Nitin, Gustavson and Aaron2023; see also Chapters 25 and 39), the latter of which disproportionately affect children with HL (Delage and Tuller, Reference Delage and Tuller2007; Friedmann and Szterman, Reference Friedmann and Szterman2006; Kawar, Reference Kawar2021; Moeller et al., Reference Moeller, Tomblin, Yoshinaga-Itano, Connor and Jerger2007; Nittrouer and Lowenstein, Reference Nittrouer and Lowenstein2021; Szterman and Friedmann, Reference Szterman and Friedmann2014; but see also Briscoe et al., Reference Briscoe, Bishop and Norbury2001).
Critics of the proposed link between HL and sequential learning, however, counter that the “sequence learning deficit in children with CI may be closely tied to the nature of the task” (Torkildsen et al., Reference von Koss Torkildsen, Arciuli, Haukedal and Wie2018:127). For example, some paradigms employing visual stimuli may nonetheless lend themselves to verbal rehearsal, thereby conflating domain-general sequential learning with phonological processing. In a sequential learning task deliberately designed to avoid this, children with CIs did not differ in performance from typically hearing peers (Torkildsen et al., Reference von Koss Torkildsen, Arciuli, Haukedal and Wie2018). This accords with evidence that children with CIs show selective impairments in verbal, compared with spatial, working memory tasks (Davidson et al., Reference Davidson, Geers, Uchanski and Firszt2019). An additional source of ambiguity stems from the challenge of decoupling auditory from linguistic exposure. For instance, children with HL who acquire sign language from birth perform similarly to typically hearing children in tests of sustained attention (Dye and Hauser, Reference Dye and Hauser2014). Early experience with sign language also appears to benefit linguistic development in children with CIs who go on to communicate orally in later life (Lyness et al., Reference Lyness, Woll, Campbell and Cardin2013). Linking back to rhythm, Deaf adults who use sign language can synchronize to a visual rhythmic stimulus as well as hearing adults do to an auditory rhythmic stimulus, with hearing adults performing comparatively poorly in visual-motor rhythmic synchronization tasks (Iversen et al., Reference Iversen, Patel, Nicodemus and Emmorey2015). Hence, active engagement with temporally dynamic patterns, whether visual or auditory, may in fact be key to the development of sequential, and possibly also rhythmic, processing. Further work is needed to disentangle modality-specific from domain-general deficits, as well as rhythm from possibly related cognitive faculties, such as verbal working memory (Saito, Reference Saito2001; Saito and Ishio, Reference Saito and Ishio1998).
42.3.2.1.4 Rhythm-Based Interventions in Speech Therapy
If rhythm processing is not strictly reliant on auditory experience, it may be that difficulties with speech rhythm observed in children with HL are more indicative of the limitations of speech prosody rehabilitation (Hargrove et al., Reference Hargrove, Anderson and Jones2009; Peppé, Reference Peppé2009) than of some innate deficit per se. Specifically, it is unclear whether children “do not know the stress patterns of the language or whether they do not have the articulatory co-ordination that would permit them to produce an intended pattern correctly” (Tye-Murray and Folkins, Reference Tye‐Murray and Folkins1990:2675). Some researchers propose that deficits in speech “flow” experienced by children with HL may be partly due to the “teaching of articulation of individual isolated elements rather than longer, more meaningful units of speech” (Gold, Reference Gold1980:408; see also Howard, Reference Howard2007; Tye-Murray and Woodworth, Reference Tye‐Murray and Woodworth1989). This could also help to explain why coarticulation – how the production of consecutive speech segments contextually affect each other – is reduced in HL, in comparison with typically hearing children (Okalidou and Harris, Reference Okalidou and Harris1999; Waldstein and Baum, Reference Waldstein and Baum1991). Yet, speech rhythm rehabilitation for children with HL may be facilitated by nonauditory sensorimotor experience – for instance, by demonstrating model prosodic patterns via finger tapping (Tye-Murray and Folkins, Reference Tye‐Murray and Folkins1990). Some research groups have explored this idea by using active rhythm training to enhance sensitivity to speech temporal features (see León Méndez et al., Reference del Carmen León Méndez, Fernández García and Daza González2023, and Pesnot Lerousseau et al., Reference Pesnot Lerousseau, Hidalgo and Schön2020, for reviews). For instance, it was found that a short, rhythm-focused training session, rather than the usual speech therapy session, was associated with more rhythmically stable speech production during a subsequent verbal turn-taking task for children with HAs and/or CIs (5–9 years old) (Hidalgo et al., Reference Hidalgo, Falk and Schön2017). The authors further demonstrated using electroencephalography (EEG) that the mismatch negativity, a neural correlate of prediction error, in response to rhythmically irregular trials was modulated by rhythmic training in children with HL (Hidalgo et al., Reference Hidalgo, Pesnot-Lerousseau, Marquis, Roman and Schön2019). These speech-focused studies concur with a similar experiment using musical stimuli, which found that active motor engagement with temporal structure enhanced the ability to identify songs in children with CIs (4–12 years old) (Vongpaisal et al., Reference Vongpaisal, Caruso and Yuan2016). Finally, in another speech rhythm priming study, the authors report that the presentation of a nonverbal rhythm, the structure of which was shared with the following utterance to be produced, improved phonological accuracy in children with various hearing devices (mean age 8.72 years old, SD 2.19; Cason et al., Reference Cason, Hidalgo, Isoard, Roman and Schön2015). It remains to be seen whether the benefits observed in any of these interventions persist over time or are rather limited to the immediate post-intervention period, and the results should be replicated with enhanced control conditions and blinding procedures – including, where possible, blinding of therapists and others administering the rhythm training (Hargrove et al., Reference Hargrove, Anderson and Jones2009; McKay, Reference McKay2021).
42.3.2.1.5 Musical Rhythm Perception and Production by Children with HL
In one study, children with CIs demonstrated poorer musical pitch and timbre perception, but not rhythm perception, when compared to typically hearing children; however, many of the children with CIs (9–13 years old) were involved in external music lessons (Innes-Brown et al., Reference Innes-Brown, Marozeau, Storey and Blamey2013). Another study reports marginally better performance for temporal, rather than pitch-based, recognition of melodies in children with CIs (5–7 years old) (Volkova et al., Reference Volkova, Trehub, Schellenberg, Papsin and Gordon2014). Other authors emphasize that children with HL who wear CIs and/or HAs (5–10 years old) are less accurate than typically hearing children when tapping in synchrony to both simple and complex rhythms, despite similar levels of motor variability when freely tapping (Hidalgo et al., Reference Hidalgo, Zécri and Pesnot-Lerousseau2021). In another study, children with CIs (median age 11.25 years old, IQR 5.58) performed at chance in a nonverbal rhythm discrimination task on average, but individual scores did not correlate with other auditory skills, nonverbal intelligence, nor grades in music class (Stabej et al., Reference Stabej, Smid and Gros2012). Taken together, it appears that temporal cues in nonverbal rhythmic stimuli are at least partially available to children with HL but that aspects of rhythmic processing are nonetheless impeded or delayed in comparison to typically hearing children – in particular, when rhythmic stimuli are more complex or incorporate changes in pitch (Darrow, Reference Darrow1984; Roy et al., Reference Roy, Scattergood-Keepper and Carver2014). Data from EEG reveal a similar pattern of results: When perceiving temporal deviations in musical rhythmic stimuli, children and adolescents with CIs produce a mismatch negativity, but it is reduced in magnitude and appears at a slower latency in comparison with typically hearing children (Petersen et al., Reference Petersen, Weed and Sandmann2015; Torppa et al., Reference Torppa, Salo and Makkonen2012). The heterogeneity of the samples in these studies and the varying levels of complexity and naturalism of the stimuli employed, however, preclude firm conclusions, and I will discuss further evidence regarding nonverbal rhythm perception in adults with HL in the next section (see also Shukor et al., Reference Shukor, Han, Lee and Seo2021, for a systematic review of musical interventions for children with HL).
42.3.2.2 Post-Lingual HL and Speech Rhythm Perception and Production in Adults
This section focuses on adults who primarily lost their hearing post-lingually, meaning that they had at least some experience with audible prosody and speech rhythm prior to their HL onset; however, acquired HL is associated with extensive changes to the neural architecture that supports central auditory processing, both within canonical auditory areas and across the entire brain (Glick and Sharma, Reference Glick and Sharma2017; Griffiths et al., Reference Griffiths, Lad and Kumar2020; Peelle and Wingfield, Reference Peelle and Wingfield2016). Hence, the availability and/or perceptual weighting of previously salient speech rhythm cues is likely to change over the lifetime of an individual with HL, and this probably also varies between individuals who differ in their hearing history and underlying etiology. I begin, as in the previous section, by discussing the perception of speech rhythm in adults with HL. I continue with production, although relatively few studies have investigated changes in vocal prosody or rhythm in adults with post-lingual HL. I then discuss the complementary literature on musical rhythm perception in HL, before turning to the final section, which deals with ageing-specific aspects of speech rhythm processing.
42.3.2.2.1 Perception of Speech Rhythm by Adults with Post-lingual HL
In an experiment using controlled, isolated stress contrasts, adult CI users did not differ from typically hearing controls when discriminating syllabic stress based on duration cues; however, they did perform significantly worse when stress cues consisted of changes in intonation or intensity (Meister et al., Reference Meister, Landwehr, Pyschny, Wagner and Walger2011). In the same study, the participants then completed a stress identification task with naturally produced stimuli in sentence form – for this more complex and realistic task, the CI users showed significantly poorer discrimination than the controls, and sensitivity to duration did not correlate with performance; however, there was an advantage for CI listeners who could better discriminate stress using a combination of intonation and intensity cues (Meister et al., Reference Meister, Landwehr, Pyschny, Wagner and Walger2011). Another study found that adults with CIs require substantially larger changes in fundamental frequency (F0) to identify pitch accents in multisyllabic words in comparison to adults with typical hearing (Dincer D’Alessandro and Mancini, Reference Dincer D’Alessandro and Mancini2019). Device-related differences in intonation perception are similar to those seen in children, such that adults who use both a CI and an HA have better perception of intonation, syllable stress, and emphatic stress than when using a CI alone (Most et al., Reference Most, Harel, Shpak and Luntz2011). Finally, adults with uncorrected HL (i.e., no CI or HA use) were found to better identify syllabic stress when the target was surrounded by speech envelope-modulated noise, as opposed to in the original, unprocessed sentence or when presented in isolation (Barac-Cikoja and Revoile, Reference Barac-Cikoja and Revoile1996). This study is intriguing in that it suggests a facilitating effect of prosodic context, but a cost of fine-grained phonetic and/or semantic processing, for adults with HL when they explicitly attend to speech rhythm cues.
The effects of specific properties of rhythm in speech have also been studied in the context of HL. For instance, similarly to typically hearing control participants, adults with CIs benefit from predictable patterns of syllabic stress, known in the literature as metrical patterns, during speech in noise perception (Perry and Kwon, Reference Perry and Kwon2015; Spitzer et al., Reference Spitzer, Liss, Spahr, Dorman and Lansford2009). Older adults with and without HL appear to use metrical rhythmic patterns similarly to younger adults when perceiving speech in noise (Woodfield and Akeroyd, Reference Woodfield and Akeroyd2010). There may be a relationship, at least in typically hearing adults, between the ability to perceive speech in noise and sensitivity to metric cues (Borrie et al., Reference Borrie, Baese-Berk, Van Engen and Bent2017). Awareness of pitch-based syllabic stress is associated with greater speech in noise perception among CI and typically hearing listeners (Dincer D’Alessandro and Mancini, Reference Dincer D’Alessandro and Mancini2019). In fact, sensitivity to rhythm more generally has also been linked with speech in noise perception. For instance, nonverbal rhythm – but not pitch or timbre – processing ability significantly predicted speech in noise perception by adults with CIs (Leal et al., Reference Leal, Shin and Laborde2003). Nonverbal rhythm perceptual ability also correlates with superior speech in noise understanding in typically hearing adults (Slater et al., Reference Slater, Kraus and Woodruff Carr2018; Slater and Kraus, Reference Slater and Kraus2016; Yates et al., Reference Yates, Moore, Amitay and Barry2019).
42.3.2.2.2 Production of Speech Rhythm by Adults with Post-lingual HL
Adults who acquire HL later in life may show changes in produced speech rhythm that include excessive lengthening of vowels and abnormal stress patterns (Cowie and Douglas-Cowie, Reference Cowie, Douglas-Cowie, Lutman and Haggard1983; Lane and Webster, Reference Lane and Webster1991; Plant and Hammarberg, Reference Plant and Hammarberg1983). Case studies suggest that, after cochlear implantation, objectively measured and listener-rated speech rhythmic patterns, such as the differentiation between stressed and unstressed syllables, improve (Economou et al., Reference Economou, Tartter, Chute and Hellman1992; Leder et al., Reference Leder, Spitzer and Milner1986; Tartter et al., Reference Tartter, Chute and Hellman1989). The basis of these changes is poorly understood, but CI users with post-lingual HL provide a unique opportunity for researchers to understand the role of auditory feedback in speech production more generally (Perkell, Reference Perkell2012; Ubrig et al., Reference Ubrig, Goffi-Gomez and Weber2011). Although CIs are thought to transmit temporal, in comparison to spectral, speech properties quite well, post-lingual CI listeners still have poorer temporal resolution, as measured via gap detection, than typically hearing listeners (Blankenship et al., Reference Blankenship, Zhang and Keith2016; Cesur and Derinsu, Reference Cesur and Derinsu2020; Duarte et al., Reference Duarte, Gresele and Pinheiro2016) – if better temporal resolution than pre-lingual CI recipients (Wei et al., Reference Wei, Cao, Jin, Chen and Zeng2007). Hence, CI speakers’ variable articulatory strategies may reflect not only the absence of auditory feedback during the pre-implantation period but the temporally degraded nature of the information provided by their device post-implantation.
42.3.2.2.3 Musical Rhythm Perception and Production by Adults with HL
As with the literature concerning children, the evidence for a general deficit in musical rhythm perception in adults with HL is mixed. Some studies report similar performance in rhythm perception across typically hearing and CI-listener groups (e.g., Brockmeier et al., Reference Brockmeier, Fitzgerald and Searle2011; Limb et al., Reference Limb, Molloy, Jiradejvong and Braun2010). A study comparing adults with post-lingual HL who used either HAs or CIs found them to perform similarly in a rhythm discrimination task (Looi et al., Reference Looi, McDermott, McKay and Hickson2008). Another experiment compared CI listeners to typically hearing listeners, although the control group in this case was much younger and more musically trained than the CI group (Kong et al., Reference Kong, Cruz, Jones and Zeng2004). There was no difference in performance on the basis of tempo discrimination but mixed results for CI listeners in rhythmic pattern identification, with particular difficulty associated with complex melodic rhythmic stimuli (Kong et al., Reference Kong, Cruz, Jones and Zeng2004). In another study examining production, adult CI users were observed to synchronize their movements to non-pitched drum rhythms as well as they could to a simpler, recurring tone, as well as a visual timing cue (Phillips-Silver et al., Reference Phillips-Silver, Toiviainen and Gosselin2015). Their sensorimotor timing was poorer than typically hearing controls, however, and they also tended to struggle more with more complex rhythmic stimuli that also incorporated changes in pitch.
It is sometimes assumed that timing information (e.g., the beat in music) is adequately available to most CI listeners (McDermott, Reference McDermott2004); however, an important limitation to the current literature on nonverbal rhythm perception with CIs is that stimuli complexity is rarely parameterized or formally compared within experiments. As a recent review notes, “many of the studies regarding CIs and rhythm perception utilize relatively simple perception tasks, such as basic pattern reproduction or sound stimuli that isolate rhythmic information” (Jiam and Limb, Reference Jiam and Limb2019:26). The few studies that have included both simple as well as more complex stimuli report performance differences across these conditions within CI listeners (Gfeller et al., Reference Gfeller, Woodworth, Robin, Witt and Knutson1997; Kong et al., Reference Kong, Cruz, Jones and Zeng2004; Phillips-Silver et al., Reference Phillips-Silver, Toiviainen and Gosselin2015). For example, adult CI listeners achieved scores similar to typically hearing controls in a simple rhythm perception task, but had significantly poorer performance in a more complex temporal patterning task that the authors suggest may better convey “aspects of musical experience such as phrasing, accent, and complex rhythmic structure” (Gfeller et al., Reference Gfeller, Woodworth, Robin, Witt and Knutson1997:259). Unlike the studies of young CI users mentioned previously, EEG studies investigating musical rhythm perception in adult CI users have not found a robust mismatch negativity to temporal deviants (Sandmann et al., Reference Sandmann, Kegel and Eichele2010; Timm et al., Reference Timm, Vuust and Brattico2014). Given this measure was present in children and adolescents, albeit reduced in comparison to typically hearing participants, the adult data may reflect differences in device advances and age of implantation, musical enjoyment and listening habits, or life experience as a result of changing attitudes towards HL and musical participation. Together, these inconsistencies across studies make it difficult to assess which temporal cues are or aren’t well conveyed by CIs, let alone what commonalities may be found across verbal versus nonverbal rhythm perception. Future work should clarify whether and how musical rhythm perception changes after HL, but if rhythm sensitivity can be trained in adults, it may serve as a specific, operationalizable target for speech-focused interventions as an alternative to general music or dance training, for example.
42.3.2.3 Speech Rhythm and Ageing
Age-related HL is associated with changes throughout the auditory system, from the periphery to the cortex (Peelle and Wingfield, Reference Peelle and Wingfield2016), and some aspects of temporal processing may diminish independently of sensorineural HL (Ajith Kumar and Sangamanatha, Reference Ajith Kumar and Sangamanatha2011; Füllgrabe et al., Reference Füllgrabe, Moore and Stone2015). As many older adults are likely to have some degree of HL, especially at higher frequencies, it is important to consider the interaction between HL and changes to central temporal processing associated with older populations. For example, psychophysical experiments show that temporal resolution – typically measured using silent gap detection tasks – as well as sensitivity to dynamic temporal properties, such as amplitude modulation, decline with age (see Walton, Reference Walton2010, for a review).
There are also ageing-related differences in the neural response to nonverbal rhythmic or temporally structured sounds (Henry et al., Reference Henry, Herrmann, Kunke and Obleser2017; Herrmann et al., Reference Herrmann, Maess and Johnsrude2023; Herrmann and Butler, Reference Herrmann and Butler2021; Parthasarathy et al., Reference Parthasarathy, Bartlett and Kujawa2019). For instance, the neural response in older, typically hearing adults to temporally regular, amplitude-modulated noise (4 Hz) is heightened in comparison to younger adults (Herrmann et al., Reference Herrmann, Buckland and Johnsrude2019). This hyper-responsiveness was observed in cortical areas associated with low-level auditory processing; yet, when examining a different neural marker that is associated with the higher-order representation of temporal structure, the authors found it to be suppressed in comparison to younger adults (Herrmann et al., Reference Herrmann, Buckland and Johnsrude2019). This may suggest that difficulties with temporal processing in ageing are not strictly related to a degraded representation of the target temporal structure but rather the susceptibility to irrelevant or intrusive maskers. For example, older and younger typically hearing adults performed a task wherein a prosody-like rhythmic pattern of modulated noise was identified against randomly modulated distractor noise with matched temporal statistics; despite typical hearing, the older adults were less accurate than the younger adults, except for in a control condition wherein targets were presented in isolation without distracting noise – in which case, the groups did not differ (Divenyi and Brandmeyer, Reference Divenyi and Brandmeyer2003). Hence, one possible function of speech rhythm for enhancing speech intelligibility in noise could be related to the top-down inhibition of unwanted noise related to over-responsive sensory processing, as opposed to the allocation of neural-cognitive resources towards anticipated, perceptually salient events in the future (Henry and Herrmann, Reference Henry and Herrmann2014). The balance between these listening strategies could shift with ageing and/or the temporal predictability of the target versus distractor. Explicit attention and task demands likely also have some part to play (Henry et al., Reference Henry, Herrmann, Kunke and Obleser2017). For example, a recent study using EEG found that the neural response to deviations in syllabic stress was larger in older in comparison to younger adults; however, in that study, participants were instructed to ignore auditory stimuli and were visually distracted by an unrelated, silent film (Giroud et al., Reference Giroud, Keller, Hirsiger, Dellwo and Meyer2019). More work is needed to understand how target-tracking versus noise-inhibiting mechanisms may interact, and whether they can be experimentally dissociated in the context of attentive speech rhythm perception.
42.4 Discussion
42.4.1 Future Directions for the Rehabilitation of Speech Rhythm in HL
For individuals with HL, engagement with speech rhythm depends on complex, intersecting factors. These factors can include the underlying cause of the HL, its onset relative to language acquisition, and which forms of amplification are used (if any), in addition to numerous other determinants. What is clear is that there is no singular experience of HL. As this chapter has discussed, HL may affect the perception of speech timing from the inner ear to the cortex, and individual differences in sensitivity to temporal cues and rhythmic structure could also help to explain why outcomes in functional hearing differ so much across listeners, irrespective of audiometry. Finally, HL can be further complicated by ageing, which is accompanied by cascading changes throughout the central nervous system that also affect temporal processing. In general, the empirical evidence surveyed in this chapter suggests that many people with HL encounter difficulties involving the perception and production of speech rhythm, but this is unlikely to be caused by the complete absence of timing information, particularly for listeners with CIs. Rather, in accordance with the literature examining musical rhythm perception, it appears that the greatest challenges for people with HL arise under ecological conditions, namely, when temporal patterning and structure represent just one facet of a multidimensional, sensorially rich percept – take, for example, multi-party conversational speech in a busy café. Hence, it is crucial that researchers make every effort to incorporate realism into experimental stimuli moving forward.
42.4.1.1 Multimodal Cues for Speech Rhythm
Despite the aforementioned complexity and inter-individual differences in HL, preliminary interventions involving rhythm training show some promise for improving speech and language in children and adolescents (León Méndez et al., Reference del Carmen León Méndez, Fernández García and Daza González2023), and the observed correlations between nonverbal rhythm processing and speech in noise perception in adults (e.g., Slater et al., Reference Slater, Kraus and Woodruff Carr2018) invite further study. The possibility that improving rhythm skills could support functional hearing is exciting, but substantial work is needed to clarify the nature of incidentally found relationships, to establish causality, and to show that rhythm-based interventions will generalize across untrained materials as well as over time. Other forms of rehabilitation could also incorporate speech timing more explicitly; for instance, co-speech gesture may represent a rich source to facilitate speech perception in HL, especially for speech in noise understanding. Although gesture has been identified as a potential target for training or intervention (Sparrow et al., Reference Sparrow, Lind and van Steenbrugge2020), rhythm and prosodic aspects of gesture have received comparatively less attention than their semantic or representational content (Prieto et al., Reference Prieto, Cravotta, Kushch, Rohrer and Vilà-Giménez2018). Previous research into audiovisual speech suggests prosodic processing, as well as speech intelligibility, is enhanced when rhythmic visual cues are also present (Dohen and Loevenbruck, Reference Dohen and Loevenbruck2009; Krahmer and Swerts, Reference Krahmer and Swerts2007; Peña et al., Reference Peña, Langus, Gutiérrez, Huepe-Artigas and Nespor2016). Moreover, adults with CIs benefit more than typically hearing adults from visual input in intonation perception (Lasfargues-Delannoy et al., Reference Lasfargues-Delannoy, Strelnikov, Deguine, Marx and Barone2021). This suggests that, together with lip-reading and other dynamic visual cues, gesture awareness could possibly benefit listeners by, for instance, guiding auditory rhythmic expectations, or providing additional prosodic information to repair misunderstandings. Similarly, vibro-tactile stimulation may also enhance access to speech rhythm as a complementary perceptual representation (Guilleminot and Reichenbach, Reference Guilleminot and Reichenbach2022).
42.5 Conclusion
HL is one of the most common disabilities globally, and with the proportion of older adults increasing, its incidence is projected to grow (Haile et al., Reference Haile, Kamenov and Briant2021). As scientific research into speech rhythm develops and expands, the scholarly community should account for differences – audiological, cognitive, and developmental, among others – across listeners, as well as inequalities in access to hearing care and devices. This is critical when we consider that rhythm is cross-modal and multisensory, making it amenable, as a domain, to different forms of sensory disability and underscoring its promise as a target or tool for rehabilitation. Yet, much more can be done to understand how rhythm is processed and experienced across individuals, and how best to rehabilitate speech rhythm perception and production remains an open question. The challenge, therefore, lies with speech rhythm researchers to consider the full breadth of human hearing ability and individual listener experiences when planning our programs for the future.
Summary
HL affects many people worldwide, but how speech rhythm is perceived and produced by people with HL remains poorly understood. This chapter draws from an interdisciplinary literature spanning psychology, neuroscience, linguistics, and clinical studies to form a broad overview of speech rhythm in auditory disability.
Implications
Empirical evidence suggests that expressive and receptive experiences of speech rhythm differ on the basis of hearing status, and this variation may also interact with other changes in central temporal processing, such as those associated with ageing. Researchers should, therefore, account for these forms of disability and difference when modelling and developing new theories concerning speech rhythm.
Gains
Speech rhythm in the context of HL has received relatively little attention, and the relevant empirical evidence is yet to be fully integrated and synthesized across disciplinary boundaries. Understanding how speech rhythm is perceived and produced by people with HL will help speech and hearing scientists to better understand its functional role, and may pave the way for future rhythm-based clinical interventions.