22.1 Introduction
Linguistic rhythm is classically viewed as the binary alternation of stressed and unstressed syllables. Selkirk (Reference Selkirk1984) calls this the principle of rhythmic alternation (PRA), which arises in trochaic languages, such as German or English, from the underlying trochaic structure of the lexicon (see also Sweet, Reference Sweet1876; Féry, Reference Féry1998). (1) shows an example of alternation of stressed and unstressed syllables.
(1) Péter wóllte héute láufen.
Peter wanted today (to) run.
“Peter wanted to run today.”
In the present chapter it will be argued that a balance measure is better suited than binary alternation to formalize speech rhythm since multiple violations naturally occur in everyday speech as a result of the clustering of stressed or unstressed syllables. The juxtaposition of two or more stressed syllables produces a stress clash, whereas the clash of two or more unstressed syllables produces a stress lapse (Selkirk, Reference Selkirk1984; Hayes, Reference Hayes1995). Correspondingly, *LAPSE denotes a constraint according to which more than one unstressed syllable in the sequence is to be avoided. *CLASH denotes a constraint that a sequence of more than one stressed syllable is to be avoided. The example in (2) shows a sentence with three stress lapses (the relevant syllables are underlined).
(2) Róbert ságte, dass Nádja ihn bekláuen sóllte.
Robert said that Nadja him steal (from) should.
“Robert said that Nadja should steal from him.”
While no system currently exists to formalize the rhythmicity of an entire sentence, Shih et al. (Reference Shih, Grafmiller, Futrell, Bresnan, Vogel and Vijver2015) propose a measure that makes formalization possible, at least for smaller phrases. Here, a 1 is subtracted from the number of unstressed syllables that line up between two stressed syllables (see 3a). (3b) symbolizes a rhythmic alternating structure, where stressed syllables are represented by X and unstressed syllables by x. (3c) shows a LAPSE and (3d) a CLASH. In the course of rhythmic alternation, the value should therefore approach 0, as in (3b). *LAPSE (3c) and *CLASH (3d) are treated as equivalent according to Shih et al. (Reference Shih, Grafmiller, Futrell, Bresnan, Vogel and Vijver2015). They are each given the value 1.
(3a) | Number of unstressed syllables between two stressed syllables -1|
(3b) x X x X x; |1-1| = 0
(3c) x X x x X x; |2-1| = 1
(3d) x XX x; |0-1| = 1
(3a–d) show that the rhythmic structure of sequences with two stressed syllables can be formalized quite well. However, if the number of stressed syllables increases to three or more, the measure operates only to a limited extent; that is, it only refers to the space between two stressed syllables without relating the values to one another in longer structures. Thus, it is not possible to systematically distinguish between (4a), (4b), and (4c), although (4a) apparently has a much more balanced rhythm than (4b) and (4c).
(4a) X xx X xx X xx X xx
(4b) X xxxx X x X X
(4c) X xxxx X x X xx X
The examples already suggest that the binary description of rhythm is not sufficient for more complex linguistic units. In (4a), for example, a dactylic rhythm is indeed shown, which, in addition to the trochee, may represent a rhythmic unit perceived as regular for everyday German (see, among others, Hanna, Reference Hanna2003; Vogel et al., Reference Vogel, van de Vijver, Kotz, Kutscher, Wagner, Vogel and Vijver2015). Therefore, the goal of this chapter is to develop a rhythmic measure that goes beyond a binary alternation. To achieve this, spoken sentences from a study on the position and prominence of the German object pronoun are reanalyzed.
22.1.1 Syllable Prominence and the German Pronoun
Syllable prominence is not conceived as an absolute but as a quantity dependent on the neighboring syllables: “the prominence of a syllable is always defined relative to the prominence of other syllables in the same phrase” (Gollrad, Reference Gollrad2013: 11; see also Cangemi and Baumann, Reference Cangemi and Baumann2020). To illustrate this, consider the examples in Figure 22.1 used several times in this chapter. The examples here are visualized in metrical grids (Liberman and Prince, Reference Liberman and Prince1977; Hayes, Reference Hayes1980; Halle and Vergnaud, Reference Halle and Vergnaud1987). In such a grid, the degree of prominence of a syllable is visible from the set of crosses accumulated vertically above the syllable. Syllable prominence is therefore revealed by the height of the grid. Finally, rhythmic structure can be determined by the horizontal spacing of similarly prominent syllables.
Grid with four levels.
Selection of rhythmic and nonrhythmic structures in the metrical grid (with the four levels 1: unstressable, 2: unstressed, 3: stressed, and 4: accented). The upper part shows the (relatively) rhythmic sentences (dactylic on the left and approximately trochaic on the right). The lower part shows the corresponding (relatively) unrhythmic structures. The arrows visualize the spacing of the prominent syllables.

Following the structured preparation of the concept of prominence by Wagner et al. (Reference Wagner, Origlia and Avezani2015), the aspects that are particularly relevant for the definition of prominence will now be captured. For the sake of clarity, these aspects are listed:
Linguistic entity: As a linguistic entity that can be more or less prominent, here for the study of rhythm I am concerned with the syllable.
Stands out: Standout is defined in the context – it is lower for prominent syllables on the left and right than on the syllable itself. Furthermore, the perceptual evaluation of syllable prominence is used as a measure here.
Prosodic characteristics: Prosodic characteristics are defined in terms of lexical stress and accentuality. A distinction is made between stressable, non-stressable, stressed, and accented syllables (see Figure 22.1). Furthermore, rhythmic processes, for example in the form of the addition of prominence, are discussed for three consecutive unaccented syllables (see Figure 22.4).
Its environment: The environment refers to the linguistic material surrounding the syllable, specifically the surrounding syllables and words within a complement clause, as in (2).
Finally, we can consider the RhythmRule. Liberman and Prince (Reference Liberman and Prince1977) provide a systematization of the PRA (Sweet, Reference Sweet1876; Selkirk, Reference Selkirk1984). They show repair mechanisms or rhythmic processes that are activated when violations occur. These are summarized under the term RhythmRule (see Henrich, Reference Henrich2015). Assuming that the RhythmRule is active across domains (Henrich, Reference Henrich2015; Vogel et al., Reference Vogel, van de Vijver, Kotz, Kutscher, Wagner, Vogel and Vijver2015), it is worth recalling Example (2): On the one hand, the pronoun is not accented and, as a function word, is usually considered unstressed; however, it can gain prominence if we assume that the RhythmRule takes effect here. Thus, the sentence in (2) could be considered both rhythmic and nonrhythmic.
22.1.2 The Position of the German Pronoun in (Dis)rhythmic Sentences
In German, both subject-object (SO) and object-subject (OS) sequences are possible. Therefore, serialization is subject to various restrictions (Bader, Reference Bader2020). Uszkoreit (Reference Uszkoreit1986) formulated tendencies to produce the agent first – but also to prefix pronominal elements. For the object pronoun, which tends to fill the thematic role PATIENS, this implies variation in placement. Furthermore, the position has implications for the relative prominence of the pronoun, as, for example, Kügler (Reference Kügler2018) and Zerbian and Böttcher (Reference Zerbian, Böttcher, Calhoun, Escudero, Tabain and Warren2019) have shown for pronouns in the dative case.
Of interest for the study here is the position and prominence of the object pronoun, typically considered unaccented, in the midfield. In German, the midfield begins after the left clause bracket, which in verb-second clauses consists of the conjugated verb. In German complement clauses, the midfield begins after the complementizer. In the midfield, pronouns precede non-pronominal elements (Hoberg, Reference Hoberg1981).
Franz (Reference Franz2022) investigated the influence of rhythm on the position of the accusative object pronoun and the pronominal adverb. One of the questions was whether a manipulation of the stress structure of the neighboring words can influence the preferred placement of the pronoun in silent reading. In a questionnaire study, sentences were presented in pairs, each of which differed in terms of word position: Thus, the unaccented element of interest was either prefixed or in a position that could be considered canonical. Furthermore, sentence pairs differed in rhythmicity: Thus, in one of the two serializations, they were each more rhythmic, depending on the stress structure of the surrounding elements (see Figure 22.1).
The participants preferred the rear position of the unaccented elements in both structures studied. Furthermore, they preferred the prefixation of the elements more often when this led to rhythmic alternation. The results on the pronominal adverb are in line with those of Vogel et al. (Reference Vogel, van de Vijver, Kotz, Kutscher, Wagner, Vogel and Vijver2015).
Moreover, the findings illustrate that rhythmic well-formedness is not necessarily to be understood as a binary alternation of stressed and unstressed syllables. In the studied sentences with pronominal adverbs, rhythmic alternation indicates a dactylic structure, whereas in those with object pronouns, it is approximately trochaic (see Figure 22.1).
The sentences in Figure 22.1 illustrate the rhythmic structure of the sentences within the metrical grid. The sentence on the upper left shows a rhythmic, dactylic structure, and the one on the upper right shows a rhythmic, approximately trochaic structure. The lower part of the figure shows the variants that are unrhythmic due to a change in the word sequence. The spacing of the prominent syllables visualizes the rhythmicity in the structures as balanced (top) and unbalanced (bottom).
22.2 The Present Study
The current investigation is based on a study on the serialization of the German object pronoun in the accusative. The initial question of the study was whether a manipulation of the stress structure of the neighboring words can influence the placement of the German object pronoun in speech. This was done to shed light on the workings of rhythm in psycholinguistic models of speech production (for details, see Franz, Reference Franz2022). In this post hoc analysis, the respective materials and datasets are used in order to develop a measure of sentence rhythm.
In the respective picture-based study, participants produced complement sentences (as in 5) with the placement of the object pronoun (ihn, “him,” in bold) as the dependent variable (SO in 5a; OS in 5b).
(5a) Der Junge sagt, dass Markus ihn auslacht.
The boy says that Markus him laughs (at).
(5b) Der Junge sagt, dass ihn Markus auslacht.
The boy says that him Markus laughs (at).
“The boy says that Markus is laughing at him.”
The materials consisted of 32 sentences. Varying factors in the target sentences were the stress pattern of the embedded subject (iambic, trochaic) and of the embedded verb (initial stress, no initial stress). An overview of the rhythmic structure of the sentences is given in Figure 22.3. It was predicted that trochaic embedded subjects (see Figure 22.3, condition a and c) and verbs with no initial stress (see Figure 22.3, condition a and b) promote sentences with a fronted pronoun (OS). Additionally, animacy varied as a between-items factor with the matrix sentence subject being human or nonhuman.
Stimuli pictures were 64 black and white drawings developed with a professional illustrator. Each picture consisted of a left part symbolizing the matrix sentence (Der Junge sagt, “the boy says”) and a right part symbolizing the embedded sentence (dass Markus ihn auslacht, “that Markus is laughing at him”). The right parts of the pictures were mirrored (yielding 128 stimuli) to avoid word order effects due to spatial order (Figure 22.2). Finally, the variation in the degree of animacy was realized by exchanging Der Junge, “the boy,” with a nonhuman referent der Hase, “the rabbit.”
Visual stimuli.
Mirrored stimulus example for the target sentence Der Junge sagt, dass Markus ihn auslacht (SO)/Der Junge sagt, dass ihn Markus auslacht (OS), “The boy says that Markus is laughing at him.”

The results had shown that participants were more likely to produce OS when the embedded subject was of trochaic structure – however, this became significant only in the human subset (for a detailed analysis, see Franz, Reference Franz2022). The stress structure of the verb had no clear effect on sequencing. Furthermore, the analysis of perceptual syllable prominence had shown different processes of rhythmic accommodation. This is further differentiated in the current post hoc analysis.
22.2.1 Development of the Metric for Longer Structures Based on Shih et al. (Reference Shih, Grafmiller, Futrell, Bresnan, Vogel and Vijver2015)
For a clarification of the predictions, a closer look at the rhythmic structures of the target sentences is now taken due to the respective opposite predictions of verb and subject. Figure 22.3 shows the number of rhythmic violations according to Shih et al. (Reference Shih, Grafmiller, Futrell, Bresnan, Vogel and Vijver2015) before and after the embedded subject in all four conditions and for both variants (SO/OS). The problem with this “local” definition of rhythm will be discussed in the following.
Development of the balance measure.
Example sentences in conditions (a–d) and the two serialization options (SO and OS). The bold portions indicate the syllables with lexical accent. The numbers below the text represent the number of rhythmic violations (stress clash/stress lapse) according to Shih et al. (Reference Shih, Grafmiller, Futrell, Bresnan, Vogel and Vijver2015). The digits to the right of each sentence show the sum of rhythmic violations. The arrow in condition (d) shows the change that was made here to represent the quality of lapse and clash (this value is used for the balance measure). The balance measure results from a subtraction of the two values of a sentence; the ranking index (R) shows the respective strength for a preference of OS or SO (see the text for details).

Shih et al. (Reference Shih, Grafmiller, Futrell, Bresnan, Vogel and Vijver2015) restrict their analysis to genitive constructions that contain at most two accents – a single or local determination of *LAPSE and *CLASH is thus sufficient. But what happens when the rhythmic structure of an entire sentence is to be determined? With about three accents, two measures could be summed up, and the one containing fewer violations in the sum would be more rhythmic. To test this, let’s take a closer look at the target sentences of the experiment. Figure 22.3 shows the summed violations to the right of each sentence. It turns out that summation does not make clear predictions. With the exception of condition (d), the two variants are predicted to be equivalent in all conditions. In contrast, there is an apparent varying distance between the respective accented syllables, as, for example, in the comparison of OS and SO in condition (a) – so far, however, there is no measure that can capture this balance. A first attempt consists in subtracting the two values calculated according to Shih et al. (Reference Shih, Grafmiller, Futrell, Bresnan, Vogel and Vijver2015) in the sentence. For this purpose, the conditions are examined in more detail.
22.2.1.1 Condition (a)
Based on the local rhythmic predictions, it has already been established above that in condition (a) the prefixation of the pronoun is predicted because of the iambic verb and trochaic subject. Looking at the sentence as a whole, however, this prediction seems less clear (Figure 22.3): The prefixation of ihn, “him,” generates a *LAPSE with the unstressed complementizer dass, “that” (*dass ihn, formalized in 6a); at the same time, the encounter of the unstressed final syllable of the trochaic subject and the unstressed initial syllable of the iambic verb generates another *LAPSE (*-kus be-, formalized in 6b).
(6a) |2-1| = 1
(6b) |2-1| =1
In the SO structure less predicted on the basis of local rhythmicity, a double *LAPSE arises to the right of the subject by the unstressed final syllable of the trochaic name, the unstressed pronoun, and the initial syllable of the iambic verb (-**kus ihn be-, formalized in 7b). To the left of the subject, no violation occurs with an alternation of stressed matrix verb, unstressed complementizer, and stressed initial syllable of the embedded subject (sagt, dass Mar-, formalized in 7a).
(7a) |1-1| = 0
(7b) |3-1| = 2
Intuitively, it seems clear that a “balanced” distribution of violations should be preferred here, and thus an OS structure. Also, it has already been shown that both trochaic and dactylic rhythms can influence word order preferences, and thus a pure alternation of stressed and unstressed syllables is not mandatory for rhythmic well-formedness. However, the sentences in this study have neither a pure trochee nor a pure dactyl, and formal addition generates the same number of violations (exactly two each) in condition (a) for OS and SO.
If we now assume that the spacing of the accents should be balanced, one value can be subtracted from the other for a balance measure, and an optimal balance measure is at a value of 0 (formalized in 8a). In this way, the formula of Shih et al. (Reference Shih, Grafmiller, Futrell, Bresnan, Vogel and Vijver2015) can be extended for longer structures. When applied to condition (a), a 0 (formalized in 8b) is obtained for the OS structure, and a 2 (formalized in 8c) for the SO structure. In this example, therefore, the OS variant formalized in (8b) is predicted with a balance measure of 0.
(8a) | value left – value right |
(8b) | 1-1 | = 0
(8c) | 0-2 | = 2
22.2.1.2 Condition (b)
Based on the local rhythmic predictions, it has already been established above that in condition (b), because of the iambic verb, the prefixation of the pronoun is predicted (OS; see Figure 22.3). At the same time, the iambic subject favors an SO structure, leaving the prediction unclear without the total sentence. The two variants in condition (b) each have two violations: In the SO structure, they are evenly distributed, respectively, with the two unstressed syllables of the complementizer and the first syllable of the iambic subject (*dass Mar-), and the pronoun together with the first syllable of the iambic verb (*ihn be-). In the OS structure, on the other hand, the violations of *LAPSE gather with the complementizer, the pronoun, and the first syllable of the iambic subject in a double *LAPSE structure (**dass ihn Mar-). Transferred to the formula in (9a), this implies an ideal balance measure for the SO structure in (9b), but not for the OS structure in (9c). SO is therefore predicted for condition (b).
(9a) | value left – value right |
(9b) | 1-1 | = 0
(9c) | 0-2 | = 2
22.2.1.3 Condition (c)
The local rhythmic predictions favor an OS structure due to the trochaic subject and an SO structure due to the trochaic verb. As in condition (b), the prediction due to local rhythmicity is unclear for condition (c). Unlike in condition (b), it remains so when looking at the complete sentence: The pronoun produces a *LAPSE with one of the adjacent syllables in both positions (SO and OS), with the second syllable of the trochaic subject (*-kus ihn) in the OS variant, and with the complementizer (*dass ihn) in the SO variant (see Figure 22.3). This is illustrated by the formulas in (10b) and (10c). For condition (c), therefore, there is no clear prediction by the rhythmic structure.
(10a) | value left – value right |
(10b) | 0-1 | = 1
(10c) | 1-0 | = 1
22.2.1.4 Condition (d)
In the OS example of condition (d), we obtain two violations of *LAPSE before the embedded subject (Marcel), resulting from the unstressed complementizer, the pronoun, and the initial syllable of the iambic subject (**dass ihn Mar-; see Figure 22.3). (11a) formalizes the *LAPSE structure according to Shih et al. (Reference Shih, Grafmiller, Futrell, Bresnan, Vogel and Vijver2015) (|number of unstressed syllables - 1|). Furthermore, the coincidence of the stressed second syllable of the iambic subject and the initial syllable of the trochaic verb (*-cel aus-), which is also stressed, creates a *CLASH (marked by 1), which is formalized in (11b) according to Shih et al. (Reference Shih, Grafmiller, Futrell, Bresnan, Vogel and Vijver2015).
(11a) |3-1| = 2
(11b) |0-1| = 1
In the SO example of condition (d), the (local) rhythmic structures look more balanced: No rhythmic violation occurs to the right of the embedded subject because the unstressed pronoun stands alone between the stressed final syllable of the subject and the stressed initial syllable of the verb (-cel ihn aus-). (12a) formalizes the alternating structure. To the left of the embedded subject, the unstressed complementizer and the unstressed initial syllable of the iambic subject create a *LAPSE (dass Mar-); this is formalized in (12b).
(12a) |1-1| = 0
(12b) |2-1| = 1
Thus, in condition (d), an SO structure is more rhythmic than an OS structure. The strength of this preference will be discussed in the following.
22.2.1.5 Reflections on the Quality of Lapse and Clash
Transferring the values of condition (d) to the balance measure introduced above yields a counterintuitive result: Although the OS structure has significantly less balanced spacing of accents compared to the SO variant, the balance measure (13a) yields the same result for both variants (OS in 13b, SO in 16c).
(13a) | value left – value right |
(13b) | 2-1 | = 1
(13c) | 0-1 | = 1
This is explained by the fact that Shih et al. (Reference Shih, Grafmiller, Futrell, Bresnan, Vogel and Vijver2015) are only concerned with local rhythmicity, not with the balance of larger sections, and the sentences considered here so far have only had lapses as violations. Apparently, however, *LAPSE and *CLASH have different effects on the spacing of accents in a sentence. While lapses move them gradually apart, a clash has the opposite effect: The stressed syllables of the accented words meet directly, preventing rhythmic alternation.
To account for this difference, the modulus proposed in Shih et al. (Reference Shih, Grafmiller, Futrell, Bresnan, Vogel and Vijver2015) is omitted. In this way, the values at a *LAPSE and a *CLASH receive different signs, producing values that relate not only to the number of rhythmic violations but also to their quality (see Figure 22.3). (14a) illustrates the corresponding rhythmic measure. (14b) shows the resulting balance measure, and (14c) and (14d) formalize the balance measures for the SO variant (14c) and the OS variant (14d) in condition (d).
(14a) (Number of unstressed syllables between two stressed syllables -1)
(14b) (value left) – (value right)
(14c) (1-1) – (2-1) = -1
(14d) (2-1) – (0-1) = 2
It turns out that (14c) and thus the SO variant is closer to 0 than the OS variant (14d). Consequently, the new rhythm and balance measure predicts the rhythmicity of the two word order variants according to the visual impression: SO is more rhythmic than OS in condition (d). Figure 22.3 summarizes the discussed values per condition for OS and SO, respectively. The optimal rhythmic structure is indicated by a balance measure (= 0). It can be seen, as in the analysis of local rhythmicity, that in condition (a) OS is predicted, in condition (c) there is again no clear prediction, and conditions (b) and (d) each prefer SO. The extent to which the development of the balance measure goes further than the previous analysis will be shown below.
22.2.1.6 Development of a Rhythmical Ranking
In the previous section, a measure was developed that maps the rhythmic structure within three accented syllables, taking into account not only the number of rhythmic violations but also their quality. This was done by qualitatively distinguishing *LAPSE and *CLASH and integrating them into the measure. In this way, for each possible utterance within one of the experimental conditions, there now exists a balance measure that expresses the rhythmic well-formedness of the sentence: The farther the value is from 0, the more unrhythmic the sentence is. Thus, the structure with a value closer to 0 is predicted. According to this measure, it can be stated that in condition (a) OS is predicted, in condition (c) there is no clear prediction, and conditions (b) and (d) both predict SO.
The gain of the balance measures now is that they are meaningful not only within conditions but also between conditions. If we first relate the balance measures within each condition to each other by adding them up, we obtain a value that is called ranking index (R) (see Figure 22.3). This value represents the strength of the predictions – the larger its distance from 0, the stronger is the prediction.
Accordingly, the prediction for SO is strongest in condition (d) with an R of 4, followed by condition (b) with an R of 2. Not further specified is the probability for SO in condition (c) with an R of 0, and it is finally least likely in condition (a) with an R of -2. The strength of the predictions for SO is consequently distributed between the conditions, as in (15a). Accordingly, the predictions for OS per condition are exactly opposite: OS is most likely in condition (a), followed by conditions (c), (b), and (d). This is summarized in (15b).
(15a) prediction SO: d >> b >> c >> a
(15b) prediction OS: a >> c >> b >> d
Accordingly, the balance measures allow us to map a graded predictive power and also show that the experimental design allows for rhythmically well-formed SO sentences rather than correspondingly rhythmic OS sentences. This might explain the rather weak rhythmic effects on sequence preferences (see Franz, Reference Franz2022, for details).
Moreover, the gradations separated into OS and SO in (15a) and (15b) can be summarized into a single four-level ranking. Thus, OS sentences of condition (a) and SO sentences of condition (d) are the most likely – these form level 1 in the ranking. Accordingly, OS sentences of condition (d) and SO sentences of condition (a) are the least likely – these form level 4 in the ranking. The prediction is simply that level 1 sentences should be the most frequent, followed by levels 2, 3, and 4. The complete four-level ranking is summarized in (16a–d). In each case, the conditions are in parentheses after the preferred structure. The corresponding example sentences are shown in Figure 22.3.
(16a) Level 1: SO(d), OS(a)
(16b) Level 2: SO(b), OS(c)
(16c) Level 3: SO(c), OS(b)
(16d) Level 4: SO(a), OS(d)
22.2.2 Methods
22.2.2.1 Participants
Fifty-one experimental participants (age: 19–81 years; M = 41.8; SD = 18) with German as (one of) their first language(s) took part in the study. Thirty-one of them were female. Three of the participants reported being bilingual (German/Polish, German/Vietnamese, and German/French). The participants were recruited through the subject pool of the Max Planck Institute for Empirical Aesthetics in Frankfurt, Germany. All reported normal or corrected-to-normal vision and no severe hearing, vision, speech, or neurological impairments. All participants gave their written consent for voice recordings to be made and for their data to be processed pseudonymously. Each participant received an expense allowance of 15 euros. All participants had the opportunity to discontinue the experiment at any time without giving reasons.
22.2.2.2 Materials
Materials used are given in Sections 22.2 and 22.2.1.
22.2.2.3 Procedure
The 128 stimuli were distributed over four different lists. Each list contained 32 stimuli, balanced by item and condition. Each of the lists was subjected to the same pseudorandomization using MIX (Casteren and Davis, 2006).
In addition, 32 stimuli were added to each sequence that contained the target sentences in written form so that they could be read aloud by participants. In a final step, pictorial and written stimuli were mixed in such a way that four blocks of 16 stimuli each were arranged consecutively. The first block consisted of written sentences, the second of the corresponding pictorial stimuli, the third of written stimuli, and the fourth again of the corresponding pictorial stimuli. (17) shows a schematic representation of a sequence.
(17) read out 16 sentences >> name 16 pictures >> read out 16 sentences >> name 16 pictures
The complete experiment took place in the laboratory of the Max Planck Institute for Empirical Aesthetics, in a soundproof room where the participant sat alone at a desk in front of a computer screen. The experimenter (the author of this chapter) controlled the experiment from an adjoining room, and contact with the participant was via an intercom system. The experiment was preceded by a distinct familiarization phase with the target items and target structures (for details, see Franz, Reference Franz2022).
The presentation of the sequences explained above and the recording of the individual utterances were performed in Matlab. Accordingly, 64 WAV files were created per participant – in the context of this chapter, only those that were created in response to the pictorial stimuli will be discussed (32 recordings per participant). The further parts of Experiment 1 (runs 2–4) were performed according to the variable speed of each participant only with those who had at least 10 minutes of their 60-minute session left. These further parts corresponded to the test phase of Experiment 1. Experiment 2 (a writing task) will not be discussed here.
22.2.2.4 Analysis
All participants’ utterances were subsequently transcribed and coded by three independent student assistants. First, all valid utterances were coded as to whether the pronoun in the complement clause was produced before (1: dass ihn Marcel massiert, “that Marcel massages him”) or after the embedded subject (0: dass Marcel ihn massiert, “that Marcel massages him”). An utterance was considered valid if it contained one of the two required matrix sentences (der Hase träumt, “the rabbit dreams”/Der Junge sagt, “the boy says”) as well as a complement sentence in the present active tense with bisyllabic verb, bisyllabic subject, and the object pronoun ihn, “him.” The factors of stress structure of embedded verbs and subjects (2: iambic/1: trochaic) were made on the basis of the utterance actually produced.
Additionally, the degree of animacy of the matrix subject was annotated (1: human/0: nonhuman), as well as the spatial arrangement (left/right) of the figures representing the subject and object of the complement clause in the stimulus (referent of object on the left: 1; referent of object on the right: 0; see Figure 22.2).
Further, the fluency of the utterances was coded on a three-point scale (1: The utterance contained no pauses, filler words, self-corrections, repetitions, or the like; 0.5: The utterance was interrupted only before the complement clause by one or more of the above-mentioned fluencies; 0: The utterance was [also] interrupted in the complement clause by one or more of the above-mentioned fluencies). Only the utterances with fluent complement clauses (fluency at least 0.5) were integrated into the analysis (n = 1819) in order to allow for an audible comparison between the prominence of the pronoun and that of the adjacent syllables during annotation.
Finally, the syllable prominence of the pronoun was coded. A student assistant with German as one of their first languages annotated the utterances. A distinction was made here between unstressed (n = 1355), stressed (n = 196), and reduced (n = 268) syllables, following Vogel et al. (Reference Vogel, van de Vijver, Kotz, Kutscher, Wagner, Vogel and Vijver2015).
22.2.3 Results
22.2.3.1 Proportions of Sentences
Table 22.1 shows the proportions of utterances produced according to the rhythmic ranking presented. The whole dataset shows a tendency according to the prediction (stage 1 is the most frequent, stage 4 the least frequent), with an irregularity in stages 2 and 3 (the latter proportion is larger).
| Ranking and word order | ||||
|---|---|---|---|---|
| Ranking index | ||||
| Dataset fluent sentences | 1 | 2 | 3 | 4 |
| OS and SO n = 1819 | 26.8 | 24.6 | 25.2 | 23.4 |
| subset OS n = 546 | 23.8 | 28 | 24 | 24.2 |
| subset SO n = 1273 | 28.7 | 23.1 | 25.5 | 22.7 |
| subset human n = 969 | 27 | 25,9 | 24,3 | 22,8 |
In addition, two subsets were formed, and these are also shown in Table 22.1 in terms of the rhythmic ranking. In the subset with SO structures, a similar picture appears as in the entire dataset, but here even more pronounced in its expression. In the subset with OS structures, on the other hand, no comprehensible pattern is discernible. It is noticeable here that level 2 is represented relatively frequently; this is taken up again below.
In summary, the descriptive analysis of the utterances reveals frequencies that tend to match the predictions of the ranking. Thus, sentences assigned to level 1 occurred most frequently, while those assigned to level 4 occurred least frequently. It is also notable that level 3 sentences were chosen more frequently than those assigned to level 2. However, this pattern only applies to the more frequent SO sentences; no such systematic pattern is apparent for the OS sentences. Moreover, variation in animacy seems to affect the results: The subset that does not vary with respect to animacy (the human subset) adheres to the predictions (also in relation to levels 2 and 3), albeit very slightly.
For statistical analysis, a generalized linear mixed-effects regression model (GLMER; Bates et al., Reference Bates, Mächler, Bolker and Walker2015) was computed in R statistical software (version 4.0.2; R Core Team, 2020). The selected model computed successive distances (frequencies) between levels of the rhythmic ranking (taking into account sentence structures OS and SO). Covariates included were, in addition to fluency of utterance, spatial arrangement (mirror) and animacy of the antecedent of the pronoun. Both item and participant were integrated as random effects. In addition, participant and run were coupled as embedded factors (Common Extensions | Mixed Models with R [m-clark.github.io]). Participant and run were thus considered embedded, as there were potentially four runs for each participant.
The covariates did not assume a significant magnitude, except for the spatial ordering. Regarding the ranking, the model for the whole set (OS and SO) shows a highly significant distance between levels 1 and 2 in the predicted direction. The negative sign (z = -1.7) between levels 2 and 3 illustrates the tendency against the prediction (there were more level 3 sentences than level 2 sentences), this difference taking on a marginally significant size (see Table 22.2).
| Model (GLMER) ranking | ||||
|---|---|---|---|---|
| Estimate | Std. Error | z value | Pr(>|z|) | |
| (Intercept) | −0.79269 | 0.14425 | −5.495 | <0.0001 *** |
| Ranking 2-1 | 0.49570 | 0.15097 | 3.283 | 0.00103 ** |
| Ranking 3-2 | −0.26226 | 0.15114 | −1.735 | 0.08271 † |
| Ranking 4-3 | 0.18580 | 0.15308 | 1.214 | 0.22484 |
| Mirror | 0.11064 | 0.05353 | 2.067 | 0.03875 * |
| Animacy | −0.06936 | 0.05361 | −1.294 | 0.19573 |
| Fluent 1 | −0.18068 | 0.14335 | −1.260 | 0.20754 |
Signif. codes: 0 “***” 0.001 “**” 0.01 “*” 0.05 “.” 0.1 “†”
22.2.3.2 Syllable Prominence
Based on the previous analysis, a conflation of syllable prominence and ranking will now be performed. It is predicted that the prominence level of the pronoun will vary systematically with the rhythmic ranking. In particular, the pronoun should be classified as stressed more often in the sense of accommodation in stage 4, which is classified as unrhythmic, than in the more rhythmic stages 1, 2, and 3.
To illustrate, the prominence levels defined at the beginning of this chapter are presented here. Syllable prominence was defined as a four-level quantity consisting of accented syllables, stressed syllables, stressable syllables, and unstressable syllables. Figure 22.4 shows a sentence in the context of this model. Here, the pronoun as a function word is initially considered unstressed. The predicted gain or reduction of prominence is marked by parentheses around the corresponding cross (lower panel) or its deletion (upper panel).
Varying prominence on the pronoun.
Upper panel: Prominence reduction on the pronoun. Lower panel: Addition of prominence on the pronoun.

Figure 22.4 Long description
The words are der Junge sagt, dass Markus ihn belügt. Each word has a series of "X" marks in a grid. The columns of "X" marks correspond to different accentuation categories: accented, stressed, unstressed, and unstressable.
For the analysis, all (semi-)fluent sentences were analyzed in terms of rhythmic ranking (1–4) and annotated syllable prominence (reduced, unstressed, stressed). Table 22.3 shows the merging of the categories in absolute numbers (n = 1819). The pronouns rated as unstressed are proportionally represented with similar frequency in the levels – with the exception of levels 2 and 3. Thus, the pronoun in level 3 was rated as unstressed with striking frequency, followed by levels 1, 4, and 2. For the other two degrees of prominence, a clear pattern emerges: According to the prediction, the proportion of stressed pronouns successively increases with decreasing rhythmicity. In the opposite direction, the proportion of reduced pronouns successively decreases with decreasing rhythmicity. Figure 22.5 visualizes the results.
| Ranking and prominence degree | ||||
|---|---|---|---|---|
| Ranking index | ||||
| Prominence degree | 1 | 2 | 3 | 4 |
| Reduced | 103 | 90 | 50 | 25 |
| Unstressed | 350 | 312 | 359 | 336 |
| Stressed | 35 | 44 | 51 | 64 |
Rhythmic ranking and degrees of prominence.

Figure 22.5 Long description
The category wise major to minor sectors representing the levels of rhythmic ranking are as follows. Reduced. 1, 2, 3, and 4. Stressed. 4, 3, 2, and 1. Unstressed. 3, 4, 1, and 2.
It should be particularly emphasized at this point that these systematics are also evident in levels 2 and 3, because the frequency distribution of these two levels had occurred in reverse order in the analysis of the previous chapter, contrary to prediction.
In addition, pronouns referring to a nonhuman referent (n = 850) were rated as stressed (n = 100) more often than those referring to a human referent (n = 969; n = 94 rated as stressed), with the latter being rated as reduced (n = 159) more often than the former (n = 109).
Finally, the relationship between annotated syllable prominence and ranking was statistically tested. For the analysis, a GLMER model (Bates et al., Reference Bates, Mächler, Bolker and Walker2015) was computed in R statistical software (version 4.0.2; R Core Team, 2020). The chosen model computes successive distances (frequencies) between levels of rhythmic ranking, here in the context of annotated syllable prominence. To do justice to the binomial character of the model, it refers to a rescaled prominence. This means that the previously three-level prominence in the model distinguishes only between those pronouns that were scored as reduced and all others. Consequently, in this model, the reduced syllables (coded as -1) were set apart from all others (coded as 0).
Included covariates were, in addition to fluency of utterance, spatial arrangement (mirror) and animacy of the antecedent of the pronoun. Both item and participant were integrated as random effects. In addition, participant and run were coupled as embedded factors, as in the previous section.
Table 22.4 summarizes the results of the model. It revealed a significant difference between ranking levels 3 and 4, and a highly significant difference between levels 2 and 3. Furthermore, animacy significantly affected the perceived syllable prominence of the pronoun.
| Model (GLMER) ranking | ||||
|---|---|---|---|---|
| Estimate | Std. Error | z value | Pr(>|z|) | |
| (Intercept) | 2.11668 | 0.18864 | 11.221 | <0.0001 *** |
| Ranking 2-1 | 0.03067 | 0.17347 | 0.177 | 0.859650 |
| Ranking 3-2 | 0.74253 | 0.20053 | 3.703 | 0.000213 *** |
| Ranking 4-3 | 0.65330 | 0.26054 | 2.508 | 0.012158 * |
| Mirror | −0.02387 | 0.07188 | −0.332 | 0.739850 |
| Animacy | −0.14641 | 0.07183 | −2.038 | 0.041521 * |
| Mirror:Animacy | 0.09343 | 0.07178 | 1.302 | 0.193072 |
Signif. codes: 0 “***” 0.001 “**” 0.01 “*” 0.05 “.” 0.1 “†”
22.3 Discussion
In the present chapter, a new metric for the evaluation of the rhythmic well-formedness of sentences was developed. This measure is based on the assumption that not only is the local rhythm relevant for the rhythmicity of a structure (stress lapse or stress clash at a defined position) but also the rhythm of the entire structure. Following the results on the relevance of trochaic and dactylic rhythm (Franz, Reference Franz2022), and the work of Hanna (Reference Hanna2003) and Vogel et al. (Reference Vogel, van de Vijver, Kotz, Kutscher, Wagner, Vogel and Vijver2015), the hypothesis was developed that the balanced spacing of accented syllables is the relevant measure. Under this assumption, the measure of Shih et al. (Reference Shih, Grafmiller, Futrell, Bresnan, Vogel and Vijver2015) of local rhythm was used and further extended. The result covers structures with up to three accented syllables and is shown schematically in (18). An ideal sentence rhythm lies at the value zero in (18b). Deviations are equally weighted in the positive and negative range.
(18a) Local rhythm: (n unstressed syllables between two stressed syllables -1)
(18b) Balance measure: (value left) - (value right)
The metric was developed during a post hoc analysis of a picture-based production study in Franz (Reference Franz2022). One of the questions of this study was whether a manipulation of the stress structure of the neighboring words (iambic or trochaic) can influence the preferred placement of the German object pronoun (SO/OS). The rhythmic quality of the target sentences of the respective study was then re-evaluated as predicted by the new metric, resulting in a rhythmic ranking consisting of four levels (descending in rhythmicity).
In order to strengthen the new metric empirically, the chapter presented two analyses. The first analysis evaluated how the structures produced by the participants were distributed on the rhythmic ranking. According to the prediction, level 1 sentences should occur more frequently than those of levels 2, 3, and 4. The results partially confirmed this prediction: Overall, the frequencies were distributed according to the ranking with an inverted order of levels 2 and 3. However, only the difference between levels 1 and 2 was significant. The evaluation showed that the ranking applies primarily to the SO sentences, not to the OS sentences. Consequently, the results indicate that the participants construct the sentences in such a way that the preferred SO structure can be realized as rhythmically as possible. Accordingly, the frequencies of SO sentences tended to follow the rhythmic ranking, but the frequencies and proportions of OS sentences were unsystematic (with respect to the ranking).
Finally, the numbers in the human subset were found to be consistent with the predictions. Although the influence of animacy did not reach significance in the statistical model, the descriptive results do suggest that it is a confounding factor here. Respectively, in the study in Franz (Reference Franz2022), rhythmic influences on pronoun placement were found to reach the significance level only when animacy did not vary (see also McDonald et al., Reference McDonald, Bock and Kelly1993; Franz et al., Reference Franz, Kentner and Domahs2021). Therefore, in a future application and further development of the metric, care should be taken to avoid animacy differences. This was also relevant for the analysis of syllable prominence.
In the second analysis, the rhythmic ranking was reviewed in terms of perceived prominence of the pronoun. The results strengthened the predictions of the measure in the sense that syllable prominence varied systematically with predicted rhythmicity. Thus, pronouns were perceived as stressed especially in sentences classified as unrhythmic, and as reduced in rhythmic sentences. Only the latter result could be statistically tested and reached the significance level. Furthermore, animacy affected perceived syllable prominence; that is, pronouns referring to human referents were perceived as less prominent than those referring to nonhuman referents. However, the results should be interpreted cautiously, since the second analysis was based on perceived prominence by one person only (for problems with this kind of evidence, see Bruggeman et al., Reference Bruggeman, Schade, Włodarczak and Wagner2022). Further, the statistical model could only be done with the reduced pronouns (the systematics were shown with the stressed pronouns in the descriptive analysis).
Nevertheless, the results concerning the proportions of the produced structures in combination with the results on syllable prominence allow for a first empirical support of the new metric. In the context of the aim of this study, the introduced metric made it possible to systematize the rhythmicity visible in the metrical grid of the studied sentences. In its existing form, the metric works for structures with three accented (or otherwise prominent) syllables and, as such, is also applicable to structures not studied in this chapter.
Besides a phonetic validation, a next step would be to apply the metric to further structures. In doing so, a promising continuation of the project would be to consider aspects of prosodic boundaries. Thus, it should be investigated whether and to what extent the occurrence of prosodic boundaries between the prominent syllables affects sentence rhythmicity. One of the many possibilities is to change the meter (the spacing of prominent syllables) within an utterance (Vogel et al., Reference Vogel, van de Vijver, Kotz, Kutscher, Wagner, Vogel and Vijver2015). In the work presented here, a new metric was developed as an extension to Shih et al. (Reference Shih, Grafmiller, Futrell, Bresnan, Vogel and Vijver2015), providing a further basis for the study of sentence rhythm.
Summary
The present chapter presents a metric for the formal computation of the rhythmic soundness of a sentence. The measure is developed in a post hoc analysis of a speech production experiment on the influence of rhythm on the placement and prominence of the German object pronoun ihn [English “him”].
Implications
Future studies in the field of sentence rhythm can use the metric to formally capture rhythmic coherence when it goes beyond a binary alternation of stressed and unstressed syllables. The developed formal distinctions of stress lapse and stress clash and their relationship within a sentence could be evolved to make them useful for longer structures.
Gains
The present elaboration demonstrates that rhythmic soundness is not necessarily equivalent to a binary alternation of stressed and unstressed syllables. This should be kept in mind when developing linguistic stimuli. Rhythmicity rather depends on the composition of stress lapses and clashes – which can be formally grasped using the metric.




