Highlights
-
• Multisensory strategies were applied to foreign language (FL) word learning.
-
• FL concepts were activated via pictures and three sound conditions.
-
• Sound presence improved performance on semantically mediated tasks over silence.
-
• Environmental sounds did not outperform neutral tones.
-
• Lexically mediated tasks were insensitive to sound conditions.
1. Introduction
If a picture is worth a thousand words, what is a sound worth? Environmental sounds, or everyday noises that we encounter in our surroundings, such as the rhythmic patter of rain against a windowpane, or the sharp ring of an alarm clock, are not merely registered in the brain as isolated auditory events or background noise: they convey meaning. Much like pictures, sounds function as powerful sensory cues that enrich the environment, evoking mental representations of associated events (e.g., rainfall; an alarm clock). The simultaneous engagement of perception and meaning plays a critical role in subsequent memory and learning, which in turn can play a role in actions (e.g., bringing an umbrella; getting up or hitting snooze). It is, therefore, worth considering how such sensory stimuli might be leveraged for practical applications. In this paper, we will explore how environmental sounds could support foreign language (FL) vocabulary learning.
In today’s globalized world, multilingualism is a highly valued skill, offering both social and professional benefits (Adsera & Pytlikova, Reference Adsera and Pytlikova2015; Chiswick & Miller, Reference Chiswick and Miller2002; Fieles-Ahmad & Huber, Reference Fieles-Ahmad and Huber2022; Stöhr, Reference Stöhr2015). Despite clear personal and societal benefits of FL skills, linguistic distance between languages and the cognitive, temporal and economic costs associated with learning continue to pose significant barriers to acquisition (Sprenger, Reference Sprenger2024). Not all adults had the social and economic opportunity to learn languages as children, when language is acquired intuitively through immersive, multisensory experiences. For adults, acquiring a new language often presents formidable challenges: mastering unfamiliar vocabulary, grammatical rules, phonemic repertoires or even new writing conventions. It is possible that conventional, memorization-heavy pedagogical models of language learning (e.g., two-column vocabulary lists and written materials) lack the sensory engagement which would be enriching and lead to experientially driven learning (for reviews on adult FL learning, see Plonsky, Reference Plonsky2011; Rasouli & Jafari, Reference Rasouli and Jafari2016).
Addressing these challenges calls for new, cognitively grounded approaches to adult language learning – approaches that align more closely with how humans naturally acquire meaning through perception and sensory experience. There is a substantial body of research indicating that FL learning is improved when relevant information is delivered through multiple sensory channels rather than relying on just one modality (Li & Deng, Reference Li and Deng2023; Rahmanu & Molnár, Reference Rahmanu and Molnár2024). This has been examined for FL vocabulary learning across various input channels including visual (Altarriba & Knickerbocker, Reference Altarriba and Knickerbocker2011; Bates & Son, Reference Bates and Son2020; Yu & Liu, Reference Yu and Liu2022), gestural (García-Gámez & Macizo, Reference García-Gámez and Macizo2019, Reference García-Gámez and Macizo2020, Reference García-Gámez and Macizo2022; Sweller et al., Reference Sweller, Shinooka-Phelan and Austin2020), olfactory (Honda et al., Reference Honda, Tanizawa, Masai, Nakamura, Choi and Fukushima2025; Xia et al., Reference Xia, Qin and Fan2024) and auditory (Bellegarda et al., Reference Bellegarda, García-Gámez and Macizo2026; Kaplan-Rakowski & Loranc-Paszylk, Reference Kaplan-Rakowski and Loranc-Paszylk2019) modalities. A multimodal, or multisensory, learning environment may be particularly effective because it mirrors how real-world experiences naturally integrate diverse sensory inputs (Ellis, Reference Ellis2019; Shams & Seitz, Reference Shams and Seitz2008).
Several theoretical frameworks help explain why multisensory learning can strengthen memory and support vocabulary acquisition. According to Dual Coding Theory (Paivio, Reference Paivio1969), encoding information in both verbal and nonverbal (i.e., sensory) formats creates parallel memory traces, increasing the likelihood of retention and retrieval of said information. Closely related, Multimedia Learning Theory (Mayer, Reference Mayer2012) posits that learning is enhanced when information is presented across multiple modalities (e.g., visual and auditory), as learners can process information through partially independent channels. Specifically, according to the principle of coherence (see Mayer, Reference Mayer2024, for a review), information that aligns with the task objective (FL learning, in our case) would facilitate knowledge acquisition (images and sounds that share the same meaning), whereas learning would be hindered by the presence of information irrelevant to the task goals (no irrelevant semantic contents were added in our study). Other redundant information, such as descriptive text accompanying learning materials (see Mayer & Johnson, Reference Mayer and Johnson2008), would hinder learning. In this study, we prioritize coherent information within the framework of Multimedia Learning Theory, by the inclusion of images and sounds that convey the same meaning, while avoiding redundant information, such as text or additional explanations during FL learning.
From another view, the Scaffolding Assumption (Eitel et al., Reference Eitel, Scheiter and Schüler2013) further suggests that being exposed to a sensory stimulus (such as a picture) before corresponding linguistic input can preactivate global conceptual structures that serve as mental scaffolds for later comprehension. While these original formulations emphasized visual imagery, their underlying principles extend to other sensory modalities, which can likewise foster meaningful connections. Complementary to these perspectives, embodied cognition (Barsalou, Reference Barsalou2008; Smith & Gasser, Reference Smith and Gasser2005) posits that cognition is grounded in sensorimotor experiences, such that understanding arises from interacting with the environment. From this standpoint, meaning is not merely conceived as abstract representation (i.e., the semantic category a concept belongs to); instead, semantic knowledge inherently involves sensory components (i.e., visual, auditory, olfactory, etc.) derived from individuals’ interactions with the world. Thus, understanding a concept such as “dog” involves integrating sensory-perceptual features such as its appearance, movements and barking. By engaging multiple senses during learning, learners can trigger these perceptual-conceptual systems by mentally simulating real-world contexts, producing richer and more durable memory representations. Together, these frameworks buttress the idea that multimodal approaches do more than present information in varied formats: they actively link linguistic forms to conceptual and experiential knowledge, enhancing comprehension and long-term learning.
To understand how multisensory language learning fits within existing theoretical frameworks for FL learning, we can turn to the Bilingual Interactive Developmental Model (BIA-d, Grainger et al., Reference Grainger, Midgley, Holcomb, Kail and Hickmann2010) which builds on its earlier iterations including the BIA (Grainger & Dijkstra, Reference Grainger, Dijkstra and Harris1992), the BIA+ (Dijkstra & van Heuven, Reference Dijkstra and Van Heuven2002), as well as the Revised Hierarchical Model (RHM, Kroll & Stewart, Reference Kroll and Stewart1994). According to this framework, beginner learners initially depend heavily on their first language (L1) to mediate the understanding of new words in the FL. With continued exposure, however, learners begin to form direct associations between FL lexical forms and their meanings, gradually reducing their reliance on L1 translations. These emerging direct links are reinforced through Hebbian learning, the neural process summarized by the phrase “neurons that fire together wire together” (Hebb, Reference Hebb2005). In other words, each coactivation (i.e., “fire together”) of an FL form and its associated conceptual representation strengthens the connection between the two (i.e., “wire together”). Within this language learning framework, multisensory inputs could facilitate learning by providing additional conceptual activation. Following Hebbian learning principles, each coactivation of an FL word and its conceptual referent strengthens their reciprocal mapping, with additional reinforcement when the concept is supported by sensory input (e.g., sounds or images). This process may accelerate the formation of direct FL-concept connections, reducing reliance on L1 mediation.
Among the cited sensory modalities, environmental sounds are a promising yet underexplored avenue for FL vocabulary learning. As Vanderveer (Reference Vanderveer1979) notes, environmental sounds are the opposite of arbitrary: they directly indicate the events or objects that produce them, automatically evoking mental representations of those referents. The strong, direct associations make environmental sounds a uniquely rich source of sensory information that could support FL vocabulary acquisition (Ballas & Howard, Reference Ballas and Howard1987). Although environmental sounds are nonlinguistic, they nonetheless carry semantic content, as demonstrated in multiple experiments. For instance, they facilitate the recognition of both spoken and written words related to the sound stimulus (Frey et al., Reference Frey, Aramaki and Besson2014; Orgs et al., Reference Orgs, Lange, Dombrowski and Heil2006, Reference Orgs, Lange, Dombrowski and Heil2007, Reference Orgs, Lange, Dombrowski and Heil2008; Van Petten & Rheinfelder, Reference Van Petten and Rheinfelder1995). Moreover, environmental sounds engage similar cognitive and neural mechanisms as words: event-related potential (ERP) studies show that meaningful sounds elicit N400 responses comparable to those elicited by words, whereas nonmeaningful sounds do not (Cummings et al., Reference Cummings, Ceponiene, Koyama, Saygin, Townsend and Dick2006, Reference Cummings, Ceponiene, Dick, Saygin and Townsend2008; Dick et al., Reference Dick, Saygin, Galati, Pitzalis, Bentrovato, D’Amico, Wilson, Bates and Pizzamiglio2007). Functional magnetic resonance imaging (fMRI) and lesion studies further indicate that the semantic processing of environmental sounds and words relies on largely overlapping neural networks (Beauchamp et al., Reference Beauchamp, Lee, Argall and Martin2004; Saygin et al., Reference Saygin, Dick, Wilson, Dronkers and Bates2003, Reference Saygin, Dick and Bates2005). Importantly, environmental sounds have also been shown to contribute to embodied cognition by linking perception and meaning: certain action-related sounds (e.g., claps or whistles) integrate at the neural level with action words, suggesting that hearing these sounds can activate corresponding motor representations (Grisoni et al., Reference Grisoni, Dreyer and Pulvermüller2016, Reference Grisoni, Mohr and Pulvermüller2019).
To date and to our knowledge, there exist only two studies that directly examine how sounds affect FL vocabulary acquisition. Kaplan-Rakowski and Loranc-Paszylk (Reference Kaplan-Rakowski and Loranc-Paszylk2019) researched how Polish university students learned low-frequency onomatopoeic English nouns (e.g., lash), which phonetically resemble the sounds they represent (e.g., the crack of a whip). Onomatopoeic sounds can be thought of as a subset of environmental sounds (e.g., rain would be an environmental sound, but patter would be an onomatopoeic sound because when said out loud, the pronunciation of the word resembles the meaning). In a classroom setting, participants learned the new words under four sound conditions: onomatopoeic sound only, spoken pronunciation of the new word only, both and no-audio. During learning, a Microsoft PowerPoint slide was displayed for 15 seconds and contained five components: the written FL word, its Polish translation, a definition in English, a corresponding image and the corresponding audio condition. The results indicated that presenting the onomatopoeic sound alone during the learning improved both the immediate and the delayed (seven-day) recall compared to the no-audio condition. In contrast, the spoken pronunciation condition and the condition in which both sounds were present (onomatopoeic and pronunciation) did not yield advantages over the no-audio condition. The authors proposed that the simultaneous presentation of both sound inputs may have exceeded the learners’ working memory capacity, thereby limiting potential benefits from the dual auditory input. They explained from the perspective of the Cognitive Load Theory (Sweller, Reference Sweller1994) that the dual input could have introduced competing information, which divided attention across signals rather than integrate it. Nonetheless, the findings highlight that sounds that are contextually relevant (such as onomatopoeic sounds) can act as effective learning cues that support encoding and retrieval of word meanings.
More recently, Bellegarda et al. (Reference Bellegarda, García-Gámez and Macizo2026) extended this line of inquiry by examining how environmental sounds affect FL word learning in Spanish-speaking university students. In their study, participants learned high-frequency Spanish-FL word pairs under three auditory conditions: environmental sounds (e.g., the sound of rain paired with the word rain), neutral tones or silence. Participants were exposed solely to the written word pairs and the sound condition, with no additional inputs. Subsequent testing revealed task-dependent effects. For tasks that relied heavily on semantic knowledge, both sound conditions (environmental sounds and neutral tones) tended to show higher accuracy than the silence condition. However, for tasks where form-based recognition was most important, the opposite pattern was observed. The absence of differences between the environmental sounds and neutral tones suggests that the observed benefit was driven by the general presence of auditory input that supported semantic access, rather than by deeper semantic enrichment. From an embodied cognition perspective, this pattern may indicate that the environmental sounds used in the study did not sufficiently strengthen the multisensory memory trace to active the perceptual-semantic representations fully.
The present study directly addresses this issue. To determine whether the advantage observed by Bellegarda et al. (Reference Bellegarda, García-Gámez and Macizo2026) was purely due to auditory input or whether richer sensory information could further support semantic learning, we will introduce a critical methodological change: the inclusion of a visual image as an additional, novel source of conceptual input. The present design combines pictures and sounds to provide dual-sensory conceptual activation. This richer, semantically aligned input will allow us to test whether engaging multiple sensory modalities effectively supports embodied semantic systems and strengthens vocabulary learning.
1.1. The current study
Taken together, this emerging evidence points to the promise yet complexity of integrating environmental sounds into multimodal FL learning. For the present study, we designed an experiment to examine whether conceptual information could be activated solely through sensory input, while minimizing reliance on L1 mediation. Specifically, we asked whether visual and auditory cues (in the form of pictures and environmental sounds) presented before an FL word can jointly enhance conceptual representations and thereby facilitate FL word learning. Bellegarda et al. (Reference Bellegarda, García-Gámez and Macizo2026) employed pairs of L1 and FL words in FL vocabulary learning, accompanied by various auditory stimuli. In this study, we replaced the L1 words with images representing the meanings of the FL words. This change was made for two reasons. Firstly, to promote semantic processing, and secondly, to reduce the influence of L1 during FL word learning. As an anonymous reviewer of an earlier version of this article noted, processing a picture can trigger associated lexical information (e.g., phonology and orthography; see the Independent Network Model, Caramazza, Reference Caramazza1997). However, if we compare learning using L1–FL word pairs (Bellegarda et al., Reference Bellegarda, García-Gámez and Macizo2026) with learning using pictures–FL words (the current study), semantic processing would be greater and L1 lexical processing would be minimized in this study. Support for these predictions comes from classical studies. For example, Glaser and Glaser (Reference Glaser and Glaser1989) concluded through six experiments that words have privileged access to the lexicon, whereas pictures have privileged access to the semantic network.
Prior work has shown that pictures can serve as scaffolds that preactivate relevant schemas and improve subsequent comprehension (García-Gámez & Macizo, Reference García-Gámez and Macizo2020; Yu & Liu, Reference Yu and Liu2022), and that environmental sounds can reinforce links between referents and labels in some cases (Bellegarda et al., Reference Bellegarda, García-Gámez and Macizo2026; Kaplan-Rakowski & Loranc-Paszylk, Reference Kaplan-Rakowski and Loranc-Paszylk2019). Yet it remains unclear (a) whether FL-concept associations can be formed effectively when learning relies primarily on nonverbal (sensory) input, and (b) to what extent the combination of relevant environmental sounds and visual cues leads to stronger FL-concept associations.
To address these questions, we adopted a novel, dual-sensory conceptual activation paradigm tailored to FL learning and designed to amplify the multisensory memory trace. We selected high-frequency concrete nouns typical of early FL instruction and paired each item with a written FLFootnote 1 target, a picture and a sound. During learning, each concept was activated via a picture and one of three sound conditions: a relevant environmental sound “congruent” with the picture, a neutral sound or silence. In the congruent condition, the environmental sound (e.g., a bee buzzing) was semantically related to the picture (e.g., a bee), providing coherent dual-sensory activation. The neutral condition used a single-frequency tone that was acoustically salient but semantically irrelevant; this condition controlled for the general contribution of auditory presence and tested whether nonsemantic auditory input could support learning by promoting increased learner engagement during encoding (self-involvement theory, Helstrup, Reference Helstrup1987). Our third condition, the silence condition, served as a baseline for FL learning with only visual activation. We intentionally did not include incongruent pairings because our aim was to model naturalistic learning contexts (where mismatches are uncommon) and to test facilitation as our primary goal, rather than interference.
Participants performed four evaluation tasks designed to assess different aspects of vocabulary learning: forward and backward translation tasks, a picture naming task, and a lexical decision task. While all four tasks were completed after the entire learning phase, the translation tasks were additionally completed halfway through the learning. This allowed us to track how learning progressed over time (see García-Gámez & Macizo, Reference García-Gámez and Macizo2023), and specifically, to examine how lexical and semantic representations evolved under different sound conditions. We selected the four tasks because they are widely used in FL learning research and provide evidence for formation of semantic and lexical links between new FL words, L1 words and conceptual knowledge (Grainger et al., Reference Grainger, Midgley, Holcomb, Kail and Hickmann2010; Lindsay & Gaskell, Reference Lindsay and Gaskell2013; Mestres-Missé et al., Reference Mestres-Missé, Rodriguez-Fornells and Münte2007).
Additionally, using these four tasks enables us to eliminate the possibility that differences in participants’ performance are due to the specific demands of each task rather than the fact that two of them are semantically mediated (forward translation and picture naming) and two are lexically mediated (backward translation and lexical decision).
The forward translation and picture naming tasks required participants to engage in explicit semantic processing and are thus considered conceptually mediated (Kroll & Stewart, Reference Kroll and Stewart1994; Sholl et al., Reference Sholl, Sankaranarayanan and Kroll1995). In the forward translation task, participants saw a Spanish item and produced its FL equivalent, while in the picture naming task they named pictures representing the learned concepts in the FL. Successful performance on both tasks relied on access to the semantic network: the forward translation task via activation of the conceptual representation associated with the L1 word and the picture naming task through conceptual activation from the image (de Groot & van Hell, Reference de Groot and van Hell2005; Vigliocco et al., Reference Vigliocco, Vinson, Damian and Levelt2002). We predicted that the enriched congruent condition (learning with a semantically relevant visual and auditory cues) would enhance performance on these tasks, as the meaningful inputs should strengthen conceptual representations, rendering them more vivid and accessible.
In contrast, the backward translation and lexical decision tasks do not rely on explicit semantic access, but can still reveal the organization and consolidation of the FL lexicon (Goldinger, Reference Goldinger1996; Jusezyk & Luce, Reference Jusezyk and Luce2002; Strauß et al., Reference Strauß, Wu, McQueen, Scharenborg and Hintz2022; Vitevitch, Reference Vitevitch2003). In the backward translation task, participants saw a word in the FL and provided its Spanish equivalent, while in the lexical decision task they decided whether a visually presented FL word had been learned or was novel. Prior findings indicate that learning methods emphasizing semantic mediation not only support the development of conceptual links but also facilitate lexical consolidation (García-Gámez & Macizo, Reference García-Gámez and Macizo2019; García-Gámez et al., Reference García-Gámez, Cervilla, Casado and Macizo2024). Therefore, it might be possible that the congruent sound condition would also promote more robust lexical organization, leading to better performance on these lexically mediated tasks.
With regard to the other conditions, we anticipated a better performance for words learned with a neutral tone relative to silence. In both these conditions, conceptual information was provided largely by the picture. However, although the tone lacked semantic meaning, its presence may still enrich memory representations by enhancing engagement and learner self-involvement during encoding (Helstrup, Reference Helstrup1987).
We are aware that there is evidence that previous language experience (LEX) (particularly multilingualism) and socioeconomic status (SES) are both factors that could play a role in language learning. LEX has been linked to enhanced cognitive flexibility and metalinguistic awareness with regard to language tasks (Bonnet & Siemund, Reference Bonnet and Siemund2018; Cenoz, Reference Cenoz2013; Hirosh & Degani, Reference Hirosh and Degani2018; Kaiser et al., Reference Kaiser, Eppenberger, Smieskova, Borgwardt, Kuenzli, Radue, Nitsch and Bendfeldt2015; Kaushanskaya & Marian, Reference Kaushanskaya and Marian2009). Similarly, SES, a multidimensional construct that includes parental education, income and access to resources, has been shown to influence both cognitive and linguistic development. Although its effects are more pronounced in childhood (Abo Hamza et al., Reference Abo Hamza, Tindle, Pawlak, Bedewy and Moustafa2024), SES-related disparities can persist into adulthood (Dong, Reference Dong2024). Therefore, we propose measuring LEX and SES with the Language Experience and Proficiency Questionnaire (LEAP-Q; Marian et al., Reference Marian, Blumenfeld and Kaushanskaya2007) and the MacArthur Scale of Subjective Social Status (Adler et al., Reference Adler, Epel, Castellazzo and Ickovics2000), respectively, to account for their potential impact on FL acquisition. As part of an exploratory analysis, these variables will be included to examine how they may modulate learning outcomes, although no specific predictions were made.
2. Methods
2.1. Participants
A total of 40 native Spanish speakers (25 women; M age = 21.25, SD age = 3.32) from the University of Granada took part in the study in return for course credit. None of the participants reported visual, auditory, neurological or language-related impairments. Written informed consent was obtained before participation. The study procedures followed national and institutional guidelines, and were consistent with the principles on human experimentation outlined in the Helsinki Declaration (1964; latest revision 2024). The experimental protocol was reviewed and approved by the University of Granada Ethics Committee (approval number: 957/CEIH/2019).
The sample size required for this study was estimated using G*Power (Faul et al., Reference Faul, Erdfelder, Buchner and Lang2009) based on a repeated-measures ANOVA with one within-subject factor (sound condition, 3 levels). This approach was used as a pragmatic and conservative approximation, as simulation-based power estimation for mixed-effects models requires a priori specification of variance components and random-effects structures, which are often unavailable before data collection (Kumle et al., Reference Kumle, Võ and Draschkow2021). A total of N = 36 participants was needed to achieve 95% statistical power with α = .05 and a medium effect size (Cohen’s f = 0.275, η 2p = .07).
Before beginning the experiment, participants completed two questionnaires to provide demographic information. SES was assessed with the MacArthur’s Scale of Subjective Social Status (Adler et al., Reference Adler, Epel, Castellazzo and Ickovics2000), in which individuals placed themselves on a ladder with 10 rungs, representing their relative standing with regards to income, education and occupation. The participants’ mean subjective rating for SES was M = 6.15, SD = 1.33, range 3–8.
Language background and experience (LEX) were evaluated using a modified version of the LEAP-Q (Marian et al., Reference Marian, Blumenfeld and Kaushanskaya2007). Participants reported the languages they knew, ordered by acquisition, fluency and frequency of exposure. They also indicated what percentage of time they would typically read or speak in each language and provided self-ratings of reading, writing, listening and speaking abilities (1–10 scale) for all languages in which they had experience. In addition, participants indicated the cultural groups they self-identified with. To quantify LEX, we calculated average second language (L2) proficiency across the four skill domains (reading, writing, speaking, listening), since all participants reported knowledge of at least two languages. This was relevant given evidence that prior LEX can influence subsequent language learning (Hirosh & Degani, Reference Hirosh and Degani2018). On average, participants reported speaking 2.68 languages (SD = 0.92, range 2–6). Participants self-reported their L2 proficiency skills as M = 7.23, SD = 1.33, range 3.5–10. Summary statistics for these language measures are presented in Table 1.
Language use and proficiency by order of acquisition as well as cultural identities self-reported by participants

Table 1. Long description
Spanish (National): 94.87% identifying; strength of identification 6.97 (SD 2.03). Spanish (Regional or Local): 84.62% identifying; strength 7.86 (SD 2.32). Includes Andalusian, Asturian, Canary Islander, Castilian, Catalonian, Extremaduran, and Granadan. Anglophone: 25.64% identifying; strength 4.20 (SD 2.94). Includes English. Mediterranean: 15.38% identifying; strength 8.17 (SD 1.94). Non-Iberian European: 10.26% identifying; strength 3.25 (SD 1.50). Includes French, German, and Italian. Latin American: 7.69% identifying; strength 6.83 (SD 3.55). Includes Chilean, Cuban, and Mexican. Arab/Moroccan: 7.69% identifying; strength 5.00 (SD 4.36).
Section 3: Cultural Identity. Reports the percentage of participants identifying with each cultural group and the mean strength of identification (scale 1–10). Participants could indicate multiple identities.
L1: reading 9.75 (SD 0.49); writing 9.68 (SD 0.53); speaking 9.65 (SD 0.66); listening 9.85 (SD 0.36). L2: reading 7.60 (SD 1.24); writing 6.83 (SD 1.48); speaking 6.98 (SD 1.75); listening 7.50 (SD 1.77). L3: reading 5.28 (SD 2.44); writing 3.83 (SD 1.95); speaking 4.17 (SD 2.26); listening 4.94 (SD 2.26). L4: reading 3.00 (SD 3.03); writing 2.33 (SD 1.97); speaking 2.67 (SD 1.75); listening 4.00 (SD 3.58). L5: reading 1.00 (SD 0.00); writing 1.00 (SD 0.00); speaking 1.00 (SD 0.00); listening 1.00 (SD 0.00). L6: reading 1.00 (SD 0.00); writing 1.00 (SD 0.00); speaking 1.00 (SD 0.00); listening 1.00 (SD 0.00).
Section 2: Language Proficiency. Reports mean self-rated proficiency (scale 1–10) across four skills (reading, writing, speaking, listening) for each language (L1–L6).
L6: Daily exposure 5.00 (SD 0.00); reading 0.00 (SD 0.00); speaking 0.00 (SD 0.00).
L5: Daily exposure 0.00 (SD 0.00); reading 0.00 (SD 0.00); speaking 0.00 (SD 0.00).
L4: Daily exposure 2.43 (SD 1.99); reading 0.71 (SD 1.50); speaking 0.86 (SD 1.57).
L3: Daily exposure 3.44 (SD 3.50); reading 2.78 (SD 6.28); speaking 4.00 (SD 8.67).
L2: Daily exposure 17.20 (SD 9.87); reading 19.65 (SD 14.43); speaking 20.73 (SD 26.44).
L1: Daily exposure 80.95 (SD 11.07); reading 79.23 (SD 15.45); speaking 77.43 (SD 27.57).
Section 1: Language Use. Reports mean percentage of daily exposure in each language (L1–L6) and preferences for reading and speaking (i.e., percentage of time participants reported choosing to read or speak in each language). Languages are listed in order of acquisition. In all, participants reported experience with 2–6 languages: the sample was composed of 22 bilinguals, 11 trilinguals, and 7 multilinguals with 4 or more languages.
Table 1 is divided into three sections: Language Use, Language Proficiency, and Cultural Identity. Values are means with standard deviations (SD) in parentheses.
Note: Language use measures reported mean daily exposure (i.e., percentage of time participants reported exposure to each language) and preferences for reading and speaking (i.e., percentage of time participants reported choosing to read or speak in each language). Participants reported experience with 2–6 languages: there were 22 bilinguals, 11 trilinguals and 7 multilinguals with experience with 4 or more languages. Language proficiency reports mean proficiency across 4 dimensions (reading, speaking, writing, listening), on a scale of 1–10. Cultural identity measures the percentage of participants identifying with each cultural group, along with the mean strength of identification (scale of 1–10). Participants could indicate multiple cultural identities. Standard deviations are shown in parentheses.
a Includes Andalusian, Asturian, Canary Islander, Castilian, Catalonian, Extremaduran, and Granadan.
b Includes English.
c Includes French, German, and Italian.
d Includes Chilean, Cuban, and Mexican.
2.2. Design and materials
The data, analyses, study materials and supplementary files supporting the findings of this research are openly available in the Open Science Framework (OSF) repository at: https://doi.org/10.17605/OSF.IO/6QAZ2. In this study, participants learned 42 FL words. Each word was presented visually within a sequence that began with a picture representing the word, which was accompanied either by silence, a congruent sound or a neutral tone. Using a within-subject design, 14 words were assigned to each condition, totaling 42 new words. Previous research suggests that this number of items can be learned comfortably in one session with a comparable number of repetitions (Bellegarda et al., Reference Bellegarda, García-Gámez and Macizo2026; García-Gámez & Macizo, Reference García-Gámez and Macizo2019, Reference García-Gámez and Macizo2023; García-Gámez et al., Reference García-Gámez, Cervilla, Casado and Macizo2024). The assignment of words to the different sound conditions was counterbalanced across participants.
2.3. Materials
2.3.1. L1 and FL words
A total of 42 Spanish nouns were chosen, with half representing natural entities (e.g., “gato,” or “cat” in English) and half representing man-made entities (e.g., “alarma,” or “alarm” in English). All nouns were familiar words (M = 5.83, SD = 0.51, scale 1–7 with 7 indicating maximum familiarity) and had an average lexical frequency of 22.12 (SD = 51.46) occurrences per million words (SUBTLEX-ES corpus, Cuetos et al., Reference Cuetos, González-Nosti, Barbón and Brysbaert2011). Eighty-four FL words were created conforming to Spanish phonological and orthographical patterns. The FL items had few Spanish orthographic neighbors (M = 0.13, SD = 0.46). Half of the FL words were used in the learning task, while the remaining half of the FL word served as pseudowords (“no” responses) in the lexical decision task, as they were never presented to the learners. In the learning task, an FL word was randomly paired with one of the Spanish nouns (although creating pairings that began with the same initial phoneme was avoided). The 42-word pairs were grouped into 3 sets of 14 words, which were matched on multiple lexical-semantic variables for the Spanish words (lexical frequency, number of syllables, orthographic length, number of orthographic neighbors, number of shared graphemes with the FL pair, familiarity, imageability and concreteness) and lexical variables for the FL words (orthographic length, number of orthographicFootnote 2 neighbors and number of syllables). Data on number of orthographic neighbors, familiarity, imageability and concreteness were consulted in EsPal database (Duchon et al., Reference Duchon, Perea, Sebastián-Gallés, Martí and Carreiras2013). For the lexical decision task, “learned” (FL real words) and “new” items (nonpracticed FL words that functioned as pseudowords) were matched in terms of number of syllables, number of orthographic neighbors and orthographic length. The full set of linguistic materials is available in the Supplementary Materials, Table S1, with descriptive statistics for the learning and lexical decision tasks provided in Supplementary Tables S2 and S3, respectively.
2.3.2. Sounds
Forty-two environmental sounds corresponding to the 42 selected Spanish nouns were adapted from Bellegarda et al. (Reference Bellegarda, García-Gámez and Macizo2026). These sounds belonged to various semantic categories: alarms/alerts (n = 4, e.g., alarm clock ringing), household sounds (n = 6, e.g., someone pulling a curtain), sports (n = 1, e.g., someone bowling), tools (n = 7, e.g., someone hammering), transport (n = 3, e.g., a car starting), animal vocalizations (n = 11, e.g., horse whinnying), environmental phenomena (n = 8, e.g., wind whistling), and insect sounds (n = 2, e.g., bee buzzing). The neutral sound condition consisted of 14 monochromatic unique sinusoidal tones ranging from 525 to 1000 Hz, created using Praat software (Boersma & Weenink, Reference Boersma and Weenink2020). Both the environmental sounds and the tones lasted 4 seconds, were digitized at 48 kHz with a 16-bit sampling rate, and were normalized to an RMS level of 70 dB SPL (Dittinger et al., Reference Dittinger, Barbaroux, D’Imperio, Jäncke, Elmer and Besson2016).
2.3.3. Pictures
Forty-two colored pictures representing the meaning of the Spanish-FL word pairs were selected from the MultiPic database (Duñabeitia et al., Reference Duñabeitia, Crepaldi, Meyer, New, Pliatsikas, Smolka and Brysbaert2018). For items not available in that database, additional stimuli were sourced from the PicPsy database (Martínez et al., Reference Martínez, Matute and Goikoetxea2020) and other databases with public domain licenses. These pictures were used in a picture familiarization task, the learning task and the picture naming task.
To verify that the selected sounds and pictures accurately represented the meaning of the Spanish nouns and were easily recognizable, we conducted a norming study with 23 participants who were not involved in the main experiment. Participants completed three rating tasks. In the first task, they heard an auditory stimulus while a Spanish word appeared on the screen and rated how well the word’s meaning matched the sound on a 7-point Likert scale (1 = “high mismatch,” 7 = “high match”). In the second and third tasks, the procedure was the same, but participants judged either picture-word or picture-sound pairings. To prevent fatigue and avoid excessive repetition (since the 42 original stimuli would otherwise only appear in congruent combinations), each stimulus was shown twice per task: once in a congruent pairing and once in an incongruent one. In incongruent trials, the mismatched item always came from the same semantic category (e.g., an oinking sound was paired with “marrano” [hog in English] in the congruent condition, and with “caballo” [horse in English] in the incongruent condition). In total, participants rated 252 stimuli pairs (84 per task) in a session lasting roughly 20 minutes. Results confirmed that participants consistently distinguished congruency across all modalities. Sound-word pairs received higher ratings when congruent (M = 6.46, SD = 1.30) than when incongruent (M = 1.62, SD = 1.45), t(22) = 34.55, p < .001. Picture-word pairs were also rated higher in the congruent (M = 6.84, SD = 0.82) than the incongruent (M = 1.34, SD = 1.10) condition, t(22) = 45.89, p < .001. Finally, picture-sound pairs showed the same pattern (congruent: M = 6.63, SD = 1.23; incongruent: M = 1.38, SD = 1.23), t(22) = 45.12, p < .001.
2.4. Familiarization task
Participants first completed an initial familiarization phase, which was included to ensure that the pictures were consistently associated with the intended conceptual meanings of the words to be learned. Participants viewed 42 pictures along with their corresponding Spanish labels (as described in Materials). Each label corresponded to the noun most commonly associated with the depicted concept.Footnote 3 The pictures were presented one by one in random order, and participants were instructed to take their time to study the images and words, advancing at their own pace by pressing the spacebar.
2.5. Learning task
The three sets of Spanish-FL word pairs described in the Materials section were each assigned to one of three sound conditions (silence, congruent sound or neutral tone). The assignment of sets to conditions was counterbalanced across the participants. Learning took place in a blocked manner, with one block per sound condition, see Figure 1. Blocked presentation, rather than interleaving sound trials, was used to avoid introducing additional attentional or carry-over effects between trials. One round of learning included all 3 blocks, providing exposure to all 42 FL words and all 3 sound conditions. Participants completed 10 rounds of learning in total, resulting in 10 exposures to each item (420 trials in all). They were given a short break after every two rounds (approximately every 14 minutes). Each trial began with a central fixation point (500 ms), followed by the display of a picture accompanied by the sound condition (silence, congruent sound or neutral tone) (4000 ms). The corresponding written FL word then appeared centrally (4000 ms), followed by an intertrial interval (500 ms). Participants were told that each picture represented the meaning of the upcoming FL word and were instructed to focus on learning the word meanings. Within each sound block, the order of the FL word-picture pairs was randomized.
Description of one round of foreign language (FL) word learning. Participants learned 42 items in a blocked design, with 14 items in each block, and each block associated to a sound condition (congruent environmental sound, neutral tone or silence).

Figure 1. Long description
A vertical flowchart represents the trial events during one round of foreign language word learning. Example stimuli from three sequential learning blocks are shown. Block 1 (congruent sound): A picture of a bird is displayed, and a speaker icon indicates that a congruent environmental sound (bird chirping) plays during the presentation of the stimulus. Then the foreign language word “argu” is displayed. The text “… 13 more FL words” indicates this trial structure repeats for the remaining 13 items in the block. Block 2 (neutral tone): A picture of a curtain is displayed, and a speaker icon indicates that a neutral tone plays during the presentation of the stimulus. Then the foreign language word “sorlato” is displayed. The text “… 13 more FL words” follows. Block 3 (silence): A picture of a rain cloud is shown, and a speaker icon with a cross indicates silence during this block. Then the foreign language word “valdri” is displayed. The text “… 13 more FL words” follows.
2.6. Evaluation tasks
2.6.1. Forward translation task
Participants completed an oral forward translation task where the written target was presented in Spanish and an oral response was given in FL. This task was administered at two timepoints: after 5 exposures to the items in the learning task and after 10 exposures. At the first timepoint, participants translated half of the words (21 items) and at the second timepoint, participants translated the remaining items (21 words).
2.6.2. Backward translation task
Participants also completed an oral backward translation task where the written target was presented in FL and an oral response was given in Spanish. Similarly, this task was administered at two timepoints during the experiment: after 5 exposures to the items in the learning task and again after 10 exposures. At the first timepoint, participants translated half of the words (21 items) and at the second timepoint, participants translated the remaining items (21 words).
The order of the performance of the translation task (forward or backward) at each of the two timepoints was counterbalanced across the participants. The assignment of items to Timepoint 1 or Timepoint 2 was counterbalanced across participants, so that each word appeared equally often at each timepoint across the sample. Across both timepoints, all 84 items (42 Spanish and 42 FL) were translated once, ensuring that no item was repeated across tasks. Within each translation task, items were presented in a random order.
Both translation tasks had similar trial events. Each trial began with a centrally presented fixation point (500 ms), followed by the written display of a Spanish or FL stimulus (500 ms). A question mark (?) then appeared and remained on the screen until the participants provided their oral response, which triggered the voice-activated key. The trial then concluded with a blank screen (450 ms). Participants were instructed to respond loudly and clearly, focusing on both speed and accuracy and to avoid stuttering.
2.6.3. Picture naming task
Participants saw the 42 pictures from the learning task and were instructed to name each aloud using the FL vocabulary they had learned previously, focusing on speed and accuracy. Each trial began with a centrally displayed fixation point (300 ms), followed by an image (e.g., the picture of a door, 10 × 10 cm) appearing centrally (500 ms). A question mark (?) signaled that the participant could provide his or her spoken response. This prompt remained on screen until the response was made and the voice-key triggered, after which a blank screen was shown (450 ms). The pictures were presented in random order.
2.6.4. Lexical decision task
Participants were asked to indicate, using the M and Z keys, whether a presented item was a learned, real FL word or an FL pseudoword. The assignment of the keys to “yes” or “no” responses was counterbalanced across participants. Each trial began with a centrally presented fixation point (300 ms), followed by the target word presented visually, which remained on screen until a response was made. After the response, a blank screen was displayed (450 ms) before the next trial. The task included a total of 84 trials: 42 trials with previously learned words, and 42 new, never-seen words. All stimuli were presented in random order.
2.7. Procedure
Before coming to the laboratory, participants completed online questionnaires (LEAP-Q and SES) via a link provided on the recruitment platform. In the laboratory, they performed the learning and evaluation tasks in a single session lasting approximately 110 minutes, though the duration varied depending on the individual performance. The session took place in a quiet room to minimize distractions, and participants wore a Trust Mauro USB headset to listen to the binaural stimuli. Initially, participants completed 48 trials of a visual Stroop task (Stroop, Reference Stroop1935), which lasted about 5 minutes; the materials were from Bellegarda and Macizo (Reference Bellegarda and Macizo2021). Data from this task were collected for a different purpose and are not analyzed in this paper. Next, they completed the familiarization task (2 minutes), followed by the first half of the learning task (35 minutes), the translation tasks (5 minutes), then the second half of the learning task (35 minutes) and again the translation tasks (5 minutes). After a short break, participants performed the picture naming task (10 minutes) and the lexical decision task (10 minutes). The task order was fixed, with picture naming preceding the lexical decision task. The lexical decision task was intentionally placed last to avoid giving only a subset of participants extra exposure to the target FL words (counterbalancing would have resulted in unequal input across groups). All stimuli were presented using E-Prime 2.0 (Schneider et al., Reference Schneider, Eschman and Zuccolotto2012) on a 19-inch Captiva E1903 monitor. Phonation initiation during the translation and picture naming tasks was recorded using an ATR20 Audio-Technica Cardioid Low Impedance microphone connected to a PST serial response box (Psychology Software Tools Inc.), with responses also captured on a Sony ICD PX-440 voice recorder.
2.8. Data analysis
The primary index of FL word learning was task accuracy (translations, picture naming and lexical decision) and discrimination ability (d′ scores), as successful performance presupposes FL learning. However, reaction times (RTs) were also examined.
Statistical analyses were conducted using linear mixed-effects models for continuous data and general linear mixed-effects models for binomial data using the lme4 package (Bates, Maechler, et al., Reference Bates, Maechler, Bolker and Walker2015) in R version 4.2.1 (R Core Team, 2022). Statistical models included random intercepts for participants and items. Random slopes were initially considered but could not be retained in the final models due to convergence issues; therefore, parsimonious random structures were adopted in line with recommendations to capture data variability without overfitting (Bates, Kliegl, et al., Reference Bates, Kliegl, Vasishth and Baayen2015). For the RT analyses, only correct responses were considered. Univariate outliers exceeding ±3 SD from each participant’s mean were removed, and RTs were log-transformed (natural log). The significance of fixed-effects and interactions was evaluated using likelihood ratio tests, and estimated marginal means were obtained via emmeans package (Lenth, Reference Lenth2023). For binomial mixed-effects models, pairwise comparisons of estimated marginal means were exponentiated to obtain odds ratios (ORs), which quantify the differences in the odds of a correct response between conditions. ORs > 1 indicate higher odds in the first condition of the contrast relative to the second, whereas ORs < 1 indicate lower odds. Ninety-five percent confidence intervals (CIs) are also reported on the OR scale. Tukey’s correction (Tukey, Reference Tukey1953) was applied to adjust for multiple comparisons in post hoc analyses. Full model structures are provided in Table 2.
Summary of mixed-effects models by task and dependent variable (DV)

Table 2. Long description
Models for Forward Translation and Backward Translation included Sound and Timepoint as interacting fixed effects, as both were measured at two timepoints. Models for Picture Naming and Lexical Decision included Sound only. Accuracy was modelled using generalized linear mixed-effects models (glmer); response time (logRT) used linear mixed-effects models (lmer). All tasks included a model with ACC and a model with logRT. Lexical Decision additionally included a model for d′. All models included random intercepts for participants and items.
Table 2 has three columns: Task, Dependent Variable (DV), and Model Formula. Tasks are divided into semantically mediated (Forward Translation, Picture Naming) and lexically mediated (Backward Translation, Lexical Decision).
Note: ACC = accuracy; RT = response time (ms); logRT = natural logarithm of response time (ms); glmer = generalized linear mixed-effects model; lmer = linear mixed-effects model. Fixed-effect predictors included Timepoint (levels: 1, 2) and/or Sound (levels: congruent, silence, tone), depending on task. All models included random intercepts for participants and items.
For the binary coding of verbal response data in the translation and picture naming tasks, responses were scored as 1 (correct) if they met the criteria for correctness and given a 0 (incorrect) otherwise. Minor errors were tolerated depending on the target word’s length: (a) for monosyllabic words, a vowel substitution; (b) for disyllabic words, a vowel or consonant substitution, but not both; (c) for words with three or more syllables, a vowel-consonant inversion or the replacement of a vowel or consonant. These criteria follow those used in prior research employing verbal response data in FL word learning (García-Gámez & Macizo, Reference García-Gámez and Macizo2019; García-Gámez et al., Reference García-Gámez, Cervilla, Casado and Macizo2024). Semantically equivalent responses (e.g., responding “cerdo” instead of “marrano,” comparable to “pig” instead of “hog”) were also treated as correct (score = 1). For RT analyses, data points were excluded if: (a) the voice key was triggered by nonverbal sounds or (b) the participant stuttered or hesitated during the response.
2.8.1. Questionnaires
In an exploratory analysis, we extended our primary model to include two additional fixed-effect predictors (SES and LEX), given previous findings suggesting their potential influence on learning outcomes. Both predictors were scaled and centered before analysis. For each task and each dependent variable (accuracy, RTs, d′), two additional models were tested, one including SES and one including LEX. To reduce the likelihood of false positives due to multiple comparisons, the Benjamini-Hochberg procedure (Benjamini & Hochberg, Reference Benjamini and Hochberg1995) was applied to control the false discovery rate (FDR) across all exploratory analyses.
2.9. Semantically mediated tasks
2.9.1. Forward translation task
A generalized linear mixed-effects model was fitted to the binomial accuracy data (correct and incorrect responses), and a linear mixed-effects model was fitted to the RT data. The models included Timepoint (1, 2) and Sound (congruent, silence, neutral) as fixed-effects predictors.
2.9.2. Picture naming task
A generalized linear mixed-effects model was used to assess accuracy (correct and incorrect responses), and a linear mixed-effects model was used for the RTs. In both models, Sound (congruent, silence, neutral) was entered as the fixed-effect predictor.
2.10. Lexically mediated tasks
2.10.1. Backward translation task
A generalized linear mixed-effects model was fitted to the binomial accuracy data (correct and incorrect responses) and a linear mixed-effects model was fitted to the RT data. The models included Timepoint (1, 2) and Sound (congruent, silence, neutral) as fixed-effects predictors.
2.10.2. Lexical decision task
Binomial accuracy data (correct and incorrect responses) were analyzed with a generalized linear mixed-effects model. To account for the two types of errors, misses (or, responding “new” to an old item) and false alarms (or, responding “old” to a new item), we calculated d′ values by z-transforming hit and false alarm rates for each participant (MacMillan & Creelman, Reference MacMillan and Creelman2005). These d′ values were analyzed with a linear mixed-effects model. An additional linear mixed-effects model was fitted to the RT data. Sound (congruent, silence, neutral) served as the fixed-effect predictor in all models.
3. Results
Participants whose accuracy on any evaluation task was less than 10% after completing all 10 exposures to the FL items were excluded from the analyses for that specific task, in line with the criteria established by García-Gámez and Macizo (Reference García-Gámez and Macizo2020). Accordingly, one participant was removed from the translation task analyses at the second timepoint, and two participants were removed from the picture naming task analysis. The resultsFootnote 4 presented below exclude these participants in the indicated tasks. Task means are provided in Table 3.
Mean accuracy (%) and response times (RT in ms and logRT) by evaluation task and sound condition

Table 3. Long description
Lexical Decision. Congruent: accuracy 94.82% (SD 22.18) [CI 92.98, 96.66]; RT 1029 ms (SD 643) [CI 973, 1084]; logRT 6.82 (SD 0.43) [CI 6.78, 6.86]. Silence: accuracy 93.04% (SD 25.48) [CI 90.93, 95.15]; RT 1049 ms (SD 712) [CI 987, 1110]; logRT 6.84 (SD 0.43) [CI 6.80, 6.87]. Tone: accuracy 94.11% (SD 23.57) [CI 92.15, 96.06]; RT 1016 ms (SD 532) [CI 970, 1062]; logRT 6.83 (SD 0.41) [CI 6.79, 6.86].
Picture Naming. Congruent: accuracy 76.88% (SD 42.20) [CI 73.29, 80.47]; RT 2113 ms (SD 1819) [CI 1930, 2296]; logRT 7.38 (SD 0.72) [CI 7.31, 7.45]. Silence: accuracy 73.31% (SD 44.28) [CI 69.54, 77.07]; RT 2320 ms (SD 2364) [CI 2079, 2562]; logRT 7.42 (SD 0.77) [CI 7.34, 7.50]. Tone: accuracy 78.57% (SD 41.07) [CI 75.08, 82.06]; RT 2199 ms (SD 2074) [CI 1994, 2404]; logRT 7.37 (SD 0.78) [CI 7.29, 7.45].
Backward Translation, Timepoint 2. Congruent: accuracy 84.25% (SD 36.49) [CI 79.92, 88.58]; RT 2184 ms (SD 1822) [CI 1942, 2425]; logRT 7.43 (SD 0.69) [CI 7.34, 7.52]. Silence: accuracy 84.25% (SD 36.49) [CI 79.92, 88.58]; RT 2160 ms (SD 1764) [CI 1928, 2393]; logRT 7.44 (SD 0.66) [CI 7.36, 7.53]. Tone: accuracy 87.55% (SD 33.08) [CI 83.62, 91.47]; RT 2295 ms (SD 2481) [CI 1976, 2615]; logRT 7.43 (SD 0.73) [CI 7.33, 7.52].
Backward Translation, Timepoint 1. Congruent: accuracy 69.60% (SD 46.08) [CI 64.13, 75.06]; RT 2439 ms (SD 2024) [CI 2134, 2743]; logRT 7.52 (SD 0.73) [CI 7.41, 7.63]. Silence: accuracy 58.97% (SD 49.28) [CI 53.13, 64.82]; RT 2363 ms (SD 1953) [CI 2044, 2682]; logRT 7.52 (SD 0.70) [CI 7.41, 7.63]. Tone: accuracy 61.90% (SD 48.65) [CI 56.13, 67.68]; RT 2770 ms (SD 2536) [CI 2370, 3169]; logRT 7.65 (SD 0.70) [CI 7.54, 7.76].
Forward Translation, Timepoint 2. Congruent: accuracy 76.92% (SD 42.21) [CI 71.92, 81.93]; RT 2856 ms (SD 2856) [CI 2527, 3185]; logRT 7.68 (SD 0.75) [CI 7.58, 7.79]. Silence: accuracy 70.33% (SD 45.77) [CI 64.90, 75.76]; RT 2751 ms (SD 2751) [CI 2457, 3046]; logRT 7.69 (SD 0.68) [CI 7.59, 7.78]. Tone: accuracy 73.99% (SD 43.95) [CI 68.78, 79.21]; RT 2891 ms (SD 2891) [CI 2550, 3233]; logRT 7.70 (SD 0.71) [CI 7.60, 7.80].
Forward Translation, Timepoint 1. Congruent: accuracy 47.86% (SD 50.04) [CI 42.00, 53.72]; RT 3188 ms (SD 2604) [CI 2721, 3656]; logRT 7.79 (SD 0.74) [CI 7.66, 7.92]. Silence: accuracy 37.86% (SD 48.59) [CI 32.17, 43.55]; RT 3638 ms (SD 2950) [CI 3051, 4224]; logRT 7.92 (SD 0.74) [CI 7.77, 8.07]. Tone: accuracy 47.50% (SD 50.03) [CI 41.64, 53.36]; RT 2862 ms (SD 2528) [CI 2402, 3322]; logRT 7.70 (SD 0.69) [CI 7.57, 7.82].
Table 3 shows mean accuracy (%), response time (RT in ms), and log-transformed response time (logRT) reported for 4 tasks (forward translation at timepoints 1 and 2, backward translation at timepoints 1 and 2, picture naming, and lexical decision) for 3 sound conditions (congruent, silence, tone). Standard deviations (SD) and 95% confidence intervals (CI) are also shown.
Note: Values are presented as means (M), with standard deviations (SD) in parentheses and 95% confidence intervals (CI) indicated in brackets with lower and upper limits. Accuracy values are reported as percentages. RT = response time (ms); logRT = natural logarithm of response time (ms); T1 = Timepoint 1; T2 = Timepoint 2.
3.1. Questionnaires
For the additional models performed, we found that neither LEX (all uncorrected ps > .02) nor SES (all uncorrected ps > .12) provided a significant effect on performance for any task or measure when p-correction was performed (FDR-corrected ps > .31).
3.2. Semantically mediated tasks
3.2.1. Forward translation task
Accuracy. Participants showed a general accuracy rate of M = 59%, SE = 1.21. There was a main effect of Timepoint, (χ2) (1) = 184.19, p < .001, with accuracy increasing from 44% (SE = 1.72) at Timepoint 1 to 74% (SE = 1.54) at Timepoint 2. There was also a main effect of Sound, (χ2) (2) = 15.97, p < .001. The Timepoint × Sound interaction was not significant, (χ2) (2) = 1.59, p = .45. The post hoc for Sound detected differences between the tone (M = 61%, SE = 2.08) and silence conditions (M = 54%, SE = 2.12) (OR = 1.68, 95% CI [1.13, 2.48], p = .006), between congruent (M = 62%, SE = 2.06) and silence conditions (OR = 1.84, 95% CI [1.24, 2.74], p < .001), but not between tone and congruent conditions, p = .84. Given the absence of a Timepoint × Sound interaction, accuracy is shown collapsed across timepoints in Figure 2; however, descriptive statistics by timepoint are available in Table 3.
Observed mean accuracy (%) in both semantically mediated tasks: (A) Forward translation task and (B) Picture naming task by sound conditions (congruent sound, silence, neutral tone). Vertical lines indicate standard error, and mean accuracy is indicated for each condition.

Figure 2. Long description
Panel B: Picture Naming Task. Accuracy was high and relatively similar across all conditions: neutral tone was highest (78.6%), followed by congruent sound (76.9%) and silence (73.3%). A significance bracket indicates neutral tone was significantly more accurate than silence (p = .022). No other significant differences were indicated.
Panel A: Forward Translation Task. Congruent sound yielded the highest mean accuracy (62.2%), followed by neutral tone (60.6%) and silence (53.9%). Significance brackets indicate that congruent sound was significantly more accurate than silence (p < .001), and neutral tone was significantly more accurate than silence (p = .006).
Figure 2 contains two grouped bar charts showing mean accuracy (%) for semantically mediated tasks. Each chart has three bars representing the three sound conditions: congruent sound (orange), silence (blue), and neutral tone (yellow). The y-axis shows mean accuracy (%). Error bars indicate standard errors.
Reaction times. The percentage of data eliminated as outliers was 2.07%. There was a main effect of Timepoint, with participants responding faster at Timepoint 2 (M log = 7.69, SE log = 0.03) in relation to the first (M log = 7.80, SE log = 0.04), (χ2) (1) = 16.11, p < .001. There was no Sound main effect, (χ2) (2) = 1.48, p = .48 or Sound × Timepoint interaction, (χ2) (2) = 3.32, p = .19.
3.2.2. Picture naming task
Accuracy. Participants achieved an overall accuracy of M = 76% (SE = 1.07). The model revealed a main effect of Sound, (χ2) (2) = 7.15, p = .028. Post hoc tests indicated that participants were more accurate in the tone condition (M = 79%, SE = 1.78) than in the silence condition (M = 73%, SE = 1.92) (OR = 1.61, 95% CI [1.06, 2.45], p = .022), see Figure 2. No significant differences were found between the congruent sound (M = 77%, SE = 1.83) and tone conditions, p = .53, or between the congruent and silence conditions, p = .26.
Reaction times. Outlier removal excluded 2.22% of the data. Results from our model showed that Sound had no effect, (χ2) (2) = 1.02, p = .60.
3.3. Lexically mediated tasks
3.3.1. Backward translation task
Accuracy. Participants showed a general accuracy rate of M = 74%, SE = 1.08. There was a main effect of Timepoint, (χ2) (1) = 110.11, p < .001, with accuracy increasing from 63% (SE = 1.68) at Timepoint 1 to 85% (SE = 1.24) at Timepoint 2. There was also a main effect of Sound, (χ2) (2) = 6.62, p = .037. The Timepoint × Sound interaction was significant, (χ2) (2) = 6.94, p = .031. The post hoc of the interaction showed that differences were detected at Timepoint 1 between congruent (M = 70%, SE = 2.79) and silence (M = 59%, SE = 2.98) conditions (OR = 2.07, 95% CI [1.23, 3.49], p = .003), and between congruent and tone (M = 62%, SE = 2.94) conditions (OR = 1.69, 95% CI [1.00, 2.85], p = .049). The difference at Timepoint 1 between silence and tone conditions was not significant, p = .61, nor were any of the comparisons at Timepoint 2, ps > .35.
Reaction times. The percentage of data eliminated as outliers was 2.06%. There was a main effect of Timepoint, with participants responding faster at Timepoint 2 (M log = 7.43, SE log = 0.03) in relation to the first (M log = 7.56, SE log = 0.03), (χ2) (1) = 38.01, p < .001. There was no Sound main effect, (χ2) (2) = 1.55, p = .46 or Sound × Timepoint interaction, (χ2) (2) = 4.49, p = .11.
3.3.2. Lexical decision task
Accuracy. Participants were highly accurate (M = 95%, SE = 0.37). Statistical analysis revealed that Sound had no effect, (χ2) (2) = 1.83, p = .40, see Figure 3.
Observed mean accuracy (%) in both lexically mediated tasks: (A) Backward translation task and (B) Lexical decision task by sound conditions (congruent sound, silence, neutral tone). Vertical lines indicate standard error, and mean accuracy is indicated for each condition.

Figure 3. Long description
Panel B: Lexical Decision Task. Accuracy was high and near-identical across conditions: congruent sound (94.8%), silence (93.0%), and neutral tone (94.1%). No significant differences between conditions were indicated.
Panel A: Backward Translation Task, shown at two timepoints. At Timepoint 1, congruent sound yielded the highest accuracy (69.6%), followed by neutral tone (61.9%) and silence (59.0%). Significance brackets indicate congruent sound was significantly higher than silence (p = .003) and than neutral tone (p = .049). At Timepoint 2, accuracy increased for all conditions and no significant differences were indicated: congruent sound (84.2%), silence (84.2%), and neutral tone (87.5%).
Figure 3 contains two bar charts showing mean accuracy (%) for lexically mediated tasks. Bars represent three sound conditions: congruent sound (orange), silence (blue), and neutral tone (yellow). The y-axis shows mean accuracy (%). Error bars indicate standard errors.
D′. The mean sensitivity score was M = 9.43 (SE = 0.49). All d′ values were above 0, showing that participants’ performance was better than chance. Higher d′ values reflect greater sensitivity and ability to discriminate items. No significant main effect of Sound was found for this measure, (χ2) (2) = 2.72, p = .26.
Reaction times. A small proportion of trials (2.34%) were excluded as outliers. Analysis showed that RTs were not influenced by Sound (χ2) (2) = 0.57, p = .75.
4. Discussion
This study was the first to investigate dual-sensory input applied to the realm of FL learning by examining whether the combination of semantically meaningful environmental sounds and visual cues (pictures) could enhance recall. Previous literature has established the general utility of environmental sounds for enhancing learning (Heikkilä et al., Reference Heikkilä, Alho, Hyvönen and Tiippana2015) and specifically, for FL acquisition (Bellegarda et al., Reference Bellegarda, García-Gámez and Macizo2026; Kaplan-Rakowski & Loranc-Paszylk, Reference Kaplan-Rakowski and Loranc-Paszylk2019), and it is well established that pictures are beneficial for FL learning (Altarriba & Knickerbocker, Reference Altarriba and Knickerbocker2011; Bates & Son, Reference Bates and Son2020; Yu & Liu, Reference Yu and Liu2022). Importantly, unlike previous studies that presented written linguistic cues alongside new words, our paradigm relied on conceptual activation of the target. This design aimed to strengthen the conceptual node and the FL-concept connection by engaging learners’ semantic networks through sensory stimulation while minimizing L1 mediation. To assess how the amount and quality of sensory input influenced learning, we manipulated the number of sensory inputs (single visual input or dual visual/auditory input) and the characteristic of the auditory input. Participants were exposed to three auditory conditions: semantically meaningful environmental sound, neutral tone and silence. Across tasks, a coherent pattern emerged in accuracy performance depending on if tasks were semantically (forward translation and picture naming) or lexically (backward translation and lexical decision) mediated. Accuracy results for the forward translation task indicated that the presence of auditory input in general – regardless of semantic relevance – was associated with better learning relative to the silence condition. However, the environmental sounds did not appear to provide additional benefits beyond the presence of sound alone. In contrast, lexically mediated tasks showed no difference between sound conditions once learning had taken place. RTs showed no significant differences across conditions. This pattern may reflect the intensive repetition during the learning phase, which could have reduced variability in retrieval speed across conditions. This also highlights accuracy as the most reliable indicator of learning in our paradigm. The implications of the findings, and in particular a possible redundancy due to the dual-sensory input, are discussed below.
Contrary to our initial hypothesis, which predicted that two channels of semantically relevant input (e.g., pictures and environmental sounds) would enhance semantic processing and lead to superior learning outcomes, the environmental sound condition did not outperform the other neutral sound condition. Instead, both types of sounds yielded similar performance during testing and outperformed the silence condition. This pattern can be understood through several complementary, though not mutually exclusive, mechanisms. One possibility is that the presence of auditory input in general, independent of semantic content or task relevance, may have enriched the encoding process and supported the formation of more elaborate semantic memory representations. This would be consistent with the Dual Coding Theory, which posits that an additional sensory trace enhances the probability of later correct retrieval (Paivio, Reference Paivio1969). In addition, both types of sounds may have increased participants’ involvement or engagement in the task, in line with the self-involvement theory (Helstrup, Reference Helstrup1987). Consistent with our pattern, Bellegarda et al. (Reference Bellegarda, García-Gámez and Macizo2026) also found that both environmental sounds and neutral tones performed similarly in an FL learning paradigm. Thus, even though the present study enriched the sensory trace through two modalities (visual and auditory) instead of one (auditory only, Bellegarda et al., Reference Bellegarda, García-Gámez and Macizo2026), this enrichment did not yield additional semantic benefits for the congruent condition. Instead, both studies converge on the finding that semantic relevance of the auditory input did not affect the learning outcomes, suggesting that the mere presence of sound, rather than its informational content, may be the critical factor shaping encoding.
A second explanation derives from the redundancy or semantic overload effect. In multisensory learning, redundancy occurs when multiple modalities provide overlapping information, diminishing the additive value of each, while semantic overload may occur when multiple inputs activate conceptual information simultaneously, exceeding optimal processing capacity (Sweller, Reference Sweller1994). According to Multimedia Learning Theory (Mayer, Reference Mayer2012), additional information through separate input channels does not necessarily enhance learning unless it contributes uniquely to the representation. In these situations, learners may inefficiently allocate cognitive resources to processing the duplicated information across different inputs. In the present design, methodological steps were taken to reduce redundancy by using a sequential presentation, thereby minimizing disruption to orthographic encoding. At the same time, the materials were designed to maintain coherence, ensuring that auditory and visual information were not extraneous or distracting. Nevertheless, during the simultaneous multisensory presentation, because pictures already provided rich conceptual access, additional semantic information from environmental sounds may have offered limited incremental benefit. Under these conditions, both environmental sounds and neutral tones may have supported learning to a similar extent, as neither provided uniquely diagnostic information beyond what was already available from visual input to foster further integration of the FL word into the mental conceptual model created.
This interpretation would align with evidence suggesting that visual stimuli may present certain advantages for conceptual activation compared to environmental sounds. Pictures offer immediate perceptual correspondence through structural and idiosyncratic features (i.e., shape, spatial configuration), allowing fast access to stored conceptual knowledge (Paivio, Reference Paivio1969). This view is further supported by evidence from picture-word interference paradigms showing strong semantic involvement in picture naming (Bajo et al., Reference Bajo, Puerta-Melguizo and Macizo2003; Macizo & Bajo, Reference Macizo and Bajo2004). In contrast, environmental sounds may provide less specific conceptual information and can be more ambiguous with respect to their source. For example, a picture of a horse directly specifies the concept “horse,” whereas a whinny requires listeners to infer the sound source before retrieving the associated concept. So, although environmental sounds can activate semantic representations, pictures may provide more precise and diagnostically rich conceptual cues. Altogether, the pattern in our data can be interpreted as the conceptual activation being largely driven during the learning phase through the visual modality. Nevertheless, the presence of both environmental sound and neutral tone conditions still yielded an additional benefit, likely through general engagement-related mechanisms.
From a broader theoretical perspective, these findings align with the BIA-d framework (Grainger et al., Reference Grainger, Midgley, Holcomb, Kail and Hickmann2010), which proposes that early FL vocabulary learning is driven by repeated coactivation of the FL form and its conceptual representation. In our paradigm, the picture was intended to activate the conceptual node, and the presence of sound (regardless of semantic content) likely boosted this activation by increasing engagement. Strong conceptual activation at the moment the FL word appeared would therefore strengthen the emerging FL-concept link, accounting for the advantage of any sound over silence. At the same time, the absence of an advantage for environmental sounds over neutral tones is consistent with the possibility of a ceiling or boundary condition on conceptual activation when visual cues are already highly informative: once the concept was fully activated by the picture, the possibility of congruent sounds providing additional semantic value (viz., conceptual activation) was limited.
Given the BIA-d assumption that conceptual support primarily benefits meaning-based tasks, the absence of sound effects in the two lexically mediated tasks is not surprising. Visual lexical decision relies largely on word form recognition and does not require accessing the conceptual system, making it less sensitive to the conceptual boosting that occurred during learning. Similarly, backward translation requires mapping newly learned FL forms onto highly dominant L1 lexical representations, such that performance does not hinge on explicit semantic access (Kroll & Stewart, Reference Kroll and Stewart1994). These processing characteristics also make both tasks less demanding compared to the semantically mediated tasks that required active item retrieval, form-to-meaning mapping and the resolution of conflict arising from competing L1 representations, which are dominant and well-established in the lexicon. In the case of lexical decision in particular, the task required only a binary keypress (indicating whether a word was learned or new). Consistent with this, overall accuracy in both lexically mediated tasks approached ceiling (95% in lexical decision and 85% in backward translation at the second timepoint), suggesting that participants were able to recognize or produce almost all studied words regardless of the auditory condition; this ceiling effect likely obscured any subtle differences that might otherwise have emerged. Indeed, even participants who struggled in the more demanding meaning-based tasks performed well above chance in lexical decision.
An interesting observation is that in backward translation at the first timepoint, congruent sounds provided a modest advantage over silence and neutral tones, suggesting that early in learning, auditory semantic cues may transiently facilitate FL form-meaning mapping. By the second timepoint, however, this advantage disappeared. This pattern aligns with behavioral and neurophysiological evidence showing that even very novice learners can access semantic representations of newly learned FL words (Comesaña et al., Reference Comesaña, Perea, Piñeiro and Fraga2009; García-Gámez & Macizo, Reference García-Gámez and Macizo2022; Pu et al., Reference Pu, Holcomb and Midgley2016). As learning progressed and the FL forms became more established, participants appeared to rely primarily on direct FL–L1 lexical connections to perform lexically mediated tasks, reducing the contribution of explicit semantic mediation.
An important note of comparison is that Bellegarda et al. (Reference Bellegarda, García-Gámez and Macizo2026) reported reduced performance from both congruent and neutral sounds in their lexical decision task, whereas in the present study we observed a null effect. Part of this difference may reflect methodological variations: Bellegarda et al. used more items and provided written Spanish forms, while our design relied exclusively on conceptual activation through pictures, which may have promoted especially stable memory traces (Comesaña et al., Reference Comesaña, Perea, Piñeiro and Fraga2009). This divergence can also be interpreted in light of the different sensory configurations: their study presented sounds alone, without visual input, and the observed decrease in performance was attributed to increased cognitive load hindering lexical access. In contrast, in the present study, beneficial effects from pictures and potential difficulty introduced by concurrent sounds may have acted in opposite directions, effectively canceling each other out and producing the null effect observed here. These findings reinforce that multisensory contributions are not inherently additive: its effects can interact nonlinearly, particularly in tasks driven by rapid form-based processing such as lexical decision and backward translation.
It is also interesting to understand how previous LEX or SES could impact FL vocabulary acquisition. Although both have been linked to differences in linguistic development, our null result in our exploratory analysis suggests that their influence may be limited in the context of short-term FL vocabulary learning. One possibility is that the associative nature of the task relied primarily on domain-general memory mechanisms rather than on preexisting linguistic systems or linguistic transfer processes. Additionally, multilingualism effects often appear when learning involves interference between known and new languages (e.g., cognates, grammatical overlap). In this experiment, participants learned isolated FL nouns and there were neither cognates nor grammatical elements. Regarding SES, given that our sample primarily consisted of university students, this lack of variability may have attenuated differences. Future research may probe these variables under conditions involving greater linguistic interference and greater SES variability.
4.1. Future directions
Environmental sounds and neutral tones differ acoustically and have been proposed to engage partially different processing routes. In acoustic terms, neutral (or pure) tones are stimuli with predictable spectral and temporal structures. In contrast, environmental sounds are acoustically dense signals. They feature rapid, asymmetric modulations and an intrinsic statistical structure that leads to a different processing route compared to neutral tones (Theunissen & Elie, Reference Theunissen and Elie2014). However, we did not find behavioral evidence for this distinction. This finding is informative in itself, as it suggests that the contribution of auditory input to FL learning may depend on a specific boundary condition that remains to be specified in future research.
Additionally, to better understand the processing of environmental sounds vs. neutral sounds, neuroimaging techniques could clarify how these cues recruit semantic networks differently at a neural level, even when behavioral performance may appear unchanged. Word meaning is not stored in a single, encapsulated module, but is instead supported by distributed neuronal assemblies spanning multiple cortical regions (Pulvermüller, Reference Pulvermüller2005). From this perspective, environmental sounds may function as sensory “anchors” that facilitate the automatic recruitment of topographically specific semantic networks. Neurophysiological methods such as electroencephalography (EEG) or magnetoencephalography (MEG) are particularly well suited to testing this possibility. By examining the spatiotemporal dynamics of brain activity, future studies could determine whether environmental sounds promote stronger grounding of lexical representations in action and perception systems compared to neutral cues, potentially leading to more durable memory traces through the engagement of distributed neurocognitive networks.
Another important direction for future work concerns consolidation over time. Because our study tested participants’ learning on the same day, it remains possible that the benefits of environmental sounds may emerge more clearly over time. Semantically richer encoding could produce more durable memory traces, which have greater resistance to forgetting. Indeed, Kaplan-Rakowski and Loranc-Paszylk (Reference Kaplan-Rakowski and Loranc-Paszylk2019) observed sustained advantages for sound-enhanced words on a seven-day delayed recall test when compared to the silence control condition. Notably, this advantage was not present for their pronunciation or sound and pronunciation conditions. Future research incorporating delayed assessments would therefore be valuable for examining whether environmental sounds support longer-term retention.
Finally, expanding the lexical scope of training materials to include a greater diversity of lexical items (i.e., verbs) may further clarify how multisensory input supports vocabulary acquisition. Expanding the learning set could also help determine whether the effects generalize across a broader range of linguistic contexts. Verbs, in particular, are often more “embodied” than nouns, as they engage sensorimotor systems during learning and retrieval (Grisoni et al., Reference Grisoni, Dreyer and Pulvermüller2016, Reference Grisoni, Mohr and Pulvermüller2019).
5. Conclusions
To conclude, this study examined whether relevant multisensory stimuli enhance adult FL vocabulary learning. The results showed that the presence of sound during learning (regardless of semantic relevance) led to improved performance compared to silence, particularly in tasks that required semantic mediation. Although environmental sounds did not provide an advantage over neutral tones for learning, the overall benefit of sound suggests that sensory engagement plays a meaningful role during early FL learning. Importantly, the absence of sound effects in lexically mediated tasks highlights that multisensory contributions are not inherently additive: combined influences may interact in task-dependent ways, sometimes resulting in null behavioral effects. Rather than indicating that environmental sounds lack pedagogical value, our results point to a possible boundary condition under which multisensory enrichment is most effective. When visual input already provides strong conceptual activation, additional semantic cues may yield diminishing returns. Understanding when auditory meaning contributes uniquely and when it becomes redundant remains an important direction for future research.
Supplementary material
The supplementary material for this article can be found at http://doi.org/10.1017/S1366728926101692.
Data availability statement
The data, analyses, study materials and supplementary files supporting the findings of this research are openly available in the Open Science Framework (OSF) repository at: https://doi.org/10.17605/OSF.IO/6QAZ2
Acknowledgments
This research was supported by the Spanish Ministry of Science and Innovation through an FPU University Teacher Training grant (grant number FPU19/04616) awarded to Melodie Bellegarda and the project (reference number PID2019-111359GB-I00) awarded to Pedro Macizo. Additional funding was provided to the Mind, Brain and Behavior Research Center (CIMCYC), University of Granada through CEX2023-001312-M/MICIU/AEI/10.13039/501100011033 and UCE-PP2023-11/UGR. Funding for open access charge: University of Granada / CBUA. We are grateful to Daniel Navarrete Sancho for his valuable assistance in data collection. All procedures performed in this study involving human participants were in accordance with the ethical standards of the research ethical committee at the University of Granada (number issued by the Ethical Committee: 957/CEIH/2019) as well as APA ethical standards. We are grateful to two anonymous reviewers for their help in the review process.
Competing interests
The authors declare none.