Document Type : Research Article
Authors
1 Department of English Language and Literature, Faculty of Foreign Languages and Literatures, University of Tehran, Tehran, Iran
2 Department of Teaching English as a Foreign Language (TEFL), Faculty of Humanities, University of Hormozgan, Bandar Abbas, Iran
Abstract
Keywords
Main Subjects
Introduction
The meaning of “comprehension” seems to differ among individuals (Lee et al., 2025). What researchers seem to agree on about the meaning of comprehension might be a range of abilities and processes involved in the comprehension process (Kendeou et al., 2007). Various types of comprehension share basic processes, including interpreting content, utilizing existing knowledge to help this interpretation, and eventually generating a coherent mental representation of the processed content (Kendeou et al., 2007). Recently, the majority of input comprehension studies have focused on reading (Choi & Zhang, 2021; Tong et al., 2023; Zhang & Zhang, 2022), listening (Du & Man, 2022; Vafaee & Suzuki, 2020; Zhang & Zhang, 2022) and reading-while-listening (Hui, 2024; Pellicer-Sánchez et al., 2020; Serrano & Pellicer-Sánchez, 2022), but this trend has started to change during the last decade given the popularity of audio-visual materials. This change seems to be timely because research indicates that nowadays L2 learners tend to receive more audio-visual input than other input types, such as reading or listening materials (Muñoz et al., 2023). To illustrate, in a study undertaken by Ahrabi Fakhr et al. (2021), the questionnaire results showed that Iranian EFL students majoring in English watched English videos more frequently than they read English books or listened to English materials. Thus, it can also be summarized that technology could promote L2 comprehension dramatically.
Viewing comprehension may pose a challenge for L2 learners (Schroeders et al., 2010). Various techniques have been devised to facilitate comprehension, with captioning being one of the most widely used ones (Teng, 2025). Some research evidence suggests that captions can enhance viewing comprehension (Sidorov et al., 2020). For example, Pujadas and Muñoz (2024) found that captioned viewing led to superior comprehension scores compared to uncaptioned viewing. In a study with a between-subjects design, Teng (2025) found that viewing captioned videos resulted in improvements in comprehension scores.
Few studies can be found on the effect of textual enhancement (TE) of vocabulary items in captions on viewing comprehension. TE has been defined as manipulating a text to make certain linguistic targets more perceptually salient (Lee, 2021). In vocabulary learning studies, this technique has been employed to draw L2 learners’ attention to target lexical items and promote their acquisition. However, the impact of implementing this technique on students’ viewing experience and viewing comprehension is yet unresolved. In Hsieh’s (2020) study, the results showed that uncaptioned viewing, fully-captioned viewing, and fully captioned viewing plus textually enhanced target single words did not differ significantly from each other in terms of comprehension scores. However, the fully-captioned viewing condition resulted in non-significantly better scores than the other conditions. The writer pointed out that the presence of textually enhanced words in the captioning line might have directed learners’ attention to vocabulary rather than video content.
Majuddin et al. (2021) conducted a study with Malaysian L2 English learners randomly allocated to one of the experimental conditions. These conditions differed in their caption condition (no captions, unenhanced captions, captions + TE on target multiword items). The results showed that learners in the unenhanced captioned viewing condition significantly outperformed those in the uncaptioned condition. However, rather surprisingly, enhanced captioned viewing did not result in significantly higher comprehension scores compared to the uncaptioned condition, suggesting a negative effect of TE on video comprehension.
Despite many advances in TE and captioning, very few studies have focused on this aspect of EFL education in Iranian settings (Roohani et al., 2013). Existing findings are inconsistent, making it difficult for teachers and materials developers to determine whether textual enhancement should be incorporated into captioned materials for Iranian EFL learners. So, the question remains as to whether directing Iranian EFL learners’ attention toward enhanced lexical items facilitates or interferes with content comprehension. Considering the abundance of audiovisual input received by many EFL learners and the changing conditions of EFL learning that increase opportunities for language exposure, the paucity of research in this area needs to be addressed. Applying this approach in research can also pave the way for practical applications in EFL education.
Literature Review
Theoretical Background
Two theories that are closely related to the present investigation are the Noticing Hypothesis (Schmidt, 1990) and the Cognitive Theory of Multimedia Learning (Mayer, 2021). The Noticing Hypothesis posits that conscious attention to linguistic forms is a necessary condition for language learning. According to this hypothesis, language input cannot become intake unless learners consciously notice specific linguistic features. Textual Enhancement (TE) was developed as an instructional technique to increase the perceptual salience of target linguistic forms and thereby promote noticing. By making target items more visually prominent, TE is expected to increase the likelihood that learners notice and subsequently acquire them.
According to the Cognitive Theory of Multimedia Learning (Mayer, 2021), meaningful learning occurs when learners actively select relevant verbal and visual information, organize it into coherent mental representations, and integrate these representations with prior knowledge. Consequently, presenting words together with corresponding images generally promotes deeper learning than presenting words alone. However, this theory also assumes that working memory has a limited processing capacity. Learners can process only a finite amount of information in the verbal and visual channels at any given time (Baddeley, 2010). When instructional materials exceed this limited capacity, cognitive overload may occur, reducing learning effectiveness. This assumption has important implications for captioned video viewing. While captions may facilitate comprehension and vocabulary learning by providing additional verbal input, excessive on-screen text may increase extraneous cognitive load and distract learners' attention from the video content.
Captions and Viewing Comprehension
Captions refer to on-screen text in the same language (Sydorenko, 2010; Vanderplank, 2016a; 2016b). They were initially used to enable deaf and hard-of-hearing individuals to understand videos, but were later exploited to benefit those with normal hearing who nevertheless had difficulties with L2 listening comprehension (Vanderplank, 2016b; 2016a). Recently, there has been a proliferation of studies investigating the potential benefits of captions for language learning purposes. Although many studies have aimed to explore the efficacy of captions for vocabulary learning, other language areas have not been neglected. In fact, research suggests that captioning videos can boost learning of different language components, including grammar (e.g., Pattemore, 2022), pronunciation (e.g., Mohsen & Mahdi, 2021), and vocabulary (e.g., Fievez et al., 2023; Teng, 2022a; 2022b).
In Winke et al.’s (2010) study, second- and fourth-year learners of Arabic, Chinese, Spanish, and Russian watched three short L2 videos twice, once with captioning and once without. For each target language, one group saw captions during the first viewing of the video and another saw captions during the second viewing. Learners of Spanish included two extra groups, one which viewed the videos twice without captions and one which viewed them twice with captions. After the second viewing, a comprehension test was given. Results demonstrated that captioned videos were more beneficial than uncaptioned videos for overall video comprehension.
Pujadas and Muñoz (2024) examined the effects of L2 proficiency (A1 to C2, measured by the OPT) and vocabulary knowledge on the comprehension of captioned and uncaptioned English TV series. The results revealed that captioned viewing led to superior comprehension scores compared to uncaptioned viewing. Additionally, both L2 proficiency and vocabulary knowledge were positively correlated with comprehension scores.
Apart from studies that have probed the effect of various on-screen texts on viewing comprehension, a few studies have compared comprehension of captioned videos and comprehension of other modes of input. To exemplify, in a study with 80 pre-intermediate Vietnamese EFL learners, Vu et al. (2023) found that reading-while-listening resulted in better comprehension scores than captioned TV viewing.
Textual Enhancement and Content Comprehension
Typographic or Textual Enhancement (TE) is a technique designed to help learners attend to specific linguistic items in the input, using methods such as underlining and/or bolding of those items (Lee, 2021). TE has been investigated mostly in the area of grammar learning through reading and more recently, through captioned viewing (Della Putta, 2019; Jahan & Kormos, 2015; LaBrozzi, 2016; Lee, 2021; Lee & Jung, 2024; Lee & Révész, 2018; Lee, 2007; Meguro, 2019; Pattemore, 2022; Rassaei, 2015; Simard, 2009). The findings of this line of research, however, indicate mixed patterns. For instance, the findings of Lee's (2007) research showed that TE facilitated the learning of target grammatical forms but had unfavorable effects on reading comprehension. Contrary to Lee's (2007) findings, in Meguro's (2019) study, TE of target grammatical constructions did not hinder L2 English learners’ reading comprehension. LaBrozzi (2016) compared the effects of different types of TE on recognition of L2 grammatical forms and reading comprehension. Different types of TE were explored, including underlining, italicizing, bolding, using capital letters, increasing font size, and changing the font. The results demonstrated that reading comprehension was unaffected by the type of TE used.
The effect of TE on vocabulary items on content comprehension in different input modes has also been investigated. As for reading, Puimege et al. (2024) asked L2 learners of English to read 10 English texts adapted from a popular science book and TED-ED scripts. Post-intervention results indicated that participants primarily focused on content comprehension, and TE did not cause learners to pay conscious attention to the form of L2 collocations.
Taking the contradictory findings of previous studies into consideration, the present study aims to examine the impact of captions and TE of multiword items and single words on EFL learners’ viewing comprehension. Additionally, learners’ attitudes and viewing experiences are scrutinized. From a theoretical perspective, whether TE facilitates or hinders learning during captioned viewing may depend on the balance between increased noticing (Schmidt, 1990) and increased processing demands (Mayer, 2021). If the enhancement successfully directs learners' attention without overwhelming working memory, vocabulary learning and comprehension may improve simultaneously. Conversely, if the additional visual salience imposes excessive cognitive load, learners may struggle to process both the message and the highlighted lexical items. Comparing different caption conditions, therefore, provides an opportunity to examine how lexical salience and cognitive load interact during L2 viewing. Although previous studies have explored these issues, their findings remain inconsistent, and evidence from Iranian EFL contexts is scarce. Accordingly, empirical investigation is required to determine whether textual enhancement of single words and multiword items in captions facilitates learning without compromising viewing comprehension. The present study, therefore, addresses the following two research questions:
Methods
This study tried to explore the effect of captions and TE of multiword items (MWIs) and single words in captions on L2 learners’ viewing comprehension. The viewing conditions included the following, randomly distributed among four intact classes: 1) captioned viewing with textually enhanced MWIs (MWI-enhanced captioned viewing), 2) captioned viewing with textually enhanced single words (single-word enhanced captioned viewing), 3) uncaptioned viewing, and 4) captioned viewing with no TE (unenhanced captioned viewing).
Participants
Sixty-two Iranian undergraduates majoring in English as a Foreign Language (TEFL) from four intact classes participated in this study. All participants knew the most frequent 3,000 English words according to the results of Webb et al.'s (2017) Updated Vocabulary Levels Test (UVLT). The UVLT was administered to ensure that participants possessed sufficient lexical knowledge to understand the documentary. Participants’ ages ranged from 18 to 23 (mean age = 20.46, SD = 1.30), and the sample included 18 male and 44 female students. Table 1 displays participants’ background information.
Table 1. Descriptive Statistics for Background Information of Participants
|
|
Age |
Gender |
||||||
|
N |
Minimum |
Maximum |
Mean |
Standard Deviation |
Male |
Female |
||
|
N |
N |
|||||||
|
Group |
Group 1 |
16 |
18.00 |
21.00 |
19.19 |
.98 |
7 |
9 |
|
Group 2 |
16 |
19.00 |
22.00 |
20.19 |
.83 |
5 |
11 |
|
|
Group 3 |
14 |
20.00 |
23.00 |
20.57 |
.94 |
3 |
11 |
|
|
Group 4 |
16 |
21.00 |
23.00 |
21.94 |
.57 |
3 |
13 |
|
Materials and Instruments
The first 32 minutes of the first episode of The World’s Sneakiest Animals, a BBC documentary series, was used as the video material. The episode titled ‘Staying alive’ was about visual trickery. The duration of the video was 32 minutes and 15 seconds, and its script consisted of 2,923 words. Analyzing the documentary script using RANGE (Nation & Heatley, 2002) revealed that 91/5% of the words came from the 3,000 most frequent English word families. In a study on the relationship between vocabulary and viewing comprehension, Durbahn et al. (2020) found that a lexical coverage of about 90% is enough for adequate viewing comprehension without external help. Additionally, van Zeeland and Schmitt (2013) found no significant difference in comprehension of spoken input between 90 % and 95% text coverage. Following Durbahn et al. (2020) and van Zeeland and Schmitt (2013), this project deemed 90% coverage as appropriate. According to the participants’ VLT scores and the results of the pilot study, the participants of the main study were believed to be able to follow the video with relative ease.
As mentioned earlier in the literature review section, some evidence from prior research indicates that vocabulary knowledge and L2 proficiency are correlated with L2 input comprehension (Pujadas & Muñoz, 2020, 2024). Taking this into account, and to ensure inter-group homogeneity, we decided to measure not only participants’ vocabulary knowledge but also their L2 proficiency. Participants’ vocabulary knowledge was assessed using Webb et al.'s (2017) Updated Vocabulary Levels Test (UVLT) and Nation and Beglar’s (2007) Vocabulary Size Test (VST). The UVLT scores showed that all participants knew the most frequent 3,000 English words. Participants’ mean VST score was 76.11 out of 140
(SD = 14.37), and their score range indicated they knew between 4,000-11,000 most frequent English word families receptively. Another test that was given to the participants was the Quick Oxford Placement Test (OPT) (Allan, 2004). The result of this test demonstrated that students’ general English proficiency ranged from 30 to 54 (mean= 40.62, SD= 6.91). These results indicated that the participants’ English proficiency level ranged from pre-intermediate or B1 (according to the Common European Framework of Reference) to advanced or C1. Three one-way analyses of variance (ANOVA) were run, showing no significant differences between the groups’ UVLT, VST, and OPT scores (p. > 0.05 in all cases).
Participants completed a 10-item true/false test designed to assess their comprehension of the BBC documentary. This test did not include any of the textually-enhanced items (which were all unknown to the participants). The questions tapped into skills such as discerning the main idea, identifying supporting details, and drawing inferences from the context. The internal consistency of the comprehension test was examined using the Kuder-Richardson formula 20 (KR-20), which is the appropriate reliability estimate for dichotomously scored items. The resulting coefficient (KR-20= .75) indicated acceptable reliability for the test. A post-viewing questionnaire, adapted from Puimège and Peters (2020), Peters and Webb (2018), and Puimège and Peters (2019), was given to the participants as well. The questionnaire was translated into Persian (participants’ native language) by the first researcher. Content validity was confirmed using judgments of two experts in Applied Linguistics. Internal consistency of the questionnaire in the current study was acceptable (Cronbach's α = .78). Both the comprehension test and the questionnaire were scored by the first researcher.
Procedure
The data were gathered within three weeks. In the initial week, participants were briefed on the research (without revealing its specific aims) before they gave consent to take part in the study. They were given the OPT test during this week to measure their English language proficiency. In the second week, they completed the UVLT and the VST. One week later, participants watched the video under their assigned conditions, using the multimedia system (computer, projector, and speakers) in their classrooms. Group 1 watched the documentary with captions that included 16 textually-enhanced unknown MWIs. Group 2 watched the documentary with captions that involved 18 textually-enhanced unknown single words. Textual enhancement was in the form of bolding and underlining. Group 3 watched the documentary without captions, and Group 4 watched it with unenhanced captions. Before exposure to the learning materials, participants were told they would later answer comprehension questions. Immediately after the treatment, they took the comprehension test (reliability = .75), followed by the post-viewing questionnaire. After completion of the questionnaire, participants were debriefed about the real purposes of the study.
Data Analysis
Responses were scored using a binary method, according to which correct responses received 1 and incorrect responses received 0. Normality was confirmed by checking the skewness and kurtosis values. All statistical analyses were implemented using SPSS software (version 26.0). Given the normal distribution of the data, a one-way ANOVA was used to answer RQ1. RQ2 was answered through descriptive analyses of the questionnaire data.
Results and Discussion
Results
Comprehension Test
Participants’ comprehension test scores showed that, irrespective of their viewing condition, all four groups comprehended the content of the video well. The average comprehension score was 7.95 (standard deviation = 1.36). Table 2 summarizes the results of the comprehension test for each group. Levene’s test for homogeneity of variances (Table 3) revealed that the variance in scores was the same for the four groups (p = .615, p > 0.05).
Table 2. Descriptive Statistics for the Comprehension Test per Group
|
|
N |
Mean |
Std. Deviation |
|
Group 1 |
16 |
8.6875 |
1.07819 |
|
Group 2 |
16 |
8.3750 |
1.20416 |
|
Group 3 |
14 |
7.9286 |
1.49174 |
|
Group 4 |
16 |
6.8125 |
.91059 |
|
Total |
62 |
7.9516 |
1.36018 |
Note: The maximum possible score was 10.
Table 3. Results of the Levene’s Test for Homogeneity of Variances
|
|
Levene Statistic |
df1 |
df2 |
Sig. |
|
|
Comp. |
Based on Mean |
.604 |
3 |
58 |
.615 |
|
Based on Median |
.603 |
3 |
58 |
.616 |
|
|
Based on Median and with adjusted df |
.603 |
3 |
51.565 |
.616 |
|
|
Based on trimmed mean |
.566 |
3 |
58 |
.640 |
|
ANOVA results (Table 4) showed that there was a significant difference between participants’ comprehension scores, F (3, 58) = 7.752, p = .000. Post-hoc comparisons using the Tukey HSD test (Table 5) indicated that the mean score for Group 4 (M = 6.81, SD = .91) was significantly lower than that for Group 1 (M = 8.68, SD = 1.07) and Group 2 (M = 8.37, SD = 1.2). Other mean scores were not significantly different from each other. According to Cohen’s (1988) eta squared guidelines (as cited in Pallant, 2016, p. 212), 0.01 is considered a small effect, 0.06 a medium effect, and 0.14 a large effect. Here, the effect size, calculated using eta squared, was 0.286, which indicates that the difference in comprehension test mean scores between the groups was large.
Table 4. ANOVA Results
|
|
Sum of Squares |
df |
Mean Square |
F |
Sig. |
|
Between Groups |
32.301 |
3 |
10.767 |
7.752 |
.000 |
|
Within Groups |
80.554 |
58 |
1.389 |
|
|
|
Total |
112.855 |
61 |
|
|
|
Table 5. Multiple Comparisons for Comprehension Scores
|
(I) Group |
(J) Group |
Mean Difference (I-J) |
Std. Error |
Sig. |
95% Confidence Interval |
|
|
Lower Bound |
Upper Bound |
|||||
|
Group 1 |
Group 2 |
.31250 |
.41666 |
.876 |
-.7896 |
1.4146 |
|
Group 3 |
.75893 |
.43129 |
.303 |
-.3819 |
1.8997 |
|
|
Group 4 |
1.87500* |
.41666 |
.000 |
.7729 |
2.9771 |
|
|
Group 2 |
Group 1 |
-.31250 |
.41666 |
.876 |
-1.4146 |
.7896 |
|
Group 3 |
.44643 |
.43129 |
.730 |
-.6944 |
1.5872 |
|
|
Group 4 |
1.56250* |
.41666 |
.002 |
.4604 |
2.6646 |
|
|
Group 3 |
Group 1 |
-.75893 |
.43129 |
.303 |
-1.8997 |
.3819 |
|
Group 2 |
-.44643 |
.43129 |
.730 |
-1.5872 |
.6944 |
|
|
Group 4 |
1.11607 |
.43129 |
.057 |
-.0247 |
2.2569 |
|
|
Group 4 |
Group 1 |
-1.87500* |
.41666 |
.000 |
-2.9771 |
-.7729 |
|
Group 2 |
-1.56250* |
.41666 |
.002 |
-2.6646 |
-.4604 |
|
|
Group 3 |
-1.11607 |
.43129 |
.057 |
-2.2569 |
.0247 |
|
|
*. The mean difference is significant at the 0.05 level. |
||||||
Questionnaire
Viewing habits: The results of the first section of the questionnaire showed that 98.3% of the respondents (out of 59) watch English videos (including films, series, TV programs, YouTube, etc.). Of 57 respondents, 34 (59.6%) reported that they watch English videos every day, 20 (35.1%) claimed that they watch English videos once a week, and 3 (5.3%) stated that they watch English videos once a month. Of 58 respondents, 12 (20.7%) reported that they watch English videos without any captions or subtitles, 14 (24.1%) reported that they use Persian subtitles while watching English videos, 24 (41.4%) stated that they watch English videos with English captions, 7 (12.1%) claimed that they watch English videos sometimes with Persian subtitles and sometimes with English captions, and 1 person (1.7%) reported that she watches English videos sometimes with and sometimes without English captions.
Affective response and content comprehension: According to the first item of the second section of the questionnaire, 46 participants (74.2%) found the topic of the video interesting (agree or strongly agree), 12 (19.4%) were uncertain, and a few did not find the topic interesting. Regarding the length of the video, 31 participants (50%) reported that it was appropriate, 16 (25.8%) were uncertain, and others believed that the length was not appropriate (it was long). As for the understandability of the video, 38 participants (61.3 %) stated that its content was easily understandable, 13 (21%) participants were uncertain, and 10 (16.2%) did not find the content easily understandable. Forty-six participants (74.2%) reported that they mainly focused on the content of the video; 8 participants (12.9%) were uncertain, and 7 (11.3%) either disagreed or strongly disagreed with the claim.
Vocabulary awareness: Thirty-two participants (51.6%) reported that they paid attention to the vocabulary items that were used in the video, 19 (30.6%) were uncertain, and 10 (16.1%) stated that they did not pay attention to the vocabulary items that were used in the video.
Attitudes towards viewing L2 videos and language learning: Answers to the sixth question of the second section of the questionnaire showed that 54 participants (88.5%) either agreed or strongly agreed that watching videos is a good way to improve one’s English proficiency. Answers to the other three questions of this section were also illuminating.
It was shown that 56 students (91.8%) either agreed or strongly agreed that watching English videos is a good way to improve L2 learners’ listening skills. Fifty-two participants (85.3%) either agreed or strongly agreed that watching videos is a good way to improve one’s English vocabulary. Participants appeared less certain about the effectiveness of English videos for improving L2 learners’ English grammar; 32 of them (53.4%) selected either agree or strongly agree, and 14 students (23.3%) rated this item 3, indicating their lack of certainty.
Open-ended questions: The other section of the questionnaire tapped into learners’ comprehension of the documentary. The answers to this section indicated that participants had no trouble understanding the gist of the content. Specifically, the questionnaire item “what have you learned in terms of content?” elicited references to central topics in the documentary. For example, participants mentioned “I have learned about the camouflage of animals,” “how animals stay alive in nature, including strategies and solutions,” “secrets behind how animals manage to stay alive,” “survival strategies in different animals,” “camouflage of sea animals,” and “I have learned that animals can be sneaky, too”. Some participants mentioned the details they remembered from the documentary, for example:
Several participants mentioned that they enjoyed the pictures and the visual effects of the documentary. One participant reported that because the presenter used casual language to speak, he was encouraged to listen to the content more carefully.
The questionnaire item “What have you learned in terms of language?” elicited illuminating answers. Two participants mentioned that although they could not remember the words and expressions from the video, they had picked up some items subconsciously. A few participants generally stated that they learned new content and new words and expressions without mentioning any examples. However, many participants named expressions and words they had learned from the documentary.
In the first group, in which the students watched the video with captions that were highlighted for target MWIs, one participant mentioned that although she comprehended the gist of the message, she did not learn the meaning of words and expressions because we were not allowed to look them up. The same participant stated that because they were tested on the words and expressions before, she was familiar with words and expressions and understood them in context, but was not able to remember their meanings. This answer is enlightening as the participant indicated her awareness about the importance of dictionary look-up and the effect of pretesting on later treatment. A participant from the third group, who watched the video without captions, said, “If I could, I would like to re-watch the video several times to take notes of important words and expressions and expand my vocabulary”. Another participant from the fourth group, who watched the video with normal captions, mentioned, “I remember venomous, predator, and gecko because they were repeated several times throughout the video”. This response is interesting because it not only suggests that the frequency of occurrence may boost vocabulary learning from captioned viewing but also shows that some learners pay attention to repetitions consciously.
As for the questionnaire item about whether or not participants wanted to re-watch the video, several participants mentioned that now that they understood the content of the video, they wanted to re-watch the video to learn new words and expressions, while others claimed that they wanted to re-watch it because they found the content interesting and wanted to pay attention to details as they re-watched the video. Still, others mentioned that because they understood all the content, they preferred not to re-watch the video. Several participants mentioned that they did not like to re-watch the video because if one understands the topics introduced in a video, re-watching that video would be a boring activity. One participant claimed that “it is not necessary to re-watch what I have understood, but if I intend to learn its vocabulary, re-watching would be a good idea”. These answers are valuable as they suggest that there may be individual differences between students regarding how they view a repeated activity.
Discussion
The primary purpose of the current study was to examine the effect of captioning and TE of multiword items and single words on Iranian EFL learners’ viewing comprehension. The secondary purpose of this study was to gain additional insights into the students’ viewing experience. The descriptive statistics demonstrated that the performance of the four groups on the comprehension test was as follows: Group 1 > Group 2 > Group 3 > Group 4.
Regarding the first research question, the ANOVA results revealed a significant difference between the comprehension scores of the four groups. Post-hoc analyses revealed that Group 1 and Group 2 significantly outperformed Group 4 on the comprehension test. This finding suggests that in this study, TE of captions had a positive impact on students’ viewing comprehension. The significant impact of caption condition on viewing comprehension scores differs from Montero Perez et al. (2014a) and Hsieh (2020) but is similar to Majuddin et al. (2021). It appears that in the current study, the presence of TE motivated learners to pay closer attention to the video content.
Interestingly, rather than distracting from the overall message, TE of lexical forms appears to have supported comprehension. From the perspective of Mayer's (2021) Cognitive Theory of Multimedia Learning, this outcome implies that the additional visual salience introduced by TE did not exceed learners' available cognitive capacity. Rather than imposing extraneous cognitive load, which would be expected to impair learning by overwhelming the verbal and visual processing channels, TE appears to have guided learners toward meaningful selection and organization of relevant linguistic and content information. This suggests that when TE is applied judiciously, it can assist learners in integrating verbal input with the visual and auditory channels of the video without causing cognitive overload.
As for TE of single words, the significantly greater viewing comprehension of Group 2 compared to Group 4 contradicts Hsieh's (2020) study, in which no significant difference was observed between the viewing comprehension scores of the single-word enhanced captioned viewing and unenhanced captioned viewing conditions. Interestingly, even the descriptive statistics of that study and the current study do not align, because in that study, unenhanced fully captioned viewing led to higher scores than both uncaptioned and single-word enhanced captioned viewing conditions.
A theoretical account of why single-word TE aided comprehension in the present study, but not in Hsieh's (2020), may relate to the interaction between noticing and learner proficiency. For higher-proficiency learners, as in the present sample, the act of noticing a highlighted word may require relatively less cognitive effort, freeing up working memory resources for simultaneous meaning construction. This is consistent with Mayer's (2021) theory, which posits that meaningful learning depends on the learner's ability to actively select and integrate relevant information without exceeding processing capacity. For lower-proficiency learners, conversely, the same TE might generate extraneous cognitive load by pulling attention away from the overall message at a point when available cognitive resources are already heavily engaged in basic decoding.
With respect to MWIs, the significantly higher viewing comprehension scores of Group 1 compared to Group 4 are in contrast with Majuddin et al. (2021), who found no significant disparity between the comprehension scores of the MWI-enhanced captioned viewing and the uncaptioned viewing conditions. However, this result receives partial support from one of the findings of Majuddin et al. (2021). According to that finding, when watching the video twice, the MWI-enhanced captioned viewing group non-significantly outscored the unenhanced captioned group in terms of viewing comprehension score.
A plausible theoretical explanation for why TE aided comprehension, rather than distracting from it, may lie in the Noticing Hypothesis (Schmidt, 1990) and its interaction with learners' proficiency level. For high proficiency learners (as in the current study), the additional processing load imposed by TE may be within cognitive capacity, allowing them to notice highlighted items without sacrificing comprehension of the overall message. The present results, therefore, suggest that TE may be beneficial specifically when learner proficiency is sufficient to accommodate the dual demands of comprehension and form-focused attention.
The lack of a significant difference between the comprehension score of the uncaptioned viewing group compared to the unenhanced captioned viewing group aligns with Montero Perez et al. (2014b) and Hsieh (2020) but differs from Majuddin et al. (2021), who observed that learners in the unenhanced captioned viewing condition significantly outperformed those in the uncaptioned condition. Moreover, this finding is in contrast with Montero Perez et al. (2013), who reported that captioned viewing had a significantly large effect size on listening comprehension.
Although not reaching statistical significance, the higher mean comprehension score of the uncaptioned viewing group compared to the unenhanced captioned viewing group is similar to Sydorenko (2010), whose findings indicated that uncaptioned viewing developed listening comprehension. However, this finding conflicts with Winke et al. (2010) and Pujadas and Muñoz (2024), who found that captioned viewing led to superior comprehension scores compared to uncaptioned viewing. One plausible explanation for this finding can be the high proficiency level of the participants. As stated earlier, Pujadas and Muñoz (2024) discovered that learner reliance on captions can vary according to their proficiency level. That is, lower-level learners may rely on captions for listening comprehension more than higher-level learners. Additionally, some research findings indicate that captions may cause frustration or hinder comprehension in higher-level learners (Leveridge & Yang, 2013).
The lack of a significant difference between the comprehension score of Group 3 and that of groups 1 and 2 is in line with Majuddin et al. (2021) and Hsieh (2020), respectively. Finally, our questionnaire data indicate that a great proportion of participants regularly watched English-language videos either with or without captions/subtitles, a finding that confirms the popularity of L2 audio-visual materials among Iranian English-major EFL learners. The finding that most participants used L2 captions while watching English videos (41.4%) is consistent with Ahrabi Fakhr et al. (2021), who found that Iranian EFL students watched English-captioned videos for almost three hours each week. However, in that study, students reported watching uncaptioned English videos more than L1-subtitled videos. This pattern differs from the viewing habits of the participants in the current study, where more participants reported using Persian subtitles (24.1%) compared to no captions/subtitles (20.7%).
Another interesting insight from the questionnaire is that while research suggests repeated viewing can enhance comprehension (Majuddin et al., 2021), not all individuals enjoy this repetition; therefore, personal preferences should be taken into consideration. Instead of watching the same video repeatedly, students could be encouraged to watch TV series or YouTube vlogs with similar topics. This way, relevant content is repeated without causing learners unnecessary boredom and frustration.
Conclusions and Implications
Given the inconsistent findings of previous studies regarding the effect of TE and caption condition on L2 viewing comprehension, the present study aimed to clarify the effect of four captioning conditions on Iranian EFL learners’ viewing comprehension. It contributes to a growing body of research on L2 viewing comprehension by demonstrating that TE of lexical items in captions does not necessarily compromise, and can in fact support, viewing comprehension, at least among high-proficiency EFL learners. This finding reframes a concern that has persisted in the literature (e.g., Hsieh, 2020; Majuddin et al., 2021), namely that drawing attention to lexical items may come at the cost of meaning comprehension. The results suggest this trade-off is not inevitable and that proficiency may act as a moderating variable. Another notable finding was that unenhanced captioned and uncaptioned viewing did not lead to significantly different comprehension outcomes, a finding that challenges the assumption that captions always aid viewing comprehension. Finally, the questionnaire results indicated that participants were used to watching English videos and most of them (41.4%) watched English videos with L2 captions.
The findings of the present investigation have some pedagogical implications. First, the presence of TE in captioned videos to draw learners’ attention to certain lexical items does not necessarily lead to a reduction in content comprehension. Conversely, it may positively impact their comprehension if students are of high language proficiency. Therefore, provided that learners’ L2 proficiency is high, teachers and materials developers can textually enhance target lexical items in the captions without worrying about a trade-off between attention to textually-enhanced items and video comprehension. Second, the lack of a significant difference in the mean comprehension scores of the unenhanced captioned viewing group and the uncaptioned group indicates that L2 teachers can encourage high-proficiency learners to turn the captions off while watching English videos, so that the learners become more independent and more confident viewers.
Like any study, this investigation had several limitations. One limitation was the fact that, because of using intact classes, the study had a quasi-experimental, rather than true experimental, design. In addition, although the sample size (n = 62 across four groups) was moderate, it was not particularly large for between-group comparisons, which may have limited the statistical power to detect smaller effects. Another limitation was that, due to time restrictions, it was not extended longitudinally. Furthermore, participants consisted of Iranian university students from a single university. To be able to generalize the results, future studies involving learners from other academic levels and cultural backgrounds are required because the comprehension patterns for other learner groups might be different. Finally, eye-tracking data are required to provide us with helpful insights into the way learners process textually-enhanced captions.
Acknowledgements
We extend our sincere gratitude to the students and colleagues who provided invaluable assistance during the data collection phase.