Skip to main content
PLOS One logoLink to PLOS One
. 2025 Dec 1;20(12):e0337195. doi: 10.1371/journal.pone.0337195

Cognitive load and visual attention assessment using physiological eye tracking measures in multimedia learning

Fatemeh Shahnabati 1, Atefeh Sabourifard 1, S Hamid Amiri 1, Alireza Bosaghzadeh 1, Reza Ebrahimpour 2,3,*
Editor: Vishal Bharmauria4
PMCID: PMC12668483  PMID: 41325323

Abstract

Effective multimedia content design can boost performance, capture visual attention, and optimize cognitive load. The current study employs eye-tracking technology to establish metrics to measure cognitive load, analyze visual attention allocation, and evaluate learners’ performance in English language learning. The study focuses on creating and comparing two different multimedia presentations. The differentiation between them lies in their adherence to or deviation from Mayer’s educational multimedia design principles: coherence, signaling, and spatial contiguity. participants were randomly assigned to two groups. The first group viewed with principles version, while the second group viewed without principles version, during which their eye movement data were collected. Subsequently, both groups participated in a recall test and completed the NASA-TLX questionnaire. The research establishes connections between specific eye-tracking parameters, subjective cognitive load scores, and recall test results through regression models and analyzes fixation distributions. The study also delves into microsaccades rate and changes in pupil size, each analyzed within times of interest. The study’s findings indicate that the examined metrics can significantly help distinguish between the two conditions: principles and no principles. These metrics are pertinent for assessing individuals’ cognitive load and visual attention and serve as beneficial indicators for gauging the efficacy of the designed multimedia content.

1. Introduction

In contrast to traditional educational methods reliant solely on textual information, educational multimedia offers an immersive learning environment that seamlessly integrates visual and verbal cues to enhance the learning experience. For instance, research indicates that university students who combined online instructional videos with traditional classroom teaching achieved superior outcomes compared to those exposed solely to conventional face-to-face instruction through an approach known as blended learning [1]. The effectiveness of educational multimedia is most pronounced when multimedia design and the integration of visual and verbal elements align harmoniously with the functioning of the human brain [2]. Consequently, the impact of educational multimedia can vary significantly based on its design and its influence on learning improvement [2]. Given the limitations of learners’ cognitive capacities, employing multimedia design strategies that effectively engage learners with educational content while minimizing the unnecessary cognitive load on working memory becomes paramount, ultimately leading to more meaningful learning experiences [3].

Cognitive load is a multifaceted concept reflecting the cognitive demands imposed by specific tasks on a learner’s cognitive system [4]. The Cognitive Load Theory (CLT) categorizes cognitive load into three types: intrinsic, extraneous, and germane. Intrinsic cognitive load, shaped by the interaction between the material and learners’ expertise, is influenced by the complexity of material interactivity [5]. This cannot be directly modified or minimized through instructional design. The more the number of elements that are related to each other in the educational materials, the more their complexity will increase; therefore, the intrinsic cognitive load will increase. Extraneous cognitive load results from poorly designed instructions, adding unnecessary burdens beyond intrinsic cognitive load. This additional load hinders learning and should be minimized to optimize cognitive resource allocation [6,7]. Germane cognitive load pertains to processes contributing to schema construction and automation, directly fostering learning. Both extraneous and germane loads are modifiable by instructional designers [6,7]. Educational multimedia must be meticulously crafted to alleviate unnecessary cognitive load, maximize cognitive resource allocation, and enhance working memory capacity, bolstering germane load and improving learning outcomes. Mayer’s work encapsulates 12 principles for effective multimedia presentations that mitigate undue cognitive load [8]. Three of these principles include: 1. Coherence Principle, 2. Signaling Principle, 3. Spatial Contiguity Principle. Three primary methodologies are employed to measure cognitive load: subjective rating, performance-based measures, and physiological measures [4,9,10]. Subjective rating methods assume that people can estimate their cognitive processes and report their mental efforts [11]. These methods typically hinge on questionnaires spanning either one-dimensional or multidimensional formats. While unidimensional scales capture overall cognitive load, multidimensional scales encompass diverse components like fatigue, mental effort, and frustration [12]. Notably, the NASA-TLX scale, assessing six dimensions, finds widespread application. Nevertheless, subjective rating methods exhibit limitations, including potential response exaggeration, individual biases, and the inability to track cognitive load fluctuations during testing. Performance-based measurement quantifies cognitive load based on individual task performance or behavior. For instance, the dual-task paradigm involves concurrently executing a secondary task alongside the primary task, with secondary task performance as a gauge for changes in cognitive load within the primary task. However, while performance-based methods facilitate real-time tracking of cognitive load variations, they often prove better suited to controlled laboratory environments, proving practically challenging or unfeasible for real-world application. Physiological methods, constituting the third category of cognitive load measurement, offer the capability to discern cognitive shifts via specific physiological variables (such as fMRI, electroencephalogram, and heart rate variability) [1315], showcasing real-time sensitivity and resolution [10]. EEG stands as a prominent tool leveraging brain activity to gauge cognitive load [16], through feature extraction from EEG signals, established a viable technique for assessing cognitive load in educational multimedia [16]. Farkish and colleagues by measuring the cognitive load in educational multimedia using EEG signals, showed that following Mayer’s multimedia design principles significantly reduces the cognitive load of users [17]. This underscores the potential of physiological measures to provide nuanced insights into cognitive load dynamics during multimedia engagement.

Eye-tracking technology has emerged as a vital tool across numerous studies for observing cognitive processes and gauging attention in multimedia education experiments [15]. Various aspects of eye movement behavior are categorized into distinct groups, including:

  1. Fixational and Saccadic Measures: This category encompasses metrics like fixation count, fixation duration, microsaccadic measures, saccade amplitude, saccade velocity, and saccade latency.

  2. Blink Measures: Blink rate and blink amplitude provide insights into visual engagement.

  3. Visual Search Measures: Parameters such as time to first fixation, gaze transition, scan path similarity, and dwell time shed light on visual exploration patterns.

  4. Pupil Measurements: Pupil size variation and The Index of Cognitive Activity (ICA) contribute to understanding cognitive processing.

Research has shown that these criteria offer valuable information about visual information processing, learning strategies, and visual search patterns and can even serve as indicators for assessing cognitive load. For instance, Chen and colleagues found that longer and slower fixations correlate with higher attentional effort demands [18]. The relationship between cognitive load and fixation duration can vary with the task; fixation duration might increase [19] or decrease [20,21] as mental workload intensifies. Similarly, the frequency of saccadic movements depends strongly on the nature of the task at hand [22]. Microsaccades, covering less than 1° of the visual angle, are crucial in preventing currently viewed visual information from fading [23]. Recent studies suggest a link between microsaccade occurrence and cognitive load. According to [24], greater task complexity corresponds to a decrease in microsaccade rate and an increase in microsaccade amplitude. Additionally, [25] revealed that the microsaccadic rate drops under high-load conditions of a memory task compared to low-load conditions. Conversely, other evidence indicates that microsaccadic frequency increases with the visual complexity of a task [26]. This underscores the potential of microsaccades as a window into cognitive processing during various cognitive tasks. The findings of [27] highlight a noteworthy pattern: Individuals tend to maintain their gaze on a specific area of an image or scene for an extended duration to enhance attention and minimize distractions. Consequently, a reduced blink rate emerges as an informative cue indicative of heightened attention and increased concentration on the scene. Another significant gauge linked to visual information processing is the frequency of gaze transitions between Areas of Interest (AOIs). This metric offers insights into whether learners effectively establish connections between different informational components or encounter challenges integrating diverse elements [28]. Diverse interpretations of this metric have emerged across studies [29]. It sometimes benefits cognitive processes such as organization and integration [30,31]. However, in other scenarios, excessive fixation transitions are viewed as detrimental to learning, signaling split attention [32]. The fluctuation of pupil diameter, believed to mirror changes in brain activity and human cognition, represents another fascinating avenue of research. Scholars have examined the correlation between pupil size, cognitive load, and varied tasks, encompassing activities like driving while engaged in conversation, solving mathematical problems, recalling numbers, and processing visual stimuli [3335]. This exploration underscores the dynamic relationship between pupil changes and cognitive engagement across diverse cognitive endeavors.

2. This study

Two primary objectives guided our study. Firstly, we aimed to uncover the impact of deviating from Mayer’s multimedia design principles on the cognitive load experienced by students’ working memory. Secondly, we wanted to have a systematic and methodological approach to measure cognitive load and processing and visual resource allocation eye tracking tools. We wanted to know if the eye behaviors differ between the two groups when the cognitive load changes due to watching different multimedia versions.

While numerous investigations have delved into multimedia design principles, firstly, their focus has centered on learning outcomes. We intended to explore the cognitive load resulting from principle violations using physiological eye tracking data. Secondly, within Mayer’s 12 proposed principles, most attention in eye-tracking studies has been directed toward modality and signaling. In contrast, certain principles, such as spatial contiguity and coherence, have received relatively limited scrutiny. For instance, a solitary study delved into gaze behavior during multimedia learning to probe the spatial contiguity principle [36]. However, this study failed to establish a significant contiguity effect on learning outcomes like retention or transfer test scores [37]. Furthermore, the effectiveness of video and animation-based training and dynamic content compared to static training content necessitates further investigation [37]. Equally pivotal is the investigation into the impact of adhering to or violating principles of educational multimedia design on eye behaviors. While eye behavioral measures have been applied in diverse studies, spanning mathematical calculations to human-computer interaction question comprehension, comprehending how observing or violating multimedia design principles influences eye behaviors warrants deeper exploration. For instance, while probing microsaccades or pupil responses could serve as avenues for understanding shifts in cognitive load, few studies have systematically addressed alterations in microsaccade-related metrics or pupil size during multimedia engagement. Moreover, the literature underscores contradictory findings regarding pattern changes and the susceptibility of many of these metrics to the task context. Thus, an enhanced understanding of these metrics, particularly within video-oriented multimedia education, can shed light on these intricacies and provide valuable insights moving forward.

3. Materials and methods

3.1. Participants

A total of 34 university students voluntarily participated in the study with an average age of 22.8 years (SD=±2.5). All of the participants were male. The experiment encompassed two distinct sessions. All 34 participants participated in the first session of the experiment, and 26 in the second session. Data from participants who had participated in one session were also used in the analyses. So, in total, we had 28 participants in the P condition and 28 participants in the NP condition. All participants have normal or corrected to normal vision and hearing. It’s worth noting that two participants were excluded due to calibration issues. Participants were eligible for inclusion in the study if they met the following conditions: they were between the ages of 20 and 30, had normal or corrected-to-normal vision, were native Farsi speakers, had no prior familiarity with the experimental materials, were right-handed, and obtained a score of 6–7 on the IELTS Simulator Test (their English listening skills were evaluated using the listening component of the IELTS -International English Language Testing System- simulation test). Participants were excluded from the study if they had prior experience or knowledge of the study material, had recently used alcohol or drugs that could impact cognitive performance, exhibited excessive eye blinking or head movements, showed signs of fatigue or sleep deprivation, or obtained an IELTS Simulator Test score lower than 6 or higher than 7. In adherence to ethical guidelines, all participants provided written consent before the experiment. Notably, the experimental protocols received approval from the ethics committee of the Iran University of Medical Sciences.

3.2. Materials

We utilized Lessons 6 and 11 from the Oxford Open Forum 3 [38] textbook to create 4 multimedia versions. The initial two versions, each lasting 290 seconds, were based on Lesson 6. One of these versions adhered comprehensively to five educational multimedia design principles (coherence, signaling, redundancy, spatial contiguity, and temporal contiguity) – referred to as the “P: with principles” condition. Conversely, the other version intentionally deviated from these principles – referred to as the “NP: without principles” condition. Similarly, the 2 subsequent versions spanned 342 seconds and corresponded to Lesson 11. In one version, the same five principles mentioned earlier were dutifully upheld (P). Conversely, the other version disregarded these principles (NP).

3.3. Experimental design

The experiment has 6 main stages, the full description of each stage is mentioned in the caption of Fig 1. The participants were divided randomly into two groups of equal numbers (G1 and G2). During the first session, G1 viewed the “P” version of Lesson 6 (P6), while G2 watched the “NP” version of Lesson 6 (NP6). Subsequently, in the second session of the experiment, G1, watched the “NP” version of Lesson 11 (NP11), and G2, viewed the “P” version of Lesson 11 (P11). Before commencing the experiment, all participants received a comprehensive explanation of the experimental procedure. Additionally, they engaged in a simulated experiment to acquaint them with the experimental conditions and methodologies. Once prepared, participants positioned themselves behind the monitor. Before the multimedia’s initiation, participants focused on a black circle positioned at the center of a gray screen for 10 seconds. This preliminary step facilitated the collection of eye data about the baseline state. Following this, the designated version from the pool of four multimedia was automatically presented. Afterward, an automatic recall test was administered, wherein participants answered test questions within 420 seconds. Subsequently, participants completed a paper-based NASA-TLX workload questionnaire.

Fig 1. Experimental design.

Fig 1

The experiment utilized two audio narrations to generate four video variations: P6, NP6, P11, and NP11. For each narration, two versions were produced—one adhering to multimedia design principles (labeled P6 and P11, marked in green) and one violating those principles (labeled NP6 and NP11, marked in purple). Participants were randomly assigned to two groups: Group 1 (G1) viewed the principled version of the first narration (P6) and the non-principled version of the second (NP11) across two sessions. Group 2 (G2) watched the non-principled version of the first narration (NP6) and the principled version of the second (P11). Experiment steps: First, English language level test (IELTS simulation test); second, randomly dividing the participants who got scores ranging from 6 to 7 into two equal groups; third, looking at a black-filled circle for recording baseline data; fourth, watching the multimedia (no interaction); fifth, taking part in the recall test (via mouse interface); and finally completing the NASA-TLX questionnaire (paper-based version). From the third step, all the steps are completely identical in both sessions. In the third and fourth steps, eye-tracking data are collected. All of these steps were carried out completely consecutively and immediately.

According to the cognitive theory of multimedia learning, five strategies are employed to manage cognitive overload and minimize redundant processing. In our eye-tracking data analysis, we concentrated on three of these strategies within the context of multimedia content. By deliberately violating these principles in the “NP” multimedia version, we aimed to introduce an increased extraneous cognitive load. The three principles are as follows:

  1. Spatial Contiguity Principle: This principle underscores the importance of arranging text and related graphical content closely to minimize the visual search and the cognitive effort required for scanning back and forth between the text and graphics. In the “NP” multimedia version, we deliberately spatially separated corresponding images and text descriptions within specific frames (see Fig 2-A). This manipulation is anticipated to prompt subjects to continually shift their focus between the graphical and text components as they strive to comprehend the content presented within the scene within a limited timeframe.

  2. Coherence Principle: This principle advocates eliminating extra content that could distract learners. Only the pertinent information required for mastering the main subject should be included. The coherence principle optimizes cognitive resources by eradicating redundant information [2]. Accordingly, in the “NP” multimedia version, we deliberately introduced additional and unnecessary content in various parts of the scenes within certain frames (refer to Fig 2B). We anticipate that learners’ cognitive resources will be distributed among the diverse types of information presented on the screen during multimedia viewing. Consequently, some cognitive resources will be expended to ignore extraneous details [39].

  3. Signaling Principle: This principle dictates that crucial information should be highlighted to guide learners’ attention. By emphasizing significant content, the signaling principle steers learners toward focal points, allowing them to disregard irrelevant details and allocate cognitive resources more effectively to process vital information [2]. For instance, a recent study underscored the enhanced learning outcomes achieved by integrating visual cues in non-procedural activities within educational videos [40]. Fig. 2C illustrates the signaling principle’s application or violation within frames of both the “P” and “NP” multimedia versions.

Fig 2. Educational multimedia.

Fig 2

Three Examples of application (i.e., right panels)/violation (i.e., left panels) of multimedia design principles. Panels A, B, and C are respectively related to the application or violation of the principles of spatial contiguity, coherence, and signaling.

To ensure the validity of extraneous load manipulation, the NASA-TLX workload index was employed to gauge the workload of participants in both G1 and G2 groups. Renowned for its reliability, this index comprehensively evaluates subjective workload. It consists of six facets: mental pressure, physical pressure, time pressure, performance, effort, and frustration. Given that each facet contributes differently to workload during a specific task, the raw score of each facet is multiplied by its corresponding weight. These weights are determined through 15 pairwise comparisons completed by the participants. The cumulative scores of all six facets are then combined and divided by the number of pairwise comparisons (15) to yield the total workload score [41]. This comprehensive approach guarantees a reliable assessment of the workload experienced by participants in each experimental group.

3.4. Eye movement measures

Violating coherence and spatial contiguity principles within multimedia design leads to visual and textual content dispersion throughout the scene. To comprehensively explore the impact of these violations on eye behaviors, we employed the Total Number of Fixations (TNF), Mean Fixation Duration (MFD), and Mean Saccade Amplitude (MSA) measures within specific time of interest (TOI) segments. These measures collectively provide insights into factors like visual effort exertion, the efficacy of visual search, and shifts in additional cognitive load. Below, we provide detailed explanations for each of these metrics:

  1. TNF: This metric quantifies the total fixations occurring across the screen within a designated TOI [42].

  2. MFD: MFD represents the average duration of all fixations observed within the TOI across the complete screen [39].

  3. MSA: MSA denotes the average length of saccades within the TOI across the entire screen [39].

In evaluating concepts such as the findability and noticeability of vital content, as well as the extent of attention on specific screen areas, we calculated the following metrics within AOIs containing crucial learning content:

  1. Time to First Fixation (TFF): TFF gauges the time between stimulus presentation and the initial gaze directed toward that stimulus. This metric quantifies bottom-up attention originating from stimuli and assesses visual search’s effectiveness in evaluating user interfaces [4344].

  2. Dwell Time (DWT): DWT measures the cumulative duration participants fixate on a specific AOI. It offers insights into the sustained attention dedicated to particular content within the AOI.

Fixation Distribution Map These visual representations showcase the distribution of gaze across various AOIs, providing an intuitive depiction of participants’ focus and engagement. This map is constructed by smoothing fixation distributions with Gaussian kernels. It comprises a 3D representation generated from 2D fixation maps, relying on x and y coordinate locations for fixations, while the third dimension represents fixation intensity (weighted by fixation count). This map provides a comprehensive depiction of fixations across the AOIs [42,45].

By leveraging these metrics, we can comprehensively understand participants’ visual behavior, the attention directed toward significant content, and the cognitive processes accompanying multimedia engagement.

In addition to these metrics, we employed changes in pupil size and microsaccade rate within the designated TOIs corresponding to the observation or violation of the three principles mentioned earlier.

Pupillary Responses to Cognitive Load Pupillary responses denote changes in pupil size in response to cognitive demands or mental effort. An increased mental effort typically leads to pupil dilation. Pupillary responses are utilized in research to assess mental workload and attention and are sensitive to diverse cognitive tasks and mental states [4647].

Microsaccades Responses to Cognitive Load: Microsaccades, akin to pupillary dilation, can be investigated from an information processing perspective. Numerous studies indicate an inverse relationship between microsaccade rate and cognitive load. Siegenthaler and colleagues, for example, established that heightened task complexity correlates with a reduction in microsaccade rate [24].

These comprehensive metrics collectively facilitate a deeper understanding of participants’ cognitive processes, attention allocation, and engagement dynamics during multimedia interactions.

3.5. Recall test

Given the likelihood of learners’ performance being influenced by the design of the educational video they watched, we employed a recall test after multimedia viewing to gauge and compare the learning outcomes between the “P” and “NP” conditions.

Upon completion of the multimedia engagement, the software integrated into the laboratory computers automatically initiated a recall test. This test comprised 12 multiple-choice questions about the content covered within the multimedia. The questions were meticulously crafted to encompass crucial points from the material, facilitating a comprehensive evaluation of participants’ grasp of the subject matter. Respondents were required to provide answers using a mouse, navigating between questions while being restricted to viewing one question at a time. A time frame of 420 seconds was allotted for answering the questions. Participants could leave questions unanswered or modify their responses during this duration. This recall test methodology assessed participants’ comprehension and learning progress.

3.6. Data analysis

We employed a video-based monocular eye tracker, specifically the Eyelink-1000 Plus model, to track eye position and pupil size. Stimuli were presented on a 17-inch LCD panel with a 1280 × 1024 pixels resolution. The sampling frequency of the eye tracker was set to 1000 Hz and data were obtained from both eyes of the participants. The experiment was conducted in a dimly lit room with no direct light sources. We used a standard 9-point calibration. It required participants to fixate on a series of known points, allowing us to map eye positions to specific locations.

In line with established practice, fixations of less than 90 ms were deemed too brief for meaningful information processing and, thus, were excluded from the analysis. because they were more likely to be artifacts, noise, or involuntary eye movements. Following existing research methodologies [4851], we employed linear interpolation scipy’s interpolate module to replace pupil values during blink episodes in which we had some missing data. Furthermore, the raw pupil signals underwent smoothing using a low-pass digital filter to enhance signal quality. We used Savitzky-Golay filter, that fits a polynomial to a local window of data using least square regression. We set the degree of the polynomial to 3 which provides better results in rapid changes while still avoiding overfitting. Regarding the pupillary response to changes in luminance, it’s noteworthy that only a single fixed frame was presented within all the selected TOIs. Moreover, the brightness of these frames remained consistent between the P and NP conditions. This brightness uniformity was supported by the outcome of two-tailed paired t-tests conducted on the brightness of the frames. The result of this test demonstrated no significant difference (t = 0.243, p = 0.83) between the two conditions’ frame brightness.

It’s also crucial to acknowledge that the frames preceding the TOIs’ onset were identical across both the P and NP conditions. Consequently, considering the uniformity of the frames and the lack of significant difference in brightness, the expectation was that the variations in pupil size attributable to lighting conditions between the two conditions would not exhibit substantial differences. For the detection of microsaccades, we adapted Engbert and Kliegl’s algorithm [52]. Our choice of parameters included a velocity threshold of λ = 5 and a minimum duration threshold set at 6 ms. To evaluate pupillary responses and microsaccade rates, we executed paired t-tests to test our hypothesis that manipulating cognitive load and allocating attentional resources trigger pupillary and microsaccadic responses. We conducted an FDR (False Discovery Rate) correction to mitigate the possibility of false discoveries, ensuring a controlled probability of such errors. By adhering to these rigorous methodological steps, we ensured a comprehensive and reliable analysis of pupillary responses and microsaccade rates in response to changes in cognitive load and attentional allocation strategies.

4. Results

This section comprehensively analyzes participants’ performance, subjective cognitive load estimation, and collected gaze data from two experimental sessions.

4.1. NASA-TLX and recall test

To assess the differences between the P and NP conditions, we conducted an independent samples t-test on the students’ NASA scores and recall performance. To assess the normality of the data distribution visually, Q-Q (Quantile-Quantile) plots were generated. The points closely followed the reference line, suggesting no substantial deviation from normality. For the NASA-TLX scores, skewness (0.45) and kurtosis (0.37) fell within the acceptable range (−1 to +1), while for the Recall scores, skewness (−0.47) and kurtosis (0.6) were also within this range. These results indicate approximate symmetry and a mesokurtic distribution. The normality of the data was also assessed using the Kolmogorov-Smirnov test. The results indicated that the data were consistent with a normal distribution (the tests have more than 0.05 p-value suggesting that the data are consistent with normality). Levene’s test for homogeneity of variance indicated that the variances were equal across groups, suggesting that the assumption of homogeneity of variance was met. We anticipated observing meaningful distinctions between these two conditions. Illustrated in Fig 3, the participants in the NP condition indicated a notably higher workload than those in the P condition (t(54) = −6.223, p = 1.06E-08). Furthermore, the P condition performed significantly better on the recall test than the NP condition (t(54) = 5.325, p = 2E-06). To compare the two conditions, the non-parametric Mann–Whitney U test was also applied. For the NASA scores, the test yielded U = 74, p < 0.001, indicating a statistically significant difference between the conditions. Similarly, for the Recall scores, U = 674.5, p < 0.001, also indicating a significant difference.

Fig 3. Result of NASA-TLX and Recall test.

Fig 3

The scores are in the range of 0 to 100. Middle line represents the median (50th percentile) of the data. Whiskers (error bars) extend to the minimum and maximum values within 1.5 × IQR from Q1 and Q3.

4.2. Regression analysis

In the initial eye movement data analysis, we identified the TOIs which were pertained to observing or violating coherence and spatial contiguity principles. Subsequently, for each participant, we quantified three metrics: TNF, MFD, and MSA within each TOI.

To delve deeper, we employed linear regression analysis to explore potential connections between these metrics and participants’ performance in the test and their perceived cognitive load. In order to assess the assumption of normality for the residuals in our linear regression model, we conducted the Shapiro-Wilk test. The results showed that the residuals did not significantly deviate from a normal distribution (all p-values > 0.05). Thus, we conclude that the residuals of the model are approximately normally distributed. The complete results of regression analysis are comprehensively presented in Table 1 and Table 2.

Table 1. Summary Table of linear regression analysis results for predicting NASA-TLX.

Eye Movement Measures NASA – TLX R2
TNFa 5.275** 0.162
MFDb −4.658** 0.133
MSAc 7.617*** 0.356

aconstant: 37.450

bconstant: 49.281

cconstant: 49.281

Table 2. Summary table of linear regression analysis results for predicting recall scores.

Eye Movement Measures Recall Scores R2
TNFa −8.294*** 0.246
MFDb 8.655*** 0.283
MSAc −8.315*** 0.261

aconstant: 91.519

bconstant: 72.916

cconstant: 72.916

Table 1 illustrates a significant correlation between NASA-TLX scores and all three measures: TNF, MFD, and MSA. This correlation suggests that as the number of fixations increases and saccade sizes become larger, while fixation durations decrease within TOIs, participants experience heightened cognitive and mental load on their working memory. This underscores the sensitivity of the cognitive load perceived by individuals to all three metrics.

Furthermore, as indicated in Table 2, learners with lower TNF and MSA measures achieved notably higher scores on the performance test. Conversely, learners with higher MFD measures achieved significantly lower scores. This implies that an increased number of fixations within limited TOI and extensive movement across distant screen areas can have a relationship with decreasing fixation duration. This, in turn, can impact on overall cognitive load, resulting in poorer performance on the test.

Our subsequent analysis involves eye movement data from TOIs associated with the observation or violation of the signaling principle. In these TOIs, essential graphical content for learning is highlighted through techniques such as motion, color changes, and graphical cues to attract users’ attention. We identified AOIs within each frame to pinpoint these critical components. We measured TFF and DWT for each TOI. We performed a logistic regression analysis to understand the relationship between these measures and recall scores.

Concerning the content within each TOI, a corresponding question was included in the recall test for Lesson 6. Correct responses to these questions were assigned a score of one, while incorrect answers received a score of zero. The outcomes of this analysis are documented in Table 3. The regression model employed as follows:

Table 3. Summary table of logistic regression analysis results for predicting recall scores.

Eye Movement Measures Recall Scores
DWT 1.668*
TFF −1.565**

aintercept: −1.307

b* P < .05 ** P < .01 *** P < .001

Test score = B0 + B1 × DWT + B2 × TFF (1)

Based on the outcomes of this analysis, where the pseudo-R2 value is 0.072, noteworthy relationships have been identified. A significant positive correlation emerged between learning test scores and DWT, while a significant negative correlation was observed between learning test scores and TFF. This suggests that students who quickly focus on AOIs and allocate more time to observe them tend to perform better in the final learning test.

Continuing our exploration, we undertook a linear regression analysis to delve into the connections between DWT and TFF regarding compliance/violation of the signaling principle. The results of the regression analysis unveiled substantial associations. Specifically, a positive and significant relationship was established between DWT and the Observing signaling principle (F(1,166) = 27.99, R2 = 0.144, B = 0.185, p < 0.001). Conversely, TFF exhibited a negative relationship with the Observing signaling principle (F(1,166) = 165.4, R2 = 0.499, B = −0.443, p < 0.001).

These findings highlight the predictive nature of observing or violating the signaling principle about DWT and TFF among students. Moreover, these results provide compelling evidence of the significant influence of the “cuing effect” on learning within educational content.

4.3 Microsaccadic rate

Initially, we validated the accuracy of the algorithm in detecting microsaccades. To accomplish this, we executed a correlation analysis involving peak velocity and amplitude, a measure referred to as the “main sequence.” This analysis is grounded in the established understanding that microsaccades exhibit a positive correlation between these two parameters, like regular saccades.

Our investigation yielded a notably strong positive correlation (r = 0.913, p < 0.001), affirming the correct identification of microsaccades. This finding provides confidence in the algorithm’s ability to detect microsaccades, as depicted in Fig 4, accurately.

Fig 4. Microsaccadic peak velocity vs. magnitude (main sequence).

Fig 4

This correlation established that microsaccades were correctly detected.

We investigated the impact of compliance/violation of principles on microsaccade responses. We calculated the microsaccadic rate within a 100-ms temporal interval, commencing from initiating the TOIs. As part of this analysis, we established a baseline rate by averaging the rates within the epoch spanning −100 to +100 ms relative to the onset of the TOIs. This involved the individual computation of microsaccadic rates for each participant, followed by the averaging of results across all participants.

For a more comprehensive understanding of microsaccade behavior, we conducted multiple comparisons, corrected for false discovery rate (FDR), between the mean microsaccadic rates observed in the P and NP conditions. These comparisons were performed within a 300-ms moving window, initiated at the onset of the TOIs. This methodology aligns with similar approaches documented in references [5354].

The results revealed that the rate of microsaccades during TOIs was notably lower in the NP condition compared to the P condition. This observation underscores distinct differences in microsaccade patterns between the two conditions in response to the manipulation of principles. As depicted in Fig 5, specifically in panel A, comparing baseline and initial microsaccade rates following the commencement of the TOIs revealed no significant differences between the two conditions. However, as we progressed within this time interval, the rate of microsaccades experienced a decline. It decreased from approximately 2.2 to around 1 within the NP condition. Notably, throughout the TOIs, the microsaccade rate in the NP condition consistently remained significantly lower than in the P condition. This pattern persisted until the conclusion of the TOIs, marked by the gray vertical line. It’s worth noting that all FDR-corrected comparisons indicated significant differences between the two conditions (ps < 0.02 and ts > 7.36), except for the first two 300-ms time windows (0–300 ms and 300–600 ms), where the comparisons showed non-significant results (ps > 0.18 and ts < 1.41), as represented by the red bars beneath the graphs.

Fig 5. Microsaccadic rate calculated within the TOIs.

Fig 5

Shaded areas indicate the standard error of the mean. Green and red rectangles above the x-axis indicate the 300-ms time windows (FDR-corrected) used for statistical testing. Asterisk denotes a significant difference between the two experimental conditions, while “ns” means the difference was non-significant.

Moving to panel B, during the baseline state, the microsaccade rate of the NP condition was initially higher than that of the P condition. However, after the initiation of the TOIs, the microsaccade rate of the NP condition experienced a decline, evident around the 800-ms mark. At the onset of this suppression, the microsaccade rates in both condition were equal. Subsequently, the microsaccade rate of the NP condition continued to decrease. Consequently, the frequency of microsaccades in the NP condition gradually diminished, sustaining this gradual decline until the conclusion of the TOIs.

In terms of statistical significance, within the NP condition and P condition comparisons, there were three 300-ms time windows (600–900 ms, 900–1200 ms, and 1200–1500 ms) where the FDR-corrected comparisons did not demonstrate significant differences (indicated by the red bars in Fig 5B; ts < 2.10, ps > 0.17). Conversely, outside these time windows, all comparisons reached canonical levels of significance (indicated by the green bars in Fig 5B; ts > 15.92, ps < 0.02). In general, within TOIs where the principles of multimedia design are not adhered to, there is a noteworthy reduction in the microsaccade rate compared to the condition following the principles. This decrease is observed concerning the baseline level. Notably, the TOIs depicted in Fig 5A and Fig 5B pertain to coherence and spatial contiguity principles. Conversely, no significant difference is detected in the microsaccade rate between the conditions when considering TOIs associated with the signaling principle.

Fig 5C is dedicated to evaluating compliance or violation of the signaling principle in one of the testing and observation conditions (TOIs). As the illustration reveals, there is no discernible pattern in the fluctuations of microsaccade rates within the two conditions. Additionally, no statistically significant distinction is observable between these rates.

4.4. Pupil size

An increase in pupil size could be a sensitive indicator of mental effort and cognitive load. To test this hypothesis, the following approach was employed:

  1. Baseline Calculation: A baseline value was established by averaging the pupil size data from −200 ms to +100 ms relative to initiating the TOIs. This baseline value was used as a reference point for subsequent analysis.

  2. Baseline Correction: Each data point was subtracted from the corresponding baseline value, following the application of the baseline correction procedure. This procedure was executed independently for each participant regarding individual differences, allowing for accurate comparison.

Continuing the analysis, pupil size within the TOIs underwent a similar procedure as that of the microsaccade rate. Specifically, ten FDR-corrected comparisons were conducted between the mean pupil size changes for the P and NP conditions. These comparisons were carried out using a 300-ms moving window, commencing at the onset of the TOIs.

In Fig 6A, the pupil size changes observed in the NP condition exhibited a distinct pattern. Roughly 1500 ms following the initiation of the TOI, a pronounced expansion in pupil size peaked at around 2000 ms, representing a notable increase of 300 units compared to the baseline pupil size. Following this expansion, there was a subsequent drop, leading to a suppression of pupil size until approximately 3800 ms. This suppression eventually reached a relatively stable state. In contrast, the pupil size changes in the P condition did not demonstrate significant variations compared to the NP condition. Across all 300-ms windows, the results of FDR correction consistently demonstrated significant differences between the two conditions. When the pupil size of the NP condition was smaller than that of the P condition, p-values were less than 0.001, and t-scores were greater than 71.43. Conversely, when the pupil size of the NP condition was larger than that of the P condition, p-values were less than 0.001, and t-scores were less than −10.

Fig 6. Change in pupil size calculated within the TOIs.

Fig 6

Shaded areas indicate the standard error of the mean. Green and red rectangles above the x-axis indicate the 300-ms time windows (FDR-corrected) used for statistical testing. Asterisk denotes a significant difference between the two experimental conditions, while “ns” means the difference was non-significant.

Both conditions’ graphs in Fig 6B displayed a rapid expansion following a minor contraction. This expansion reached its zenith approximately 1000 ms after the commencement of the TOI. Subsequently, there was a gradual decrease, around 2500−3000 ms, after the TOI began, leading to an approximately stabilized state. Notably, the size of changes in the NP condition was more substantial, featuring a more pronounced expansion than the P condition. Throughout the TOI, significant differences between the two conditions were consistently observed. Specifically, within the first 300-ms window, FDR correction results indicated p-values below 0.001 and t-scores exceeding 8.80. Beyond the initial window, for the remaining time windows, p-values were below 0.001, and t-scores were below −83.06. In panel C of the analysis, distinct patterns in pupil size changes were observed. In the NP condition, a momentary suppression was noticeable, succeeded by a rebound that occurred around 950 ms and peaked approximately 1500 ms after the onset of the TOI. Conversely, in the P condition, the pupil size changes increased from the start of the TOI, reaching its peak at roughly 1000 milliseconds. Subsequently, there was a gradual decline, leading to a stable state at around 2000 milliseconds.

The FDR correction results for various time windows revealed significant differences between the NP and P conditions. During the time windows when pupil size changes were smaller in the NP condition than in the P condition, p-values were below 0.001, and t-scores exceeded 9.20. For time windows where pupil size changes were larger in the NP condition compared to the P condition, p-values remained below 0.001, and t-scores fell below −55.23.

These observations collectively indicate that larger changes in pupil size were consistently observed in the NP condition compared to the P condition across all TOIs associated with the three principles of contiguity, coherence, and signaling.

4.5. Fixation distribution map

We analyzed fixation distributions in signaling TOIs for all participants, creating fixation distribution maps. The outcomes are illustrated in Fig 7, where fixation distribution maps for both participant conditions are displayed across panels A, B, C, and D. To illuminate disparities in attention allocation, we subtracted the gaze distribution maps of the P condition from those of the NP condition in the third column of the first row. The second row showcases frames associated with TOIs, denoted by black boxes marking AOIs. Notably, certain AOIs linked to the signaling principle in this study exhibit dynamic attributes, such as motion or color changes, which may need to be discerned in static images.

Fig 7. Fixation Distribution Maps.

Fig 7

Axes values represent number of fixations. Each panel shows a TOI. The first and second images in each panel are related to the fixation distributions of the two P and NP conditions. Higher fixation probability is represented by yellow. The third picture shows the difference probability maps of the P-NP, with yellow (positive values) indicating a higher fixation probability for P than NP. The fourth and fifth pictures (two pictures of the second row in each panel) show Multimedia frames presented in TOIs for P and NP conditions.

In Fig 7A, during the discussion of India’s economic growth, the arrow in the image gradually transitions to red within the P condition’s multimedia. Evident from the 2D density map of condition P, this multimedia element effectively captures learners’ attention. Fig 7B involves the explanation of the geographical locations of Indian states. Within the P condition’s multimedia, this concept gains emphasis by placing a compass on the map, which intermittently flashes toward the northwest direction. In contrast, the NP condition’s multimedia solely displays the country’s map.

Fig 7C entails the speaker’s account of their journey to the subcontinent, wherein the P condition’s multimedia highlights this episode through a red marking, accompanied by the appearance of the word “subcontinent” within the image. Fig 7D discusses local winds known as monsoons. In condition P, the key content’s prominence is achieved through the dynamic movement of curves representing the winds, accompanied by displaying the wind’s name adjacent to these curves.

As anticipated, across all the TOIs examined in this analysis, the count of fixations within the designated AOI was notably higher for the P condition. In this condition, the main content was specifically highlighted to capture learners’ attention. The P condition consistently demonstrated this superiority in focused attention on crucial screen areas compared to the NP condition. Conversely, participants in the NP condition exhibited more scattered fixations, likely due to their inability to identify the primary content.

4.6. Visualizing gaze behavior

To gain insights into players’ visual attention and interaction patterns during watching multimedia, we analyzed eye-tracking data using scan paths and fixation distribution map. Scan paths illustrate the sequential movement of gaze, highlighting the order and transitions between fixations, in Fig 8, you can see an example of a scan path of a TOI in which the spatial contiguity principle observed/violated.

Fig 8. Scan the path of two subjects while watching multimedia with (left) and without (right) the application of the Spatial Contiguity Principle.

Fig 8

Fixation distribution maps provide an aggregated view of fixation density across different regions of interest. These visualizations allow us to identify key areas that captured user’s attention, common gaze patterns, and variations in attentional focus based on dynamics of Compliance and non-compliance with multimedia design principles. In Fig 9 we present the result of fixation distribution map analyses, demonstrating how users visually engage with the multimedia.

Fig 9. Attention distribution of two subjects while watching multimedia with (left) and without (right) the application of the Coherence Principle.

Fig 9

Cronbach’s alpha was computed to assess the internal consistency of the eye movement metrics utilized in this study. Given the expected variability in eye movements between the P and NP conditions, separate analyses were performed for each. As shown in Table 4, the results indicate a high reliability level across most tested eye movement metrics.

Table 4. Cronbach’s Alpha values for the eye movement measures.

Eye Movement Measures Cronbach’s alpha
P NP
TNF 0.750 0.890
MFD 0.711 0.620
MSA 0.710 0.618
TFF 0.844 0.813
DWT 0.266 0.201

5. Discussion

The primary objective of this study was to examine the impact of adherence to or divergence from Mayer’s multimedia design principles on learning effectiveness, cognitive load experienced by learners, their eye behavior, and, subsequently, the allocation of attention to various presented content. This goal was accomplished by analyzing participants’ eye-tracking data, performance test outcomes, and responses to the NASA workload questionnaire.

The NASA-TLX questionnaire was initially employed to gauge subjective cognitive load during multimedia and the findings indicate that non-compliance with multimedia design principles significantly heightened the perceived cognitive load among individuals.

Likewise, in examining recall test results, we assert that the educational content’s design influences student performance on the recall test. Upon evaluating the test scores, a distinct pattern emerged and the results underscore the robust connection between recall test scores and subjective cognitive load as measured by the questionnaire; both were closely tied to adherence or non-adherence to multimedia design principles. Subsequently, we employed linear regression analysis to assess the relationship between eye movement metrics and Recall and NASA test scores.

While prior studies [19,5558] have proposed that fixation-related metrics can indicate mental workload, attention, and search efficiency, divergent outcomes have been observed in different research findings. Nevertheless, the sensitivity of fixation-related metrics to cognitive load hinges upon factors such as design methodology and content complexity [59] As noted by Jacob and Karen [42,60], the greater the number of fixations becomes, the less the viewer’s search for information on the computer screen is assumed to be.

The elevated count of fixations observed within the NP condition can signify diminished target findability or reduced target accessibility, attributed to numerous unrelated distractor elements where coherence principles are violated [6162]. The elevated count of fixations also arises from frequently shifting gaze between disparate visual and textual information sources. A depiction of this can be found in Fig 8, representing the scan paths of two subjects from both P and NP conditions concerning the spatial contiguity principle.

Findings from various studies indicate divergent patterns in how fixation durations change under conditions of high cognitive load. Some studies have shown that increased cognitive load leads to longer fixation durations as individuals spend more time processing information [63]. This is typically observed when individuals are faced with complex or unfamiliar stimuli that require additional cognitive resources for understanding or decision-making. In such scenarios, the brain’s need for greater information processing may manifest as longer fixations, particularly on critical aspects of the task.

On the other hand, other research suggests that higher cognitive load can lead to shorter fixation durations in some contexts [57,64]. This is often seen in tasks where individuals are overwhelmed by excessive information or are under time pressure. In such cases, participants might demonstrate a more superficial level of processing with shorter, more frequent fixations, attempting to process information rapidly without delving deeply into any one element.

Several factors can contribute to these varying patterns of fixation duration under high cognitive load like:

Task Complexity: More complex tasks that require deep cognitive processing tend to lead to longer fixations, while tasks with simpler or familiar components might result in shorter fixations.

Information Overload: When cognitive resources are taxed by excessive or overwhelming information, the individual may reduce the duration of each fixation to quickly scan and process multiple elements.

Our experiment’s dynamic, video-format educational content imposed temporal constraints on frame and content display. Consequently, the anticipated increase in fixations within the NP condition was expected to lead to shorter fixation durations, aligning with the outcomes observed. The reduction in fixation duration in the NP condition should not be misconstrued as an indication of content simplicity or ease of processing. Rather, it points toward shallower and more superficial processing [39].

Hyona [65] proposes one approach to assess users’ visual search efficiency and ability to identify crucial on-screen information: analyzing the dispersion of fixations. In the NP condition, the fixations were widely dispersed and distanced from each other due to the diffusion of distinct content elements across the screen. The elongated saccades in the NP condition stemmed from the spatial dispersion of various content components, as illustrated in Fig 9, depicting the Attention distribution heatmap of two subjects during the multimedia presentation. The violation of the spatial contiguity principle led to a separation between graphical and textual content, confounding learners and necessitating transitions between different elements.

It’s noteworthy that, contrary to general trends, an increase in cognitive load results in shorter saccade lengths [6667] and increased average fixation durations [68]. However, as explained earlier, our counterintuitive results within the NP condition don’t signify simplified processing, reduced mental load, or efficient visual searches. Drawing upon references [8,69], two pivotal strategies for reducing extraneous cognitive processing, thereby augmenting resources for essential and generative processing, are the spatial contiguity principle and the coherence principle. It is posited that when content is presented disparately, greater visual searching becomes imperative, necessitating cognitive resources for upholding individual elements within working memory before their mental integration. This influx of extraneous processing consumes resources that would otherwise be available for crucial and generative processing.

Furthermore, incorporating seductive details within multimedia can divert students’ attention, complicating the organization and integration of vital learning content. Consequently, introducing extraneous material amplifies extraneous processing, depleting cognitive working memory resources indispensable for essential and generative processing.

Consequently, the violation of these multimedia design principles subjects users to many challenges, including working memory overload, disregarding seductive details and managing visual search, and the endeavor to sustain focus on primary content while processing information. We also employed logistic regression analysis to explore the influence of signaling principles on individuals’ learning and performance. As elucidated in the literature [42], TFF metric enables the assessment of AOI noteworthiness and findability, indicating how swiftly an AOI captures attention.

Given that AOIs were observed faster in the P condition’s multimedia than in the NP condition, we can infer that adherence to the signaling principle rendered vital areas conspicuous. These zones’ accentuation and prompt detection translated into enhanced test performance. These outcomes resonate with findings from earlier investigations [8,7072].

According to the result obtained for DWT metric users are drawn towards AOIs through the signaling principle, comprehending that these areas furnish significant information. The prolonged time dedicated to AOIs arises from the processing time allocated to these pivotal regions, ultimately culminating in superior performance in the learning test. These findings align with previous research [8,7072].

This utilization of logistic regression analysis enables a deeper understanding of the impact of signaling principles on individuals’ learning and performance, affirming the significance of these principles in shaping effective multimedia design.

To explore the connection between mental effort and cognitive load with the count of microsaccades and pupil size variations compared to baseline levels, we assessed these parameters within TOIs related to spatial contiguity, coherence, and signaling principles. The findings are summarized below:

Microsaccade Rate: In TOIs corresponding to spatial contiguity and coherence principles, the microsaccade rate exhibited a gradual decrease within the NP condition. Prior studies have demonstrated that task load exerts influence over the microsaccade rate. Specifically, increased visual load correlates with an elevated microsaccade rate, while augmented mental load is associated with a diminished microsaccade rate [26,53,73]. Also, in the TOIs related to violating the signaling principles for the NP condition, we do not observe significant changes. Consequently, further research is needed to validate these findings based on the microsaccade rate in the context of multimedia learning and complex video-based representations.

Changes in pupil size: Across all analyzed TOIs, a significant dilation of pupil size was observed in the NP condition, occurring approximately after ~1000 ms from initiating the TOIs. It is worth noting that, as we mentioned before, the frames preceding the TOI start are the same in both conditions. Additionally, the frames associated with TOI demonstrated no significant disparity between the two conditions concerning brightness levels. Consequently, the pronounced difference in pupil size alterations between the two conditions stemmed from an alternate factor, namely cognitive load and the participants’ mental engagement. These outcomes align seamlessly with prior research findings [7477].

As indicated by the results concerning the 2D density map, our expectations were fulfilled. Using the signaling principle successfully directed learners’ attention to vital content within all TOIs. The elevated count of fixations within the AOI within the P condition, relative to the NP condition, underscored the conspicuousness and ease of locating the designated AOI in the P multimedia. Consequently, it reduced the necessity for extended visual searching to identify them.

In a broader context, it is plausible to argue that under the premise of dual coding and the Cognitive Theory of Multimedia Learning (CTML), applying the signaling principle helps students eliminate extraneous content. Simultaneously, the signaling principle adeptly organizes pertinent graphical and textual information, diminishing the need for exhaustive scene exploration and enhancing focus on principal themes. This strategic approach optimizes working memory capacity and significantly alleviates cognitive load during multimedia consumption.

6. Conclusion

In this study, we investigated the effectiveness of multimedia educational content design principles using eye behavior data as a physiological measure, the NASA-TLX test as a self-reporting measure, and the recall test as a performance measure. The possibility of evaluating the cognitive load imposed on people’s working memory and the amount of visual attention was investigated using different analyses. First, in order to investigate the effect of violation/observance of the design principles, we extracted various features from the eye movement data to measure various concepts such as visual attention, cognitive load, findability, and visual search strategy of users in the scene. Then, by applying linear and logistic regression models, we investigated the relationship between these eye-tracking measures and NASA-TLX scores and recall tests. This study also delves into microsaccade rates and changes in pupil size, dissecting learners’ eye behavior in response to compliance with or violating the principles. Each of these analyses was performed separately within distinct TOIs. We analyzed fixation distributions to show the concentration of fixations in certain areas of the scene. These analyses can significantly help educational multimedia designers in creating multimedia by imposing the optimal amount of cognitive load on learners’ working memory and drawing more attention to critical content.

7. Limitations and areas for future research

This study possesses several limitations that warrant consideration. Notably, the potential influence of visual stimuli on microsaccade rate alteration should be noticed, given its significance alongside cognitive processes. The observed variations in microsaccade rates between TOIs related to coherence or spatial proximity and the lack of disparities in TOIs connected to signaling might predominantly stem from differences in the visual complexity of stimuli rather than the workings of the working memory system. Addressing this issue necessitates comprehensive and in-depth investigations in the future.

In addition, the sample of participants was composed entirely of males, which may restrict the generalizability of the findings across genders. Another factor that might have affected the results is possible session fatigue, as the experimental tasks required sustained attention over time. Both aspects should be taken into account when interpreting the outcomes.

Moreover, an overarching challenge arises from the multifaceted nature of some metrics, entailing the need for a deeper understanding of how eye movements can effectively reflect cognitive processes. Consequently, forthcoming research endeavors should diligently explore this intricate relationship.

Furthermore, it is recommended that future investigations account for individuals’ distinct learning styles, thereby enhancing the comprehensiveness of findings. Lastly, to attain a more comprehensive understanding of multimedia learning dynamics, future studies should investigate the various forms of cognitive load (intrinsic, extraneous, and germane) and employ eye tracker metrics to measure them individually. By undertaking these proposed avenues of inquiry, multimedia learning can advance toward a more nuanced comprehension of its underlying mechanisms.

Data Availability

The data supporting the conclusions of this article is made available without undue reservation in the link below: https://doi.org/10.5281/zenodo.15101629.

Funding Statement

This work was supported by the Iranian National Science Foundation (INSF) (under proposal number of 4015666 to Reza Ebrahimpour).

References

  • 1.Expósito A, Sánchez-Rivas J, Gómez-Calero MP, Pablo-Romero MP. Examining the use of instructional video clips for teaching macroeconomics. Computers & Education. 2020;144:103709. doi: 10.1016/j.compedu.2019.103709 [DOI] [Google Scholar]
  • 2.Mayer RE. Cognitive theory of multimedia learning. The Cambridge handbook of multimedia learning. 2005. [Google Scholar]
  • 3.Mayer RE, Moreno R. Nine Ways to Reduce Cognitive Load in Multimedia Learning. Educational Psychologist. 2003;38(1):43–52. doi: 10.1207/s15326985ep3801_6 [DOI] [Google Scholar]
  • 4.Paas F, Tuovinen JE, Tabbers H, Van Gerven PWM. Cognitive Load Measurement as a Means to Advance Cognitive Load Theory. Educational Psychologist. 2003;38(1):63–71. doi: 10.1207/s15326985ep3801_8 [DOI] [Google Scholar]
  • 5.Sweller J, Chandler P. Why some material is difficult to learn. Cogn Instr. 1994;12(3):185–233. [Google Scholar]
  • 6.Paas F, Ayres P, Pachman M. Assessment of cognitive load in multimedia learning. Recent innovations in educational technology that facilitate student learning. Charlotte, NC: Information Age Publishing Inc. 11–35. [Google Scholar]
  • 7.Sweller J. Cognitive load during problem solving: effects on learning. Cogn Sci. 1988;12(2):257–85. [Google Scholar]
  • 8.Mayer RE. Multimedia learning. Psychology of Learning and Motivation. Elsevier. 2002. 85–139. doi: 10.1016/s0079-7421(02)80005-6 [DOI] [Google Scholar]
  • 9.Brunken R, Plass JL, Leutner D. Direct Measurement of Cognitive Load in Multimedia Learning. Educational Psychologist. 2003;38(1):53–61. doi: 10.1207/s15326985ep3801_7 [DOI] [Google Scholar]
  • 10.Galy E, Cariou M, Mélan C. What is the relationship between mental workload factors and cognitive load types?. Int J Psychophysiol. 2012;83(3):269–75. doi: 10.1016/j.ijpsycho.2011.09.023 [DOI] [PubMed] [Google Scholar]
  • 11.Gopher D, Braune R. On the Psychophysics of Workload: Why Bother with Subjective Measures?. Hum Factors. 1984;26(5):519–32. doi: 10.1177/001872088402600504 [DOI] [Google Scholar]
  • 12.Paas FGWC. Training strategies for attaining transfer of problem-solving skill in statistics: A cognitive-load approach. Journal of Educational Psychology. 1992;84(4):429–34. doi: 10.1037/0022-0663.84.4.429 [DOI] [Google Scholar]
  • 13.Tomasi D, Ernst T, Caparelli EC, Chang L. Common deactivation patterns during working memory and visual attention tasks: an intra-subject fMRI study at 4 Tesla. Hum Brain Mapp. 2006;27(8):694–705. doi: 10.1002/hbm.20211 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Urrestilla N, St-Onge D. Measuring Cognitive Load: Heart-rate Variability and Pupillometry Assessment. In: Companion Publication of the 2020 International Conference on Multimodal Interaction, 2020. 405–10. doi: 10.1145/3395035.3425203 [DOI] [Google Scholar]
  • 15.Wang J, Antonenko P, Celepkolu M, Jimenez Y, Fieldman E, Fieldman A. Exploring relationships between eye tracking and traditional usability testing data. Int J Hum Comput Int. 2018;35(6):483–94. [Google Scholar]
  • 16.Sarailoo R, Latifzadeh K, Amiri SH, Bosaghzadeh A, Ebrahimpour R. Assessment of instantaneous cognitive load imposed by educational multimedia using electroencephalography signals. Frontiers in Neuroscience. 2022;16. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Farkish A, Bosaghzadeh A, Amiri SH, Ebrahimpour R. Evaluating the Effects of Educational Multimedia Design Principles on Cognitive Load Using EEG Signal Analysis. Educ Inf Technol. 2022;28(3):2827–43. doi: 10.1007/s10639-022-11283-2 [DOI] [Google Scholar]
  • 18.Chen S, Epps J, Ruiz N, Chen F. Eye activity as a measure of human mental effort in HCI. In: Proceedings of the 16th international conference on Intelligent user interfaces, 2011. 315–8. doi: 10.1145/1943403.1943454 [DOI] [Google Scholar]
  • 19.Di Stasi LL, Antolí A, Cañas JJ. Evaluating mental workload while interacting with computer-generated artificial environments. Entertainment Computing. 2013;4(1):63–9. doi: 10.1016/j.entcom.2011.03.005 [DOI] [Google Scholar]
  • 20.De Rivecourt M, Kuperus MN, Post WJ, Mulder LJM. Cardiovascular and eye activity measures as indices for momentary changes in mental effort during simulated flight. Ergonomics. 2008;51(9):1295–319. doi: 10.1080/00140130802120267 [DOI] [PubMed] [Google Scholar]
  • 21.Foy HJ, Chapman P. Mental workload is reflected in driver behaviour, physiology, eye movements and prefrontal cortex activation. Appl Ergon. 2018;73:90–9. doi: 10.1016/j.apergo.2018.06.006 [DOI] [PubMed] [Google Scholar]
  • 22.Sevcenko N, Appel T, Ninaus M, Moeller K, Gerjets P. Theory-based approach for assessing cognitive load during time-critical resource-managing human–computer interactions: an eye-tracking study. J Multimodal User Interfaces. 2022;17(1):1–19. doi: 10.1007/s12193-022-00398-y [DOI] [Google Scholar]
  • 23.Pouget P. Introduction to the Study of Eye Movements. Studies in Neuroscience, Psychology and Behavioral Economics. 2019. 3–10. [Google Scholar]
  • 24.Siegenthaler E, Costela FM, McCamy MB, Di Stasi LL, Otero-Millan J, Sonderegger A. Task difficulty in mental arithmetic affects microsaccadic rates and magnitudes. Eur J Neurosci. 2013;39(2):287–94. [DOI] [PubMed] [Google Scholar]
  • 25.Dalmaso M, Castelli L, Scatturin P, Galfano G. Working memory load modulates microsaccadic rate. J Vis. 2017;17(3):6. doi: 10.1167/17.3.6 [DOI] [PubMed] [Google Scholar]
  • 26.Benedetto S, Pedrotti M, Bridgeman B. Microsaccades and Exploratory Saccades in a Naturalistic Environment. JEMR. 2011;4(2). doi: 10.16910/jemr.4.2.2 [DOI] [Google Scholar]
  • 27.Maffei A, Angrilli A. Spontaneous blink rate as an index of attention and emotion during film clips viewing. Physiol Behav. 2019;204:256–63. doi: 10.1016/j.physbeh.2019.02.037 [DOI] [PubMed] [Google Scholar]
  • 28.Alemdag E, Cagiltay K. A systematic review of eye tracking research on multimedia learning. Computers & Education. 2018;125:413–28. doi: 10.1016/j.compedu.2018.06.023 [DOI] [Google Scholar]
  • 29.Deng R, Gao Y. A review of eye tracking research on video-based learning. Educ Inf Technol (Dordr). 2023;28(6):7671–702. doi: 10.1007/s10639-022-11486-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Krebs M-C, Schüler A, Scheiter K. Just follow my eyes: The influence of model-observer similarity on Eye Movement Modeling Examples. Learning and Instruction. 2019;61:126–37. doi: 10.1016/j.learninstruc.2018.10.005 [DOI] [Google Scholar]
  • 31.Wang F, Zhao T, Mayer RE, Wang Y. Guiding the learner’s cognitive processing of a narrated animation. Learn Instr. 2020;69:101357. [Google Scholar]
  • 32.Wang J, Antonenko P, Dawson K. Does visual attention to the instructor in online video affect learning and learner perceptions? An eye-tracking analysis. Comput Educ. 2020;146:103779. [Google Scholar]
  • 33.Beatty J. Task-evoked pupillary responses, processing load, and the structure of processing resources. Psychological Bulletin. 1982;91(2):276–92. doi: 10.1037/0033-2909.91.2.276 [DOI] [PubMed] [Google Scholar]
  • 34.Chen S, Epps J, Chen F. A comparison of four methods for cognitive load measurement. In: Proceedings of the 23rd Australian Computer-Human Interaction Conference, 2011. 76–9. doi: 10.1145/2071536.2071547 [DOI] [Google Scholar]
  • 35.Lallé S, Toker D, Conati C, Carenini G. Prediction of Users’ Learning Curves for Adaptation while Using an Information Visualization. In: Proceedings of the 20th International Conference on Intelligent User Interfaces, 2015. 357–68. doi: 10.1145/2678025.2701376 [DOI] [Google Scholar]
  • 36.Schmidt-Weigand F, Kohnert A, Glowalla U. Explaining the modality and contiguity effects: New insights from investigating students’ viewing behaviour. Applied Cognitive Psychology. 2010;24(2):226–37. [Google Scholar]
  • 37.Coskun A, Cagiltay K. A systematic review ofeye‐tracking‐basedresearch on animated multimedia learning. Computer Assisted Learning. 2021;38(2):581–98. doi: 10.1111/jcal.12629 [DOI] [Google Scholar]
  • 38.Duncan J, Parker A. Open Forum: Academic Listening and Speaking. Oxford University Press. 2007. [Google Scholar]
  • 39.Zu T, Hutson J, Loschky LC, Rebello NS. Using eye movements to measure intrinsic, extraneous, and germane load in a multimedia learning environment. J Educ Psychol. 2019. [Google Scholar]
  • 40.Xie H, Zhao T, Deng S, Peng J, Wang F, Zhou Z. Using eye movement modelling examples to guide visual attention and foster cognitive performance: A meta‐analysis. J Comput Assist Learn. 2021;37(4):1194–206. [Google Scholar]
  • 41.Development of NASA-TLX (Task Load Index): Results of Empirical and Theoretical Research. Advances in Psychology. Elsevier. 1988. 139–83. doi: 10.1016/s0166-4115(08)62386-9 [DOI] [Google Scholar]
  • 42.Molina AI, Navarro Ó, Ortega M, Lacruz M. Evaluating multimedia learning materials in primary education using eye tracking. Computer Standards & Interfaces. 2018;59:45–60. doi: 10.1016/j.csi.2018.02.004 [DOI] [Google Scholar]
  • 43.Holmqvist K, Nyström M, Andersson R, Dewhurst R, Jarodzka H, Van de Weijer J. Eye tracking: A comprehensive guide to methods and measures. OUP Oxford. 2011. [Google Scholar]
  • 44.Mahanama B, Jayawardana Y, Rengarajan S, Jayawardena G, Chukoskie L, Snider J, et al. Eye Movement and Pupil Measures: A Review. Front Comput Sci. 2022;3. doi: 10.3389/fcomp.2021.733531 [DOI] [Google Scholar]
  • 45.Habibi M, Oertel WH, White BJ, Brien DC, Coe BC, Riek HC, et al. Eye tracking identifies biomarkers in α-synucleinopathies versus progressive supranuclear palsy. J Neurol. 2022;269(9):4920–38. doi: 10.1007/s00415-022-11136-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Krejtz K, Duchowski AT, Niedzielska A, Biele C, Krejtz I. Eye tracking cognitive load using pupil diameter and microsaccades with fixed gaze. PLoS One. 2018;13(9):e0203629. doi: 10.1371/journal.pone.0203629 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Mihelčič M, Podlesek A. Cognitive workload affects ocular accommodation and pupillary response. J Optom. 2023;16(2):107–15. doi: 10.1016/j.optom.2022.05.001 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Fink L, Simola J, Tavano A, Lange E, Wallot S, Laeng B. From pre-processing to advanced dynamic modeling of pupil data. Behav Res Methods. 2024;56(3):1376–412. doi: 10.3758/s13428-023-02098-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Dan EL, Dinsoreanu M, Muresan RC. Accuracy of six interpolation methods applied on pupil diameter data. In: 2020 IEEE Int Conf Autom, Qual Test, Robot (AQTR), 2020. [Google Scholar]
  • 50.Grootjen JW, Weingärtner H, Mayer S. Investigating the Effects of Eye-Tracking Interpolation Methods on Model Performance of LSTM. In: Proceedings of the 2024 Symposium on Eye Tracking Research and Applications, 2024. 1–6. doi: 10.1145/3649902.3656353 [DOI] [Google Scholar]
  • 51.Pedrotti M, Mirzaei MA, Tedesco A, Chardonnet J-R, Mérienne F, Benedetto S, et al. Automatic Stress Classification With Pupil Diameter Analysis. International Journal of Human-Computer Interaction. 2014;30(3):220–36. doi: 10.1080/10447318.2013.848320 [DOI] [Google Scholar]
  • 52.Engbert R, Kliegl R. Microsaccades uncover the orientation of covert attention. Vision Res. 2003;43(9):1035–45. doi: 10.1016/s0042-6989(03)00084-1 [DOI] [PubMed] [Google Scholar]
  • 53.Dalmaso M, Castelli L, Scatturin P, Galfano G. Working memory load modulates microsaccadic rate. J Vis. 2017;17(3):6. doi: 10.1167/17.3.6 [DOI] [PubMed] [Google Scholar]
  • 54.Dalmaso M, Castelli L, Galfano G. Microsaccadic rate and pupil size dynamics in pro-/anti-saccade preparation: the impact of intermixed vs. blocked trial administration. Psychol Res. 2020;84(5):1320–32. doi: 10.1007/s00426-018-01141-7 [DOI] [PubMed] [Google Scholar]
  • 55.Huddleston P, Behe BK, Minahan S, Fernandez RT. Seeking attention: an eye tracking study of in-store merchandise displays. International Journal of Retail & Distribution Management. 2015;43(6):561–74. doi: 10.1108/ijrdm-06-2013-0120 [DOI] [Google Scholar]
  • 56.Laeng B, Alnaes D. Pupillometry. Studies in Neuroscience, Psychology and Behavioral Economics. Springer International Publishing. 2019. 449–502. doi: 10.1007/978-3-030-20085-5_11 [DOI] [Google Scholar]
  • 57.Mallick R, Slayback D, Touryan J, Ries AJ, Lance BJ. The use of eye metrics to index cognitive workload in video games. In: 2016 IEEE Second Workshop on Eye Tracking and Visualization (ETVIS), 2016. 60–4. doi: 10.1109/etvis.2016.7851168 [DOI] [Google Scholar]
  • 58.Recarte MA, Nunes LM. Effects of verbal and spatial-imagery tasks on eye fixations while driving. J Exp Psychol Appl. 2000;6(1):31–43. doi: 10.1037//1076-898x.6.1.31 [DOI] [PubMed] [Google Scholar]
  • 59.Liu J-C, Li K-A, Yeh S-L, Chien S-Y. Assessing Perceptual Load and Cognitive Load by Fixation-Related Information of Eye Movements. Sensors (Basel). 2022;22(3):1187. doi: 10.3390/s22031187 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Jacob RJK, Karn KS. Eye Tracking in Human-Computer Interaction and Usability Research. The Mind’s Eye. Elsevier. 2003. 573–605. doi: 10.1016/b978-044451020-4/50031-1 [DOI] [Google Scholar]
  • 61.Goldberg JH, Kotval XP. Eye movement-based evaluation of the computer interface. Advances in Occupational Ergonomics and Safety. Amsterdam: ISO Press. 1998. 529–32. [Google Scholar]
  • 62.Kotval XP, Goldberg JH. Eye Movements and Interface Component Grouping: An Evaluation Method. Proceedings of the Human Factors and Ergonomics Society Annual Meeting. 1998;42(5):486–90. doi: 10.1177/154193129804200509 [DOI] [Google Scholar]
  • 63.Walter K, Bex P. Cognitive load influences oculomotor behavior in natural scenes. Sci Rep. 2021;11(1):12405. doi: 10.1038/s41598-021-91845-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Zhou Y-B, Ruan S-J, Zhang K, Bao Q, Liu H-Z. Time pressure effects on decision-making in intertemporal loss scenarios: an eye-tracking study. Front Psychol. 2024;15:1451674. doi: 10.3389/fpsyg.2024.1451674 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Hyönä J. The use of eye movements in the study of multimedia learning. Learning and Instruction. 2010;20(2):172–6. doi: 10.1016/j.learninstruc.2009.02.013 [DOI] [Google Scholar]
  • 66.Loschky LC, Ringer RV, Johnson AP, Larson AM, Neider M, Kramer AF. Blur Detection is Unaffected by Cognitive Load. Vis cogn. 2014;22(3):522–47. doi: 10.1080/13506285.2014.884203 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Reimer B, Mehler B, Wang Y, Coughlin JF. A field study on the impact of variations in shortterm memory demands on drivers’ visual attention and driving performance across three age groups. Hum Factors. 2012;54(3):454–68. doi: 10.1177/0018720812437274 [DOI] [PubMed] [Google Scholar]
  • 68.Nuthmann A, Henderson JM. Using CRISP to model global characteristics of fixation durations in scene viewing and reading with a common mechanism. Visual Cognition. 2012;20(4–5):457–94. doi: 10.1080/13506285.2012.670142 [DOI] [Google Scholar]
  • 69.Krüger JM, Bodemer D. Application and Investigation of Multimedia Design Principles in Augmented Reality Learning Environments. Information. 2022;13(2):74. doi: 10.3390/info13020074 [DOI] [Google Scholar]
  • 70.Li W, Wang F, Mayer RE, Liu H. Getting the point: Which kinds of gestures by pedagogical agents improve multimedia learning?. Journal of Educational Psychology. 2019;111(8):1382–95. doi: 10.1037/edu0000352 [DOI] [Google Scholar]
  • 71.Skuballa IT, Schwonke R, Renkl A. Learning from narrated animations with different support procedures: working memory capacity matters. Appl Cognit Psychol. 2012;26(6):840–7. [Google Scholar]
  • 72.Xie H, Mayer RE, Wang F, Zhou Z. Coordinating visual and auditory cueing in multimedia learning. J Educ Psychol. 2019;111(2). [Google Scholar]
  • 73.Gao X, Yan H, Sun H-J. Modulation of microsaccade rate by task difficulty revealed through between- and within-trial comparisons. J Vis. 2015;15(3):3. doi: 10.1167/15.3.3 [DOI] [PubMed] [Google Scholar]
  • 74.Fukuda K, Stern JA, Brown TB, Russo MB. Cognition, blinks, eye-movements, and pupillary movements during performance of a running memory task. Aviat Space Environ Med. 2005;76(7 Suppl):C75-85. [PubMed] [Google Scholar]
  • 75.Klingner J, Tversky B, Hanrahan P. Effects of visual and verbal presentation on cognitive load in vigilance, memory, and arithmetic tasks. Psychophysiology. 2011;48(3):323–32. doi: 10.1111/j.1469-8986.2010.01069.x [DOI] [PubMed] [Google Scholar]
  • 76.Nakayama M, Takahashi K, Shimizu Y. The act of task difficulty and eye-movement frequency for the “Oculo-motor indices”. In: Proceedings of the symposium on Eye tracking research & applications - ETRA ’02, 2002. 37. doi: 10.1145/507072.507080 [DOI] [Google Scholar]
  • 77.Unsworth N, Robison MK. Tracking working memory maintenance with pupillometry. Attention, Perception, & Psychophysics. 2017;80(2):461–84. [DOI] [PubMed] [Google Scholar]

Decision Letter 0

Vishal Bharmauria

10 Jul 2025

Dear Dr. Ebrahimpour,

Language Confound : Reviewer 2 raised concerns about using English materials for Persian-speaking participants, which may have introduced extraneous cognitive load. Please discuss this limitation more fully.

Statistical Rigor : Given the small sample size, consider using non-parametric tests and add visual assessments of normality. Include an a priori power analysis.

Figure Quality : Improve figure resolution (≥300 dpi) and enhance AOI labeling for clarity.

Limitations : Add a brief discussion of other limitations, such as the all-male sample and session fatigue.

publication criteria  and not, for example, on novelty or perceived impact.

For Lab, Study and Registered Report Protocols: These article types are not expected to include results but may include pilot data. 

==============================

Please submit your revised manuscript by Aug 24 2025 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org . When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

  • A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'.

  • A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.

  • An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols . Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols .

We look forward to receiving your revised manuscript.

Kind regards,

Vishal Bharmauria

Academic Editor

PLOS ONE

Journal requirements:

When submitting your revision, we need you to address these additional requirements.

1. Please ensure that your manuscript meets PLOS ONE's style requirements, including those for file naming. The PLOS ONE style templates can be found at

https://journals.plos.org/plosone/s/file?id=wjVg/PLOSOne_formatting_sample_main_body.pdf and

https://journals.plos.org/plosone/s/file?id=ba62/PLOSOne_formatting_sample_title_authors_affiliations.pdf

2. When completing the data availability statement of the submission form, you indicated that you will make your data available on acceptance. We strongly recommend all authors decide on a data sharing plan before acceptance, as the process can be lengthy and hold up publication timelines. Please note that, though access restrictions are acceptable now, your entire data will need to be made freely accessible if your manuscript is accepted for publication. This policy applies to all data except where public deposition would breach compliance with the protocol approved by your research ethics board. If you are unable to adhere to our open data policy, please kindly revise your statement to explain your reasoning and we will seek the editor's input on an exemption. Please be assured that, once you have provided your new statement, the assessment of your exemption will not hold up the peer review process.

Additional Editor Comments (if provided):

[Note: HTML markup is below. Please do not edit.]

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. Is the manuscript technically sound, and do the data support the conclusions?

Reviewer #1: Yes

Reviewer #2: Yes

**********

2. Has the statistical analysis been performed appropriately and rigorously? -->?>

Reviewer #1: Yes

Reviewer #2: No

**********

3. Have the authors made all data underlying the findings in their manuscript fully available??>

The PLOS Data policy

Reviewer #1: No

Reviewer #2: No

**********

4. Is the manuscript presented in an intelligible fashion and written in standard English??>

Reviewer #1: Yes

Reviewer #2: Yes

**********

Reviewer #1: I already reviewed a previous version of this manuscript. The revised manuscript is much improved. The authors have addressed the major concerns I raised previously, particularly around English language proficiency, the clarity and validity of their multimedia manipulations, and the use of eye-tracking measures. I appreciate the added detail in your methods and the more careful interpretation of the results, especially regarding fixation patterns, pupil size, and microsaccades.

A few minor points remain. Please consider improving the labeling and readability of AOIs in the scan path and fixation figures. Also, the overall quality of the figures is quite low. I’m not sure if that’s due to journal compression, but please ensure that the final version includes clear, high-resolution images.

Lastly, a brief mention of limitations (e.g., all-male sample, possible session fatigue) in the discussion would be helpful.

Reviewer #2: The paper titled "Cognitive Load and Visual Attention Assessment Using Physiological Eye-Tracking Measures in Multimedia Learning" presents an interesting investigation into how multimedia instructional design affects learners' cognitive load and visual attention patterns. The researchers employed eye-tracking technology to analyze ocular behaviors and their relationship with cognitive load and learning performance. While the study offers valuable insights, several critical issues need to be addressed before publication.

1- One notable limitation of this study is the use of English-language instructional materials for Persian-speaking participants, despite their native language being Persian. This issue may have influenced the results in the many ways. Particularly, it can increase the cognitive load. Even though participants had intermediate English proficiency, processing educational content in a non-native language requires additional cognitive effort. This could introduce extraneous cognitive load unrelated to multimedia design. Accordingly, some of the reported cognitive load differences between the two groups (P and NP) might stem from individual variations in English proficiency rather than the experimental manipulation. Also, although participants' language skills were assessed using an IELTS simulator test, individual differences in listening comprehension, processing speed, or vocabulary familiarity could affect their interaction with the content. For instance, participants with better comprehension of specific terms might experience lower cognitive load, even when exposed to the non-principled (NP) version. On the other hand, in can cause potential confounding effects on eye-tracking metrics and eye movements (e.g., fixation counts or durations) might reflect difficulty in understanding English text rather than multimedia design flaws. To address this limitation, I recommend using instructional materials in participants' native language to eliminate language-related cognitive load.

2- With 34 participants divided into two groups (approximately 14 per group after exclusions), the sample size in each group is relatively small. In such cases, normality tests like Kolmogorov-Smirnov or Shapiro-Wilk may have reduced sensitivity in detecting deviations from normality. Studies show these tests can produce false negatives with small samples (<30), potentially misidentifying non-normal data as normal. Instead, using non-parametric tests (e.g., Mann-Whitney U) could be preferable as they don't require normality assumptions. Visual methods like Q-Q plots or examining skewness/kurtosis could qualitatively assess normality (though not mentioned in the paper).

3- The lack of a priori power analysis represents a significant limitation because the sample size may be inadequate (i.e., the study might be underpowered to detect true effects). Also, the risk of false negatives (Type II errors) increases (where real significant differences are incorrectly deemed non-significant). Authors should report detailed power analysis (e.g., using G*Power) to justify sample sizes.

4- The resolution of the figures is currently insufficient for proper visualization. All figures must be rendered at a minimum of 300 dpi (dots per inch) to meet publication standards.

**********

what does this mean? ). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy

Reviewer #1: No

Reviewer #2: No

**********

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/ . PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at figures@plos.org . Please note that Supporting Information files do not need this step.

PLoS One. 2025 Dec 1;20(12):e0337195. doi: 10.1371/journal.pone.0337195.r002

Author response to Decision Letter 1


22 Sep 2025

Comments to the Author

We would like to express our sincere gratitude to the reviewers for their valuable, insightful, and constructive feedback on our manuscript. Your comments have been extremely helpful in refining the quality, clarity, and overall contribution of this work. We greatly appreciate the time and effort you devoted to a careful evaluation of our study.

We have carefully addressed each of the points raised and revised the manuscript accordingly. In the following pages, we provide a detailed, point-by-point response to all the comments and suggestions. We hope that the revisions made will meet your expectations and improve the manuscript to a level suitable for publication.

1. Is the manuscript technically sound, and do the data support the conclusions? The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #1: Yes

Reviewer #2: Yes

We greatly appreciate your careful review and valuable feedback.

2. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #1: Yes

Reviewer #2: No

We recognize the concerns raised by Reviewer #2. To address these issues, we have carefully considered all the points and suggestions you raised, and the necessary clarifications and revisions have been incorporated into the manuscript. Details of the changes are provided in the responses to the reviewer’s comments below.

3. Have the authors made all data underlying the findings in their manuscript fully available? The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #1: No

Reviewer #2: No

We have made the raw data and statistical analysis data available via a link mentioned in data availability statement for transparency and further verification.

4. Is the manuscript presented in an intelligible fashion and written in standard English? PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #1:

I already reviewed a previous version of this manuscript. The revised manuscript is much improved. The authors have addressed the major concerns I raised previously, particularly around English language proficiency, the clarity and validity of their multimedia manipulations, and the use of eye-tracking measures. I appreciate the added detail in your methods and the more careful interpretation of the results, especially regarding fixation patterns, pupil size, and microsaccades.

A few minor points remain. Please consider improving the labeling and readability of AOIs in the scan path and fixation figures. Also, the overall quality of the figures is quite low. I’m not sure if that’s due to journal compression, but please ensure that the final version includes clear, high-resolution images.

Lastly, a brief mention of limitations (e.g., all-male sample, possible session fatigue) in the discussion would be helpful.

We sincerely thank the reviewer for the valuable suggestions regarding the figures. In response, the AOIs in Figure 7 have been clearly highlighted using distinct colors to improve visibility. In Figure 8, both the AOIs and their corresponding labels have been made more explicit. Additionally, the overall quality of all figures has been enhanced, and each now meets a minimum resolution of 300 dpi to ensure clear readability.

The mentioned limitations, including the all-male sample and possible session fatigue, have been added to the manuscript.

Reviewer #2:

The paper titled "Cognitive Load and Visual Attention Assessment Using Physiological Eye-Tracking Measures in Multimedia Learning" presents an interesting investigation into how multimedia instructional design affects learners' cognitive load and visual attention patterns. The researchers employed eye-tracking technology to analyze ocular behaviors and their relationship with cognitive load and learning performance. While the study offers valuable insights, several critical issues need to be addressed

before publication.

1- One notable limitation of this study is the use of English-language instructional materials for Persian-speaking participants, despite their native language being Persian. This issue may have influenced the results in the many ways. Particularly, it can increase the cognitive load. Even though participants had intermediate English proficiency, processing educational content in a non-native language requires additional cognitive effort. This could introduce extraneous cognitive load unrelated to multimedia design. Accordingly, some of the reported cognitive load differences between the two groups (P and NP) might stem from individual variations in English proficiency rather than the experimental manipulation. Also, although participants' language skills were assessed using an IELTS simulator test, individual differences in listening comprehension, processing speed, or vocabulary familiarity could affect their interaction with the content. For instance, participants with better comprehension of specific terms might experience lower cognitive load, even when exposed to the non-principled (NP) version. On the other hand, in can cause potential confounding effects on eye-tracking metrics and eye movements (e.g., fixation counts or durations) might reflect difficulty in understanding English text rather than multimedia design flaws. To address this limitation, I recommend using instructional materials in participants' native language to eliminate language-related cognitive load.

As previously mentioned, prior to the experiment all participants completed the standardized IELTS test, and their scores ranged between 6-7. This indicates that they all possessed a relatively high and comparable level of English proficiency. Participants were then randomly assigned to the P and NP conditions, which means that exposure to the second language was a common factor across both groups. Since there were no significant differences in language proficiency between the two conditions, it can be assumed that, on average, a similar level of cognitive load stemming from second language processing was imposed on both groups.

If the observed eye-movement behaviors and cognitive load were entirely attributable to the use of a second language, then no meaningful differences should have emerged between the P and NP conditions in terms of cognitive load, performance, or related measures. Furthermore, the instructional material employed was a standardized resource specifically designed for non-native English learners, which ensured a high level of control. The material is typically aimed at learners with proficiency levels equivalent to IELTS scores of approximately 5–6.5. Given that our participants’ proficiency scores ranged from 6-7, the vocabulary and content were appropriate for their level and unlikely to have posed substantial difficulty, as most participants were already familiar with the majority of the lexical items.

2- With 34 participants divided into two groups (approximately 14 per group after exclusions), the sample size in each group is relatively small. In such cases, normality tests like Kolmogorov-Smirnov or Shapiro-Wilk may have reduced sensitivity in detecting deviations from normality. Studies show these tests can produce false negatives with small samples (<30), potentially misidentifying non-normal data as normal. Instead, using non parametric tests (e.g., Mann-Whitney U) could be preferable as they don't require normality assumptions. Visual methods like Q-Q plots or examining skewness/kurtosis could qualitatively assess normality (though not mentioned in the paper).

To compare the two conditions, the non-parametric Mann–Whitney U test was applied. For the NASA data, the test yielded U = 74, p < 0.001, indicating a statistically significant difference between the conditions. Similarly, for the Recall data, U = 674.5, p < 0.001, also indicating a significant difference.

To assess the normality of the data distribution visually, Q-Q (Quantile-Quantile) plots were generated. The points closely followed the reference line, suggesting no substantial deviation from normality. For the NASA data, skewness (0.45) and kurtosis (0.37) fell within the acceptable range (−1 to +1), while for the Recall data, skewness (−0.47) and kurtosis (0.6) were also within this range. These results indicate approximate symmetry and a mesokurtic distribution.

The relevant descriptions have been added to the manuscript.

Below, you can see the corresponding Q-Q plot charts for the NASA-TLX and recall scores.

NASA-TLX:

Recall:

3- The lack of a priori power analysis represents a significant limitation because the sample size may be inadequate (i.e., the study might be underpowered to detect true effects). Also, the risk of false negatives (Type II errors) increases (where real significant differences are incorrectly deemed non-significant). Authors should report detailed power analysis (e.g., using G*Power) to justify sample sizes.

As you can see below, a priori power analysis was performed using G*Power to estimate the required sample size for the planned comparison. The analysis specified a two-tailed independent-samples t-test with an expected effect size of d= 0.7, a significance level of α=0.05, and desired statistical power of 1-β = 0.8 . The calculation indicated that 34 participants per group, corresponding to a total sample size of 68, would be sufficient to achieve the targeted power.

we believe it is reasonable to treat the two sessions per participant as independent data points, allowing us to double the total sample size from 34 individuals to 68 observations. Although each participant took part in two experimental sessions, we argue that these should be treated as independent observations. Each session was conducted on a different day and involved completely different video content. The participants were not exposed to the same material in both sessions. In one session, they were assigned to P condition, and in the other to NP condition. There was no overlap in content, no learning or transfer effects were expected, and the sessions were separated in time to minimize any carryover influence. Moreover, the order of conditions was randomized across participants to control for any potential order effects. That is, some participants started with P version, while others began with NP version. This consideration is also supported by the independence of stimuli and the randomization procedures applied.

4- The resolution of the figures is currently insufficient for proper visualization. All figures must be rendered at a minimum of 300

dpi (dots per inch) to meet publication standards.

We sincerely thank the reviewer for the valuable suggestions regarding the figures. The overall quality of all figures has been enhanced, and each now meets a minimum resolution of 300 dpi to ensure clear readability.

Attachment

Submitted filename: Response to Reviewers.docx

pone.0337195.s002.docx (289.6KB, docx)

Decision Letter 1

Vishal Bharmauria

6 Nov 2025

<p>Cognitive load and visual attention assessment using physiological eye tracking measures in multimedia learning

PONE-D-25-22353R1

Dear Dr. Ebrahimpour,

We’re pleased to inform you that your manuscript has been judged scientifically suitable for publication and will be formally accepted for publication once it meets all outstanding technical requirements.

Within one week, you’ll receive an e-mail detailing the required amendments. When these have been addressed, you’ll receive a formal acceptance letter and your manuscript will be scheduled for publication.

An invoice will be generated when your article is formally accepted. Please note, if your institution has a publishing partnership with PLOS and your article meets the relevant criteria, all or part of your publication costs will be covered. Please make sure your user information is up-to-date by logging into Editorial Manager at Editorial Manager®  and clicking the ‘Update My Information' link at the top of the page. For questions related to billing, please contact billing support .

If your institution or institutions have a press office, please notify them about your upcoming paper to help maximize its impact. If they’ll be preparing press materials, please inform our press team as soon as possible -- no later than 48 hours after receiving the formal acceptance. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

Kind regards,

Vishal Bharmauria

Academic Editor

PLOS ONE

Additional Editor Comments (optional):

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

Reviewer #1: All comments have been addressed

Reviewer #2: All comments have been addressed

**********

2. Is the manuscript technically sound, and do the data support the conclusions??>

Reviewer #1: Yes

Reviewer #2: Yes

**********

3. Has the statistical analysis been performed appropriately and rigorously? -->?>

Reviewer #1: Yes

Reviewer #2: Yes

**********

4. Have the authors made all data underlying the findings in their manuscript fully available??>

The PLOS Data policy

Reviewer #1: (No Response)

Reviewer #2: No

**********

5. Is the manuscript presented in an intelligible fashion and written in standard English??>

Reviewer #1: Yes

Reviewer #2: Yes

**********

Reviewer #1: I don't have any more comments. The authors have sufficiently responded to all of my concerns. The revised version is ready for publication.

Reviewer #2: (No Response)

**********

what does this mean? ). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy

Reviewer #1: No

Reviewer #2: Yes:  Amirhossein Ghaderi

**********

Acceptance letter

Vishal Bharmauria

PONE-D-25-22353R1

PLOS ONE

Dear Dr. Ebrahimpour,

I'm pleased to inform you that your manuscript has been deemed suitable for publication in PLOS ONE. Congratulations! Your manuscript is now being handed over to our production team.

At this stage, our production department will prepare your paper for publication. This includes ensuring the following:

* All references, tables, and figures are properly cited

* All relevant supporting information is included in the manuscript submission,

* There are no issues that prevent the paper from being properly typeset

You will receive further instructions from the production team, including instructions on how to review your proof when it is ready. Please keep in mind that we are working through a large volume of accepted articles, so please give us a few days to review your paper and let you know the next and final steps.

Lastly, if your institution or institutions have a press office, please let them know about your upcoming paper now to help maximize its impact. If they'll be preparing press materials, please inform our press team within the next 48 hours. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

You will receive an invoice from PLOS for your publication fee after your manuscript has reached the completed accept phase. If you receive an email requesting payment before acceptance or for any other service, this may be a phishing scheme. Learn how to identify phishing emails and protect your accounts at https://explore.plos.org/phishing.

If we can help with anything else, please email us at customercare@plos.org.

Thank you for submitting your work to PLOS ONE and supporting open access.

Kind regards,

PLOS ONE Editorial Office Staff

on behalf of

Dr. Vishal Bharmauria

Academic Editor

PLOS ONE

Associated Data

    This section collects any data citations, data availability statements, or supplementary materials included in this article.

    Supplementary Materials

    Attachment

    Submitted filename: Response to Reviewers.docx

    pone.0337195.s002.docx (289.6KB, docx)

    Data Availability Statement

    The data supporting the conclusions of this article is made available without undue reservation in the link below: https://doi.org/10.5281/zenodo.15101629.


    Articles from PLOS One are provided here courtesy of PLOS

    RESOURCES