Skip to main content
PLOS One logoLink to PLOS One
. 2026 May 18;21(5):e0348069. doi: 10.1371/journal.pone.0348069

The effect of tempi and mode on the rating of the perceived emotion in music

Ulvhild Helena Færøvik 1,*, Karsten Specht 1,2
Editor: Giulia Prete3
PMCID: PMC13183231  PMID: 42149877

Abstract

In this study, we investigated the impact of tempi variations (60, 100, 120, and 150 beats per minute) and major and minor modes on the perceived emotional content of music. To explore this, we created eight versions of five original compositions, resulting in 40 musical stimuli. Control stimuli included variants of white, brown, and pink noise, as well as six human voice recordings. A total of 1280 participants took part in an online survey. Participants were diverse, comprising 262 musicians (defined as individuals with over six years of instrument-playing experience who identified as amateur, semi-professional, or professional). Participants were asked to rate the perceived emotional content of the stimuli, using the second-order factors of the GEMS-9 framework: ‘sublimity’, ‘unease’, and ‘vitality’. We employed a mixed between-within-subjects analysis of variance to analyse the data. Specifically examining the influence of the four tempi, the two modes, while controlling for musicianship, gender, age, and education level. Results showed that as tempi increased, compositions were consistently rated as less sublimity and more vitality, in both major and minor modes. A higher tempi in the major mode resulted in lower ratings of uneasiness. There were also several significant findings for musicianship, gender, and education level, although these mostly had small effect sizes. For age groups, there were two substantial differences between the groups with larger effect sizes. Overall, our findings suggest that tempi and mode significantly influence how compositions are rated for sublimity, unease, and vitality, as they interact.

Introduction

Listening to music is a ubiquitous and inherently emotional experience, shaped by a myriad of factors. We wanted to examine tempi and mode while controlling for musicianship, gender, age, and educational level. Yet, many theories exist about measuring and conceptualising emotion in music. Despite research on the impact of musical structures, key constructs like tempi and mode are examined using heterogeneous emotion models, stimuli and analytical frameworks [1].

In the following, we briefly introduce the main approaches used to conceptualise and measure emotions in music, to motivate the choice of emotion model used in the present study. Additionally, we will also give a brief overview of previous research on emotional perception of tempi, mode in music, and some covariates like age, gender, and musicianship.

Research on music and emotion typically relies on discrete, dimensional, or domain-specific emotion models, which differ in their theoretical assumptions and measurement focus. It is essential to distinguish between emotion theories (emotivist, cognitivist, and constructionist, emotion measurement models (discrete and dimensional) and phenomenological target of measurement (felt and perceived emotion).

Discrete emotion models are rooted in basic emotion theory and assume that a limited set of universal emotions can be evoked or perceived through music [2,3]. These models typically use categorical emotion labels such as happiness, sadness, anger, and fear [4] and assume that music evokes emotions in a manner comparable to other affective stimuli, such as Ekman’s facial expressions [5].

Discrete emotion approaches are prevalent in music and emotion research because basic emotions are considered relatively universal and listeners tend to identify them reliably in musical stimuli [6,7]. However, a central limitation of basic emotion theory is the lack of consensus regarding which emotions should be considered “basic” [8]. As a measurement framework, discrete emotion models can be applied to both felt and perceived emotions, as the felt–perceived distinction primarily concerns task instructions rather than the underlying emotion model itself [4].

An alternative approach is the dimensional model of emotion, often grounded in appraisal and constructionist theories [9]. Emotions are considered continuous along dimensions such as valence and arousal and can be applied to both felt and perceived experiences [2]. In this context, listeners typically rate music in terms of valence, arousal, or related constructs such as tension [10]. For example, both scary and happy music are often associated with high arousal, whereas pleasant and sad music are associated with low arousal [11]. While dimensional models allow for graded and continuous emotion ratings, they may provide a relatively coarse description of emotional experience and may not fully capture the aesthetic qualities of musical emotions.

Studying emotions in music presents additional challenges, as emotional responses to music are often less intense than those elicited by life events or autobiographical memories [12]. Scherer, therefore, argues that music should be studied as a conscious feeling that can be measured cognitively and physiologically [13]. This perspective is reflected in the component process model of emotion (CPM), which explains emotional induction and perception in music based on musical structure, stimulus characteristics, and listeners’ subjective and physiological responses [14,15]. Similarly, it has been shown that regardless of basic or dimensional models, emotions in music can be both evoked and perceived with comparable emotional qualities [16]. In line with this dynamic view of emotion, the construction organisation dynamics appraisal (CODA) model emphasises that emotions are constructed from basic psychological components and evolve through continuous appraisal and reappraisal during music listening [17].

In response to the limitations of general emotion models, domain-specific emotion models have been developed to capture emotions that are particularly relevant to music. A literature review of music and emotion research indicates that only a minority of studies employ music-specific emotion models [18]. One model is the Geneva Emotional Music Scale-9 (GEMS-9) [19], which was specifically developed to assess aesthetic emotions in music. The GEMS-9 consists of nine first-order factors: wonder, transcendence, tenderness, nostalgia, peacefulness, joyful activation, power, tension, and sadness. Each first-order factor represents multiple emotional words. The GEMS-9 can also be collapsed into three meta-factors, or second-order factors: sublimity, vitality, and unease [19]. Importantly, the GEMS-9 was developed to test both felt and perceived emotions in music.

Although the distinction between felt and perceived emotions remains debated, most music stimuli are used to assess perceived rather than felt emotions [4]. By focusing on aesthetic emotions that are characteristic of musical experiences, the GEMS-9 provides a music-specific framework for studying emotional responses to music that are less directly tied to utilitarian action tendencies [20].

One strength of the GEMS-9 is that it has been validated across a wide range of music genres, from classical to techno. The GEMS can also generalise well to music styles beyond those included in its original development, such as neoclassical, ambient, new age, punk, folk rock, electronica, electro-medieval, disco, heavy metal, and opera [21–23]. This broad applicability supports the use of the GEMS-9 as a domain-specific emotion model capturing aesthetic emotions across diverse musical contexts.

At the same time, the GEMS-9 has been described as less efficient than discrete and dimensional emotion models [24,25]. Despite this lower efficiency, the GEMS-9 contains both the strongest and weakest items when compared to discrete and dimensional models, indicating substantial variability in item-level performance [24,25]. In contrast, the authors of the GEMS-9 model reported that GEMS-9 outperformed discrete and dimensional models [19]. Together, these findings suggest that while the GEMS-9 may be more complex than general emotion models, it offers advantages in sensitivity and domain relevance.

The current study was part of a larger project on music and emotion, which aimed to measure both the emotional induction, through physiological data, and the perception of emotion through self-report ratings. Hence, we chose to use the thoroughly tested GEMS-9 scale, which was developed to measure both evoked and perceived emotion [19]. To balance domain specificity with measurement efficiency, we focused on the second-order structure of the GEMS-9. Halpern proposed that the GEMS could be condensed into three or four overarching categories, effectively capturing higher-level dimensions of musical emotion while retaining the model’s aesthetic focus [26]. Accordingly, the GEMS-9 can be collapsed into three second-order factors—sublimity, vitality, and unease, which summarise the covariance among the nine first-order factors [19]. Specifically, sublimity shows strong correlations with wonder (1.00), tenderness (0.75), nostalgia (0.89), transcendence (0.66), and peacefulness (0.65). Vitality is defined by near-perfect correlation with power (1.00) and joyful activation (0.95), indicating high internal reliability. Unease has a perfect correlation with tension (1.00), but not with sadness (0.27). Reflecting both a dominant core component and meaningful differentiation within the factor. Together, these strong within-factor associations and the clear separation between second-order factors support collapsing the nine first-order GEMS dimensions into three reliable and interpretable higher-order factors [19,26].

Using the second-order factors allows for a more parsimonious representation of emotional responses to music while preserving the domain-specific conceptualisation of aesthetic emotions embedded in the GEMS framework. Although these second-order factors resemble dimensional emotion spaces, they differ from generic valence–arousal models by reflecting emotion categories derived specifically from musical experience rather than from general affective theory. Thus, the use of GEMS-9 second-order factors provides a theoretically grounded and psychometrically supported compromise between expressive richness and analytical efficiency. For these reasons, the GEMS-9 second-order factors were selected as the emotion model in the present study.

Tempi, defined as the frequency rates of underlying pulse structures [27], is commonly expressed in beats per minute (BPM) [28]. Fast tempi are associated with activation, happiness, tension, arousal, anger, and fear [29–32]. Whereas slow tempi tend to evoke emotions of calmness, peace, sadness, tenderness, longing, boredom, and disgust [29,32]. Further, higher tempi and valence ratings are associated with higher arousal [33].

Early research [34] postulates that elements like tempi and rhythm profoundly influence bodily movement and motor response, supporting the idea that motion in music plays a significant role in emotional expression. This is corroborated in another study [28], showing that small children master tempi perception before mode perception. Similarly, patients with cochlear implants also rely on tempi for emotional ratings when listening with implant-ear alone [35]. Another study found that arousal in music was determined by tempi, while valence was determined by mode [36].

The intricate relationship between tempi and emotional states not only highlights its role as a fundamental aspect of music expression but also underscores its influence on our physiological and emotional responses.

Major and minor triads (not being named thus) became the most general trichords in Renaissance music [37]. In Western music, emotional differences between major and minor tonalities might have emerged at the same time, between the 14th and 17th centuries, when the tonal system emerged [38]. A recent systematic review posits that the major-minor dichotomy represents the primary emotional feature in Western music [39].

A musical key is defined by tonic and scale, commonly referred to as mode, typically categorised as either minor or major [40]. Mode can affect valence and arousal ratings in music [41,42]. Major modes typically convey a sense of happiness, while minor modes are linked to feelings of sadness [32]. Despite minor modes being linked with negative valence, this association is modulated by the tempi at which music is played [34]. Such nuances emphasise the complex interplay of musical elements in shaping emotional experiences for listeners.

For example, listeners understand emotions in Hindustani ragas, even though they are unfamiliar with the tonal system [43]. The authors report that listeners relied on tempi and mode to perceive joy, sadness, and anger. Further, Chinese students rated music in a major key as higher on pleasure, arousal, and dominance compared to a minor key [44]. Interestingly, regardless of mode, music in slow tempi has been rated shorter than music in moderate and fast tempi, suggesting that tempi affects the subjective time perception [45].

Another example is the study by Chubb and coworkers (2013) across three experiments on tone scrambles in either major or minor across 275 listeners. They found that listeners tend to be sorted into two subpopulations. One group is sensitive to major and minor, and the second group is not. Years of music training showed only a modest correlation of around 0.5, and thus a preexisting listening sensitivity explains why some are better at distinguishing major from minor [46].

Previous research has not found a difference in the perception of emotion in music regarding gender [47], and gender and expertise [48]. However, children perceived more happiness and less anger in music compared to adults [48]. Therefore, it is important to control for gender, age, education level, and expertise when doing studies on the perception of music. A systematic review found that perception of major and minor modes is influenced by age, culture, personality and health [49].

Building on age differences, they have been noted in several studies on music perception [25,50,51]. Interestingly, older participants reported stronger emotions toward happy music and less sadness and scary music [52]. Furthermore, adolescents and older adults tend to perceive music more positively than young and middle-aged adults [50]. However, it has been shown that middle-aged adults rate sadness and fear in music lower than young adults [53]. Taken together, these studies show quite varying results for age differences and music perception.

Another important feature of our current study was to have pilot-tested control stimuli. Some argue for acoustical control stimuli, emphasising their importance in contrast with music conditions [54]. Others posit that control stimuli should be non-musical, highlighting the inconsistency of the optimal choice of control conditions in music-related research [55].

It is also common to compare music to control conditions, such as silence and unfamiliar music [56]. However, silence as a control condition could increase the potential for a placebo effect, particularly when other conditions have an auditory component [57]. In two studies, the authors discovered the importance of an active control condition. They observed an effect of music in the first study (silence as control) but not in the second (radio show as control), leading to a discussion on the impact of active versus passive control conditions [58].

A passive control stimulus is something that does not require much attention from the participant, such as silence. An active control stimulus is something that requires attention from the participant, such as giving a subjective or physiological response. While researchers often try to use a “neutral” stimulus as a control condition, the stimuli have typically not been validated beforehand, which may lead to biases and false conclusions [59]. As such, we opted for various control stimuli, using both voice and noise controls.

To contribute to a more nuanced understanding of the perception of emotion in music, our study aimed to explore how alterations in tempi and mode influence ratings, using multiple control stimuli and a large data sample. We used self-composed music, unknown to the participants, since familiarity with music can confound the results of studies [4]. Additionally, we sought to control the impact of musicianship, gender, age, and education level on ratings of perceived emotion.

We acknowledge that there has been a lot of research on the effect of tempi and mode on the perception of music and emotion. While tempi and mode effects are established, their robustness across large samples and within a music-specific aesthetic framework remains underexplored.

Although the effects of tempi and mode on emotional ratings are well established, much of the existing literature relies on small samples, heterogeneous stimuli, and varying emotion models. High-powered conceptual replications using controlled, unfamiliar musical material remain comparatively rare. Without going into the discussion on the replication crisis, it is crucial to note the importance of robustness and generalisability. Particularly in domains where effect sizes are often small, and stimuli vary widely. By combining a large sample, original and pilot-tested compositions, and multiple control stimuli, the present study aims to provide a rigorous confirmatory test of tempi and mode effects within a music-specific emotion framework. As such, the current study is not a replication study, but a conceptual replication study.

Hypothesis

Building upon prior research, our four-point hypothesis is as follows:

H1. Tempo will positively predict vitality across both modes.

H2. Mode will predict unease independent of tempo.

H3. Tempo and mode will interact such that tempo attenuates minor–unease effects.

H4. There will be differences in emotional ratings based on gender, age, musicianship, and education.

Method

Stimuli

Five compositions were chosen from a selection of original pieces. We edited each composition in Logic Pro X [60] to adjust the tempi to 60, 100, 120, and 150 beats per minute (BPM). Each composition lasted 21 seconds. The stimuli were composed with simple melodies, without melodic changes, harmonic tirades, or elements of surprise that could alter the perceived tempi [61]. We also checked that the tempi beat in the compositions were perceived correctly in a small, unpublished pilot study (N = 82). We specifically composed the stimuli to control the musical parameters of tempi and mode. In a previous but similar pilot study, we also investigated dynamics in addition to tempi and mode, and we found that tempi and mode were the parameters that influenced emotional ratings the most, available as a preprint [62].

Additionally, we altered the compositions into a major mode and a minor mode. Consequently, each of the five compositions was played once in every tempi, in both major and minor modes, resulting in each composition being played a total of eight times across all conditions. The survey encompassed 40 compositions.

Three additional original compositions had been previously pilot-tested and available as a preprint [62] and identified as exemplary baselines for sublimity, unease, and vitality compositions. These three compositions served as test stimuli and were presented at the outset of the survey.

To control for the influence of music stimuli versus other auditory stimuli in the study, three types of noise were chosen: variants of white, pink, and brown noise. Additionally, neutral voice recordings with one female and one male speaker, each reading three different texts in an eastern Norwegian dialect, were selected. This selection resulted in a total of six piloted audio recordings and three additional noises. All music stimuli and control stimuli can be listened to here (S1 File: https://figshare.com/projects/Music_and_control_stimuli/246239). See sheet music in S1 File.

Materials

The studies were conducted using an internet survey based on l.a.m.p. (Linux/Apache/MySQL/PhP) and HTML/Javascript/CSS, which was developed in house.

The second-order factors from the GEMS-9 were used to assess the emotional validation of all stimuli [19]. This modification, aimed to align the GEMS model more closely with a dimensional approach while still ensuring comprehensive testing within the domain of music and emotion scales [26]. We translated the GEMS items test into Norwegian at our department in 2014. Two students (including the first author) translated the words to English, and a third student back translated to English. Detailed description of the words is available in the supporting information (S1 File). By using three second-order factors, we also counterbalanced the GEMS-9, as the first-order factors consist of five emotion words for sublimity (overweight) but only two for unease and vitality, respectively. This makes the GEMS-9 more scalable for participants to be able to choose three items that do not overlap, compared to five categories that do overlap (sublimity first-order factors).

To aid participants in understanding the emotional terms, a comprehensive description of the first-order factors from the GEMS-9, as well as additional synonyms, was presented alongside the sounds. See supporting information for the Norwegian synonyms we used, with English translation (S1 File).

Additionally, for the demographic section of the survey, a subset of questionnaire items from the Profile of Music Perception Skills (PROMS) [63] was incorporated. The PROMS is a questionnaire that measures the musical abilities of musicians and non-musicians.

Participants

The initial sample comprised 1280 individuals, 757 of whom identified as female, 496 as male, and 27 who chose not to disclose their gender. The participants’ ages ranged from 19 to 94 (M = 40, SD = 15.6). To be included in the data analysis, participants were required to answer more than 70% of the survey.

Regarding educational attainment, one participant had received no formal education, 22 had basic education, 218 had completed high school, 573 had higher education qualifications, 442 had a master’s degree or PhD as the highest completed level of education, and 24 participants chose not to provide their educational background.

Musicians were operationally defined as individuals who had played a musical instrument for over six years and identified as amateur, semi-professional, or professional musicians (N = 262).

Power

The target for respondents was set at 1562 to have a representative sample for the Norwegian population of 5 million people. We calculated the number with a sample size calculator (https://www.calculator.net/sample-size-calculator.html) with a 95% confidence level and a 2.48 confidence interval. With our obtained sample size of 1280 participants, this corresponds to a margin of error of ±2.68% at a 95% confidence level. Using G*Power [64,65], Post hoc power analysis indicated statistical power approaching 1.00, with a critical F value of 11.85, when the effect size was set to.2, for the sample of 1280.

Ethics

Participants were recruited through snowball sampling [66] via social media. They were eligible to join if they were over 18 and had normal or corrected hearing. This information was screened before analysis. Recruitment for the survey started on June 1, 2020, and ended on December 1, 2022.

All procedures were approved by the Regional Committee for Medical and Health Research Ethics (REK 45655) and carried out per the Code of Ethics of the World Medical Association, Declaration of Helsinki. Upon agreeing to participate, they received a username, a password, and a link to access the online survey. All participants gave electronic informed consent as part of the initial page of the survey before they could access the rest of the survey.

Procedure

Participants received a link, individual passwords, and usernames to access the survey. They were explicitly instructed to wear headphones, although this could not be verified or tested by us.

The survey, encompassing the full set of stimuli, took approximately 21 minutes to complete if participants listened to the full 21 seconds of each stimulus. However, the average completion time among participants was 16 minutes and 44 seconds, indicating that most participants spent less than 21 seconds listening to the stimulus. We excluded the times of 17 participants that exceeded 2 hours for the mean calculation of completion time. This exclusion was made under the assumption that extended times may be attributed to participants leaving the survey tab open on their computers.

The opening section of the survey requested demographic information from participants, including gender, age, education level, handedness, and details about music training. The initial part of the survey presented three examples of compositions that had undergone pilot testing and were identified as sublimity, vitality, and unease, available as a preprint [62].

After completing the test page, participants progressed to the main survey section, where one stimulus was presented at a time.

Participants were instructed to rate all compositions, noises, and voice recordings using the second-order factors from the GEMS-9 questionnaire. Each composition was to be rated on a seven-point Likert scale, ranging from 1 (indicating very little) to 7 (indicating very much). For each sound, participants were asked to provide ratings of perceived sublimity, unease, and vitality.

Participants had the flexibility to advance to the next sound as soon as they had rated the stimuli on all three emotions, allowing them the option to not necessarily listen to the full 21 seconds of the stimuli. The survey design also allowed participants the freedom to play each stimulus multiple times if desired. To maintain control and consistency, every fourth stimulus was a control sound, noise, or voice. The presentation order of stimuli was semi-randomised.

Analysis

Participants who completed less than 70% of the survey were excluded from the respective analyses. Because missing responses varied slightly across conditions, the effective sample size differs marginally between models.

The primary aim of the analysis was to examine the effects of tempo and mode, rather than differences between individual compositions. Each participant rated five structurally comparable compositions under each tempo–mode condition. To obtain robust condition-level estimates and reduce stimulus-specific variance, ratings were aggregated across the five compositions within each tempo–mode combination.

The aggregation procedure was conducted in the following steps:

For each participant, emotion ratings (sublimity, unease, vitality) were recorded separately for each composition. Within each tempo–mode condition (e.g., 60 BPM major), ratings were averaged across the five compositions.

This resulted in eight condition-level variables per emotion (4 tempi × 2 modes).

Before aggregation, descriptive statistics were inspected at the composition level. Mean ratings and variability were highly similar across the five compositions within each tempo–mode condition (S1 File), supporting aggregation. This procedure reduces idiosyncratic stimulus effects and strengthens inference at the parameter level (tempo and mode).

Control stimuli (voice and noise) were analysed separately and were not included in the factorial tempo–mode models.

Statistical model.

Data analysis was conducted in IBM SPSS Statistics (version 29.0.20). To examine the effects of tempo and mode on emotional ratings, we conducted mixed-design analyses of variance (ANOVAs) separately for sublimity, unease, and vitality.

The design included: Within-subject factors with tempi (60, 100, 120, 150 BPM) and mode (major, minor). Between-subject factors: Musicianship (musician vs non-musician) and Gender. Covariates: Age (continuous) and Education level.

The full factorial model tested main effects of tempo and mode, as well as their interaction, while controlling for demographic variables and their interactions with the within-subject factors.

Using Wilkinson–Rogers notation [67], the tested model can be expressed as:

Emotional rating ~ Tempo × Mode + Musicianship + Gender + Age + Education

Separate models were estimated for each emotional dimension.

To control the overall type 1 error rate (α), the Bonferroni-type approach proposed by Benjamini and Hochberg was applied to t-tests; the threshold was adjusted for multiple comparisons [68].

Multivariate test statistics are reported using Wilks’ lambda (Λ). In this context, smaller values of Λ indicate stronger effects of the predictor variable on the dependent variable. Partial eta squared (ηp²) is reported as effect size. Importantly, effect sizes were interpreted following conventional benchmarks defined by Cohen (1988): small effect (ηp² ≈ .01), medium effect (ηp² ≈ .06) and large effect (ηp² ≥ .14). Given the large sample size, statistical significance was interpreted alongside effect sizes to avoid overemphasising trivially small effects.

Results

For each of the four tempi and two modes, mean ratings were first computed across the five songs. Inspection of composition-level descriptives revealed highly consistent patterns across all five stimuli for each emotion and tempo (Supporting Tables, S1 File).

These condition averages were then organised into the three emotional dimensions—sublimity, unease, and vitality. Results are reported separately for each dimension.

Table 1 provides a synoptic overview of the main effects and interactions of tempo and mode across the three emotional dimensions. Across analyses, the largest effects were observed for mode on unease and tempo on vitality. The effects of sublimity were consistently small. Given the large sample size, statistical significance is interpreted alongside effect sizes.

Table 1. Summary of main effects and interactions for tempo and mode across emotional dimensions.

Effect Sublimity Unease Vitality
Tempo F(3,1035)=6.00***, ηp² = .01 F(3,1034)=9.90***, ηp² = .02 F(3,1035)=35.10***, ηp² = .09
Mode F(1,1037)=18.32***, ηp² = .01 F(1,1036)=189.51***, ηp² = .15 F(1,1037)=135.86***, ηp² = .11
Tempo × Mode F(3,1035)=4.40**, ηp² = .01 F(3,1035)=4.40**, ηp² = .01 F(3,1035)=19.10***, ηp² = .05
Tempo × Age ns F(3,1037)=6.50***, ηp² = .01 ns
Mode × Age F(1,1037)=18.18***, ηp² = .01 F(1,1036)=179.57***, ηp² = .14 ns
Other interactions† small effects small effects small effects

*** p < .001.

** p < .01.

†Includes interactions involving gender, musicianship, and education level; all had small effect sizes (ηp² < .02).

ηp² = partial eta squared.

ns = non-significant.

Sublimity

A mixed-design ANOVA revealed a significant main effect of tempi on sublimity ratings, F(3,1035)=6.00, p < .001, ηp² = .01. Slower tempi were associated with higher sublimity ratings. However, the effect size was small.

There was also a significant main effect of mode, F(1,1037)=18.32, p < .001, ηp² = .01, with major mode yielding slightly higher sublimity ratings than minor mode. Again, the effect size was small.

The tempo × mode interaction reached significance, F(3,1035)=4.40, p = .004, ηp² = .01, although the magnitude of this interaction was small. The pattern indicated that tempo-related decreases in sublimity were comparable across modes.

A significant mode × age interaction was observed, F(1,1037)=18.18, p < .001, ηp² = .01, suggesting modest age-related differences in how mode influenced sublimity. No medium or large effects were observed for sublimity.

Overall, while tempo and mode significantly influenced sublimity ratings, all observed effects were small in magnitude (Fig 1). See supporting information tables (S1 File) for details.

Fig 1. Mean sublimity ratings as a function of tempi and mode.

Fig 1

Mean sublimity ratings are shown for four tempi conditions (60, 100, 120, 150 Beats Per Minute) in major and minor modes. Ratings were averaged across five structurally matched compositions within each tempi–mode condition. Error bars represent ±1 standard error of the mean. Higher values indicate stronger perceived sublimity. N = 1289 participants (complete cases vary slightly across conditions).

Unease

For unease ratings, a robust main effect of mode was observed, F(1,1036)=189.51, p < .001, ηp² = .15, representing a large effect. Compositions in the minor mode were consistently rated higher on unease than those in the major mode.

A significant main effect of tempi was also found, F(3,1034)=9.90, p < .001, ηp² = .02, although this effect was small. Unease ratings varied modestly across tempi.

The tempi × mode interaction was significant, F(3,1035)=4.40, p = .004, ηp² = .01, indicating that the influence of tempo differed slightly between modes, though the magnitude of this interaction was small.

A substantial mode × age interaction emerged, F(1,1036)=179.57, p < .001, ηp² = .14, approaching a large effect. This interaction indicated that age moderated the relationship between mode and unease, particularly in minor mode conditions.

Additional interactions involving tempo and age reached statistical significance (e.g., tempo × age), but all were small in magnitude (ηp² ≤ .02).

Taken together, unease ratings were primarily driven by mode, with tempo and demographic interactions contributing smaller effects (Fig 2). See supporting information tables (S1 File) for more details.

Fig 2. Mean unease ratings as a function of tempi and mode.

Fig 2

Mean unease ratings are displayed for four tempi levels (60, 100, 120, 150 Beats Per Minute) in major and minor mode. Ratings were averaged across five compositions per tempi–mode condition. Error bars represent ±1 standard error of the mean. Higher values indicate greater perceived unease. N = 1289 participants. There was a pronounced main effect of mode.

Vitality

For vitality ratings, a strong main effect of tempi was observed, F(3,1035)=35.10, p < .001, ηp² = .09, representing a medium-to-large effect. Faster tempi were consistently associated with higher vitality ratings.

A substantial main effect of mode was also found, F(1,1037)=135.86, p < .001, ηp² = .11, representing a large effect. Major mode elicited higher vitality ratings than minor mode.

The tempi × mode interaction was significant, F(3,1035)=19.10, p < .001, ηp² = .05, indicating a moderate interaction. The increase in vitality with faster tempi was more pronounced in major mode.

Interactions involving age, gender, musicianship, and education reached statistical significance in several cases; however, all were small in magnitude (ηp² < .02).

Overall, vitality ratings were strongly influenced by tempo and mode, with tempo showing the clearest graded relationship (Fig 3). See supporting information tables (S1 File) for means and standard deviations.

Fig 3. Mean vitality ratings as a function of tempi and mode.

Fig 3

Mean vitality ratings are shown across four tempi conditions (60, 100, 120, 150 Beats Per Minute) in major and minor mode. Ratings were aggregated across five compositions within each condition. Error bars represent ±1 standard error of the mean. Higher values indicate greater perceived vitality. N = 1289 participants.

Age.

Given significant interactions involving age in the unease dimension, additional exploratory analyses were conducted. Participants were grouped into three age categories: 18–35 years, 36–59 years and 60 + years.

Separate one-way ANOVAs were conducted for selected tempo–mode combinations showing significant age interactions. Because 24 comparisons were examined, a Bonferroni-adjusted alpha level (α = .002) was applied.

Effect sizes (η²) are reported to contextualise statistically significant findings.

Given the pronounced mode × age interaction for unease, exploratory analyses were conducted across age groups. Significant differences were observed particularly for minor mode at 120 and 150 BPM, where younger participants reported higher unease ratings than older participants. These effects ranged from small to medium magnitude (η² up to.09).

For sublimity and vitality, age-related differences were generally small. See details of analysis in the supporting information tables (S1 File).

Visual representation of unease, 120 minor (Fig 4) and 150 minor (Fig 5), which were the only ones with a medium effect size, as defined by Cohen [69].

Fig 4. Unease ratings by age group across tempi in 120 minor mode.

Fig 4

Mean unease ratings for minor-mode stimuli are displayed separately for three age groups (18–35, 36–59, 60 + years) across four tempi levels. Ratings were averaged across five compositions per condition. Error bars represent ±1 standard error of the mean.

Fig 5. Unease ratings by age group across tempi in 150 minor mode.

Fig 5

Mean unease ratings for minor-mode stimuli are displayed separately for three age groups (18–35, 36–59, 60 + years) across four tempi levels. Ratings were averaged across five compositions per condition. Error bars represent ±1 standard error of the mean.

Control stimuli.

To verify that emotional effects observed in the music conditions were not attributable merely to participation in an auditory task, paired-sample t-tests compared aggregated music ratings with aggregated control stimulus ratings. Control stimuli consisted of:

Three noise types (variants of white, pink, brown) and six neutral voice recordings.

For each emotional dimension, control ratings were averaged across all control stimuli. Comparisons were conducted separately for major and minor music conditions.

Bonferroni-adjusted significance thresholds were applied. Effect sizes (η²) were computed for all comparisons.

Paired-sample comparisons between aggregated music stimuli and control stimuli confirmed that emotional ratings differed systematically between music and non-musical sounds.

Music stimuli in both major and minor modes were rated significantly higher on sublimity and vitality than control stimuli (all p < .001). For unease, minor-mode music elicited higher ratings than controls, whereas major-mode music stimuli elicited lower unease ratings than controls.

Effect sizes for music versus control comparisons ranged from moderate to large. [68]. See the supporting information tables (S1 File) for a detailed analysis of all the second-order factors compared to control stimuli. A visual representation is presented (Fig 6).

Fig 6. Comparison of music and control stimuli across emotional dimensions.

Fig 6

Mean ratings for musical stimuli (aggregated across tempi and mode) and non-musical control stimuli (voice and noise) are shown for sublimity, unease, and vitality. Musical ratings were averaged across all tempi–mode conditions; control ratings were averaged across all control stimuli. Error bars represent ±1 standard error of the mean. N = 1289 participants.

Discussion

Overall, the findings were consistent with our four hypotheses. Faster tempi were associated with significantly higher vitality ratings, with a medium-to-large effect size. Slower tempi were associated with higher sublimity ratings, although these effects were small in magnitude. Minor mode stimuli elicited substantially higher unease ratings than major mode stimuli, representing the largest observed effect. These patterns were robust across the aggregated stimulus conditions, although the magnitude of effects varied across emotional dimensions. This finding is consistent with earlier research, which suggests that listeners often report experiencing negative emotions in music, such as sadness and tension, which is tied to the aesthetic appreciation rather than utilitarian goals [70].

Additionally, in our prior study, available as a preprint [62], which included tempi, mode, and dynamics as variables, tempi and mode similarly emerged as the primary influences, albeit with a smaller sample size. With the high sample size of the current study, our results had a 2.68% chance of incorrectly rejecting the null hypothesis. In the past century, psychological research tended to have small effect sizes, which often comes from small sample sizes [71]. Low power reduces the likelihood of finding a true effect [72]. With our study, we confirm previous research on emotion and tempi and mode, with high power, large sample size, and rigorously tested compositions and control stimuli.

In addition, we pilot tested our compositions, which has been suggested is important to control individual elements for emotional valence in music [34]. Yet only 1% of music stimuli is composed for studies, and only 3% of music stimuli is pilot tested [4]. We hope that our study contributes to the field by building on existing knowledge.

We implemented controls to address the possibility of compositions having a disproportionate effect, receiving higher ratings than control stimuli, except for major unease, where control stimuli received higher ratings. These findings might suggest that our music stimuli in a major mode were not rated high for unease, which aligns with our hypothesis and our other findings. Although the control stimuli may have also influenced the results to some extent, the use of rigorous control stimuli appears to be a beneficial methodological procedure.

Fast tempi have consistently been linked with joyful emotions [42], evoking a sensation of vitality, joyful activation, and power. This is in line with Gagnon & Peretz´s study, where it was reported that tempi and mode significantly contribute to happy/sad judgments, with tempi being more salient [73]. Similarly, fast tempi are associated with happiness, while slow tempi are linked to peacefulness [29]. Furthermore, patients with cochlear implants found that tempi were the most important parameter for rating music as happy when listening only with the Cochlear implant ear [35]. These studies, taken together with our findings, show that fast tempi are an important parameter in the perception of happiness/joy in music.

Comparing our study to others that do not employ the second-order factors of the GEMS-9 might pose challenges due to variations in emotional wordings. In the GEMS-9, emotional words associated with sublimity include wonder, transcendence, tenderness, nostalgia, and peacefulness. Previous research has demonstrated that slow tempi can evoke feelings of peace and tenderness [32]. This is in alignment with our study, where slow tempi were associated with sublimity ratings. Furthermore, fast tempi align with happiness and activation [32], which are subcategories of vitality. In our study, we found that fast tempi evoked ratings of vitality. The consistency across our study and previous research underscores the robust impact of tempi and mode on the perception of emotional content in music.

Furthermore, the CODA model encourages listeners to construct emotional experiences in response to music elements, such as tempi and mode [17]. There is a continuous appraisal process, which, in our study, is reflected in the changes of the ratings along with the changing tempi.

In our study, we found several significant differences related to gender, age, and education, although most of these had small effect sizes. It is crucial to acknowledge that small effect sizes indicate a relatively small impact on the results. With a sufficiently large sample, statistical methods will almost always demonstrate a significant difference unless there is no effect whatsoever [74].

For the current study, we aimed to measure perceived emotion, using a scale that was designed to measure both perceived and induced emotion. The music stimuli and the study were designed as part of a larger project, where both induction and perception would be measured. Participants were simply asked to give their opinions. We acknowledge that the current study might not be a pure perception study, and it was not intended to be so.

Age

Most age differences in our study had small effects, except for unease in minor 120 and 150 BPM, which exhibited medium effects. The youngest age group (aged 18–35) had the highest ratings for compositions of unease in minor at 120 and 150 BPM, and the oldest age group (+60) had the lowest ratings for compositions of unease in minor at 120 and 150 BPM. These observations somewhat deviated from one study [26] where younger individuals showed more extreme reactions to music. Notably, our younger group demonstrated stronger ratings for unease minors at high tempi, suggesting potential age-related variations in the emotional ratings of unease. However, music likely operates at multiple levels, biological, psychological, and cultural [75], which means we may get emotional cues from tempi and mode, although there will always be individual differences, for example, between age groups.

Contrary to Peace and Halpern´s findings, our study somewhat aligns with another study [52], showing that adolescents and older adults tend to perceive music more positively compared to young and middle-aged adults. Additionally, a decrease in intense and contemporary music genre preferences with age and an increase in unpretentious and sophisticated music (genres from the MUSIC model [76]) has been shown [53]. Furthermore, older adults did not discriminate arousal differences for peaceful and threatening music compared to younger adults [77]. Similarly, compared to younger listeners, older participants rate consonant chords (typical for major mode) less pleasant and dissonant chords (typical for minor mode) more pleasant [78]. Middle-aged adults (mean age 47) rated sadness and fear lower than young adults (mean age 24) [53]. These results are somewhat like our findings, as unease represents sadness and fear.

Interestingly, other studies have found no effect on age and music perception [79,80]. These nuanced age-related differences underscore the complex interplay between age and emotional responses to music. It’s also worth noting that our age groups did not have an equal number of participants, with the oldest age group only consisting of 180 participants, compared to the youngest group with over 600 participants, and the middle age group with around 400. As most of our age-related differences had small effect sizes, it might indicate an age sensitivity for tempi and mode, but more research is needed to explore this further.

Limitations

One limitation of our study is the wording of the translated GEMS-9 scale in Norwegian. Using non-everyday words to describe emotions in Norwegian may have introduced a potential source of bias or misunderstanding. Additionally, a recent study [81] found correlations between rating music with emotional words and taste, which means we may not be constricted by language when attributing qualities to music. A valuable avenue for future research could involve introducing a dedicated potential source of bias or misunderstanding. Future research could involve a dedicated study evaluating and refining the emotional wording in the translated Norwegian version of GEMS-9, comparing it to the original English version of GEMS-9. Similar efforts have been undertaken for the German and English translations [82].

Additionally, the online nature of our study introduced limitations related to the control over participants´ environmental conditions. While participants were instructed to use headphones and take the survey in a quiet environment, the online setting made it impossible to check if the instructions were followed. Factors such as changes in volume, headphone usage, and ambient noise could have influenced participants’ experiences. Although efforts were made to ensure participants completed the survey in one sitting, we could not exclude them or the possibility that some participants left the page open on their computers for extended periods. These aspects highlight the inherent challenges and limitations associated with conducting online studies. However, another study [83] has pointed out the benefit of doing an online survey, such as higher ecological validity.

Another limitation is using the second-order factors of the GEMS-9. This is also a strength which makes our study more comparable to others that use dimensional models. However, in the confirmatory factor analysis correlation of the development of the GEMS-9, results indicated that the second-order factor unease was relatively low for the feeling of sadness, but perfect for tension [19] However, it has been shown that it is less common to feel negative emotions for music listening, such as sad, angry, tense, disgust, and guilt, compared to positive emotions such as happy, nostalgia, calm, loving, tender, and amused [84].

Using simple MIDI piano beats with constant rhythms, we reduced confounding music parameters variables, but they were less ecologically valid. Despite using a control design setting, this is also a limitation, and some participants expressed that they did not consider the stimulus music. While others expressed that they enjoyed the music very much.

Finally, the noises used for the control stimuli were only like brown, pink, and white noise, limited to a lower frequency range than the actual brown, pink, and white noise. This difference was discovered after the data had been collected and analysed. For the analysis, we combined all the noise stimuli and the voice recordings, which were control stimuli. The emotional ratings for control stimuli were significantly lower than for the music.

Conclusion

In conclusion, our study delved into the intricate interplay of tempi and mode concerning the perception of emotion in original music stimuli. Our findings were aligned with previous research and our three research hypotheses. Fast tempi result in higher ratings for vitality, slower tempi yield higher sublimity ratings, and minor mode leads to higher ratings of unease. Though this is not a novel finding, our data were based on a large sample size, with high power, many large effect sizes, self-composed pilot-tested music, and various control stimuli, all fall under a high methodological standard for research.

It is also interesting that our research aligns with the CODA model to investigate tempi and mode with an emotional construct of appraisal and basic emotion ideas. Tempi and mode are significant factors in shaping how music stimuli is rated for sublimity, unease, and vitality. Minor mode and higher tempi also emerged as important for some age-related differences in uneasiness ratings.

For future research, it would be intriguing to conduct similar large-scale studies with different types of music stimuli, another music evaluation scale, and perhaps in different parts of the world to explore whether there exists a universal trend in the perception of tempi and mode in music. Future research should also strive to pilot test music and use control stimuli, and report the analysis of the control stimuli compared to the music.

Supporting information

S1 File. Supporting information, including access to musical and control stimuli, sheet music for all stimuli, and supplementary tables.

(DOCX)

pone.0348069.s001.docx (834KB, docx)

Acknowledgments

A great thanks to Kjetil Vikene, who programmed the survey, and emphasised the importance of age differences. Thanks to Amalie Gloppen Norheim and Tina Emilie Johansen, who were excellent research assistants and helped with data collection. Thanks to the clinical psychology program students Kasper Bjørø, Elin Maria Kavlie Bryde, Kaja Kåstad Hauge, Bernhard Vestby Edvardsen, Olav Erland, Randi Vaage Nordihus, Hani Kadri Vikebø, Sofie Follestad Kverkild, and Torgeir Jensen, who assisted with the data collection. Thanks to Lydia Brunvoll Sandøy and Tobias Bashevkin for letting me steal their voices.

Data Availability

Data is uploaded to figshare: https://doi.org/10.6084/m9.figshare.30052507.

Funding Statement

This project was financed through a grant from the Research Council of Norway (Grant Number: 217932/F20) awarded to KS. The salary of UF was covered in part by a master’s student scholarship at the University of Bergen, and in part by a grant from the Research Council of Norway (Grant Number: 260576) awarded to Stefan Koelsch. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.

References

  • 1.Lamont A, Eerola T. Music and emotion: Themes and development. Musicae Scientiae. 2011;15(2):139–45. doi: 10.1177/102986491101500201 [DOI] [Google Scholar]
  • 2.Warrenburg LA. Subtle semblances of sorrow: Exploring music, emotional theory, and methodology. The Ohio State University. 2019. [Google Scholar]
  • 3.Ekman P. Basic Emotions. Handbook of Cognition and Emotion. Wiley. 1999. 45–60. doi: 10.1002/0470013494.ch3 [DOI] [Google Scholar]
  • 4.Warrenburg LA. Choosing the right tune: A review of music stimuli used in emotion research. Music Percept. 2020;37(3):240–58. [Google Scholar]
  • 5.Juslin PN, Västfjäll D. Emotional responses to music: The need to consider underlying mechanisms. Behav Brain Sci. 2008;31(5):559–75; discussion 575-621. doi: 10.1017/S0140525X08005293 [DOI] [PubMed] [Google Scholar]
  • 6.Omigie D. Basic, specific, mechanistic? Conceptualizing musical emotions in the brain. J Comp Neurol. 2016;524(8):1676–86. doi: 10.1002/cne.23854 [DOI] [PubMed] [Google Scholar]
  • 7.Laukka P, Eerola T, Thingujam NS, Yamasaki T, Beller G. Universal and culture-specific factors in the recognition and performance of musical affect expressions. Emotion. 2013;13(3):434–49. doi: 10.1037/a0031388 [DOI] [PubMed] [Google Scholar]
  • 8.Cespedes-Guevara J, Eerola T. Music communicates affects, not basic emotions - A constructionist account of attribution of emotional meanings to music. Front Psychol. 2018;9:215. doi: 10.3389/fpsyg.2018.00215 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Russell JA. A circumplex model of affect. J Pers Soc Psychol. 1980;39(6):1161–78. doi: 10.1037/h0077714 [DOI] [Google Scholar]
  • 10.Vuoskoski JK, Eerola T. Domain-specific or not? The applicability of different emotion models. Proceedings of the 10th International Conference on Music Perception and Cognition, 2010. 196–9. [Google Scholar]
  • 11.Kim J, André E. Emotion recognition based on physiological changes in music listening. IEEE Trans Pattern Anal Mach Intell. 2008;30(12):2067–83. doi: 10.1109/TPAMI.2008.26 [DOI] [PubMed] [Google Scholar]
  • 12.Konečni VJ, Brown A, Wanic RA. Comparative effects of music and recalled life-events on emotional state. Psychology of Music. 2008;36(3):289–308. doi: 10.1177/0305735607082621 [DOI] [Google Scholar]
  • 13.Scherer KR. Which emotions can be induced by music?. J New Music Res. 2004;33(3):239–51. doi: 10.1080/0929821042000317822 [DOI] [Google Scholar]
  • 14.Scherer KR. The dynamic architecture of emotion: Evidence for the component process model. Cogn Emot. 2009;23(7):1307–51. doi: 10.1080/02699930902928969 [DOI] [Google Scholar]
  • 15.Scherer KR, Coutinho E (2013) How music creates emotion: A multifactorial process approach. Cochrane T, Fantini B, Scherer KR. The emotional power of music. Oxford: Oxford University Press. 121–45. [Google Scholar]
  • 16.Song Y, Dixon S, Pearce MT, Halpern AR. Perceived and induced emotion responses to popular music. Music Perception. 2016;33(4):472–92. doi: 10.1525/mp.2016.33.4.472 [DOI] [Google Scholar]
  • 17.Lennie TM, Eerola T. The CODA model: A review and skeptical extension. Frontiers in Psychology. 2022;13:822264. doi: 10.3389/fpsyg.2022.822264 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Eerola T, Vuoskoski JK. A review of music and emotion studies. Music Percept. 2013;30(3):307–40. doi: 10.1525/mp.2012.30.3.307 [DOI] [Google Scholar]
  • 19.Zentner M, Grandjean D, Scherer KR. Emotions evoked by the sound of music. Emotion. 2008;8(4):494–521. doi: 10.1037/1528-3542.8.4.494 [DOI] [PubMed] [Google Scholar]
  • 20.Scherer KR, Zentner M. Music-evoked emotions are aesthetic rather than utilitarian. Behav Brain Sci. 2008;31(5):595–6. doi: 10.1017/S0140525X08005505 [DOI] [Google Scholar]
  • 21.Kaelen M, Barrett FS, Roseman L, Lorenz R, Family N, Bolstridge M, et al. LSD enhances the emotional response to music. Psychopharmacology (Berl). 2015;232(19):3607–14. doi: 10.1007/s00213-015-4014-y [DOI] [PubMed] [Google Scholar]
  • 22.Irrgang M, Egermann H. From motion to emotion. PLoS One. 2016;11(7):e0154360. doi: 10.1371/journal.pone.0154360 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Miu AC, Balteş FR. Empathy manipulation impacts music-induced emotions. PLoS One. 2012;7(1):e30618. doi: 10.1371/journal.pone.0030618 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Eerola T, Vuoskoski JK. A comparison of discrete and dimensional models. Psychol Music. 2011;39(1):18–49. doi: 10.1177/0305735610362821 [DOI] [Google Scholar]
  • 25.Vuoskoski JK, Eerola T. Measuring music-induced emotion. Musicae Sci. 2011;15(2):159–73. doi: 10.1177/1029864911403367 [DOI] [Google Scholar]
  • 26.Pearce MT, Halpern AR. Age-related patterns in emotions evoked by music. Psychology of Aesthetics, Creativity, and the Arts. 2015;9(3):248–53. doi: 10.1037/a0039279 [DOI] [Google Scholar]
  • 27.Thaut MH, Trimarchi PD, Parsons LM. Human brain basis of musical rhythm perception. Brain Sci. 2014;4(2):428–52. doi: 10.3390/brainsci4020428 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Dalla Bella S, Peretz I, Rousseau L, Gosselin N. A developmental study of the affective value of tempo and mode in music. Cognition. 2001;80(3):B1-10. doi: 10.1016/s0010-0277(00)00136-0 [DOI] [PubMed] [Google Scholar]
  • 29.Bresin R, Friberg A. Emotion rendering in music: Range and characteristic values of seven musical variables. Cortex. 2011;47(9):1068–81. doi: 10.1016/j.cortex.2011.05.009 [DOI] [PubMed] [Google Scholar]
  • 30.Fernández-Sotos A, Fernández-Caballero A, Latorre JM. Influence of tempo and rhythmic unit in musical emotion regulation. Front Comput Neurosci. 2016;10:80. doi: 10.3389/fncom.2016.00080 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.van der Zwaag MD, Westerink JHDM, van den Broek EL. Emotional and psychophysiological responses to tempo, mode, and percussiveness. Musicae Scientiae. 2011;15(2):250–69. doi: 10.1177/1029864911403364 [DOI] [Google Scholar]
  • 32.Hallam S, Cross I, Thaut M. Oxford handbook of music psychology. Oxford: Oxford University Press. 2009. [Google Scholar]
  • 33.Hofbauer LM, Rodriguez FS. Emotional valence perception in music and subjective arousal: Experimental validation of stimuli. Int J Psychol. 2023;58(5):465–75. doi: 10.1002/ijop.12922 [DOI] [PubMed] [Google Scholar]
  • 34.Hevner K. The affective character of the major and minor modes in music. Am J Psychol. 1935;47(1):103–18. doi: 10.2307/1416710 [DOI] [Google Scholar]
  • 35.D’Onofrio KL, Caldwell M, Limb C, Smith S, Kessler DM, Gifford RH. Musical emotion perception in bimodal patients: relative weighting of musical mode and tempo cues. Front Neurosci. 2020;14:114. doi: 10.3389/fnins.2020.00114 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Morreale F, Masu R, De Angeli A, Fava P. The effect of expertise in evaluating emotions in music. Proceedings of the 3rd International Conference on Music & Emotion (ICME3), Jyväskylä, Finland, 2013. [Google Scholar]
  • 37.Parncutt R (2024) Psychoacoustic foundations of major-minor tonality. Cambridge, MA: MIT Press. [Google Scholar]
  • 38.Parncutt R. The emotional connotations of major versus minor tonality: One or more origins?. Musicae Scientiae. 2014;18(3):324–53. doi: 10.1177/1029864914542842 [DOI] [Google Scholar]
  • 39.Carraturo G, Pando-Naude V, Costa M, Vuust P, Bonetti L, Brattico E. The major-minor mode dichotomy in music perception. Phys Life Rev. 2025;52:80–106. doi: 10.1016/j.plrev.2024.11.017 [DOI] [PubMed] [Google Scholar]
  • 40.Krumhansl CL. Cognitive foundations of musical pitch. Oxford: Oxford University Press. 2001. [Google Scholar]
  • 41.Costa M, Fine P, Ricci Bitti PE. Interval distributions, mode, and tonal strength of melodies as predictors of perceived emotion. Music Perception. 2004;22(1):1–14. doi: 10.1525/mp.2004.22.1.1 [DOI] [Google Scholar]
  • 42.Smit EA, Dobrowohl FA, Schaal NK, Milne AJ, Herff SA. Perceived emotions of harmonic cadences. Music & Science. 2020;3. doi: 10.1177/2059204320938635 [DOI] [Google Scholar]
  • 43.Ramos D, Bueno JLO, Bigand E. Manipulating Greek musical modes and tempo affects perceived musical emotion. Braz J Med Biol Res. 2011;44:165–72. doi: 10.1590/S0100-879X2011007500003 [DOI] [PubMed] [Google Scholar]
  • 44.Balkwill LL, Thompson WF, Matsunaga R. Recognition of emotion in Japanese, Western, and Hindustani music by Japanese listeners. Jpn Psychol Res. 2004;46(4):337–49. [Google Scholar]
  • 45.Fang L, Shang J, Chen N. Perception of western musical modes: A Chinese study. Front Psychol. 2017;8:1905. doi: 10.3389/fpsyg.2017.01905 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Pereira LAS, Ramos D, Bueno JLO. The influence of different musical modes and tempi on time perception. Acta Psychol (Amst). 2022;229:103701. doi: 10.1016/j.actpsy.2022.103701 [DOI] [PubMed] [Google Scholar]
  • 47.Chubb C, Dickson CA, Dean T, Fagan C, Mann DS, Wright CE, et al. Bimodal distribution of performance in discriminating major/minor modes. J Acoust Soc Am. 2013;134(4):3067–78. doi: 10.1121/1.4816546 [DOI] [PubMed] [Google Scholar]
  • 48.Karageorghis CI, Jones L, Low DC. Relationship between exercise heart rate and music tempo preference. Res Q Exerc Sport. 2006;77(2):240–50. doi: 10.1080/02701367.2006.10599357 [DOI] [PubMed] [Google Scholar]
  • 49.Robazza C, Macaluso C, D’Urso V. Emotional reactions to music by gender, age, and expertise. Percept Mot Skills. 1994;79(2):939–44. doi: 10.2466/pms.1994.79.2.939 [DOI] [PubMed] [Google Scholar]
  • 50.Cohrdes C, Wrzus C, Wald-Fuhrmann M, Riediger M. The sound of affect: Age differences in perceiving valence and arousal in music. Musicae Sci. 2018;24(1):21–43. doi: 10.1177/1029864918765613 [DOI] [Google Scholar]
  • 51.Bonneville-Roussy A, Rentfrow PJ, Xu MK, Potter J. Music through the ages: Trends in musical engagement and preferences. J Pers Soc Psychol. 2013;105(4):703–17. doi: 10.1037/a0033770 [DOI] [PubMed] [Google Scholar]
  • 52.Vieillard S, Gilet AL. Age-related differences in affective responses to music. Front Psychol. 2013;4:711. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Castro SL, Lima CF. Age and musical expertise influence emotion recognition in music. Music Perception. 2014;32(2):125–42. doi: 10.1525/mp.2014.32.2.125 [DOI] [Google Scholar]
  • 54.Koelsch S, Jäncke L. Music and the heart. Eur Heart J. 2015;36(44):3043–9. doi: 10.1093/eurheartj/ehv430 [DOI] [PubMed] [Google Scholar]
  • 55.Chanda ML, Levitin DJ. The neurochemistry of music. Trends Cogn Sci. 2013;17(4):179–93. doi: 10.1016/j.tics.2013.02.007 [DOI] [PubMed] [Google Scholar]
  • 56.Hunt AM, Legge AW. Neurological research on music therapy for mental health. Music Therapy Perspectives. 2015;33(2):142–61. doi: 10.1093/mtp/miv017 [DOI] [Google Scholar]
  • 57.Mallik A, Russo FA. Effects of music and auditory beat stimulation on anxiety. PLoS One. 2022;17(3):e0259312. doi: 10.1371/journal.pone.0259312 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Groarke JM, Groarke A, Hogan MJ, Costello L, Lynch D. Does listening to music regulate negative affect?. Applied Psychology: Health and Well-Being. 2020;12(2):288–311. doi: 10.1111/aphw.12185 [DOI] [PubMed] [Google Scholar]
  • 59.Hahn L, Buttlar B, Künne R, Walther E. TUNA database. PLoS One. 2024;19(5):e0302904. doi: 10.1371/journal.pone.0302904 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Apple Inc. Logic Pro X. 2024. [Google Scholar]
  • 61.Quinn S, Watt R. The perception of tempo in music. Perception. 2006;35(2):267–80. doi: 10.1068/p5353 [DOI] [PubMed] [Google Scholar]
  • 62.Færøvik U, Specht K. Development of ecologically valid and standardized music stimuli. Research Square. 2025. 10.21203/rs.3.rs-6510451/v1 [DOI]
  • 63.Law LNC, Zentner M. Profile of music perception skills. PLoS One. 2012;7(12):e52508. doi: 10.1371/journal.pone.0052508 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Faul F, Erdfelder E, Lang AG, Buchner A. G*Power 3. Behav Res Methods. 2007;39(2):175–91. doi: 10.3758/BF03193146 [DOI] [PubMed] [Google Scholar]
  • 65.Faul F, Erdfelder E, Buchner A, Lang AG. G*Power 3.1. Behav Res Methods. 2009;41(4):1149–60. doi: 10.3758/BRM.41.4.1149 [DOI] [PubMed] [Google Scholar]
  • 66.Goodman LA. Snowball sampling. Ann Math Stat. 1961;32(1):148–70. doi: 10.1214/aoms/1177705148 [DOI] [Google Scholar]
  • 67.Wilkinson GN, Rogers CE. Symbolic description of factorial models. J R Stat Soc C. 1973;22(3):392–9. doi: 10.2307/2346786 [DOI] [Google Scholar]
  • 68.Benjamini Y, Hochberg Y. Controlling the false discovery rate. J R Stat Soc B. 1995;57(1):289–300. [Google Scholar]
  • 69.Cohen J. Statistical power analysis for the behavioral sciences. 2nd ed. Hillsdale, NJ: Erlbaum. 1988. [Google Scholar]
  • 70.Eerola T, Vuoskoski JK, Peltola HR, Putkinen V, Schäfer K. Enjoyment of sadness in music. Phys Life Rev. 2018;25:100–21. doi: 10.1016/j.plrev.2017.11.016 [DOI] [PubMed] [Google Scholar]
  • 71.Richard FD, Bond CFJ, Stokes-Zoota JJ. One hundred years of social psychology. Review of General Psychology. 2003;7(4):331–63. doi: 10.1037/1089-2680.7.4.331 [DOI] [Google Scholar]
  • 72.Button KS, Ioannidis JPA, Mokrysz C, Nosek BA, Flint J, Robinson ESJ. Power failure. Nat Rev Neurosci. 2013;14(5):365–76. doi: 10.1038/nrn3475 [DOI] [PubMed] [Google Scholar]
  • 73.Gagnon L, Peretz I. Mode and tempo contributions to emotion judgments. Cogn Emot. 2003;17(1):25–40. doi: 10.1080/02699930302279 [DOI] [PubMed] [Google Scholar]
  • 74.Sullivan GM, Feinn R. Using effect size. J Grad Med Educ. 2012;4(3):279–82. doi: 10.4300/JGME-D-12-00156.1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 75.Reybrouck M, Eerola T. Music and its inductive power. Frontiers in Psychology. 2017;8:494. doi: 10.3389/fpsyg.2017.00494 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76.Rentfrow PJ, Goldberg LR, Levitin DJ. Structure of musical preferences. J Pers Soc Psychol. 2011;100(6):1139–57. doi: 10.1037/a0022406 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77.Vieillard S, Didierjean A, Maquestiaux F. Changes in musical emotion perception with age. Exp Aging Res. 2012;38(4):422–41. doi: 10.1080/0361073X.2012.699371 [DOI] [PubMed] [Google Scholar]
  • 78.Bones O, Plack CJ. Aging and perception of musical harmony. J Neurosci. 2015;35(9):4071–80. doi: 10.1523/JNEUROSCI.3509-14.2015 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79.Lahdelma I, Eerola T. Single chords convey emotional qualities. Psychol Music. 2016;44(1):37–54. doi: 10.1177/0305735614568886 [DOI] [Google Scholar]
  • 80.Halpern AR, Bartlett JC, Dowling WJ. Perception of mode and rhythm. Music Percept. 1998;15(4):335–55. doi: 10.2307/40285779 [DOI] [Google Scholar]
  • 81.Guedes D, Prada M, Garrido MV, Lamy E. Taste & affect music database. Behav Res Methods. 2023;55(3):1121–40. doi: 10.3758/s13428-022-01879-4 [DOI] [PubMed] [Google Scholar]
  • 82.Lykartsis A, von Coler H, Lepa S. Emotionality of sonic events. Proceedings of the 3rd International Conference on Music & Emotion (ICME3), Jyväskylä, Finland, 2013. [Google Scholar]
  • 83.Honing H, Ladinig O. The potential of the internet for music perception research. 2008.
  • 84.Laukka P. Uses of music and psychological well-being among the elderly. J Happiness Stud. 2006;8(2). doi: 10.1007/s10902-006-9024-3 [DOI] [Google Scholar]

Decision Letter 0

Andrea Schiavio

14 Apr 2025

-->PONE-D-25-05889-->-->The Effect of Tempo and Mode on the Rating of the Perceived Emotion in Music-->-->PLOS ONE

Dear Dr. Færøvik,

Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process.

Please submit your revised manuscript by May 29 2025 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you’re ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the ‘Submissions Needing Revision’ folder to locate your manuscript file.

Please include the following items when submitting your revised manuscript:-->

  • A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled ‘Response to Reviewers’.

  • A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled ‘Revised Manuscript with Track Changes’.

  • An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled ‘Manuscript’.

If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols. Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols.

We look forward to receiving your revised manuscript.

Kind regards,

Andrea Schiavio

Academic Editor

PLOS ONE

Journal requirements:

When submitting your revision, we need you to address these additional requirements.

1. Please ensure that your manuscript meets PLOS ONE’s style requirements, including those for file naming. The PLOS ONE style templates can be found at

https://journals.plos.org/plosone/s/file?id=wjVg/PLOSOne_formatting_sample_main_body.pdf and

https://journals.plos.org/plosone/s/file?id=ba62/PLOSOne_formatting_sample_title_authors_affiliations.pdf.

2. We note that the grant information you provided in the ‘Funding Information’ and ‘Financial Disclosure’ sections do not match.

When you resubmit, please ensure that you provide the correct grant numbers for the awards you received for your study in the ‘Funding Information’ section.

3. Thank you for stating the following financial disclosure:

[copy in funding statement].

Please state what role the funders took in the study.  If the funders had no role, please state: ""The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.""

If this statement is not correct you must amend it as needed.

Please include this amended Role of Funder statement in your cover letter; we will change the online submission form on your behalf.

4. In the online submission form, you indicated that [Data can be included in supplementary materials upon publication, or in another format upon request.].

All PLOS journals now require all data underlying the findings described in their manuscript to be freely available to other researchers, either 1. In a public repository, 2. Within the manuscript itself, or 3. Uploaded as supplementary information.

This policy applies to all data except where public deposition would breach compliance with the protocol approved by your research ethics board. If your data cannot be made publicly available for ethical or legal reasons (e.g., public availability would compromise patient privacy), please explain your reasons on resubmission and your exemption request will be escalated for approval.

5.  Please include captions for your Supporting Information files at the end of your manuscript, and update any in-text citations to match accordingly. Please see our Supporting Information guidelines for more information: http://journals.plos.org/plosone/s/supporting-information.

[Note: HTML markup is below. Please do not edit.]

Reviewers' comments:

Reviewer's Responses to Questions

-->Comments to the Author

1. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented. -->

Reviewer #1: Yes

Reviewer #2: Partly

**********

-->2. Has the statistical analysis been performed appropriately and rigorously? -->

Reviewer #1: I Don't Know

Reviewer #2: Yes

**********

-->3. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.-->

Reviewer #1: Yes

Reviewer #2: Yes

**********

-->4. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.-->

Reviewer #1: No

Reviewer #2: Yes

**********

-->5. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)-->

Reviewer #1: Overall this is an interesting study but the manuscript needs a bit more work, I think. I have a small question about the statistical analysis, hence my response to the prior question. I see the authors have promised to include the data upon publication. My comments are bellow.

The authors present an interesting study that develops on existing foundations in emotion perception in music. They acknowledge that this is a partial replication, delving into familiar ground but with some advantages such as having a large sample size. Their method deviates somewhat from prior studies in their choice of ratings terminology, utilising bespoke compositions, and different control stimuli. While there are clear advantages to this, I feel the authors could do more to contextualise their research and emphasise their contribution to the field. The writing can be a bit vague and disjointed, and certain theoretical concepts are not very well explained. The results could be presented in a clearer and more concise way to help the reader grasp the findings. I think there is clearly a valuable source of data here, and the article can be improved with more consideration to its presentation and contextualisation in the research field. More details as follows:

Introduction

In your discussion of different theories of emotion, I think there may be some confusion between notions of discrete and dimensional models of emotion and emotivist and cognitivist theories of emotions, or at least the descriptions and distinctions of different concepts (how we conceptualise emotions in music and how we measure emotions in music) are not very clear. You may want to review this. You might also consider delving more into constructionist theories and their implications as posited by Cespedes-Guevara & Eerola, who you cite. See also Lennie & Eerola, 2022.

In your discussion of different measures, you point out that Eerola & Vuoskoski found that a dimensional model of valence, energy, and tension outperformed other models, including the GEMS-9. I think you can better explain why you still chose to use the GEMS-9 scale: why not adopt valence (sublime), energy (vital), and tension (uneasy)? Did you reflect on two-dimensional models of affect (valence and arousal)? One point you make is that the GEMS scales are designed for measuring either evoked (felt?) or perceived emotion – is that not the case for other scales?

Method

It would be nice to know a bit more about the compositions. I appreciate they may have been used in other studies and those articles may have more details, but a summary here, if only brief, would be helpful as well to give a better sense of the intentions/design, including a consideration of any other musical factors that might contribute and/or need controlling, such as dynamics, rhythm, and timbre. For example, I note having listened to some of the samples through the Dropbox link that they were MIDI piano tracks with constant crotchet beat rhythms. This could be important for considering the research against other studies with different kinds of stimuli.

I think there needs to be a bit more detail on the choice of ratings questions, as mentioned above. I would clarify the ratings scale and details of how the measure is presented in the Materials section, not in the Procedure as it is currently, to make it clear and concise. There is also mention of additional synonyms being included; what were they? In the supplementary material I can see the nine primary factors, and it strikes me that there are five for sublime but only two for the others. Was anything considered to balance this out? Was this the reason for the additional synonyms?

Analysis

If I have understood correctly, you have aggregated the ratings of individual compositions by tempo and mode groupings. Did you test if there was any effect of composition? I can see the justification for this, but it would be good to see it clearly explained in the text. This relates to the question about knowing the details of the compositions and being able to interpret or consider these aspects of the methodology.

Results

I think you can make the results section a bit clearer by perhaps summarising the numerical outputs/ANOVA results in a table. It is a bit unwieldly in its current form. You have included similar tables in the supplementary documents. I also wonder if you have sufficiently corrected for multiple testing – you include a correction in your t-tests for the control conditions, did you consider applying any corrections across your other tests?

Discussion

The first three paragraphs are a repetition of the results and shouldn’t be necessary in a Discussion section.

I would be interested to hear more about the advantages of your study (e.g., large sample size) and where you believe it has especially contributed to the field. You acknowledge early in the paper that this is something of a replication and you compare your results against previous work; what are the new implications of your study? Why are your results important? Are there any methodological recommendations you can make?

Limitations

You mention some limitations of the GEMS scale factors, and you acknowledge in your Discussion that there are challenges related to variations of emotion terms. I think you can acknowledge this further. The last paragraph of your Limitations section describes being able to compare with dimensional ratings, but while the three categories you’ve chosen may be considered relevant to valence (sublime?), energy (vitality) and tension (uneasy), for example, the specific terms may have different connotations for your participants. You have tried to control this with the addition of other terms and synonyms in the presentation of the ratings, but it could still be a limiting factor. Are there other ways your choice of terms could be a strength?

As mentioned previously, your choice of music is another important factor in this study. What are the strengths and limitations of the stimuli selection? What about ecological validity, perhaps?

Small points

Presentation of figures, tables and graphs – could be tidied up a bit to make things more presentable, for example removing the _ in the labels/text. Consider recommended formatting styles e.g., APA. Graphs could be more distinguished with better labelling or grouping – if you’ve labelled the bars, I don’t think you need to include a legend, or vice versa. You might want to consider other ways of identifying/grouping variables to make it clearer – figures 1-3 could show bars grouped by tempo, colour coded by mode, for example.

Some typographical errors throughout, for example in the Introduction; ‘Genova’ instead of ‘Geneva’; Results section: ‘See table 1. For mean and standard deviations. See figure 1. For visual representation.’ – capitalisation following the full stop after each number; missing ‘p’ when reporting the p value in most of the stat reporting. Generally, the article could be revised for clarity and precision in the writing.

I hope my comments are helpful. Thank you for your work.

Reviewer #2: This is a paper of moderate relevance. Though the paper provides a lot of data with a huge number of participants (N = 1280), there are not many new ideas that may trigger the interest of the reader. The statistical analysis seems to be sound but what is missing to some extent is a broadening of scope and some background explanation and broader positioning of the findings. The findings are not really innovative, and it can be questioned what is really new: finding that tempo and mode influence the ratings of music is quite common knowledge. A more critical elaboration and interpretation of the findings would make the paper stronger.

General remarks

• The major strength of the paper is the huge number of participants.

• The coherence of the collected data from the literature review is not totally convincing.

• Some terms are used in a rather loose way and should be defined more strictly. E.g., what is meant with “uneasy music”? Focal concepts such as sublimity, unease, and vitality should also be explained more in detail.

• The statistics must be explained somewhat more in detail. How are the values for the ratings computed?

• The reference list is substantial but more original sources could be mentioned, especially with regard to the distinction between discrete and dimensional approach to emotions. This holds in particular for the work of Scherer.

• Try to avoid references of papers that are submitted but not yet in press.

• The most important take home message of the paper is not very strong and does not add very much to already existing knowledge.

• In its current form the paper seems to focus more on effects sizes rather than on critical elaboration of the findings. Some more in-depth discussions of the findings should make the paper stronger.

• Stronger motivations could be given for some methodological choices, such as, e.g., the use of second-order factors.

• The conclusion is very short. A stronger take home message should be given. What are the major findings? What is new? What is different from current knowledge?

Detailed comments

• Page 9, last sentence: this sentence seems to be grammatically incomplete. Please reword.

• Page 10, 2nd paragraph: Please add additional basic references for the description of discrete emotion as applied to music (Eerola; Reybrouck, and others, see suggestions below)

• Page 10, last par.: reference should be made also to the so-called “aesthetic emotions”. Reference should be made to Scherer. There is a comma lacking after (GEMS-9). Please explain somewhat more in detail what is meant with a domain-specific model.

• Page 11, 1st par.: Please explain somewhat more in detail the meaning of the second-order factors and their relation to the first-order factors.

• Page 11, 3rd par.: please substitute “valence” for “valance”. This holds also for all other appearances.

• Page 13: hypotheses: the 3rd hypothesis is very generalizing; the concept of “unease” must be better explained.

• Page 13, penultimate par.: the categories of uneasy, sublime and vital must be much better specified.

• Page 14, 1st par.: it should be explained more clearly that for each of the three neutral voice recordings there was one male and one female speaker.

• Page 14: it is not easy to identify the uploaded music examples: the used abbreviations to tag the examples should be explained somewhat more in detail. For example, what is the meaning of B_N, P_N, SUBLIM, UROLIG, V-1, etc. Please try to be as clear as possible to avoid seeking efforts by the readers.

• Page 15, 1st. par.: The target was set at 1562. Please explain shortly why this number was chosen and what was the rationale behind. This can be very short but try to be intuitive.

• Page 17, results: this is a very abrupt beginning of presenting the results. Please insert a short initial text that announces that the three variables (sublime, uneasy, vital) will be described.

• Page 17, table 1. It is not directly clear how the values (ratings) for the mean have been computed. Please explain as clearly as possible.

• Page 18, Figure 1. Is it possible to insert the significance level in the figure y using *, **, ***?

• Page 25, 2nd par.: there seems to be a contradicting assertion here: an increase in both unpretentious and sophisticated music. Or these categories not opposed to each other?

Suggestions for additional rerences (not mandatory)

Eerola, T. & Vuoskoski, J. (2013) A Review of Music and Emotion Studies: Approaches, Emotion Models, and Stimuli. Music Perception, 30(3), 307–340. DOI: 10.1525/MP.2012.30.3.307

Eerola, T., Vuoskoski, J., Peltola, H.-R., Putkinen, V., Schäfer, K. (2018). An integrative review of the enjoyment of sadness associated with music. Physics of Life Reviews 25, 100–121. https://doi.org/10.1016/j.plrev.2017.11.016

Lamont, A. & Eerola, T. (2011). Music and emotion: Themes and development. Musicae Scientiae,15 (2), 139-145.

Reybrouck, M. & Eerola, T. (2017). Music and its inductive power: a psychobiological and evolutionary approach to musical emotions. Frontiers in Psychology, 8, Art.No. 494. Open Access

Reybrouck, M., Eerola, T. (2017). Music and its inductive power: a psychobiological and evolutionary approach to musical emotions. Frontiers in Psychology, 8, Art. No. 494. Open Access

Robinson, J. (2009). Aesthetic emotions (philosophical perspectives). In D. Sander & K. R. Scherer (Eds.), The Oxford companion to emotion and the affective sciences (pp. 6–9). New York: Oxford University Press.

Scherer, K. (2008). Music evoked emotions are different – more often aesthetic than utilitarian. Behavioral and Brain Sciences, 31, 5, 595.

**********

-->6. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.-->

Reviewer #1: Yes:Rory Kirk

Reviewer #2: No

**********

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at figures@plos.org. Please note that Supporting Information files do not need this step.

PLoS One. 2026 May 18;21(5):e0348069. doi: 10.1371/journal.pone.0348069.r002

Author response to Decision Letter 1


5 Sep 2025

Dear editor,

We thank you and the reviewers for the thoughtful and constructive feedback on our manuscript, titled The Effect of Tempo and Mode on the Rating of the Perceived Emotion in Music. We have revised the manuscript thoroughly in response to the comments and believe the changes have substantially improved its clarity, structure, and contribution to the field. Below, we address each point raised by the reviewers, with our responses following each comment in bold.

Reviewer #1

Overall this is an interesting study but the manuscript needs a bit more work, I think. I have a small question about the statistical analysis, hence my response to the prior question. I see the authors have promised to include the data upon publication. My comments are bellow.

We thank the reviewer for this encouraging overall assessment and for their valuable suggestions. We have addressed each point in detail below.

The authors present an interesting study that develops on existing foundations in emotion perception in music. They acknowledge that this is a partial replication, delving into familiar ground but with some advantages such as having a large sample size. Their method deviates somewhat from prior studies in their choice of ratings terminology, utilising bespoke compositions, and different control stimuli. While there are clear advantages to this, I feel the authors could do more to contextualise their research and emphasise their contribution to the field. The writing can be a bit vague and disjointed, and certain theoretical concepts are not very well explained. The results could be presented in a clearer and more concise way to help the reader grasp the findings. I think there is clearly a valuable source of data here, and the article can be improved with more consideration to its presentation and contextualisation in the research field. More details as follows:

Introduction

In your discussion of different theories of emotion, I think there may be some confusion between notions of discrete and dimensional models of emotion and emotivist and cognitivist theories of emotions, or at least the descriptions and distinctions of different concepts (how we conceptualise emotions in music and how we measure emotions in music) are not very clear. You may want to review this. You might also consider delving more into constructionist theories and their implications as posited by Cespedes-Guevara & Eerola, who you cite. See also Lennie & Eerola, 2022.

We revised the Introduction to clearly distinguish between discrete and dimensional models of emotion and included a new discussion on constructionist perspectives, citing Cespedes-Guevara & Eerola (2018) and Lennie & Eerola (2022). We now explain how these models relate to GEMS-9 and clarify the conceptual foundation of the study.

In your discussion of different measures, you point out that Eerola & Vuoskoski found that a dimensional model of valence, energy, and tension outperformed other models, including the GEMS-9. I think you can better explain why you still chose to use the GEMS-9 scale: why not adopt valence (sublime), energy (vital), and tension (uneasy)? Did you reflect on two-dimensional models of affect (valence and arousal)? One point you make is that the GEMS scales are designed for measuring either evoked (felt?) or perceived emotion – is that not the case for other scales?

We now clearly articulate the rationale for selecting GEMS-9 second-order factors. This includes their focus on domain-specific aesthetic emotions and the flexibility to compare with both dimensional and discrete approaches. We also cite Pearce & Halpern (2015) and Zentner et al. (2008) to support this choice and clarify that other scales do not all support both perceived and felt emotion as GEMS-9 does.

Method

It would be nice to know a bit more about the compositions. I appreciate they may have been used in other studies and those articles may have more details, but a summary here, if only brief, would be helpful as well to give a better sense of the intentions/design, including a consideration of any other musical factors that might contribute and/or need controlling, such as dynamics, rhythm, and timbre. For example, I note having listened to some of the samples through the Dropbox link that they were MIDI piano tracks with constant crotchet beat rhythms. This could be important for considering the research against other studies with different kinds of stimuli.

We have expanded the Methods section to describe the compositions in more detail, including the use of a MIDI piano with a constant crotchet rhythm, and our rationale for this choice. We also note that the stimuli were pilot-tested to confirm beat perception consistency.

I think there needs to be a bit more detail on the choice of ratings questions, as mentioned above. I would clarify the ratings scale and details of how the measure is presented in the Materials section, not in the Procedure as it is currently, to make it clear and concise. There is also mention of additional synonyms being included; what were they? In the supplementary material I can see the nine primary factors, and it strikes me that there are five for sublime but only two for the others. Was anything considered to balance this out? Was this the reason for the additional synonyms?

We moved the GEMS-9 scale description to the Materials section and clarified the synonyms used (now also listed in the supplementary materials). We address the asymmetry in the number of terms, especially for the Sublime factor, and reflect on how this may relate to the valence biases in music-related vocabulary.

Analysis

If I have understood correctly, you have aggregated the ratings of individual compositions by tempo and mode groupings. Did you test if there was any effect of composition? I can see the justification for this, but it would be good to see it clearly explained in the text. This relates to the question about knowing the details of the compositions and being able to interpret or consider these aspects of the methodology.

We did not test for individual composition effects, as our design aggregated ratings by tempo and mode to strengthen generalizability. We now explain this choice and rationale more explicitly in the analysis section.

Results

I think you can make the results section a bit clearer by perhaps summarising the numerical outputs/ANOVA results in a table. It is a bit unwieldly in its current form. You have included similar tables in the supplementary documents. I also wonder if you have sufficiently corrected for multiple testing – you include a correction in your t-tests for the control conditions, did you consider applying any corrections across your other tests?

We restructured the Results section, added a comprehensive ANOVA summary table, and clarified the use of Bonferroni correction. Effect sizes are also reported throughout, and all adjustments are now mentioned in the Analysis section.

Discussion

The first three paragraphs are a repetition of the results and shouldn’t be necessary in a Discussion section.

I would be interested to hear more about the advantages of your study (e.g., large sample size) and where you believe it has especially contributed to the field. You acknowledge early in the paper that this is something of a replication and you compare your results against previous work; what are the new implications of your study? Why are your results important? Are there any methodological recommendations you can make?

We revised the opening of the Discussion to avoid repeating results and rewrote sections to highlight the unique contributions of the study — including the large sample size, controlled stimuli, second-order GEMS factors, and a strong methodological framework. We also articulate the theoretical implications more clearly.

Limitations

You mention some limitations of the GEMS scale factors, and you acknowledge in your Discussion that there are challenges related to variations of emotion terms. I think you can acknowledge this further. The last paragraph of your Limitations section describes being able to compare with dimensional ratings, but while the three categories you’ve chosen may be considered relevant to valence (sublime?), energy (vitality) and tension (uneasy), for example, the specific terms may have different connotations for your participants. You have tried to control this with the addition of other terms and synonyms in the presentation of the ratings, but it could still be a limiting factor. Are there other ways your choice of terms could be a strength?

We have expanded the Limitations section to reflect more critically on the influence of translating GEMS-9 terms into Norwegian. We also argue that using less familiar, nuanced words may have encouraged thoughtful reflection (i.e., System 2 processing), which could be seen as a strength in this context.

As mentioned previously, your choice of music is another important factor in this study. What are the strengths and limitations of the stimuli selection? What about ecological validity, perhaps?

We now explicitly acknowledge that while ecological validity is limited by the use of simplified stimuli, this choice allowed us to isolate the effects of tempo and mode without other musical confounds.

Small points

Presentation of figures, tables and graphs – could be tidied up a bit to make things more presentable, for example removing the _ in the labels/text. Consider recommended formatting styles e.g., APA. Graphs could be more distinguished with better labelling or grouping – if you’ve labelled the bars, I don’t think you need to include a legend, or vice versa. You might want to consider other ways of identifying/grouping variables to make it clearer – figures 1-3 could show bars grouped by tempo, colour coded by mode, for example.

We revised Figures 1–3 for better visual clarity, grouping by tempo and color-coding by mode. Labels were cleaned up, legends clarified, and the overall style now conforms more closely to APA standards.

Some typographical errors throughout, for example in the Introduction; ‘Genova’ instead of ‘Geneva’; Results section: ‘See table 1. For mean and standard deviations. See figure 1. For visual representation.’ – capitalisation following the full stop after each number; missing ‘p’ when reporting the p value in most of the stat reporting. Generally, the article could be revised for clarity and precision in the writing.

These have been corrected throughout the manuscript.

Reviewer #2

This is a paper of moderate relevance. Though the paper provides a lot of data with a huge number of participants (N = 1280), there are not many new ideas that may trigger the interest of the reader. The statistical analysis seems to be sound but what is missing to some extent is a broadening of scope and some background explanation and broader positioning of the findings. The findings are not really innovative, and it can be questioned what is really new: finding that tempo and mode influence the ratings of music is quite common knowledge. A more critical elaboration and interpretation of the findings would make the paper stronger.

We thank Reviewer 2 for their thoughtful critique and useful suggestions. We address each point in turn below.

General remarks

• The major strength of the paper is the huge number of participants.

• The coherence of the collected data from the literature review is not totally convincing.

We have expanded the Introduction and Discussion to more clearly position the findings in the broader literature, including constructionist, evolutionary, and aesthetic perspectives. We also emphasize the value of replication and methodologically rigorous studies, especially with large samples.

• Some terms are used in a rather loose way and should be defined more strictly. E.g., what is meant with “uneasy music”? Focal concepts such as sublimity, unease, and vitality should also be explained more in detail.

These terms are now more clearly defined in relation to their underlying GEMS-9 first-order factors. The introduction has been revised to present these categories with more precision and theoretical grounding.

• The statistics must be explained somewhat more in detail. How are the values for the ratings computed?

We added an explanation of how Mean scores were calculated — specifically, averages across five stimuli per tempo/mode condition — in both the Methods and Results sections.

• The reference list is substantial but more original sources could be mentioned, especially with regard to the distinction between discrete and dimensional approach to emotions. This holds in particular for the work of Scherer.

We have added several references to Scherer’s work (2004, 2008, 2013) to support our discussion of aesthetic emotions and the Component Process Model (CPM).

• Try to avoid references of papers that are submitted but not yet in press.

Unfortunately, this is unavoidable; our other paper is still under review, but we’ve made the citation more transparent. Taking it out would mean duplicating and making this paper longer.

• The most important take home message of the paper is not very strong and does not add very much to already existing knowledge.

The Conclusion has been revised to clearly articulate the main contributions and why they matter — including the role of structure (tempo and mode) in shaping perceived musical emotion, even under tightly controlled conditions.

• In its current form the paper seems to focus more on effects sizes rather than on critical elaboration of the findings. Some more in-depth discussions of the findings should make the paper stronger.

We agree that effect sizes are important but have now added more interpretation and commentary in the Discussion to complement the statistical perspective. Effect sizes emphasise whether the findings are believable or not; they lower the chance of random occurrences. This paper emphasises methodology, which is why the discussion follows methodological approaches.

• Stronger motivations could be given for some methodological choices, such as, e.g., the use of second-order factors.

We now provide a clearer rationale for using second-order factors, including their interpretability and fit with previous studies that used GEMS-9 at multiple levels.

• The conclusion is very short. A stronger take home message should be given. What are the major findings? What is new? What is different from current knowledge?

We revised and lengthened the Conclusion to reflect key findings, their contribution to existing knowledge, and future directions.

Detailed comments

• Page 9, last sentence: this sentence seems to be grammatically incomplete. Please reword.

• Page 10, 2nd paragraph: Please add additional basic references for the description of discrete emotion as applied to music (Eerola; Reybrouck, and others, see suggestions below)

• Page 10, last par.: reference should be made also to the so-called “aesthetic emotions”. Reference should be made to Scherer. There is a comma lacking after (GEMS-9). Please explain somewhat more in detail what is meant with a domain-specific model.

• Page 11, 1st par.: Please explain somewhat more in detail the meaning of the second-order factors and their relation to the first-order factors.

• Page 11, 3rd par.: please substitute “valence” for “valance”. This holds also for all other appearances.

• Page 13: hypotheses: the 3rd hypothesis is very generalizing; the concept of “unease” must be better explained.

• Page 13, penultimate par.: the categories of uneasy, sublime and vital must be much better specified.

• Page 14, 1st par.: it should be explained more clearly that for each of the three neutral voice recordings there was one male and one female speaker.

• Page 14: it is not easy to identify the uploaded music examples: the used abbreviations to tag the examples should be explained somewhat more in detail. For example, what is the meaning of B_N, P_N, SUBLIM, UROLIG, V-1, etc. Please try to be as clear as possible to avoid seeking efforts by the readers.

• Page 15, 1st. par.: The target was set at 1562. Please explain shortly why this number was chosen and what was the rationale behind. This can be very short but try to

Attachment

Submitted filename: Response to reviewers.docx

pone.0348069.s003.docx (26.9KB, docx)

Decision Letter 1

Andrea Schiavio

21 Oct 2025

-->PONE-D-25-05889R1-->-->The Effect of Tempo and Mode on the Rating of the Perceived Emotion in Music-->-->PLOS ONE

Dear Dr. Færøvik,

Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process.-->--> -->-->Following the first round of reviews, your revised manuscript was sent back to the two original reviewers. One requested further amendments, while the other recommended rejection. To ensure a balanced evaluation, I invited a third reviewer to assess the paper, and they have now recommended major revisions. The manuscript will therefore require substantial and careful revision before it can be considered for publication.-->--> -->-->Please submit your revised manuscript by Dec 05 2025 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

Please include the following items when submitting your revised manuscript:-->

  • A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'.

  • A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.

  • An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols. Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols.

We look forward to receiving your revised manuscript.

Kind regards,

Andrea Schiavio

Academic Editor

PLOS ONE

Journal Requirements:

If the reviewer comments include a recommendation to cite specific previously published works, please review and evaluate these publications to determine whether they are relevant and should be cited. There is no requirement to cite these works unless the editor has indicated otherwise.

[Note: HTML markup is below. Please do not edit.]

Reviewers' comments:

Reviewer's Responses to Questions

-->Comments to the Author

1. If the authors have adequately addressed your comments raised in a previous round of review and you feel that this manuscript is now acceptable for publication, you may indicate that here to bypass the “Comments to the Author” section, enter your conflict of interest statement in the “Confidential to Editor” section, and submit your "Accept" recommendation.-->

Reviewer #1: All comments have been addressed

Reviewer #2: (No Response)

Reviewer #3: (No Response)

**********

-->2. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented. -->

Reviewer #1: Yes

Reviewer #2: Partly

Reviewer #3: Yes

**********

-->3. Has the statistical analysis been performed appropriately and rigorously? -->

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: I Don't Know

**********

-->4. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.-->

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: Yes

**********

-->5. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.-->

Reviewer #1: Yes

Reviewer #2: No

Reviewer #3: Yes

**********

-->6. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)-->

Reviewer #1: My thanks to the authors for their work making considerable amendments and improvements to their article. I am broadly satisfied that my comments have been addressed. I particularly appreciate the elaboration of their methods, providing more detail on the stimuli chosen and the justification of their measures.

I still have some questions about some of the contextualisation of the music and emotions literature and the positioning of the study. There still appears to be some blurring of the concepts of discrete and dimensional models (how we measure emotions) and emotivist and cognitivist theories of emotions. For example, Introduction para 3 (or 4? Formatting isn’t clear) reads like it is suggesting that dimensional models cannot be applied to felt emotions, is that correct? Generally, I would advise a thorough proofread as the writing still lacks clarity in many places. There are also some typographical errors and inconsistencies throughout (e.g., use of ‘first-‘ and ‘1st-‘).

I think you can also go further in justifying the value of this contribution. The final paragraph of the Introduction (before the Hypothesis) is quite weak; a very strong case can be made for the need for replications, for example. You do not need to dive deeply into e.g., discussion on the replication crisis, but some consideration of the value of replication generally would I think make your article more convincing as a valuable contribution.

Reviewer #2: This paper seems to delve into an area of investigation that has been studied already extensively before. As such, it should be accepted only if there are some additional new substantial findings. At first sight this seems not be the case. Given that the paper works with a very large number of subjects (N = 1289), it should at least be considered for a first review, but reading the paper is not very convincing. The methodology is not well described and there is a lack of academic standards and academic writing. The methodology, which seems to be quite elaborated (at least the statistical part), is not explained and presented in the needed format, with manh sloppy paragraphs that are difficult to read and understand. In its current form, the paper cannot be accepted for publication. I list below some general remarks and detailed comments to motivate may rather harsh decision.

General remarks

• The contents are of moderate relevance with a lot of repetition of known facts. The major question is: what is new?

• The research questions are very general, which makes them not really suitable for an empirical study.

• The style of writing is not very mature (young scholars?) and the academic standards are not very high.

• The style of referencing is rather weak. Some major references are lacking, there are also multiple references to submitted papers that are not yet accepted.

• The research questions are very general and reductionistic. More motivation is needed to explain why some selections have been made: why only tempo and mode? What about other parameters (dynamics, agogics, timbre, etc.)?

• The methodology seems to have some potential, but the methods themselves are very badly explained. Figures and tables are introduced in the text without the needed accompanying explaining text, which makes it very difficult to interpret them. Many questions can also be raised about the collection of the data, which have been collected in less standardized circumstances.

• It is not all clear how the ratings (mean values) have been computed.

• There is overall too much generalization. For example, speaking of music as a general category is very problematic to be consider as an independent variable. There are so many kinds of music, all with a different potential effect on listeners.

• The paper as whole seems to be an example of method-driven analysis, with a lot of statistical processing and computation, but with a scarcitiy of starting ideas and challenging intuitions. The theoretical background is also rather weak.

Detailed comments

The pages are not numbered, and the lines of the pages are also not numbered. I therefore refer to the page numbering of the pdf file.

• Page 10, 3rd paragraph: the distinction between discrete, categorical models and dimensional models must be elaborated somewhat more in depth. Much more references can be given here.

• Page 10, last par.: explain more in detail to what extent the dimensional models outperform.

• Page 11: explain somewhat more in depth the second-order category. How are they described? Is the description clear enough for the subjects to understand them. A term like sublime may be problematic for many common subjects.

• Page 11, 2nd par.: use “valence” rather than “valance”. This hold for the whole paper.

• Page 12, 3rd par. This is a very strong claim/statement. More references are needed to motivate such a strong position.

• Page 13: research questions are very general and reductionistic. Referring to all compositions: those of the study, or more general? Reducing music to tempo and modes, is also very reductionistic. This should be motivated more strongly. It is questionable that parameters in isolation are the real eliciting factors for some effects.

• Page 13: explain the terms uneasy, sublime and vital in more operational definitions and descriptions. Is there some guarantee that the subjects understand the terms?

• Page 13, last par.: referring to “submitted” references is weak referencing style; this holds for the whole paper.

• Page 17: the results should be presented in a synoptic table, as is mostly done in empirical papers, with all the scores and values brought together, and using *, **, *** for instance to show the significance level. There is also some ambiguity with respect to the effect size: Wilks’ lambda is to be understood as: closer to zero is more significant effect. This seems not to be the case in the results presented here? Also for the partial effects: 0.01 is a small effect size, 0.06 is a medium effect size and 0.1 is a large effect size. This is also not to be found in the presented results. This holds also for the next pages. It should also be explained how the mean values have been computed. All steps must be described much more in detail.

• Page 24, 1st par.: This text should have a better place in the introduction rather than in the discussion.

• Page 24, 2nd par.: same remark

• Pages 32 ff: The figures must be explaineD much more in depth so that they can be interpreted appropriately.

Reviewer #3: This manuscript presents a carefully designed study on how tempo and mode shape emotional responses to music, using controlled original stimuli and the GEMS-9 framework. The dataset is valuable and the analyses are competently executed. However, several conceptual and methodological issues remain insufficiently justified or under-explained. The paper still overstates novelty relative to prior work, and several key methodological choices (use of second-order GEMS factors, averaging across compositions, control stimuli, translation validation) require fuller motivation or additional analyses.

Detailed comments:

1) Rephrase or substantiate the statement that tempo and mode are “inconsistently examined across emotion models.” Either specify the precise inconsistencies the paper addresses (e.g., differing stimuli, measures, or theoretical mappings) or soften the claim to position the work as a large-scale, stimulus-controlled extension of established findings. The manuscript itself cites many prior tempo/mode studies.

2) In the introduction, I recommend streamlining the sections on emotion models to improve logic: briefly introduce discrete, dimensional, and domain-specific approaches, then clearly justify the use of GEMS-9. The current text reads as somewhat disjointed.

3) More importantly, flesh out the motivation for using GEMS second-order factors. Provide a stronger rationale for collapsing GEMS-9 to three second-order factors and report psychometric information, where available (internal reliabilities, inter-factor correlations…).

4) The sentence on emotional differences “emerging accidentally during the 14th–16th centuries” is ambiguous or unclear. Reword for clarity and cite an appropriate historical or psychoacoustic source (e.g., Parncutt, 2024).

5) Replication vs. extension – The study departs from previous methods and thus is not a strict replication. Please revise phrasing to “conceptual replication” or “extension,” and clearly distinguish replicated versus novel aspects.

6) Unclear terminology: “pure major” and “pure minor”

7) I understand the rationale behind using control stimuli, but couldn’t one raise the issue that control stimuli are auditory as well?

8) Describe the translation procedure (forward/back translation, expert review) and report internal consistency statistics.

9) Specify the hosting platform (e.g., Qualtrics, Gorilla, PsychoPy online), playback format, use of headphones...

10) Averaging all five stimuli into single scores affects stimulus variability. Please test for composition effects (e.g., mixed-effects model with composition as a random factor) or at least report variability across compositions to demonstrate generalizability.

11) Add recent and relevant sources on major-minor perception, such as: Carraturo, G. et al. (2024). The major–minor mode dichotomy in music perception. Physics of Life Reviews.Parncutt, R. (2024). Psychoacoustic Foundations of Major–Minor Tonality. MIT Press.

**********

-->7. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.-->

Reviewer #1: Yes:Rory Kirk

Reviewer #2: No

Reviewer #3: No

**********

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

To ensure your figures meet our technical requirements, please review our figure guidelines: https://journals.plos.org/plosone/s/figures

You may also use PLOS’s free figure tool, NAAS, to help you prepare publication quality figures: https://journals.plos.org/plosone/s/figures#loc-tools-for-figure-preparation.

NAAS will assess whether your figures meet our technical requirements by comparing each figure against our figure specifications.

PLoS One. 2026 May 18;21(5):e0348069. doi: 10.1371/journal.pone.0348069.r004

Author response to Decision Letter 2


2 Mar 2026

Thank you to all the reviewers for the hard work and thorough comments. We have made the changes you asked for. In this letter, we detail each change to the separate reviewers.

Reviewer #1

We thank Reviewer #1 for their careful reading of the revised manuscript and for their constructive and encouraging feedback. We are pleased that the reviewer finds the methodological revisions satisfactory and appreciates the added detail regarding stimuli and measures. Below, we respond to the remaining points raised.

Conceptual clarification of emotion models and theories

We are grateful to the reviewer for highlighting the need for clearer conceptual distinctions between discrete versus dimensional emotion models, felt versus perceived emotions, and underlying emotion theories (e.g., emotivist, cognitivist, constructionist). We agree that these distinctions were not sufficiently explicit in the previous version.

To address this, we have revised the Introduction to clearly differentiate between:

(a) emotion theories (how emotions are generated and conceptualised),

(b) emotion measurement frameworks (discrete vs. dimensional models), and

(c) the phenomenological target of measurement (felt vs. perceived emotion).

We have added an explicit clarifying paragraph early in the Introduction stating that discrete and dimensional models are measurement frameworks that can be applied to both felt and perceived emotions, and that the felt–perceived distinction primarily concerns task instructions rather than the choice of emotion model itself. In addition, we have revised wording that could be read as implying that dimensional models cannot be applied to felt emotions, as this was not our intended claim.

Positioning of the GEMS-9 and use of second-order factors

Relatedly, we have clarified our use of the GEMS-9 framework. While we employ the second-order factors of the GEMS-9, which resemble dimensional emotion spaces, we now explicitly state that we retain the domain-specific, music-focused conceptualisation of aesthetic emotions rather than treating the scale as a generic dimensional model. This clarification is now made explicit in the Introduction and Methods sections to avoid conceptual ambiguity.

Strengthening the contribution and value of the study

We appreciate the reviewer’s suggestion to strengthen the final paragraph of the Introduction by more clearly justifying the value of the present study. We have substantially revised this section to more explicitly frame the study as a high-powered, methodologically rigorous replication and extension of prior work on tempo and mode. In particular, we now emphasise the importance of replication for establishing robustness and generalizability in music–emotion research, especially given the variability of stimuli, emotion models, and sample sizes in the existing literature. This revised framing better situates the study’s contribution without engaging in an extended discussion of the replication crisis.

Clarity, proofreading, and consistency

Finally, we have conducted a thorough proofreading of the manuscript. This included correcting typographical errors, improving clarity in several passages, standardising terminology (e.g., consistent use of “first-order” and “second-order”), and addressing formatting inconsistencies noted by the reviewer.

We again thank Reviewer #1 for their thoughtful feedback, which has helped us improve the conceptual clarity, positioning, and overall quality of the manuscript.

Reviewer #2

We sincerely thank Reviewer #2 for the careful and detailed evaluation of our manuscript. We appreciate the critical feedback and have revised the manuscript substantially to improve clarity, methodological transparency, and theoretical positioning. Below, we respond to each of the reviewer’s general and detailed comments.

1. The contents are of moderate relevance with a lot of repetition of known facts. The major question is: what is new?

We appreciate this important question and have clarified the novelty of the study in the Introduction. Specifically, we now emphasise that the current study is a conceptual replication study. The novelty lies in combining a large sample size, multiple structurally matched compositions, second-order GEMS and direct comparison with non-musical control stimuli.

2. The research questions are very general, which makes them not really suitable for an empirical study.

We have reformulated and specified the research questions to clearly indicate that they concern graded effects of tempo and mode within the present stimulus set. We now avoid overly broad claims about “music” as a general category. And clarify that we examine parameter-level effects rather than holistic compositional effects.

3. The style of writing is not very mature (young scholars?) and the academic standards are not very high.

We have carefully revised the manuscript for clarity, structure, and academic tone. In particular, the Results section has been rewritten to reduce SPSS-style reporting. Effect sizes are now explicitly interpreted (small, medium, large). Redundant passages have been removed. Paragraph structure has been tightened throughout.

We hope the revised version reflects a clearer and more mature academic style.

4. The style of referencing is rather weak. Some major references are lacking, there are also multiple references to submitted papers that are not yet accepted.

We have removed references to “submitted” manuscripts wherever possible. Where necessary, such work has either been replaced with published sources or clearly identified as preprints.

We have also strengthened the theoretical background with additional references in sections discussing dimensional versus categorical models and second-order emotional constructs.

5. The research questions are very general and reductionistic. More motivation is needed to explain why some selections have been made: why only tempo and mode? What about other parameters (dynamics, agogics, timbre, etc.)?

We have clarified the rationale for focusing on tempo and mode, specifically as the pilot study included dynamics (the referenced preprint). Both parameters are among the most consistently studied structural determinants of musical emotion. They can be systematically manipulated while holding other parameters constant. The aim of the study was not to model the full complexity of music, but to examine controlled parameter-level effects.

We explicitly acknowledge in the revised manuscript that emotional responses to music are multi-determined and that tempo and mode represent experimentally isolatable contributors rather than exhaustive determinants.

6. The methodology seems to have some potential, but the methods themselves are very badly explained. Figures and tables are introduced in the text without the needed accompanying explaining text, which makes it very difficult to interpret them. Many questions can also be raised about the collection of the data, which have been collected in less standardized circumstances.

The Methods and Analysis sections have been substantially revised to improve transparency. Specifically we now provide a step-by-step description of how ratings were aggregated across compositions within tempo–mode conditions. The rationale for aggregation is explained. The statistical model is described clearly, including within- and between-subject factors. Interpretation of Wilks’ lambda and partial eta squared is clarified. Effect size benchmarks are explicitly provided.

We hope these revisions address the concern regarding methodological clarity.

7. There is overall too much generalization. For example, speaking of music as a general category is very problematic to be consider as an independent variable. There are so many kinds of music, all with a different potential effect on listeners.

We agree that music is not a homogeneous category. We have revised the manuscript to clarify that conclusions apply to the present stimulus set, the manipulated structural parameters and the operationalized emotional dimensions.

Claims have been carefully delimited.

8. The paper as whole seems to be an example of method-driven analysis, with a lot of statistical processing and computation, but with a scarcitiy of starting ideas and challenging intuitions. The theoretical background is also rather weak.

We have revised the Introduction to strengthen the theoretical framing and clarify the conceptual motivation behind the analyses. The Results section has also been restructured to emphasize interpretive coherence rather than statistical detail.

Additionally, we have added a synoptic summary table (Table 1) that highlights the overall pattern of effects across emotional dimensions, making the theoretical structure of findings more transparent.

Detailed Comments

Dimensional vs categorical models (p. 10)

We have expanded this section and added additional references to clarify the distinction and empirical support for dimensional approaches.

Dimensional models outperforming (p. 10)

The relevant paragraph has been revised to avoid overly strong claims and to provide clearer references supporting the statement.

Second-order categories (p. 11–13)

We have expanded the explanation of the second-order dimensions (sublime, uneasy, vital), including conceptual definitions, operationalisation in the GEMS framework and clarification that participants received standardised descriptors. We also included the intercorrelations among the first- and second-order factors in GEMS.

We also discuss potential interpretational variability in terms such as “sublime.”

“Valence” correction

Corrected throughout.

Strong claim (p. 12)

The statement has been moderated and additional references added.

Reductionism and isolation of parameters (p. 13)

We have strengthened the rationale for isolating tempo and mode and explicitly acknowledge the limitations of parameter-level approaches.

Submitted references

Removed or replaced as noted above.

Synoptic table and effect sizes (p. 17)

In response to this important suggestion, we have added a synoptic summary table (Table 1) presenting main effects and key interactions across all three emotional dimensions. Reported F-values, degrees of freedom, p-values, and partial eta squared consistently. Explicitly stated effect size benchmarks in the Analysis section. Clarified interpretation of Wilks’ lambda. Clarified how mean values were computed.

Full statistical tables are retained in supplementary materials for transparency.

Text in Discussion is better suited for Introduction (p. 24)

We have relocated the indicated paragraphs to the Introduction to improve structural coherence.

Figures (pp. 32 ff.)

All figures have been revised to include clearer captions and expanded explanatory text in the Results section to guide interpretation.

Concluding Remark

We are grateful for the reviewer’s detailed critique. The manuscript has undergone substantial revision to improve clarity, methodological transparency, theoretical positioning, and presentation of results. While the empirical findings remain unchanged, their structure and interpretation are now presented more clearly and concisely.

We hope the revised manuscript addresses the reviewer’s concerns and demonstrates the contribution of the study more convincingly.

Reviewer #3

Thank you for reviewing our manuscript. We appreciate the hard work and your comments.

1. Rephrase or substantiate the statement that tempo and mode are “inconsistently examined across emotion models.” Either specify the precise inconsistencies the paper addresses (e.g., differing stimuli, measures, or theoretical mappings) or soften the claim to position the work as a large-scale, stimulus-controlled extension of established findings. The manuscript itself cites many prior tempo/mode studies.

Thank you for this suggestion. We have rephrased this statement to avoid overgeneralization. Rather than implying broad inconsistency, we now position the study as a large-scale, stimulus-controlled extension of established findings, while acknowledging the substantial body of prior research on tempo and mode.

2) In the introduction, I recommend streamlining the sections on emotion models to improve logic: briefly introduce discrete, dimensional, and domain-specific approaches, then clearly justify the use of GEMS-9. The current text reads as somewhat disjointed.

We appreciate this comment. The previous version of the introduction had expanded in response to earlier reviewer requests for additional theoretical detail, which may have affected its coherence. We have now substantially streamlined the section by briefly introducing discrete, dimensional, and domain-specific approaches in a more structured sequence, followed by a clearer and more focused justification for the use of GEMS-9.

3) More importantly, flesh out the motivation for using GEMS second-order factors. Provide a stronger rationale for collapsing GEMS-9 to three second-order factors and report psychometric information, where available (internal reliabilities, inter-factor correlations…).

We agree that the rationale required further clarification. We have expanded the discussion of why the GEMS-9 second-order factors were selected, including their theoretical relevance and suitability for capturing music-specific emotional responses in large-scale designs. We now also report inter-factor correlations to increase transparency regarding the psychometric properties of the scales in our sample.

4) The sentence on emotional differences “emerging accidentally during the 14th–16th centuries” is ambiguous or unclear. Reword for clarity and cite an appropriate historical or psychoacoustic source (e.g., Parncutt, 2024).

Thank you for pointing this out. The sentence has been reworded for clarity, and we have added an appropriate reference to support the historical and psychoacoustic context.

5) Replication vs. extension – The study departs from previous methods and thus is not a strict replication. Please revise phrasing to “conceptual replication” or “extension,” and clearly distinguish replicated versus novel aspects.

We agree with this distinction. The study should indeed be described as a conceptual replication rather than a strict methodological replication. We have revised the terminology throughout the manuscript to clearly distinguish between replicated and novel aspects of the design.

6) Unclear terminology: “pure major” and “pure minor”

Thank you for noting this ambiguity. The terms “pure major” and “pure minor” have now been removed for clarity. The intention was simply to indicate that the stimuli were constructed in major and minor modes without modal mixture. As this terminology was potentially confusing and not essential, we have simplified the wording accordingly.

7) I understand the rationale behind using control stimuli, but couldn’t one raise the issue that control stimuli are auditory as well?

This is an important point. While alternative control modalities (e.g., visual stimuli) could have been used, we chose auditory control stimuli to ensure modality consistency and to isolate musical structure rather than sensory modality effects. We have now clarified this rationale in the introduction and explicitly addressed the limitations of this choice.

8) Describe the translation procedure (forward/back translation, expert review) and report internal consistency statistics.

We conducted forward and back translation procedures to ensure linguistic accuracy. Although an external expert review was not performed at the time, the translated instrument has since been used in multiple studies without issues. Unfortunately, internal consistency statistics were not computed in the original dataset; we acknowledge this as a limitation and clarify the translation procedure more explicitly in the manuscript.

9) Specify the hosting platform (e.g., Qualtrics, Gorilla, PsychoPy online), playback format, use of headphones...

We have added information about the hosting platform and relevant technical details in the Materials section. We previously mentioned headphone use in the limitations section; this has now also been clarified in the Procedure section for greater transparency.

Attachment

Submitted filename: Response_to_reviewers_auresp_2.docx

pone.0348069.s004.docx (27KB, docx)

Decision Letter 2

Giulia Prete

13 Apr 2026

The Effect of Tempi and Mode on the Rating of the Perceived Emotion in Music.

PONE-D-25-05889R2

Dear Dr. Færøvik,

We’re pleased to inform you that your manuscript has been judged scientifically suitable for publication and will be formally accepted for publication once it meets all outstanding technical requirements.

Within one week, you’ll receive an e-mail detailing the required amendments. When these have been addressed, you’ll receive a formal acceptance letter and your manuscript will be scheduled for publication.

An invoice will be generated when your article is formally accepted. Please note, if your institution has a publishing partnership with PLOS and your article meets the relevant criteria, all or part of your publication costs will be covered. Please make sure your user information is up-to-date by logging into Editorial Manager at Editorial Manager® and clicking the ‘Update My Information' link at the top of the page. For questions related to billing, please contact billing support.

If your institution or institutions have a press office, please notify them about your upcoming paper to help maximize its impact. If they’ll be preparing press materials, please inform our press team as soon as possible -- no later than 48 hours after receiving the formal acceptance. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

Kind regards,

Giulia Prete

Academic Editor

PLOS One

Additional Editor Comments (optional):

As you can see, one of the previous reviewers has agreed to review the new version of the manuscript and supports its publication in its current form. I have carefully read the revised version myself and agree that you have satisfactorily addressed the suggestions received. Therefore, I am happy to approve your manuscript for publication without further revisions.

Reviewers' comments:

Reviewer's Responses to Questions

-->Comments to the Author

1. If the authors have adequately addressed your comments raised in a previous round of review and you feel that this manuscript is now acceptable for publication, you may indicate that here to bypass the “Comments to the Author” section, enter your conflict of interest statement in the “Confidential to Editor” section, and submit your "Accept" recommendation.-->

Reviewer #2: All comments have been addressed

**********

-->2. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented. -->

Reviewer #2: Yes

**********

-->3. Has the statistical analysis been performed appropriately and rigorously? -->

Reviewer #2: Yes

**********

-->4. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.-->

Reviewer #2: Yes

**********

-->5. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.-->

Reviewer #2: Yes

**********

-->6. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)-->

Reviewer #2: (No Response)

**********

-->7. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.-->

Reviewer #2: No

**********

Acceptance letter

Giulia Prete

PONE-D-25-05889R2

PLOS One

Dear Dr. Færøvik,

I'm pleased to inform you that your manuscript has been deemed suitable for publication in PLOS One. Congratulations! Your manuscript is now being handed over to our production team.

At this stage, our production department will prepare your paper for publication. This includes ensuring the following:

* All references, tables, and figures are properly cited

* All relevant supporting information is included in the manuscript submission,

* There are no issues that prevent the paper from being properly typeset

You will receive further instructions from the production team, including instructions on how to review your proof when it is ready. Please keep in mind that we are working through a large volume of accepted articles, so please give us a few days to review your paper and let you know the next and final steps.

Lastly, if your institution or institutions have a press office, please let them know about your upcoming paper now to help maximize its impact. If they'll be preparing press materials, please inform our press team within the next 48 hours. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

You will receive an invoice from PLOS for your publication fee after your manuscript has reached the completed accept phase. If you receive an email requesting payment before acceptance or for any other service, this may be a phishing scheme. Learn how to identify phishing emails and protect your accounts at https://explore.plos.org/phishing.

If we can help with anything else, please email us at customercare@plos.org.

Thank you for submitting your work to PLOS ONE and supporting open access.

Kind regards,

PLOS ONE Editorial Office Staff

on behalf of

Dr. Giulia Prete

Academic Editor

PLOS One

Associated Data

    This section collects any data citations, data availability statements, or supplementary materials included in this article.

    Supplementary Materials

    S1 File. Supporting information, including access to musical and control stimuli, sheet music for all stimuli, and supplementary tables.

    (DOCX)

    pone.0348069.s001.docx (834KB, docx)
    Attachment

    Submitted filename: Response to reviewers.docx

    pone.0348069.s003.docx (26.9KB, docx)
    Attachment

    Submitted filename: Response_to_reviewers_auresp_2.docx

    pone.0348069.s004.docx (27KB, docx)

    Data Availability Statement

    Data is uploaded to figshare: https://doi.org/10.6084/m9.figshare.30052507.


    Articles from PLOS One are provided here courtesy of PLOS

    RESOURCES