Skip to main content
PLOS One logoLink to PLOS One
. 2023 Oct 26;18(10):e0293412. doi: 10.1371/journal.pone.0293412

Music listening evokes story-like visual imagery with both idiosyncratic and shared content

Sarah Hashim 1,*, Lauren Stewart 1, Mats B Küssner 1,2,, Diana Omigie 1,
Editor: Ioanna Markostamou3
PMCID: PMC10602345  PMID: 37883377

Abstract

There is growing evidence that music can induce a wide range of visual imagery. To date, however, there have been few thorough investigations into the specific content of music-induced visual imagery, and whether listeners exhibit consistency within themselves and with one another regarding their visual imagery content. We recruited an online sample (N = 353) who listened to three orchestral film music excerpts representing happy, tender, and fearful emotions. For each excerpt, listeners rated how much visual imagery they were experiencing and how vivid it was, their liking of and felt emotional intensity in response to the excerpt, and, finally, described the content of any visual imagery they may have been experiencing. Further, they completed items assessing a number of individual differences including musical training and general visual imagery ability. Of the initial sample, 254 respondents completed the survey again three weeks later. A thematic analysis of the content descriptions revealed three higher-order themes of prominent visual imagery experiences: Storytelling (imagined locations, characters, actions, etc.), Associations (emotional experiences, abstract thoughts, and memories), and References (origins of the visual imagery, e.g., film and TV). Although listeners demonstrated relatively low visual imagery consistency with each other, levels were higher when considering visual imagery content within individuals across timepoints. Our findings corroborate past literature regarding music’s capacity to encourage narrative engagement. It, however, extends it (a) to show that such engagement is highly visual and contains other types of imagery to a lesser extent, (b) to indicate the idiosyncratic tendencies of listeners’ imagery consistency, and (c) to reveal key factors influencing consistency levels (e.g., vividness of visual imagery and emotional intensity ratings in response to music). Further implications are discussed in relation to visual imagery’s purported involvement in music-induced emotions and aesthetic appeal.

1. Introduction

Listening to programmatic excerpts of music, such as Peter and the Wolf by Sergei Prokofiev, may elicit the imagination of several different scenes. For some listeners, it might enliven childhood memories of first encounters with the piece, while for others, it may render images of the wolf, dark forests, and feelings of threat. In some cases, particular themes represented by different instruments may even give rise to the visualisation of specific characters from the tale [1, 2]. All are plausible situations afforded by listening to the piece. However, such a plurality of possibilities raises the question of the extent to which listeners are prone to the same mental experiences when listening to a piece, or whether listeners’ life-long learned experiences are too complex to allow for such shared visual imagery.

For most listeners, forming a narrative in their mind’s eye is a way to engage with heard music [3, 4], and such narrative sequences are often reported to be vivid and multi-thematic experiences [5, 6]. A few investigations into visual imagery content during music listening have begun to shed light on this idea of imagery consistency across listeners [7] and potential influencing factors [6]. However, the extent to which listeners exhibit similarities in their own visual imagery across listening situations is still an open question.

1.1. Semantic associations and what listeners imagine

Musicologists commonly speak of a ‘narrative dimension’ [8] to music and refer to the explicitness through which this might be communicated by drawing parallels between music and language [1, 810]. Research has shown that semantic associations inform our understanding and perception of music, and that the process of narrativizing music is rooted in, and requires access to, linguistic cognitive resources. A study by Koelsch et al. [11] revealed that priming words with musical stimuli induced the same electrophysiological signature (the N400) as when priming with linguistic stimuli, suggesting that as with language, listeners are capable of extracting meaning from musical stimuli. Interestingly, though, despite links between music and language processing being repeatedly demonstrated [1216], they seem to nevertheless also be subserved by distinct cognitive mechanisms: Barraza et al. [17] used EEG to reveal that while linguistic processing appears to work on the basis of confirming participants’ prior expectations for word primes, music listening implicates a post-hoc cognitive strategy, focused more on forming meaning where, due to the often subjective nature of deriving musical meaning, there is a lack of it. Such findings suggest that visual imagery in response to music may be a loose and unstructured process.

Given the fact that semantic associations can be formed from music, an interesting question that follows is the type of content that listeners imagine in response to music. In an online survey, Küssner and Eerola [3] asked participants to provide descriptions of the visual imagery that they typically experience while listening to music. They found that the most common type of visual imagery were natural landscapes and personal memories. Additional content included images of musical performances as well as abstract types of visual imagery such as colours and geometric shapes.

Empirical research demonstrates that people’s thoughts when listening to music are generally often visual in nature, with mind wandering, or ‘daydreams’, tending to occur more often in the form of visual images than in the form of words [18, 19]. While examining the effects of music on thought content and valence, Koelsch et al. [19] found visual imagery to be a characteristic of mind wandering in response to music that evokes heroism and sadness. An investigation into the metaphors evoked by music carried out by Schaerlaeken et al. [20] revealed that these could be represented by five main attributes (Flow, Movement, Force, Interior, and Wandering) with visual imagery emerging as a prominent mode in which they are realised. Studying visual imagery in response to music promises theoretical insights into how cross-modal associations [21] and meaning making emerge [22]. Taken together, these studies provide some clues into the types of visual imagery content that may be expected to be derived from music listening.

1.2. How stable is music-induced visual imagery?

To date, visual imagery has been assumed to be a largely idiosyncratic experience, with listeners prone to conjuring mental images not only through acoustic features, but also through processes involving past life events and contextual information [23]. Yet, the notion of consistency has been sparingly addressed in music-induced visual imagery research.

Recent related work by Margulis et al. [6] has investigated the formation of imagined narratives–mental storytelling being formed with the potential (but not necessity) to appear visually–and suggests that imagined narratives may be more widely shared than previously thought. With two sets of independent samples recruited from the US and China, they used natural language processing techniques to analyse open-text reports of narratives formed in response to instrumental music. They found that descriptions written by individuals who shared an underlying culture received higher consistency levels than reports written by those who did not.

These results offer strong support for the semantic similarities that individuals can have in their imagined narratives to music. However, it is important to note that consistency there was made in relation to general narrative formation and not with respect to visual imagery. Thus, it is possible that focusing specifically on music-induced visual imagery would yield a less pronounced similarity level.

In an influential paper, Cross [12] proposes that music possesses a floating intentionality whereby different listeners derive different meanings from the ongoing events in music. In line with the work of Margulis and colleagues, one might expect that individuals will tend to show higher consistency of visual imagery within themselves than with one another. However, such a possibility has never been empirically tested. One aim of the current research was therefore to shed light on this matter by comparing the degree of similarity in listeners’ visual imagery reports within themselves (across two surveys) to the degree of similarity in visual imagery reports they show with other listeners.

1.3. Emotion and aesthetic appeal

Previous literature hints at the idea that one’s experience of visual imagery may somewhat be related to their hedonic responses to music. Music-induced visual imagery has been previously associated with aesthetic and emotional engagement with music, but the evidence of such links remains limited [24, 25]. In one study, Belfi [26] investigated factors contributing to the aesthetic appeal of classical, jazz, and electronic music. Participants were required to listen to tracks taken from each genre and to report on their experience with respect to the vividness of music-induced visual imagery, arousal, emotional valence, and liking/aesthetic appeal of the music. Across genres, it was found that emotional valence and visual imagery vividness were similar in the degree of their predictive influence over musical aesthetic appeal.

In another study, Presicce and Bailes [27] established a clear association between the continuous ratings of visual imagery that participants reported in response to a selection of piano pieces and their continuous ratings of engagement with the music. Finally, in an fMRI study, Koelsch et al. [28] showed evidence of an interaction between emotion and visual areas of the brain during music listening. Interestingly, listening to fearful, compared to joyful music, led to greater interactions between the visual cortex and the superficial amygdala, perhaps due to the heightened vigilance fearful music elicits. Thus, associations between music-induced visual imagery and emotion induction are evident, although the directionality of this relationship is still a widely debated topic supported by contrasting evidence [29, 30].

As previous studies have stated that extensive prior experience with the visual arts may be expected to influence the experiences of visual imagery that individuals have (mental imagery being regarded as a useful tool in creative processes including the visual arts, music, and dance [31, 32]), the current study also investigated how the prevalence, vividness, and consistency of music-induced imagery reported is influenced by this individual difference.

Taken together, it is apparent that visual imagery may be at least one tool with which one is ‘drawn in’ to music and by which music engages and induces emotions in a listener. However, there is a need to explore this with a larger sample than has been used in previous research and with a more nuanced approach that can discern particular types and subtypes of visual imagery during music listening.

1.4. The current research

In light of the reviewed literature, the present work sought to address four main aims:

  1. To run a thematic analysis to create a hierarchical framework highlighting the prevalent codes and overarching themes found in descriptions of music-induced visual imagery content,

  2. To ascertain the extent of the consistency of visual imagery within and across individuals during music listening,

  3. To examine potential behavioural factors and individual differences (general visual imagery ability, musical training, and participation in the visual arts) that may be driving the rates of within- and across-participant consistency levels,

  4. To test the extent to which the prevalence and vividness of visual imagery is associated with the emotional intensity and aesthetic appeal of music, as well as an array of individual differences (general visual imagery ability, musical training, and participation in the visual arts).

Importantly, we anticipated that storytelling, including action-based imagery [20] (but see also [33, 34]), may emerge as a prominent form of visual imagery (H1) [3, 4, 6], and that while, in line with Cross [12], within- and across-participant consistency would be shown to be modest, there would be higher within-participant consistency than across-participant consistency (H2), due to the reduced set of factors that can lead to variance (i.e., no differences in individual listening experience or personal memories when comparing individuals with themselves, in contrast to when comparing with different individuals).

Further, we predicted that we would be able to replicate previous reports of a relationship between visual imagery and both emotion induction [30, 3537] and aesthetic appeal [26]. We specifically predicted that the prevalence and vividness of visual imagery would both be predicted by ratings of emotional intensity (H3) and music liking (H4), in line with previous reports of positive links between these phenomena [23, 25, 26, 35].

Next, with regard to inter-individual differences, we predicted that the prevalence and vividness of visual imagery in response to music would be strongly associated with general visual imagery abilities (H5) but show a weaker relationship with musical training (H6); this is due to previous findings that musical expertise had little predictive influence over visual imagery content [5] and that while musically trained individuals may possess higher imagery abilities, this is not necessarily true of visual imagery [38].

Finally, we nevertheless predicted that the higher imaginative freedom of those who regularly engage with the visual arts may allow them to experience music-induced visual imagery more frequently and vividly than those who do not regularly engage with the visual arts (H7).

2. Methods

2.1. Participants

353 participants (153 female, 198 male, 2 prefer not to say) aged 18–66 years (Mean = 26.41, SD = 9.41) were recruited using either Prolific, an online participant recruitment platform, or word-of-mouth. The same sample was invited three weeks later to once more take part, with 254 participants (102 female, 149 male, 3 prefer not to say) aged 18–66 (Mean = 26.85, SD = 9.04) retaking the survey. See Table 1 for further demographic information.

Table 1. Countries of residence of the samples from survey 1 and survey 2.

Country of Residence Survey 1 (N = 353) Survey 2 (N = 254)
Sum Proportion Sum Proportion
United Kingdom 79 22.4% 38 15.0%
Portugal 59 16.7% 48 18.9%
Poland 48 13.6% 37 14.6%
Italy 18 5.1% 16 6.3%
Canada 18 5.1% 14 5.5%
Mexico 16 4.5% 15 5.9%
Germany 14 4.0% 3 1.2%
Greece 13 3.7% 12 4.7%
Other 88 24.9% 71 28.0%

This research has received ethical approval from the Ethics Committee of the Department of Psychology at Goldsmiths, University of London. Participants provided online written consent and received monetary compensation or course credit for their time.

2.2. Materials and stimuli

Three film music stimuli conveying happy, tender, and fearful emotions were selected from Eerola and Vuoskoski’s [39] database (see, https://www.jyu.fi/hytk/fi/laitokset/mutku/en/research/projects2/past-projects/coe/materials/emotion/soundtracks for access to the original stimuli). These excerpts were obtained from the catalogue of extended 1-min film excerpts (see Appendix of [40], or see, https://www.jyu.fi/hytk/fi/laitokset/mutku/en/research/projects2/past-projects/coe/materials/emotion/soundtracks-1min for access to the original stimuli), validated to be unfamiliar to most listeners and to still convey the intended emotions even in their shorter form. In terms of the films that the tracks were taken from, the excerpt conveying happy emotions was taken from The Untouchables soundtrack (track 6, number 071 from Eerola and Vuoskoski’s set of 110 tracks). The tender excerpt is from the Shine soundtrack (track 10, number 042 from set of 110 tracks). Finally, the fearful excerpt is from the Batman Returns soundtrack (track 5, number 011 from set of 110 tracks). In order to ensure uniformity amongst the musical excerpts, as well as to control the overall length of the survey, all excerpts were edited to last a duration of 45 seconds using Audacity (Version 2.3.2.0). These were also edited to finish with a fade-out to avoid an abrupt ending.

Visual imagery ratings were obtained using two items from Pekala’s Phenomenology of Consciousness Inventory [41], a 53-item questionnaire that measures a variety of personal perceptual experiences revolving around consciousness. Participants were asked about the prevalence of visual imagery in their experience (1 = I experienced no visual imagery at all, to 7 = I experienced a great deal of visual imagery), and the vividness of their imagery (1 = My visual imagery was so vague and diffuse, it was hard to get an image of anything, to 7 = My visual imagery was so vivid and three-dimensional, it seemed real).

The content of visual imagery was measured using an open-text question. Participants were asked to describe the content of their visual imagery (if any) with no limit to the length of their descriptions required. Liking ratings in response to the music were measured using a 5-point Likert scale, from 1 = Dislike a great deal, to 5 = Like a great deal. Emotional intensity ratings were also measured using a 5-point Likert scale, from 1 = Not at all intense to 5 = Extremely intense. Participants were also asked to report on the musical aspects they thought contributed to their visual imagery, but data from this question will be addressed elsewhere.

Finally, we measured three aspects of individual differences. The Musical Training dimension of the Goldsmiths Musical Sophistication Index (Gold-MSI) [42] was administered in order to gauge the extent of individuals’ musical experience. The Vividness of Visual Imagery Questionnaire (VVIQ) [43], which comprises 16 statements to which individuals are instructed to form a visual mental image in their minds, was administered as an independent measurement of visual imagery ability, along a 5-point Likert scale from 1 = No image at all, you only ‘know’ that you are thinking of an object to 5 = Perfectly clear and vivid as real seeing. Finally, any experience with activities associated with the visual arts (‘Do you participate in any activities associated with the visual arts?’) was also probed. This was a binary yes/no question, with an option to elaborate on the type of activity experienced if answered in the affirmative.

2.3. Procedure

Participants were first provided with the aims, instructions, and a definition of visual imagery as “the spontaneous formation of visual images or pictures in your mind’s eye. Your imagery experience is completely subjective, and it is completely acceptable not to have experienced any imagery at all”. Survey 1 took approximately 12 minutes to complete. The presentation order of musical stimuli was randomised across participants.

Participants were advised to use good-quality headphones and to have minimal outside disturbance throughout the study. For the three main trials, participants were first presented with the musical excerpt. They were instructed to listen to the whole excerpt and told that on the next page they would be presented with questions regarding their experience of the music. They were advised to pay attention to any visual imagery that they may be experiencing, and the musical characteristics that may have contributed to the imagery (question addressed elsewhere), as they listened. On a first response page, participants were asked to rate the prevalence (the amount experienced) and vividness (the clarity with which it was experienced) of their visual imagery. These ratings were followed by the open-text question asking them to describe the content of their visual imagery (if any). Finally, participants were asked to rate how much they liked the music, and how intense any felt emotional response to the music was.

After this, participants completed the VVIQ, the musical training dimension of the Gold-MSI, then indicated whether they have participated in any activities associated with the visual arts.

Three weeks later, participants completed an almost identical survey, excluding the VVIQ, Gold-MSI, and question about their experience with the visual arts. Certain demographic questions were collected once again, such as an anonymous ID, but also age, and nationality, to improve the chances of confidently being able to match participant data across both surveys. This second survey took approx. 8 minutes to complete.

3. Analysis

3.1. Thematic analysis of open-text reports of music-induced visual imagery content

Using Braun and Clarke’s [44] approach, we implemented a thematic analysis on the responses to the open-text question (‘Please describe the content of your visual imagery (if any)’) to identify the prominent themes that emerged in terms of visual imagery in response to the three musical excerpts. This thematic analysis was carried out using Microsoft Excel (Version 16.43) while further quantitative analyses were calculated using R (Version 4.2.3) [45].

The thematic analysis was carried out by two independent coders to reduce the risk of personal bias. The dataset was analysed as one large unit independent of excerpt emotion with the aim of producing a rich account of each developing theme. In general, to overcome any potential variability in language as a result of participants’ wide range of personal and musical backgrounds, we discerned meaning from the data at an explicit level (i.e., a literal interpretation of the text).

The dataset from survey 1 comprised a total of 1,059 reports across the three musical excerpts, while survey 2 comprised a total of 762 reports. Two coders independently examined the dataset and identified Level 3 codes that could encompass key aspects of each description, as well as making note of as many ideas and patterns that may benefit the final model as relevant. Commonalities between specific terms were identified from Level 3 (L3) codes, which were sorted and categorised into distinct groups comprising Level 2 (L2) codes. The suitability of the subthemes in this level was reviewed by the two coders, including whether to collapse redundant subthemes or to divide ones that were too diverse. The L2 codes were finally combined to form higher-order Level 1 (L1) themes. The final structure was discussed by the two coders, confirming that it formed an effective hierarchy that offered a parsimonious overview of the content of the free-form descriptions.

3.2. Measuring consistency within and across reports

We used the Jaccard coefficient index, a conservative measure of agreement that enables assessment of the level of overlap between any two lists (0 = no similarity at all, and 1 = perfect similarity):

J(A,B)=ABAB

In order to assess consistency, we needed a way of referring to each visual imagery code in a systematic way. Thus, we created a labelling system where all theme levels from our thematic analysis were assigned a numeric label as a unique identifier. To test the efficacy of these labels as a way to sort each participant report, the two coders first used this system to independently label a subset of 60 participant reports (5–6% of total) from the dataset, using L3 codes (the most detailed level of our visual imagery codes). Each individual report was assigned labels pertaining to the content of their descriptions. This first attempt by the coders yielded a Jaccard similarity average of 69% between Coders 1 and 2. Coders discussed any comments or issues that resulted from this effort, including minor reorganisation of the code levels. As a final check, a new subset of 60 reports was analysed, resulting in a version that was a more feasible and intuitive organisation of the themes. The codebook was reviewed and discussed once more by the research team to address any remaining structural issues. Finally, Coder 1 used the labelling system to assign labels to the rest of the datasets for survey 1 and survey 2. The original and coded framework can be found through the Open Science Framework using the following link: https://osf.io/nf4x7/?view_only=081602e5aca94788bb959b48ed8b47ef.

We assessed two types of consistency for each musical excerpt separately: within- and across-participant consistency. Within-participant consistency was computed by comparing the lists of themes emerging from each participant’s reports across surveys 1 and 2, and across-participant consistency was computed by comparing the list of themes from each individual participant with the lists from every other individual in the sample, resulting in 352 values per participant for each musical excerpt type (leading to three groups of values) that were then averaged to create a single consistency value per individual per excerpt. With regard to across-participant consistency specifically, individuals were always compared against other individuals’ responses that were within the same excerpt type.

In order to retain a precise estimation of the visual imagery types present in the data and its consistency within the sample, the descriptions of those who reported not experiencing any visual imagery or only vague imagery in one or both timepoints were excluded before running the Jaccard analyses. However, those who were coded as having no or vague imagery were included if they also reported additional coded experiences that could be used to compute their consistency (i.e., reports describing no or blurred visual imagery, but references to other relevant experiences such as emotional reactions or other imagery types). We also ran the same analyses on our codes only depicting visual experiences (i.e., those under the Storytelling higher-order theme, see section 4.1). This was to fully address one of our main aims and research questions regarding the consistency of content of visual imagery more specifically.

3.3. Comparing consistency levels and determining the behavioural measures and individual differences that influence them

We assessed the differences in Jaccard coefficient scores between within- and across-participant consistency within each code level (L2 (more broad) versus L3 (more detailed) codes) using linear mixed effects models. The purpose of this was to examine whether consistency levels differ when comparing within and across individuals. To test this, two models were defined, using the lme4 [46] and lmerTest [47] packages in R, one including L3 consistency values as the dependent variable and the other including L2 values as the dependent variable. In both models, Consistency Type was entered as a categorical fixed effect representing whether the consistency values originate from the Within-Participant or Across-Participant analysis, with participant and musical excerpt entered as random effects in both models.

The same method was used in a subsequent stage of the analysis to test the strength of any relationships between visual imagery consistency values (within- and across-participants) and data on the behavioural experiences of music as well as individual differences. Four models were run, each including within- and across-participant consistency values for each code level (L2 and L3) as dependent variables. In each model, prevalence and vividness of visual imagery, music liking, emotional intensity, VVIQ, musical training, and visual arts participation (yes or no) were entered as fixed effects, with participant and musical excerpt included as random effects.

3.4. Tests of associations and differences between behavioural measures and individual differences

Finally, linear mixed effects models and t-tests were run to ascertain the strength of associations and differences with regard to behavioural ratings (prevalence and vividness of visual imagery, music liking, and emotional intensity) and data on individual differences (VVIQ, musical training, and visual arts participation). To this end, two models were run, one with prevalence as dependent variable and the other with vividness as dependent variable, with music liking, emotional intensity, VVIQ, and musical training included as fixed effects, and participant and musical excerpt as random effects in both models. Further, independent samples t-tests were run to assess differences between those who do and do not participate in the visual arts in terms of their prevalence and vividness of visual imagery ratings.

4. Results

4.1. Thematic analysis of music-induced visual imagery

As can be observed in Table 2, the thematic analysis conducted on descriptions of visual imagery content resulted in three higher-order themes pertaining to the most prevalent topics in participant reports. As will be observed below, participant descriptions did not always include experiences that were strictly visual in content and often included descriptions related to affect, musical features, and other forms of mental imagery. This fact has been taken into consideration with regard to analyses and conclusions drawn.

Table 2. Thematic framework of prominent higher-order themes (L1) of music-induced visual imagery content (with L2 and L3 codes in descending order of prevalence).

Values included beside L1 codes represent their proportions within the entire framework, whereas the two values beside L2 codes reflect their proportions within their higher-order theme (left) and within the entire framework (right).

Level 1 Storytelling (74.3%) Associations (6.4%) References (19.3%)
Level 2 Setting & Location (38.0%, 28.2%) Characters (27.5%, 20.4%) Action (17.7%, 13.1%) Narrative (16.8%, 12.5%) Abstract & Other (59.9%, 3.8%) Feelings & Atmospheres (29.6%, 1.9%) Memories (10.9%, 0.7%) Other Media (89.5%, 17.2%) Music (10.5%, 2.0%)
Level 3 Building Imagery of Self Communicative Gestures Plot Lines Physical Reactions Emotional Reaction/Feelings Autobiographical Medium Genre
Urban Individuals Idle/Passive Hardship Other Senses Atmosphere/Mood Episodic Genres Composer
Nature Royal Mental Conflict & Resolution Cross-Modal Specific Movie/Game/Show Characteristics
Interior & Exterior Features Heroic Interaction Celebratory Features Instruments
Existing Locations Evildoer Musical Inquisitive
Imagined Setting Civilian Leisure
Time References (Social) Groups Gait
Season Musical
Weather Animals/Insects
Other Objects Characteristics

A summary of the frequencies of each L2 subtheme is compiled in Fig 1 (see Fig A in S2 File for frequencies of L2 codes across music excerpts, and Fig B in S2 File for frequencies of L3 codes). We summarise each higher-order theme and their subthemes (in descending order of prevalence) as follows:

Fig 1. Occurrence of L2 visual imagery content.

Fig 1

4.1.1. Storytelling

Storytelling, our first higher-order theme, is characterised by descriptions of narratives involving situations, actions, people, and locations, thus confirming Hypothesis 1. In most cases, descriptions referred to fictional situations involving non-existent characters, although participants sometimes also visualised themselves or other familiar individuals in imagined circumstances. This higher-order theme most explicitly encompasses details that were visualised. This higher-order theme was referred to in 74.3% of the total sample (2,868 times) throughout reports, and comprised four L2 subthemes:

  1. Setting & Location: describing the locations, interior/exterior, and temporal (and other) details in which the scenes took place. This comprised 38.0% of the Storytelling theme (28.2% of total, 1,090 times, e.g., “It was a tree on top of a hill”, “I was in a museum on a day with beautiful weather”)

  2. Character: referring to the presence of an individual, groups of people, animal, or imagery of oneself within content descriptions. This comprised 27.5% of the Storytelling theme, 20.4% of total, 788 times). This subtheme comprised (fictional or real) individuals ranging in background or societal status and includes a lower-level theme describing a range of bodily characteristics (e.g., “I feel like I’m at the coronation of a future king”, “A man in a button-down, sleeves rolled up to mid-arms, sitting in an apartment”)

  3. Action: referring to actions or activities performed by individuals contained in the descriptions. This comprised 17.7% of the Storytelling theme (13.1% of total, 507 times, e.g., “A couple walking on the beach”, “Man and woman dancing slowly”)

  4. Narratives: referring to plot transitions and prevalent themes portrayed by the content. This comprised 16.8% of the Storytelling theme (12.5% of total, 483 times, e.g., themes of celebration or conflict, “A celebration of a great event with lots of people…”, “Some scary monsters, and very dynamic action…”, “At first, I sort of had an imagery of a battlefieldthen later on it was more sort of fairy-tale like”)

4.1.2. Associations

Associations, our second higher-order theme, occurred 6.4% of the total number of reports (247 times throughout reports). This theme encompasses a mixture of perceptual or sensorial experiences within participants’ reports to the music and/or those contained in narrative descriptions of the music, as well as some mention of abstract forms of visual imagery. In all, this higher order theme groups together multimodal and affective experiences contained within the musical experience. It comprised three L2 subthemes:

  1. Abstract & Other: referring mostly to abstract forms of visual imagery, mental imagery additional to visual imagery, as well as physical experiences in relation to the music (e.g., “At first, I saw light and bright colours and laterclouds above a red colour”, “I felt my eyes tremble”). This subtheme occurred 59.5% of the Association theme (3.8% of total, 147 times)

  2. Feelings & Atmosphere: comprising 29.6% of the Associations theme (1.9% of total, 73 times). This subtheme describes any emotional reactions or moods portrayed by the scene or any felt by characters within the narrative or participant themselves (e.g., “Love and happiness in general”, “A scary mood, oppressive, a danger”). Whilst not representing visual imagery, its presence within reports was considerably prevalent and constituted a prominent attribute of descriptions of visual imagery experience

  3. Memories: describing two types of memories that participants may have recalled during their listening experience: autobiographical (e.g., “It reminds me of the old cartoons”) and episodic (e.g., “Walking to the stage on my graduation”). This theme was the least prevalent amongst the other elements in the framework, occurring only 10.9% of the already very small Associations theme (0.7% of total, 27 times, e.g., “My ballet classes”)

4.1.3. References

References, the final higher-order theme, occurred in 19.3% of the total sample (743 times) throughout reports. It is comprised of two L2 subthemes and contains codes referring to details pertaining to the origin of the participant’s visual imagery, and generally also includes a mixture of semantic associations (e.g., the music being characteristic of a particular genre) and/or visualised details (e.g., specific instruments or composers). It comprised two L2 subthemes:

  1. Media: the first and majority subtheme (89.5% of the theme, 17.2% of total, 665 times), comprised references to different types of media (e.g., television, film), references to media genres (e.g., horror, action, film noir), and mentions of any pre-existing media (e.g., “Animated Disney-style movie”, “A 40s horror movie”)

  2. Music: comprised references to musical genres (10.5% of the theme, 2.0% of total, 78 times), composers, or instruments that were visualised or described as contributing to the visual imagery (e.g., “As if I were sitting with tea in my hand next to a famous composer like Fryderyk Chopin”, “When the piano started playing I imagined a dark haired male pianist on a black piano”)

4.2. Do respondents show higher consistency with themselves than others?

Table 3 provides examples of reports at different (full framework) within-participant consistency boundaries. While reports at 0% reflect no discernible overlap, the 20–50% consistency range reflects minor content overlap in reports between timepoints, the 50–99% consistency range denotes significant content overlap between timepoints, and 100% reflects complete overlap and almost identical responses between timepoints.

Table 3. Example visual imagery excerpts at different bands of within-participant consistency L3 codes.

Survey 1 Excerpts Survey 1 Codes Survey 2 Excerpts Survey 2 Codes Consistency Level Reports at this level (% of N)
Akin to the end of a children’s fantasy movie 1.1.1, 2.3.1, 2.3.2, 1.4.7 Gathering of people for a celebration of royalty 1.1.4, 1.4.7, 1.4.3 0–19.99% 399 (55.4%)
I’m watching a theatre play, and it’s starting. 1.4.1, 1.2.5, 1.1.5 I have imagined that I’m in a big forest to explore it 1.3.1, 1.1.3, 1.2.6
I imagined a man playing a piano in an empty room 1.4.2, 1.2.5, 4.1.4, 1.3.1, 1.3.4 I felt like I was watching someone playing on a piano in a massive room. Beautiful, relaxing. . .and sort of nostalgic. 1.2.5, 1.4.2, 3.2.4, 1.3.1, 2.3.1 20–49.99% 208 (28.9%)
The celebration of a victory at the end of a war 1.1.3, 1.1.4 A hero triumphant return to its home village 1.4.4, 1.1.3, 1.3.1
WWI trenches, then galloping horses, then marching soldiers 1.1.1, 1.1.3, 1.2.6, 1.4.9, 1.2.7, 1.4.3 WWI, American civil war battles, armies marching 1.1.3, 1.4.3, 1.2.7 50–99.99% 90 (12.5%)
I feel like I’m at the coronation of a future king 1.4.1, 1.1.4, 1.4.3 I have the impression that I was at the coronation of the king / queen, or at some state event 1.4.1, 1.4.3, 1.3.1, 1.1.4
Lonely day at bar/cafeteria. Raining outside. Lonely afternoon. 1.3.1, 2.3.1, 1.3.9, 1.3.7 Lonely evening at the bar. Raining outside. 2.3.1, 1.3.7, 1.3.1, 1.3.9 100% 23 (3.2%)
I imagine someone of royalty getting married like for example Prince William and Kate. 1.4.3, 1.1.4 I imagine a wedding of someone of the royalty. 1.1.4, 1.4.3

N = 720

Fig 2 illustrates the distributions of within- and across-participant consistency (and Fig C in S2 File demonstrates these distributions as a function of music excerpt type). The graphs display marked differences in range, with the within-participant distribution spanning a wider set of consistency values and additionally revealing a higher proportion of zero levels of consistency, while across-participant consistency appears to span a much narrower value range and peaks at a relatively low level (see also Tables C and D in S1 File for precise consistency values for within- and across-participant profiles, respectively, at different Jaccard percentage ranges).

Fig 2. Comparing density distributions of consistency values of within- (across listening situations) and across-participant (across listeners) groups with mean intercepts.

Fig 2

(A) Level 2 codes. (B) Level 3 codes.

In order to test Hypothesis 2, we asked whether there were differences in consistency when comparing an individual with themselves at a later stage, on the one hand, to when comparing a given individual with the remaining sample. The model comparing Jaccard coefficient values for L2 codes between the within- (Mean = 27.4%, Median = 25%) or across-participant (Mean = 17.1%, Median = 17.2%) distributions, indeed, confirmed there to be distinguishable differences between the two groups (ß = 0.10, SE = 0.01, t = 13.00, p < 0.001), with the within-participant group displaying overall greater consistency values than the across-participant group. Additionally, the model predicting coefficient values at the higher granularity L3 codes similarly showed there to be overall differences between the within- (Mean = 20.6%, Median = 14.3%) and across-participant (Mean = 9.8%, Median = 9.6%) groups (ß = 0.11, SE = 0.01, t = 14.24, p < 0.001).

Fig D in S2 File presents within- and across-participant distributions using only Storytelling codes, since these, unlike the Associations and References codes, constitute codes that characterise purely visual imagery experiences. We once more show the two distributions to possess distinct shapes and peaks as previously described for all codes. Again, the model comparing Jaccard coefficient values for L2 codes between the within- (Mean = 33.6%, Median = 28.6%) or across-participant (Mean = 21.5%, Median = 21.2%) distributions for Storytelling codes confirmed differences between the two groups (ß = 0.12, SE = 0.01, t = 12.06, p < 0.001). Finally, the model predicting consistency values at the higher granularity L3 codes also demonstrated overall differences between the within- (Mean = 26.1%, Median = 16.6%) and across-participants (Mean = 11.9%, Median = 11.3%) groups (ß = 0.14, SE = 0.01, t = 14.35, p < 0.001).

4.3. Associations between visual imagery consistency and musical experience

In the next set of analyses, we asked to what extent across-participant consistency (a participant compared with everyone else in the sample) or within-participant consistency (a participant compared with themselves) were associated with participants’ experience of the music (visual imagery prevalence and vividness, liking, and emotional intensity) on the one hand, but also individual differences in general visual imagery ability, musical training, and participation in the visual arts on the other (see Table 4 for a full summary of the model results).

Table 4. Fixed effect estimates of behavioural and individual difference measures predicting within- and across-participant consistency.

Across Consistency L2 Across Consistency L3 Within Consistency L2 Within Consistency L3
ß SE t p ß SE t p ß SE t p ß SE t p
Intercept 0.11 0.02 6.32 < 0.001 *** 0.09 0.01 7.84 < 0.001 *** 0.09 0.06 1.40 0.163 0.10 0.06 1.61 0.110
Imagery Prevalence 0.01 0.00 2.96 0.003 ** 0.00 0.00 2.30 0.022 * 0.01 0.01 1.18 0.238 0.01 0.01 1.20 0.230
Imagery Vividness 0.00 0.00 2.49 0.013 * 0.00 0.00 1.24 0.214 0.01 0.01 1.42 0.155 0.02 0.01 2.17 0.030 *
Music Liking 0.00 0.00 0.62 0.534 –0.00 0.00 –0.16 0.869 0.01 0.01 0.94 0.356 –0.01 0.01 –1.30 0.193
Emotional Intensity 0.00 0.00 1.75 0.080 –0.00 0.00 –0.62 0.534 0.01 0.01 1.10 0.274 0.01 0.01 0.77 0.440
VVIQ 0.00 0.00 0.44 0.658 –0.00 0.00 –0.68 0.495 0.01 0.01 0.47 0.637 –0.01 0.01 –0.41 0.684
Musical Training –0.00 0.00 –1.41 0.161 –0.00 0.00 –0.56 0.576 –0.00 0.01 –0.30 0.768 0.00 0.01 0.62 0.534
Visual Arts 0.00 0.01 0.72 0.470 0.00 0.00 0.36 0.716 –0.03 0.02 –1.27 0.204 –0.01 0.02 –0.50 0.618

*** < 0.001

** < 0.01

* < 0.05. Abbreviations: L2 = Level 2, L3 = Level 3.

The model predicting L2 across-participant consistency revealed prevalence (ß = 0.01, SE = 0.00, t = 2.96, p = 0.003) and vividness (ß = 0.00, SE = 0.00, t = 2.49, p = 0.013) of visual imagery to be significant predictors of this level of consistency, whereas the model predicting L3 across-participant consistency showed only prevalence to be a significant predictor (ß = 0.00, SE = 0.00, t = 2.30, p = 0.022). No other behavioural measures or individual differences significantly predicted across-participant consistency in either level.

Conversely, the model predicting L2 within-participant consistency revealed that none of the behavioural measures possessed any predictive influence over this consistency level. However, the model predicting L3 within-participant consistency showed a significant influence of the vividness of visual imagery (ß = 0.02, SE = 0.01, t = 2.18, p = 0.030). No other behavioural measures or individual differences significantly predicted within-participant consistency in either level.

4.4. Prevalence and vividness of music-induced visual imagery and links between visual imagery, emotional intensity, liking, and individual differences

Table 5 presents a summary of visual imagery prevalence and vividness ratings divided and aggregated by musical excerpt. These values demonstrate high proportions of visual imagery prevalence and vividness across all musical excerpts, that persist even when one considers the individual excerpt types. Higher prevalence and vividness levels of visual imagery can be seen in response to the Fearful excerpt than the Happy and Tender excerpts, which both exhibit almost equal proportions.

Table 5. Proportions (and frequencies) of the prevalence and vividness of visual imagery for and across the musical excerpts.

Happy Tender Fearful Overall
No VI At least mild VI No VI At least mild VI No VI At least mild VI No VI At least mild VI
Prevalence 8.8% (31) 90.7% (320) 9.6% (34) 90.4% (319) 5.1% (18) 94.6% (334) 2.5% (9) 97.5% (344)
Vividness 10.9% (38) 89% (314) 13.6% (48) 86.1% (304) 5.9% (21) 93.2% (329) 5.4% (19) 94.6% (334)

Abbreviations: VI = visual imagery. “No VI” pertains to individuals who provided the lowest rating, whereas “At least mild VI” refers to ratings of 2 and upwards. Values in brackets indicate the sum of reports within that category. Combined sum of reports for or across excerpt types for each rating that do not equate to the total number of participants (N = 353) are due to missing ratings data.

Confirming Hypotheses 3, 4, 5, and 6, the model predicting prevalence of visual imagery demonstrates that emotional intensity (ß = 0.70, SE = 0.04, t = 15.83, p < 0.001), music liking (ß = 0.25, SE = 0.04, t = 5.53, p < 0.001), and the VVIQ (ß = 0.48, SE = 0.08, t = 5.95, p < 0.001) were all highly significant predictors. Contrastingly, musical training (ß = 0.02, SE = 0.04, t = 0.65, p = 0.517) did not significantly predict visual imagery prevalence.

Similarly confirming our hypotheses, the model predicting vividness of visual imagery similarly showed emotional intensity (ß = 0.70, SE = 0.04, t = 15.25, p < 0.001), music liking (ß = 0.21, SE = 0.04, t = 4.64, p < 0.001), and the VVIQ (ß = 0.54, SE = 0.08, t = 6.45, p < 0.001) to be highly significant predictors, whereas, once again, musical training was not (ß = –0.02, SE = 0.04, t = –0.45, p = 0.655).

In response to the question regarding experience with activities in the visual arts (VA), 20.1% (n = 71) reported that they participate in activities associated with the visual arts, which included activities such as painting, photography, and graphic design. With regard to the averaged excerpt ratings between those who do and do not (NVA) participate in the visual arts, independent samples t-tests showed that those who participate in the visual arts reported significantly more visual imagery prevalence (Mean-VA = 4.63, Mean-NVA = 3.99, t(119.4) = 3.78, p < 0.001) and more vividness (Mean-VA = 4.30, Mean-NVA = 3.74, t(124.1) = 3.37, p < 0.001) than those who do not, supporting Hypothesis 7. There were also significant differences with regard to each of the individual musical tracks, with individuals who do take part in the visual arts reporting higher prevalence and vividness than those who do not (see Tables A and B in S1 File for full results).

5. Discussion

The aims of the current study were multi-fold. Our main aim was to derive the most prominent themes present in listeners’ descriptions of their music-induced visual imagery and to investigate the extent to which listeners exhibited consistency in their visual imagery reports within themselves (within-participants) and with the whole cohort (across-participants). Further to this aim, we sought to then explore whether the prevalence and vividness of visual imagery, music liking and emotional intensity, as well as individual differences in general visual imagery ability, musical training, and participation in the visual arts may be associated with patterns of consistency levels shown. However, other important aims were to replicate previous findings of a link between visual imagery and emotion [30, 37], aesthetic appeal [26], and to explore the influence of general visual imagery ability, musical training, and participation in the visual arts on music-induced visual imagery experience.

In sum, we found (a) that storytelling is the most common form of visual imagery during music listening, (b) individuals are more consistent with themselves (in terms of their visual imagery content) than when compared to other listeners, (c) evidence for all-round modest consistency levels, (d) confirmation of links between visual imagery, emotional intensity, and aesthetic appeal, and (e) evidence for a nuanced role of individual differences on music-induced visual imagery experience in terms of prevalence, vividness, and consistency. Further to these findings, we ascertained that visual imagery experience averaged across the three music stimuli was very prevalent in our sample with about 97.5% of participants reporting experiencing at least mild levels of visual imagery, and about 94.6% reporting their visual imagery to be at least a mildly vivid experience (similar proportions were also found across the individual music excerpts).

5.1. Storytelling is the most common form of visual imagery during music listening

The thematic analysis of our open-text question revealed that (in support of H1) story-making was a very pervasive aspect of visual imagery experience. Indeed, our first higher-order theme Storytelling encompassed descriptions of visual imagery ranging from locations to individual characters and was present in over 74% of reports. This finding is in line with the idea that individuals are prone to imagining narratives in response to music [4, 6], and our findings support the notion that this may significantly occur in the visual domain. In addition to showing that visualising story-making is a prevalent aspect of music listening, we were also able to offer insights into the relative prevalence of certain types of visual imagery such as the prominence of details pertaining to setting and location, followed by characters, and the actions they carried out.

We coded two further themes that accompanied our sample’s visual imagery reports, Associations, comprising abstract visual imagery, emotion, and memories, and References, comprising codes regarding media comparisons, instruments, and composers. Codes underlying these themes were present in over 10% of reports, and support past arguments that music listening is a multi-modal experience [48, 49] that involves not just visual experiences but also physical, affective, and semantic connotations. In other words, in addition to reports suggesting that certain visual imagery was in fact ‘seen’, descriptions were also often accompanied by remarks of cross-modal aspects spanning the remaining senses, aesthetic evaluations, and music’s emotional power. These findings highlight the considerable prevalence of semantic associations found within listeners’ descriptions in response to the music, even when explicitly instructed to focus their attentions on visual imagery. Such patterns indicate that narrowing one’s research aims on just the visual components of a listeners’ experience could lead to overestimations of its occurrence as well as overshadows its potential unique links to other semantic associations formed in response to the music; one recent study by Cespedes-Guevara and Dibben [50] showed that a considerable portion of listeners’ reports of what went through their minds while listening comprised an array of semantic, personal, as well as visual experiences.

5.2. Participants show all-round low consistency levels but greater levels for within- than across-participants

We investigated visual imagery consistency using the Jaccard coefficient index, a measure of overlap between two lists. First, we assessed within-participants consistency by comparing each individuals’ content from surveys that were administered three weeks apart. Here, we found that, overall, consistency levels tended to cluster around 20%, indicating that listeners tended to refer to almost a fifth of the same visual imagery features in both instances. Our analysis of across-participant consistency yielded similarly low results when comparing individual participants with the rest of the sample: approx. 7–13% for L3 codes and 13–22% for L2 codes.

At first glance, our findings show notable differences with a recent investigation by Margulis et al. [6]. The authors reported high levels of consistency in their sample with regard to musical narrative engagement that they suggest was, for the most part, determined by cultural experience. However, it is important to point out both the differences in what was being reported on (visual imagery vs. imagined narrative) and how consistency is estimated in the two papers. Indeed, we opted to use the Jaccard coefficient index on lists of themes, a method which takes the presence of individual cases into consideration. In contrast, Margulis and colleagues’ approach focuses on the respondents’ text as a whole, using a cosine similarity as a weighted index to assess semantic ‘closeness’ between portions of text based on their orientations on a multi-dimensional space. Approaches used by other studies to address similar questions have also included frequency-based approaches [7] (i.e., summing occurrences of certain content). Critically, the current research aimed to assess consistency with a focus on the presence or absence of particular themes and topics. By allowing analysis of consistency on the basis of themes and topics, our approach provides a way to estimate consistency levels with a focus on particular aspects of content. However, it would be beneficial to consider the advantages of applying a weighted approach (as Margulis and colleagues have done) when scrutinising visual imagery content in combination with the current study’s methods of assessing discrete overlap. Incorporating both topic overlap as well as weighted similarity might provide insight on the significance of specific types of cross-modal content evidently present (as is seen within our thematic framework) in reports of music-induced visual imagery.

In any case, we were able to confirm H2 about how within- and across-participant consistency would differ from each other, supporting Cross’ [12] aforementioned idea on the subjectivity of music. Specifically, we show that participants possess greater consistency within themselves across two time-points than consistency with other listeners. Critically, this type of behaviour is not so different to what researchers have described as cross-modal correspondences. Deroy and Spence [21] explain that the sensory connections or cross-modal matching are not only pervasive in everyday life but also remain quite consistent over time, which could be down to environmental regularities or contextual associations.

In a final check, we aimed to re-assess consistency by only including Storytelling codes. We viewed this as an important next step to (a) confirm that it was indeed pure visual imagery (i.e., codes where participants report seeing images) that was mostly leading to our observed consistency levels above, and (b) aid comparison with previous work on narrative consistency during music listening [6]. We were able to show that the main conclusions we drew (regarding comparison of within- and across-participant consistency profiles across code levels) all continued to hold even when looking at this reduced set of codes.

5.3. Relating visual imagery consistency to behavioural ratings and individual differences

Next, we explored whether there was a predictive influence of the behavioural measures (prevalence, vividness, music liking, and emotional intensity) and individual differences (general imagery ability, musical training, and participation in the visual arts) on within- and across-participant consistency. We saw links between across-participant consistency and the prevalence and vividness of visual imagery (albeit only L2 for vividness). Interestingly, within-participant consistency was only predicted by vividness and only for L3 codes.

The generally positive relationships found between (across- and within-participant) consistency and prevalence and vividness of visual imagery clearly indicate that the qualitative nature of listeners’ visual imagery influences how consistent they are within themselves and with others. While it may not be appropriate to speculate on the nuanced differences seen with regard to levels of consistency, our results suggest that music that is able to induce highly vivid visual imagery tends to produce largely similar content across listeners. Given that we had explicitly chosen to present musical stimuli capable of inducing distinct types of emotion that likely vary in their compositional techniques, we propose that the unique acoustic features of the different excerpts could have influenced how vividly listeners experienced their visual imagery; an idea in line with literature that has highlighted features such as contrast in promoting narrative thinking [4], although is an area of research that is generally in need of much more insight in order to properly ascertain its influence over visual imagery formation [23].

Further, our analyses show that participation in the visual arts does not predict within- or across-participant consistency at either level of granularity, aligning with the idea that those who participate in the visual arts could experience visual imagery at a high enough frequency, vividness, and creative freedom to be inconsistent in content.

5.4. Visual imagery, emotion, and aesthetic appeal

In line with previous work which has linked emotion induction with visual imagery formation [24, 25, 29, 35, 36, 51], we saw significant associations between the prevalence and vividness of visual imagery and emotional intensity, whereby the amount of visual imagery and its vividness were positively related to the intensity of emotions felt, confirming H3.

However, our results were not able to speak to the nature of the relationship between emotion and visual imagery. Indeed, there has been increasing debate regarding the directionality of this relationship–recent studies suggest that it is the emotion that is first felt that then leads to visual imagery experience [30], whereas others propose a more complex and mutually beneficial interlink between music-induced visual imagery and emotional induction [29]. In at least one study, it has been shown that manipulating listeners’ experience of visual imagery by attempting to hinder it led to a mild suppression of reported induced emotion [35].

Even though our results cannot speak directly to the directionality of the relationship between visual imagery and emotion induction, they extend findings in a number of useful ways. Although our analyses showed that emotional intensity did not predict L3 across-participant consistency, the pattern for L2 consistency was just shy of significance, hinting at the idea that a minor part of what determines high consistency across individuals may be a shared high level of emotion. Future studies could further probe this link by assessing the extent to which emotional valence is a factor driving similarities in visual imagery reports. Secondly, with the thematic analysis, we obtained a lower-level code dedicated to outlining emotional experiences reported in visual imagery descriptions. Indeed, emotional experiences, whether felt or perceived within the imagery, became a prominent aspect of our categorisations and was highly intertwined in descriptions of different visual imagery scenarios [3]; this being further shown in the fact that it was the second most commonly occurring non-visual subtheme of the Associations higher-order theme. Again, while this does not offer evidence into the causal relationship between visual imagery and emotion induction, it speaks to the strong link visual imagery has with emotional experiences.

Finally, in line with previous work that showed visual imagery vividness to be strongly linked with music’s aesthetic appeal [26], and in support of H4, we were also able to reflect this finding between prevalence and vividness of visual imagery and music liking responses, whereby liking was positively associated with visual imagery prevalence and vividness. However, this relationship was weaker than seen between visual imagery and emotion, suggesting that it may not be a leading determinant of music-induced visual imagery experience. Here as well, we were unable to provide full insight into the causal relationship between visual imagery and liking. However, the generally weak associations found draw into doubt the idea that aesthetic appeal may play a major role in influencing visual imagery experience. With regard to what may cause a high amount of music-induced visual imagery, one might expect that aspects of acoustic and musical features like melody and harmony [23], structural features like tempo [5], or even interindividual factors like trait empathy [52], may play as important if not more of a role than liking.

5.5. The roles of individual differences on music-induced visual imagery

We further explored whether music-induced visual imagery ratings were predicted by generally occurring visual imagery levels. In contrast with previous assessments [35], we found that general visual imagery showed a moderate positive predictive association with music-induced visual imagery, supporting H5 and suggesting that general and music-induced visual imagery may be somewhat interrelated. This finding is partly in line with findings by Küssner and Eerola [3] who showed a positive, albeit small, correlation between imagery vividness and general visual imagery. Nevertheless, due to the inconsistencies found across studies, these results warrant further investigations into the processes that may underly the experience of general as well as music-induced visual imagery.

Furthermore, musical training showed no associations with prevalence and vividness of visual imagery ratings (extending H6 to show that the link is not only weak but in fact non-existent), and is in contrast with past observations that music-induced visual imagery was affected by training due to potential functional benefits [3]. This finding however may not come as such a surprise, as several examples of past literature have similarly found no differences between musicians and non-musicians in their visual imagery abilities [38], instead finding superior involvement of other mental imagery experiences in response to music listening (for example, auditory [53, 54] and kinaesthetic [55, 56]).

Further, in support of H7, we found that those who participate in activities associated with the visual arts provided significantly higher prevalence and vividness of visual imagery ratings, both in terms of aggregated and individual music track responses. This is unsurprising, as it is evident that various artists (visual, musical, etc.) use imagery in their creative processes to stimulate the creation and performance of their art [31]. In the context of music, mental imagery is even found to enhance the perceived creativity of music compositions [32]. Thus, although the general chain of causality is unclear, increased participation within various art modalities (in this case, visual) may increase visual imagery engagement with music.

5.6. Implications, limitations, and future directions

With the current research, we have presented a novel methodological approach to probing the content of music-induced visual imagery, a method that we hope will be adopted by future studies seeking to develop the knowledge on the topicality of visual imagery content. The ephemeral nature of visual imagery makes it difficult to measure and draw decisive conclusions. Using our approach of quantifying prevalent themes, we were able corroborate past notions of narrative thinking in the context of music listening and confirm that a large portion of it can be visual in nature.

Our approach may however also suffer from a couple of key limitations; namely that our visual imagery framework was developed in response to only three musical excerpts and were taken from film soundtracks [39]. Our choice of the film genre was partly to ensure that participants were free to provide rich accounts of their visual imagery in response to a programmatic selection of tracks, as well as giving us considerable power with which to analyse consistency, especially when selecting such a low number of listening stimuli. Such decisions mean that our stimuli fall short of being considered entirely comprehensive. It is further a possibility that the predominance of storytelling and media references was an artefact of musical cues present in our chosen tracks that were associated with the development and changes found in film scenes.

However, we propose that these issues could be considered negligible. Margulis [4] found that even when unprompted, individuals listening to instrumental classical music made references to film and television in their narrative open-text reports. Similarities between our study and theirs in revealing a high volume of media references may lie in the fact that our stimuli were predominantly orchestral. Nevertheless, future investigations should aim to incorporate music from genres with contrasting compositional characteristics and instrumentation. One piece of work by Markert and Küssner [57] which utilised ambient music, typically incorporating non-instrumental synthesised compositional techniques, found that while participant reports presented a significant amount of storytelling visual imagery in response to both ambient and classical music excerpts, reports in response to the ambient music tracks also predominantly featured abstract visual imagery (i.e., non-specific images, such as geometric shapes and colours). These findings emphasise the variability in visual imagery that differences in music genre could lead to, and hints at the potential for our framework to be representative of most music genres with further refinement. It would further be relevant to explore differences in the content of our framework and the proportions of themes and consistency levels should listeners be reporting imagery in response to self-selected music. It is likely that this would lead to more personal imagery, due to the higher proportion of autobiographical memories that can be evoked from familiar music [58].

A further important limitation to consider is that the visual imagery content descriptions provided in survey 2 could have been subject to demand characteristics. On the one hand, listeners may have only reported new visual imagery that was not experienced during the first listening instance; on the other, it is possible that some may have felt inclined to report similar visual imagery to that they recalled experiencing in survey 1. Further research on the topic should aim to take this issue into consideration.

Some literature has emphasised the positive effects that eye closure may have on visual imagery experienced in response to music (e.g., [35, 59, 60]), so much so that current imagery-based therapies adopt eye closure as a way to enhance the benefits of rehabilitation [61]. One study that specifically compares the facilitatory effects of eye opening or closure on participants’ reported visual imagery experience found that eye closure led to markedly higher visual imagery vividness as well as content [62]. The current design did not offer specific instructions on whether participants should listen to each musical excerpt with eyes open or closed, thus this action was free to vary across the sample. Given the dramatic influence that the change in instruction can have on visual imagery experience, especially in the context of music listening, such an instruction would be important to include in future research hoping to enhance the experience of visual imagery.

The design of the current study required participants to provide unrestricted reports on the content of their visual imagery to music, implicating their ability to verbalise their visual imagery experiences (i.e., constructing their narrative to music [4]). It is generally well evidenced that music and language are reflected by overlapping electrophysiological correlates [16, 6367], but some studies identify differences in this regard (e.g., between males and females [6871], and as a function of musicianship [7277]). Neuroscientific studies on music-induced visual imagery are only beginning to emerge (see [61, 78] for first evidence of neural signatures). However, we suggest that future studies may seek to combine approaches like those taken in our current paper with emerging insights into neural underpinnings in order to advance knowledge of both the brain and visual imagery during music listening.

What musical characteristics are more likely to result in music-induced visual imagery? While this concept has been minimally addressed, Margulis [4] found that one potential contributing factor in leading listeners to narratively engage with music is musical contrast (sudden and unexpected changes in musical events). Juslin [23] proposes that musical features, such as predictability and repetition, may be driving forces in leading listeners to form visual imagery to music. In any case, the musical features that may link to specific types of visual imagery is, to date, a vast and unanswered question.

Finally, future investigations may consider extending the time between administering the first and second survey to assess the endurance of visual imagery experience more effectively. A yet more comprehensive approach would involve presenting additional administrations of the survey: this is to assess potential modulations more systematically in terms of how levels of within-person consistency change over time.

5.7. Conclusion

We have presented a detailed investigation into the visual imagery content that listeners experience in response to music. We show that visual imagery is a highly prevalent aspect of individuals’ listening experience, with storytelling being particularly prominent. We also demonstrate the idiosyncrasies of listeners’ content consistency by showing that they were, on average, relatively consistent with themselves across timepoints, in contrast to when compared with other listeners.

The ease with which music appears to elicit visual imagery offers further support for the connection between music and language processing with regard to listeners’ inclination to derive meaning from the music. We anticipate that our research will set a precedence for further studies to develop and hone our understanding of the inherent visual imagery qualities experienced during music listening.

Supporting information

S1 File. Appendices.

(DOCX)

S2 File. Supplementary.

(DOCX)

Acknowledgments

We would like to thank Olivia Geibel for her assistance with coding the visual imagery description content during the thematic analysis.

Data Availability

All data files are available from the OSF database: https://osf.io/nf4x7/?view_only=081602e5aca94788bb959b48ed8b47ef.

Funding Statement

The author(s) received no specific funding for this work.

References

  • 1.Maus FE. Music As Narrative. Indiana Theory Rev. 1991;12: 1–34. [Google Scholar]
  • 2.Newcomb A. Schumann and Late Eighteenth-Century Narrative Strategies. 19th-Century Music. 1987;11: 164–174. doi: 10.2307/746729 [DOI] [Google Scholar]
  • 3.Küssner MB, Eerola T. The content and functions of vivid and soothing visual imagery during music listening: Findings from a survey study. Psychomusicology Music Mind Brain. 2019;29: 90–99. doi: 10.1037/pmu0000238 [DOI] [Google Scholar]
  • 4.Margulis EH. An Exploratory Study of Narrative Experiences of Music. Music Percept Interdiscip J. 2017;35: 235–248. doi: 10.1525/mp.2017.35.2.235 [DOI] [Google Scholar]
  • 5.Herff SA, Cecchetti G, Taruffi L, Déguernel K. Music influences vividness and content of imagined journeys in a directed visual imagery task. Sci Rep. 2021;11: 15990. doi: 10.1038/s41598-021-95260-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Margulis EH, Wong PCM, Turnbull C, Kubit BM, McAuley JD. Narratives imagined in response to instrumental music reveal culture-bounded intersubjectivity. Proc Natl Acad Sci. 2022;119: e2110406119. doi: 10.1073/pnas.2110406119 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Dahl S, Stella A, Bjørner T. Tell me what you see: An exploratory investigation of visual mental imagery evoked by music. Music Sci. 2022; 10298649221124862. doi: 10.1177/10298649221124862 [DOI] [Google Scholar]
  • 8.Nattiez J-J. Can One Speak of Narrativity in Music? J R Music Assoc. 1990;115: 240–257. doi: 10.1093/jrma/115.2.240 [DOI] [Google Scholar]
  • 9.Levinson J. Music as Narrative and Music as Drama. Mind Lang. 2004;19: 428–441. doi: 10.1111/j.0268-1064.2004.00267.x [DOI] [Google Scholar]
  • 10.McClary S. Narrative Agendas in “Absolute” Music: Identity and Difference in Brahms’s Third Symphony. In: Solie RA, editor. Musicology and Difference. University of California Press; 1993. pp. 326–344. doi: 10.1525/9780520916500-017 [DOI] [Google Scholar]
  • 11.Koelsch S, Kasper E, Sammler D, Schulze K, Gunter T, Friederici AD. Music, language and meaning: brain signatures of semantic processing. Nat Neurosci. 2004;7: 302–307. doi: 10.1038/nn1197 [DOI] [PubMed] [Google Scholar]
  • 12.Cross I. Music as a social and cognitive process. In: Rebuschat P, Rohmeier M, Hawkins JA, Cross I, editors. Language and Music as Cognitive Systems. Oxford University Press; 2011. pp. 315–328. doi: 10.1093/acprof:oso/9780199553426.003.0033 [DOI] [Google Scholar]
  • 13.Raffman D., Language Music, and Mind. Cambridge, Massachusetts: The MIT Press; 1993. [Google Scholar]
  • 14.Steinbeis N, Koelsch S. Shared Neural Resources between Music and Language Indicate Semantic Processing of Musical Tension-Resolution Patterns. Cereb Cortex. 2008;18: 1169–1178. doi: 10.1093/cercor/bhm149 [DOI] [PubMed] [Google Scholar]
  • 15.Sternin A, McGarry LM, Owen AM, Grahn JA. The Effect of Familiarity on Neural Representations of Music and Language. J Cogn Neurosci. 2021; 1–17. doi: 10.1162/jocn_a_01737 [DOI] [PubMed] [Google Scholar]
  • 16.Tillmann B Music and Language Perception: Expectations, Structural Integration, and Cognitive Sequencing. Top Cogn Sci. 2012;4: 568–584. doi: 10.1111/j.1756-8765.2012.01209.x [DOI] [PubMed] [Google Scholar]
  • 17.Barraza P, Chavez M, Rodríguez E. Ways of making-sense: Local gamma synchronization reveals differences between semantic processing induced by music and language. Brain Lang. 2016;152: 44–49. doi: 10.1016/j.bandl.2015.12.001 [DOI] [PubMed] [Google Scholar]
  • 18.Taruffi L, Pehrs C, Skouras S, Koelsch S. Effects of Sad and Happy Music on Mind-Wandering and the Default Mode Network. Sci Rep. 2017;7: 14396. doi: 10.1038/s41598-017-14849-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Koelsch S, Bashevkin T, Kristensen J, Tvedt J, Jentschke S. Heroic music stimulates empowering thoughts during mind-wandering. Sci Rep. 2019;9: 10317. doi: 10.1038/s41598-019-46266-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Schaerlaeken S, Glowinski D, Rappaz M-A, Grandjean D. “Hearing music as…”: Metaphors evoked by the sound of classical music. Psychomusicology Music Mind Brain. 2019;29: 100–116. doi: 10.1037/pmu0000233 [DOI] [Google Scholar]
  • 21.Deroy O, Spence C. Crossmodal Correspondences: Four Challenges. Multisensory Res. 2016;29: 29–48. doi: 10.1163/22134808-00002488 [DOI] [PubMed] [Google Scholar]
  • 22.Leman M, Maes P-J, Nijs L, Van Dyck E. What Is Embodied Music Cognition? In: Bader R, editor. Springer Handbook of Systematic Musicology. Berlin, Heidelberg: Springer Berlin Heidelberg; 2018. pp. 747–760. doi: 10.1007/978-3-662-55004-5_34 [DOI] [Google Scholar]
  • 23.Juslin PN. Musical Emotions Explained: Unlocking the Secrets of Musical Affect. Oxford, New York: Oxford University Press; 2019. [Google Scholar]
  • 24.Taruffi L, Küssner MB. A review of music-evoked visual mental imagery: Conceptual issues, relation to emotion, and functional outcome. Psychomusicology Music Mind Brain. 2019;29: 62–74. doi: 10.1037/pmu0000226 [DOI] [Google Scholar]
  • 25.Juslin PN. From everyday emotions to aesthetic emotions: Towards a unified theory of musical emotions. Phys Life Rev. 2013;10: 235–266. doi: 10.1016/j.plrev.2013.05.008 [DOI] [PubMed] [Google Scholar]
  • 26.Belfi AM. Emotional valence and vividness of imagery predict aesthetic appeal in music. Psychomusicology Music Mind Brain. 2019;29: 128–135. doi: 10.1037/pmu0000232 [DOI] [Google Scholar]
  • 27.Presicce G, Bailes F. Engagement and visual imagery in music listening: An exploratory study. Psychomusicology Music Mind Brain. 2019;29: 136–155. doi: 10.1037/pmu0000243 [DOI] [Google Scholar]
  • 28.Koelsch S, Skouras S, Fritz T, Herrera P, Bonhage C, Küssner MB, et al. The roles of superficial amygdala and auditory cortex in music-evoked fear and joy. NeuroImage. 2013;81: 49–60. doi: 10.1016/j.neuroimage.2013.05.008 [DOI] [PubMed] [Google Scholar]
  • 29.Vroegh T. Investigating the directional link between music-induced visual imagery and two different types of emotional responses. Berlin, Germany; 2018. [Google Scholar]
  • 30.Day RA, Thompson WF. Measuring the onset of experiences of emotion and imagery in response to music. Psychomusicology Music Mind Brain. 2019;29: 75–89. doi: 10.1037/pmu0000220 [DOI] [Google Scholar]
  • 31.Rosenberg HS, Trusheim W. Creative Transformations: How Visual Artists, Musicians, and Dancers Use Mental Imagery in Their Work. In: Shorr JE, Robin P, Connella JA, Wolpin M, editors. Imagery: Current Perspectives. Boston, MA: Springer US; 1989. pp. 55–75. doi: 10.1007/978-1-4899-0876-6_6 [DOI] [Google Scholar]
  • 32.Wong SSH, Lim SWH. Mental imagery boosts music compositional creativity. PLOS ONE. 2017;12: e0174009. doi: 10.1371/journal.pone.0174009 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Eitan Z, Granot RY. How Music Moves: Musical Parameters and Listeners’ Images of Motion. Music Percept. 2006;23: 221–248. doi: 10.1525/mp.2006.23.3.221 [DOI] [Google Scholar]
  • 34.Johnson ML, Larson S. “Something in the Way She Moves”—Metaphors of Musical Motion. Metaphor Symb. 2003;18: 63–84. doi: 10.1207/S15327868MS1802_1 [DOI] [Google Scholar]
  • 35.Hashim S, Stewart L, Küssner MB. Saccadic Eye-Movements Suppress Visual Mental Imagery and Partly Reduce Emotional Response During Music Listening. Music Sci. 2020;3: 205920432095958. doi: 10.1177/2059204320959580 [DOI] [Google Scholar]
  • 36.Juslin PN, Västfjäll D. Emotional responses to music: The need to consider underlying mechanisms. Behav Brain Sci. 2008;31: 559–575. doi: 10.1017/S0140525X08005293 [DOI] [PubMed] [Google Scholar]
  • 37.Vuoskoski JK, Eerola T. Extramusical information contributes to emotions induced by music. Psychol Music. 2015;43: 262–274. doi: 10.1177/0305735613502373 [DOI] [Google Scholar]
  • 38.Talamini F, Vigl J, Doerr E, Grassi M, Carretti B. Auditory and visual mental imagery in musicians and non-musicians. Music Sci. 2022; 102986492110627. doi: 10.1177/10298649211062724 [DOI] [Google Scholar]
  • 39.Eerola T, Vuoskoski JK. A comparison of the discrete and dimensional models of emotion in music. Psychol Music. 2011;39: 18–49. doi: 10.1177/0305735610362821 [DOI] [Google Scholar]
  • 40.Vuoskoski JK, Thompson WF, McIlwain D, Eerola T. Who Enjoys Listening to Sad Music and Why? Music Percept Berkeley. 2012;29: 311–317. [Google Scholar]
  • 41.Pekala RJ. The Phenomenology of Consciousness Inventory. Quantifying Consciousness: Emotions, Personality, and Psychotherapy. Boston: Springer; 1991. [Google Scholar]
  • 42.Müllensiefen D, Gingras B, Musil J, Stewart L. The Musicality of Non-Musicians: An Index for Assessing Musical Sophistication in the General Population. PLOS ONE. 2014;9: e89642. doi: 10.1371/journal.pone.0089642 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Marks DF. Visual Imagery Differences in the Recall of Pictures. Br J Psychol. 1973;64: 17–24. doi: 10.1111/j.2044-8295.1973.tb01322.x [DOI] [PubMed] [Google Scholar]
  • 44.Braun V, Clarke V. Using thematic analysis in psychology. Qual Res Psychol. 2006;3: 77–101. doi: 10.1191/1478088706qp063oa [DOI] [Google Scholar]
  • 45.R Core Team. R: A Language and Environment for Statistical Computing. Vienna, Austria: R Foundation for Statistical Computing; 2018. Available: https://www.R-project.org/ [Google Scholar]
  • 46.Bates D, Mächler M, Bolker B, Walker S. Fitting Linear Mixed-Effects Models Using lme4. J Stat Softw. 2015;67. doi: 10.18637/jss.v067.i01 [DOI] [Google Scholar]
  • 47.Kuznetsova A, Brockhoff PB, Christensen RHB. lmerTest Package: Tests in Linear Mixed Effects Models. J Stat Softw. 2017;82. doi: 10.18637/jss.v082.i13 [DOI] [Google Scholar]
  • 48.Deroy O. Evocation: How Mental Imagery Spans Across the Senses. 1st ed. In: Abraham A, editor. The Cambridge Handbook of the Imagination. 1st ed. Cambridge University Press; 2020. pp. 276–290. doi: 10.1017/9781108580298.018 [DOI] [Google Scholar]
  • 49.Nanay B. Multimodal mental imagery. Cortex. 2018;105: 125–134. doi: 10.1016/j.cortex.2017.07.006 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Cespedes-Guevara J, Dibben N. The Role of Embodied Simulation and Visual Imagery in Emotional Contagion with Music. Music Sci. 2022;5: 20592043221093836. doi: 10.1177/20592043221093836 [DOI] [Google Scholar]
  • 51.Balteș FR, Miu AC. Emotions during live music performance: Links with individual differences in empathy, visual imagery, and mood. Psychomusicology Music Mind Brain. 2014;24: 58–65. doi: 10.1037/pmu0000030 [DOI] [Google Scholar]
  • 52.Taruffi L, Skouras S, Pehrs C, Koelsch S. Trait Empathy Shapes Neural Responses Toward Sad Music. Cogn Affect Behav Neurosci. 2021. [cited 21 Jan 2021]. doi: 10.3758/s13415-020-00861-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Bishop L, Bailes F, Dean RT. Musical Imagery and the Planning of Dynamics and Articulation During Performance. Music Percept. 2013;31: 97–117. doi: 10.1525/mp.2013.31.2.97 [DOI] [Google Scholar]
  • 54.Keller PE, Appel M. Individual Differences, Auditory Imagery, and the Coordination of Body Movements and Sounds in Musical Ensembles. Music Percept. 2010;28: 27–46. doi: 10.1525/mp.2010.28.1.27 [DOI] [Google Scholar]
  • 55.Clark T, Williamon A. Imagining the music: Methods for assessing musical imagery ability. Psychol Music. 2012;40: 471–493. doi: 10.1177/0305735611401126 [DOI] [Google Scholar]
  • 56.Di Nuovo SF, Angelica A. Musical skills and perceived vividness of imagery: differences between musicians and untrained subjects. Ann Della Fac Sci Della Formazione—Univ Degli Studi Catania. 2015;14: 3–13. doi: 10.4420/unict-asdf.14.2015.1 [DOI] [Google Scholar]
  • 57.Markert N, Küssner MB. An exploratory study of visual mental imagery induced by ambient music. In: Stupacher J, Hagner S, editors. Proceedings of the 14th International Conference of Students of Systematic Musicology (SysMus21). Aarhus, Denmark: Open Science Framework; 2021. doi: 10.17605/OSF.IO/HG6RZ [DOI] [Google Scholar]
  • 58.Jakubowski K, Francini E. Differential effects of familiarity and emotional expression of musical cues on autobiographical memory properties. Q J Exp Psychol. 2022; 174702182211297. doi: 10.1177/17470218221129793 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Vredeveldt A, Hitch GJ, Baddeley AD. Eyeclosure helps memory by reducing cognitive load and enhancing visualisation. Mem Cognit. 2011;39: 1253–1263. doi: 10.3758/s13421-011-0098-8 [DOI] [PubMed] [Google Scholar]
  • 60.van den Hout MA, Engelhard IM, Beetsma D, Slofstra C, Hornsveld H, Houtveen J, et al. EMDR and mindfulness. Eye movements and attentional breathing tax working memory and reduce vividness and emotionality of aversive ideation. J Behav Ther Exp Psychiatry. 2011;42: 423–431. doi: 10.1016/j.jbtep.2011.03.004 [DOI] [PubMed] [Google Scholar]
  • 61.Fachner JC, Maidhof C, Grocke D, Nygaard Pedersen I, Trondalen G, Tucek G, et al. “Telling me not to worry…” Hyperscanning and Neural Dynamics of Emotion Processing During Guided Imagery and Music. Front Psychol. 2019;10. doi: 10.3389/fpsyg.2019.01561 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Herff SA, McConnell S, Ji JL, Prince JB. Eye Closure Interacts with Music to Influence Vividness and Content of Directed Imagery. Music Sci. 2022;5: 20592043221142711. doi: 10.1177/20592043221142711 [DOI] [Google Scholar]
  • 63.Carrus E, Pearce MT, Bhattacharya J. Melodic pitch expectation interacts with neural responses to syntactic but not semantic violations. Cortex. 2013;49: 2186–2200. doi: 10.1016/j.cortex.2012.08.024 [DOI] [PubMed] [Google Scholar]
  • 64.Koelsch S, Gunter TC, Wittfoth M, Sammler D. Interaction between Syntax Processing in Language and in Music: An ERP Study. J Cogn Neurosci. 2005;17: 1565–1577. doi: 10.1162/089892905774597290 [DOI] [PubMed] [Google Scholar]
  • 65.Koelsch S. Music-syntactic processing and auditory memory: Similarities and differences between ERAN and MMN. Psychophysiology. 2009;46: 179–190. doi: 10.1111/j.1469-8986.2008.00752.x [DOI] [PubMed] [Google Scholar]
  • 66.Kraus N, Slater J. Chapter 12—Music and language: relations and disconnections. In: Aminoff MJ, Boller F, Swaab DF, editors. Handbook of Clinical Neurology. Elsevier; 2015. pp. 207–222. doi: 10.1016/B978-0-444-62630-1.00012–3 [DOI] [PubMed] [Google Scholar]
  • 67.Maidhof C, Koelsch S. Effects of Selective Attention on Syntax Processing in Music and Language. J Cogn Neurosci. 2011;23: 2252–2267. doi: 10.1162/jocn.2010.21542 [DOI] [PubMed] [Google Scholar]
  • 68.Koelsch S, Maess B, Grossmann T, Friederici AD. Electric brain responses reveal gender differences in music processing. NeuroReport. 2003;14: 709–713. doi: 10.1097/00001756-200304150-00010 [DOI] [PubMed] [Google Scholar]
  • 69.Koelsch S, Grossmann T, Gunter TC, Hahne A, Schröger E, Friederici AD. Children Processing Music: Electric Brain Responses Reveal Musical Competence and Gender Differences. J Cogn Neurosci. 2003;15: 683–693. doi: 10.1162/089892903322307401 [DOI] [PubMed] [Google Scholar]
  • 70.Kramer JH, Delis DC, Daniel M. Sex differences in verbal learning. J Clin Psychol. 1988;44: 907–915. doi: 10.1002/1097-4679(198811)44:6&lt;907::AID-JCLP2270440610&gt;3.0.CO;2–8 [DOI] [Google Scholar]
  • 71.Steinbeis N, Koelsch S. Comparing the Processing of Music and Language Meaning Using EEG and fMRI Provides Evidence for Similar and Distinct Neural Representations. PLOS ONE. 2008;3: e2226. doi: 10.1371/journal.pone.0002226 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72.Brochard R, Dufour A, Després O. Effect of musical expertise on visuospatial abilities: Evidence from reaction times and mental imagery. Brain Cogn. 2004;54: 103–109. doi: 10.1016/S0278-2626(03)00264-1 [DOI] [PubMed] [Google Scholar]
  • 73.Trainor LJ, Shahin AJ, Roberts LE. Understanding the Benefits of Musical Training. Ann N Y Acad Sci. 2009;1169: 133–142. doi: 10.1111/j.1749-6632.2009.04589.x [DOI] [PubMed] [Google Scholar]
  • 74.Klein C, Liem F, Hänggi J, Elmer S, Jäncke L. The “silent” imprint of musical training. Hum Brain Mapp. 2015;37: 536–546. doi: 10.1002/hbm.23045 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 75.Aleman A, Nieuwenstein MR, Böcker KBE, de Haan EHF. Music training and mental imagery ability. Neuropsychologia. 2000;38: 1664–1668. doi: 10.1016/s0028-3932(00)00079-8 [DOI] [PubMed] [Google Scholar]
  • 76.Bangert M, Altenmüller EO. Mapping perception to action in piano practice: a longitudinal DC-EEG study. BMC Neurosci. 2003;4: 26. doi: 10.1186/1471-2202-4-26 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77.Lotze M. Kinesthetic imagery of musical performance. Front Hum Neurosci. 2013;7. Available: https://www.frontiersin.org/ [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 78.Hashim S, Küssner MB, Weinreich A, Omigie D. The neuro-oscillatory profiles of static and dynamic music-induced visual imagery. under review. [DOI] [PubMed] [Google Scholar]

Decision Letter 0

Ioanna Markostamou

27 Mar 2023

PONE-D-23-03570Music listening leads to predominantly story-like visual imagery that shows idiosyncratic tendencies

PLOS ONE

Dear Dr. Hashim,

Thank you for submitting your manuscript to PLOS ONE. The manuscript has been evaluated by three very knowledgeable experts in the topical area you are investigating, who all wish to be identified and are Dr. Steffen A. Herff, Dr. Julian Cespedes-Guevara, and Prof. Marta Olivetti Belardinelli.

As you will see when you read their critiques further below and in the attached documents, all three reviewers evaluate positively your work and highlight that it can make a significant contribution to the field. However, the reviewers raise certain concerns that should be addressed in order for the paper to be suitable for publication.

Specifically, all reviewers note that more clarifications should be provided in several parts of the manuscript – most notably, when outlining the formulation of the study aims and hypotheses. Reviewers also make some critical suggestions with respect to statistical analysis and results sections. I believe that addressing these comments would strengthen the clarity of the manuscript and increase its potential impact.

Therefore, I invite you to submit a revised version of the manuscript that addresses the points raised by the reviewers. When revising your manuscript, please consider all issues mentioned in the reviewers’ comments carefully – please outline every change made in response to their comments and provide suitable rebuttals for any comments not addressed. Please note that your revised submission may need to be re-reviewed. 

Please submit your revised manuscript by May 11 2023 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

Please include the following items when submitting your revised manuscript:

  • A rebuttal letter that responds to each point raised by the academic editor and reviewers. You should upload this letter as a separate file labeled 'Response to Reviewers'.

  • A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.

  • An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols. Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols.

I look forward to receiving your revised manuscript.

Kind regards,

Ioanna Markostamou, Ph.D.

Academic Editor

PLOS ONE

Journal Requirements:

When submitting your revision, we need you to address these additional requirements.

1. Please ensure that your manuscript meets PLOS ONE's style requirements, including those for file naming. The PLOS ONE style templates can be found at 

https://journals.plos.org/plosone/s/file?id=wjVg/PLOSOne_formatting_sample_main_body.pdf and 

https://journals.plos.org/plosone/s/file?id=ba62/PLOSOne_formatting_sample_title_authors_affiliations.pdf

2. Please provide additional details regarding participant consent. In the ethics statement in the Methods and online submission information, please ensure that you have specified what type you obtained (for instance, written or verbal, and if verbal, how it was documented and witnessed). If your study included minors, state whether you obtained consent from parents or guardians. If the need for consent was waived by the ethics committee, please include this information.

3. Peer review at PLOS ONE is not double-blinded (https://journals.plos.org/plosone/s/editorial-and-peer-review-process). For this reason, authors should include in the revised manuscript all the information removed for blind review.

4. Please include your full ethics statement in the ‘Methods’ section of your manuscript file. In your statement, please include the full name of the IRB or ethics committee who approved or waived your study, as well as whether or not you obtained informed written or verbal consent. If consent was waived for your study, please include this information in your statement as well. 

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: Yes

**********

2. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: Yes

**********

3. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: Yes

**********

4. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: Yes

**********

5. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)

Reviewer #1: 

Dear Dr. Markostamou, Dear Authors,

I had the great pleasure to review the article ‘Music listening leads to predominantly story-like visual imagery that shows idiosyncratic tendencies’ submitted for publication to PLOS ONE (PONE-D-23-03570; review request received: 18.02.2023, review accepted: 20.02.2023, review submitted: 22.02.2023).

This large-scale online study investigates inter- and intra-participant consistency of music-indued visual mental imagery. The manuscript is well written, the experiment is well designed to address the core research questions, and the results fill an important gap in the literature. The manuscript is a great contribution to the field and well suited to the audience of PLOS-One.

I can strongly recommend the manuscript for publication; however, I have a few minor suggestions and three larger concerns regarding formalism and statistical analysis that I recommend the authors should address first. These are detailed in the attached file PONE-D-23-03570 - REVIEW.pdf.

Reviewer #2: 

I would like to congratulate the author(s) for the rigorous and interesting investigation they carried out. I think it makes a valuable contribution to advancing our knowledge about the phenomenon of music-evoked visual imagery. That said, I have a few suggestions for improving the paper:

(Page 4 Lines 75-82). First, in the introduction, I suggest mentioning Day and Thompson's (2019) work, about how that in many cases, emotional responses to music may occur earlier than visual imagery. This finding suggests that visual imagery is not the necessarily the driving factor behind emotion, but the other way around. The paper does mention this paper in the discussion, but I think it is worth including in the review about the link between imagery and emotion.

(Pages 7 and 8: “The current research”). I found it difficult to understand the aims of the study. I had to read them several times to understand the difference between aims 1 and 4. Additionally, it is not immediately clear what the author(s) mean by “formulating a framework” in aim 2. I suggest rewriting these paragraphs so that their meaning is more readily apparent.

(Pages 8 and 9, Lines 182-204). The author(s) state that they expect to “be able to replicate previous reports of a relationship between visual imagery and both emotion induction and aesthetic appeal” (page 8, lines 182-183). I think this hypothesis needs to be outlined in more detail: what type of relationship did they expect to find? I also suggest labelling the hypotheses stated in this section with numbers and use them throughout the paper to identify them, particularly in the results section.

(Page 10, Lines 223-227). Please provide musical details about the musical excerpts that were used: mode, tempo, instrumentation, style, and the movies from where they were taken from.

(Page 11, Lines 249-250). Did the authors offer any explanation of what synesthesia consists of to the participants?

(Page 15, lines 340-347). I suggest expanding the rationale for using both an across-participants and a between-participants measures of consistency. I did not find this explanation offered by the authors to be sufficient to grasp the usefulness of using both measures: “to allow greater comparison of consistency values within and between individuals, a between-participants consistency was calculated by comparing each participant with a random other individual” (lines 345-347).

(Pages 23-24, lines 841-847; and Pages 26-27, lines 507-510). From my point of view, these sections of the paper are the ones that require revision more urgently, because the author(s) describe the content of the participants’ reported experiences in contradicting terms. For instance, on page 24, lines 485-487, they write: “Whilst not denoting an exact visual image, its presence within reports was considerably prevalent and constituted a prominent attribute of visual imagery experience”. This is confusing, because if the participants did not explicitly report seeing any visual images in their mind, then those experiences do not correspond to “visual experience(s)”. They are indeed, subjective mental experiences while listening to music, but not visual ones. I checked the authors’ original data on the Open Science website, and found that many of the quotes from the participants in the “Feelings & Atmosphere” and “Music” categories correspond to abstract terms such as “Hope and courage", "Sadness", "Haugtiness", "Love, care", "Tranquility peace well-being love", "Fear", "National anthem", "slow music", "A piano like music for study or sleep", "Elevator music", etc. There is nothing in these terms that suggest any visual dimension to them, and therefore, classifying them as “visual imagery” is misleading. From my point of view, these reports correspond to semantic associations about the culturally-shared uses of functions of music.

One important implication of not calling these experiences “visual imagery”, is that the reported percentages of “visual imagery prevalence” should be slightly adjusted.

(Pages 25, lines 507-510). Please provide examples of the terms that participants used in this category.

(Page 35, lines 680-688). I suggest that the authors should emphasize that the fact that a significant proportion of the participants’ reports did not have a “visual” quality to them, implies that by narrowing our focus on researching “visual” experiences we may be missing out an important aspect of listeners’ experiences. They could also mention that this has been found in previous research. For instance, Cespedes-Guevara and Dibben (2022) found that this is a common occurrence even when participants are provided with written narratives about the music meaning. Furthermore, those authors suggest that the power of music to evoke semantic associations may be a crucial factor behind listeners’ mind wandering and emotional experiences while listening to music. That interpretation may also help explain the relatively high levels of consistency found on level 2 in the present investigation.

(Page 42, lines 861-868). I agree with the authors that using movie soundtracks as stimuli was an important limitation of this study. I suggest also mentioning the fact that they only used 3 stimuli was also a limitation.

Reviewer #3: 

It would be important and interesting to give the distinction between Males and Females for all results presented. If the Authors decide to not perform this distinction, this decision should be indicated as a limitation of the paper.

**********

6. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #1: Yes: Steffen A. Herff

Reviewer #2: Yes: Julian Cespedes-Guevara

Reviewer #3: Yes: Marta olivetti Belardinelli

**********

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at figures@plos.org. Please note that Supporting Information files do not need this step.

Attachment

Submitted filename: PONE-D-23-03570 - REVIEW.pdf

Attachment

Submitted filename: PONS review.docx

PLoS One. 2023 Oct 26;18(10):e0293412. doi: 10.1371/journal.pone.0293412.r002

Author response to Decision Letter 0


28 Jul 2023

Dear Editor,

Thank you for the opportunity to submit a revised version of our manuscript ‘Music listening leads to predominantly story-like visual imagery that shows idiosyncratic tendencies’ (newly entitled ‘Music listening evokes story-like visual imagery with both idiosyncratic and shared content’) to PLOS ONE. We are grateful to the reviewers for their constructive comments and suggested amendments to our paper. All changes are made on a revised version of the manuscript with tracked changes implemented (entitled ‘Revised Manuscript with Tracked Changes’) as well as in an additional revised unmarked version (entitled ‘Manuscript’), as requested.

Please find below the reviewers’ comments and our responses following each one:

Comments from Reviewer 1 – Dr Steffen A. Herff

Major Comments

Comment: The manuscript in general is of very high quality, however, the results section and statistical approach in its current state is not on par with the rest of the excellent manuscript. Despite the experimental design relying on various related variables that are co-dependent, the authors opted for a very simple analytical approach and base much inference on a large number of individual correlations and t-tests, which (if sticking to the Null-Hypothesis-Significance-Testing analytical frame worked use here) not only leads to a dramatic alpha error accumulation, but also misrepresents the data structure, whilst simultaneously being unnecessarily verbose. This is a shame, as the data set and results are very exciting and constitute a great contribution to the field. A slightly more appropriate analysis that focuses on the main narrative of (in)consistency in music-induced visual mental imagery would substantially increase the quality of the paper, make greater use of the valuable data the authors collected, and streamline the results section. I’ve provided some detailed suggestions, including annotated example code in the detailed comments.

Response: We would like to thank the Reviewer for their perspective on our analytical approach, as well as their detailed annotations at the very end of their document. We agree with the suggested changes to the analysis pipeline and have decided to implement these. These especially include replacing our correlation tests run throughout the Results section with linear mixed effects models that take into consideration the full data across our three pieces.

Comment: The formal calculations of the ‘between’ consistency seem unconventional and possibly a bit off. The authors calculate three metrics of consistency. Within consistency, which is operationalised through Jaccard coefficient index between topic codes of the first and second free format responses, which makes a lot of sense. Across consistency, which is the averaged Jaccard coefficient between a given participant and each other participant in the same condition, which also makes a lot of sense. And a between consistency (largely used for the within vs between consistency comparison), which is the same as the across consistency, only that rather than exhaustively calculates, it samples a single other participant. I see no tangible advantage to the between measure over the across measure, but it does come with a lot of disadvantages (e.g., sample biases, lack of reproducibility, low statistical resolution, and reliance on a particular random seed that would need to be shared). It is not clear why this metric is needed or how it provides ‘greater comparison’ for within vs between tests. I recommend removing the between metric and instead focusing on the across consistency metric, which captures the same information, just much more reliably.

Response: We agree that the distinction between our ‘between’ and ‘across’ measures could have been reflected more convincingly. The initial goal here was to calculate a quasi ‘across’-participant measure that would, in principle, be similar to the within-participant measure (i.e., single-individual comparisons). After further consideration of the issues posed by the Reviewer, we have decided to exclude this measure from the paper and focus solely on the across-participant measure as our metric of inter-individual consistency wherever relevant.

Comment: Comparing consistency levels across levels (i.e., Level 2 vs Leve 3) is formally problematic. This is because the results reported are an innate feature of the data modelling framework. Through the authors’ thematic analysis, the data is structured as a convolutional, embedded hierarchy where each Level 3 codes projects to a point on the Level 2 space. A structure like this necessitates that consistency (as calculated but Jaccard coefficient indices) increases as you move from Level 3 to Level 2, and the factor of increase is largely determined by the distribution of converging coding labels, rather than semantically interpretable properties of participants’ responses. I provide more detail on this in the detailed comments. I strongly recommend removing these analyses from the manuscript, both in the results section as well as their interpretation in the discussion, as this finding is simply an innate feature of your annotation strategy, and instead focusing on the comparison within levels (which are very exciting).

Response: We would like to thank the Reviewer for their explanation into why the analysis between our L2 and L3 codes is problematic (as well as their more detailed example on this issue in another comment later). We agree and, as such, comparisons of the L2 and L3 codes have been removed.

Detailed Comments

Comment: p.1: Title ‘Music listening leads to predominantly story-like visual imagery that shows idiosyncratic tendencies’. – I feel the title might be a bit too strong. You report (even on the detailed Level 3) 20.6% within and 9.8% across consistency. Considering that participants had the opportunity to imagine whatever they wanted, it is remarkable how close your ‘across’ consistency measure got to your ‘within’ consistency measure. A title along the lines of the following might more closely represent your numerical findings: ‘Music listening evokes story-like visual imagery with both idiosyncratic as well as shared content’. Of course, the title is your choice, so please consider this merely a suggestion.

Response: We thank the Reviewer for their suggestion of a more appropriate manuscript title. We have taken this into consideration and have, for the most part, adopted the suggestions proposed by the Reviewer to create a new title: ‘Music listening evokes story-like visual imagery with both idiosyncratic and shared content’

Comment: p.2, l. 31: Abstract: ‘Of the initial sample, 254 respondents completed the survey again three weeks later. – The abstract is very well written and gives an engaging summary of the project. However, this particular sentence seems a bit out of place. At this point of the abstract, the reader does not know yet what the survey is. I recommend moving the sentence to a slightly later place. […]

Response: Agreed. This sentence has now been moved to a later more appropriate place, as suggested by the Reviewer: ‘Further, they completed items assessing a number of individual differences including musical training and general imagery ability. Of the initial sample, 254 respondents completed the survey again three weeks later.’.

Comment: p.3, l.63: Introduction ‘However, while it is unlikely that music-induced visual imagery solely comprises of static mental pictures, little is known about how it may develop over time.’ – Could just be me, but I found the phrasing here a bit unclear. What does ‘develop over time’ refer to? Based on the content of the paper, I guess it is about test-retest reliability, but since the first part of the sentence is about static images, one might think ‘develop over time’ refers simply to dynamic images within one imagery session. The next sentence then jumps to across listener consistency. Maybe you can rephrase this sentence to emphasize your point.

Response: We agree that this sentence could be clearer. This sentence has now been removed and the paragraph in general has been rephrased to reflect our points more concisely, as follows: ‘For most listeners, forming a narrative in their mind’s eye is a way to engage with heard music [3,4], and such narrative sequences are often reported to be vivid and multi-thematic experiences [5,6]. A few investigations into visual imagery content during music listening have begun to shed light on this idea of imagery consistency across listeners [7] and potential influencing factors [6]. However, the extent to which listeners exhibit similarities in their own visual imagery across listening situations is still an open question.’

Comment: p.3, l. 64: ‘Indeed, to date, there has been no systematic investigation of visual imagery content during music listening and it is unknown the extent to which it is consistent across listeners.’ – I think this might be a bit too strong. There are a few papers that look at across participant consistency of mental imagery in general, and some in music, and I think discussing this might be useful here. For example: […]. It might be worth discussing some of the references above and to be a bit more specific about the exact gap in the literature you are referring to here, currently the statement reads a bit too broad.

Response: We agree with the Reviewer that the phrasing of this sentence is not entirely accurate. Thus, this has been rephased after taking into consideration the suggested references and to reflect the current evidence: ‘A few investigations into visual imagery content during music listening have begun to shed light on this idea of imagery consistency across listeners [7] and potential influencing factors [6].’

Comment: p. 3. L.68-71: “It is also unclear how certain relevant individual differences (such as musical training, general visual imagery ability, synesthetic tendencies or participation in the visual arts) may be associated with prevalence, vividness and consistency of music-induced visual imagery.” – The main motivation statement here reads like ‘we don’t know this yet, so it’s worth investigating’. Personally, I absolutely agree with the sentiment, however, PLOS ONE is a generalist’s journal, so I think it would be important to strengthen the motivation here a bit. From a big picture perspective, what can we learn from exploring this topic in greater detail, what are the implications for our understanding of human cognition, or the role of culture and society that different results would suggest. I think your study is very exciting with many large implications and I think here is the place to highlight a bit further why investigating this topic is important, and why it is worth for the readers of PLOS ONE to read the rest of your exciting paper. Additionally, the unique selling point and main narrative of the paper is about within vs across consistency, so to streamline the narrative, it might be better focus on “It is also unclear how certain relevant individual differences (such as musical training, general visual imagery ability, synesthetic tendencies or participation in the visual arts), imagery vividness and prevalence may be associated with consistency of music-induced visual imagery. (That is focusing on consistency as the main dependent variable of interest, but this is just a suggestion to make your narrative flow nicer)

Response: We agree that the point of this final paragraph could be more convincing. It has therefore now been rephrased to instead emphasise the more pressing aims of the study: ‘For most listeners, forming a narrative in their mind’s eye is a way to engage with heard music [3,4], and such narrative sequences are often reported to be vivid and multi-thematic experiences [5,6]. A few investigations into visual imagery content during music listening have begun to shed light on consistency across listeners [7] and potential influencing factors [6]. However, the extent to which listeners exhibit similarities in their own visual imagery across listening situations is still an open question.’

Comment: p. 4, l. 75: ‘In a similar vein, while visual imagery has been associated with aesthetic and emotional engagement with music, the evidence of such links remains limited.’ – Maybe cite one or two of the theoretical frameworks or recent review articles that conceptualize this in detail here. For example:

Juslin, P. N. (2013). From everyday emotions to aesthetic emotions: Towards a unified theory of musical emotions. Physics of life reviews, 10(3), 235-266.

Taruffi, L., & Küssner, M. B. (2022). Visual Mental Imagery, Music, and Emotion. In Music and Mental Imagery.

Response: Agreed. The relevant articles have been cited alongside this statement, as follows: ‘Music-induced visual imagery has been previously associated with aesthetic and emotional engagement with music, but the evidence of such links remains limited [24,25].’

Comment: p. 8, l. 176 “…within, across, and between..’ – The conceptual difference between ‘across’ and ‘between’ is not really clear here. As mentioned earlier and detailed later, the formal difference is also not clear, so I suggest dropping the ‘between’ measure entirely.

Response: We would like to thank the Reviewer for their advice regarding the analysis of this dataset. As mentioned in response to other similar comments, the between-subjects measure has been excluded due to the fact that it does not provide a unique contribution to the results over the across-participant measure. The sentence now reads as follows: ‘iii. To ascertain the extent of the consistency of visual imagery within and across individuals during music listening,’

Comment: p. 9, l. 197 “Further, we predicted we would find little difference between synaesthetes and non-synaesthetes in the consistency of their visual imagery experience; this is due to previous results stating that non-synaesthetes are capable of showing consistency in a colour-picking for letters task, performance on which was predicted by scores in a visual imagery task [39].’ – I find the logical flow here not very clear. You predict no differences between two groups, but use a statistical approach that is not capable of supporting the Null-hypothesis, only rejecting it. If you are interested in supporting the Null-hypothesis, you’ll have to use Bayesian statistics, but frequentist/NHST approaches like the ones you use in your analysis lack the ability to support the null hypothesis (i.e., not rejecting the null is not equal supporting the null). Additionally, synaesthetes is a very broad term, maybe be clear about which precise conditions you are interested in (e.g., all those that have co-evoked percepts involving either sound or visuals or both?), otherwise it isn’t clear why you would consider individuals that say, have co-evoked percepts of pain and smell. Furthermore, why would you predict a lack of an effect between synaesthetes vs non synaesthetes solely based on the finding that non-synaesthetes can do a color-picking task consistency that has some degree of correlation with visual imagery task. I feel like I’m missing something, but currently the logical flow here feels a bit shacky.

Response: We agree with the Reviewer that this hypothesis is not very strongly supported and was not accompanied with the appropriate analyses. While we did collect data on the types of synaesthesia that participants experience, this was significantly varied across this subsample. Thus, this hypothesis, along with any related discussions and analyses, have now been completely removed.

Comment: p.10 l. 223. Materials and Stimuli – Most readers will be unfamiliar with the film-music corpus from Eerola & Vuoskoski and will likely assume that these songs are reasonably familiar to at least some participants. The corpus contains published film music, judged to be relatively unfamiliar by a few raters, but some of the pieces will likely sound familiar to participants (e.g., the Vertigo soundtrack is part of that corpus). I recommend providing a bit more information about the pieces you selected, maybe indicate which exact soundtracks the three selected pieces are from.

Response: We agree that there were key details regarding the musical stimuli missing. These have now been included in the Materials and Stimuli subsection, including information regarding the movies that the excerpts are from and their soundtrack numbers, as follows: ‘Three film music stimuli conveying happy, tender, and fearful emotions were selected from Eerola and Vuoskoski’s [39] database (see, https://www.jyu.fi/hytk/fi/laitokset/mutku/en/research/projects2/past-projects/coe/materials/emotion/soundtracks for access to the original stimuli). These excerpts were obtained from the catalogue of extended 1-min film excerpts (see Appendix of [40], or see, https://www.jyu.fi/hytk/fi/laitokset/mutku/en/research/projects2/past-projects/coe/materials/emotion/soundtracks-1min for access to the original stimuli), validated to be unfamiliar to most listeners and to still convey the intended emotions even in their shorter form. In terms of the films that the tracks were taken from, the excerpt conveying happy emotions was taken from The Untouchables soundtrack (track 6, number 071 from Eerola and Vuoskoski’s set of 110 tracks). The tender excerpt is from the Shine soundtrack (track 10, number 042 from set of 110 tracks). Finally, the fearful excerpt is from the Batman Returns soundtrack (track 5, number 011 from set of 110 tracks). In order to ensure uniformity amongst the musical excerpts, as well as to control the overall length of the survey, all excerpts were edited to last a duration of 45 seconds using Audacity (Version 2.3.2.0). These were also edited to finish with a fade-out to avoid an abrupt ending.’

Comment: p.12, l.265: ‘They were advised to pay attention to any visual imagery that they may be experiencing..’’ – In your instructions you very much highlight visual imagery, which makes a lot of sense given your research questions. However, there is a lot of literature pointing towards the effects of eyes open vs close in visual mental imagery, both in general: […]. Were you participants instructed to keep their eyes open, or to close them? Or was it up to them? I think this is worth mentioning, and maybe discussing in a couple of sentences in the discussion as the simultaneous visual input (or absence of it) can have dramatic effects (the effect are so large, they are used in therapeutic contexts) and this has been explored in particular in the context of – and interaction with- music-induced imagination – the present focus on visual mental imagery warrants a short discussion about this.

Response: The participants did not receive any specific instructions regarding whether they should listen to each musical excerpt with eyes open or closed. We acknowledge that this is an important detail to mention in the manuscript. We have explained this issue, citing a few relevant papers including some suggested by the Reviewer, in the Implications, limitations, and future directions subsection of the Discussion, as follows: ‘Some literature has emphasised the positive effects that eye closure may have on visual imagery experienced in response to music (e.g., [36,64,65]), so much so that current imagery-based therapies adopt eye closure as a way to enhance the benefits of rehabilitation [66]. One study that specifically compares the facilitatory effects of eye opening or closure on participants’ reported visual imagery experience found that eye closure led to markedly higher visual imagery vividness as well as content [67]. The current design did not offer specific instructions on whether participants should listen to each musical excerpt with eyes open or closed, thus this action was free to vary across the sample. Given the dramatic influence that the change in instruction can have on visual imagery experience, especially in the context of music listening, such an instruction would be important to include in future research hoping to enhance the experience of visual imagery.’

Comment: p. 15: Results ‘Across-participants was computed by comparing the list of themes from each individual participant with the lists from every other individual in the sample, resulting in values per participant that were then averaged to create a single consistency value per individual.’ – The general approach is solid, but if you have 352 values per participant (prior to averaging), then this means that you must have averaged across the three pieces. Or did you calculate the Jaccard coefficient index separately for each of the pieces? I think calculating the coefficient separately for each of the three pieces and each participant would make the most sense given your data structure, this would allow you to visualise and test the within vs across-distributions thoroughly and provide very insightful results, however, it isn’t clear from the description which approach was taken.

Response: We agree that the process of calculating the across-participant measure could have been explained more clearly. The coefficient was indeed calculated for each of the three pieces separately, and we did not calculate an across-excerpt average at any stage. The distributions that were analysed and plotted in the Results section later included all of the three pieces, meaning that we had used the full 1,059 (353 values for each excerpt) item dataframe. This explanation has now been amended to make this clearer: ‘We assessed two types of consistency for each musical excerpt separately: within- and across-participant consistency. Within-participant consistency was computed by comparing the lists of themes emerging from each participant’s reports across surveys 1 and 2, and across-participant consistency was computed by comparing the list of themes from each individual participant with the lists from every other individual in the sample, resulting in 352 values per participant for each musical excerpt type (leading to three groups of values) that were then averaged to create a single consistency value per individual per excerpt. With regard to across-participant consistency specifically, individuals were always compared against other individuals’ responses that were within the same excerpt type.’

Comment: p. 15. l, 344: ‘Finally, to allow greater comparison of consistency values within and between individuals, a between-participants consistency was calculated by comparing each participant with a random other individual. With regard to across- and between-participant consistency specifically, individuals were always compared against other individuals’ responses that were within the same excerpt type. – This part really confused me. Based on your description, you are effectively producing a low resolution, sampled metric from your across-participant measure. I do not understand in what way this allows greater comparison. If anything, it allows less reliable comparison, because your comparison would be massively influenced by the particular dyads you are sampling. In addition, this approach is effectively not reproducible unless you share the particular random seed used to pair the dyads, and there is no way of checking whether the dyads were handpicked. Of course, I know that they weren’t, but this approach would theoretically allow it, so this is more a point about general rigor and reportability. I really do not see any advantage in this metric over your across metric. Currently, it feels like you are throwing away a great amount of your precious data without gaining anything in return. Long story short, I strongly suggest ditching the ‘between’ metric and focusing on within vs across. This will provide very compelling support for your project. At the end of the detailed comments section, I’ll propose an alternative structure and analytical approach for your results, you can take as much or little of it on board as you like, but I think it is worth considering.

Response: We thank the Reviewer for their thoughts. As has been mentioned in response to previous comments, we have chosen to exclude the between-participants measure from the manuscript entirely and have taken on board much of the Reviewer’s suggestions regarding the analyses, especially those that enable us to take advantage of the full range of our data.

Comment: p. 17-19 ‘Out of all the ratings of the 353 responds collapsed across the three excepts’ – Why do you collapse across the three pieces? You are throwing away a lot of data. Additionally, the large number of Pearson correlations you are conducting afterwards result in a dramatic alpha error accumulation. If you were to correct for this, basically no result would be significant anymore (I do not suggest you do this). Additionally, based on the degrees of freedom, it seems like all your correlations were also calculated on the aggregated (across excerpts) data even though you deliberately chose those songs to be different across key dimension (such as emotional intensity, or might differ in liking), so you are again throwing away a large amount of data (in the order of magnitude factor 3!). In many instances it isn’t clear why you are correlating particular variables (say liking and emotional intensity) when there wasn’t a clear hypothesis for this. Additionally, the multiple correlations with your target variables ‘e.g., imagery vividness’ using variables that show co-linearity, such as emotional intensity and liking, makes interpreting these results effectively impossible, as all your effects may or may not be tapping into the same variance.

Long story short, I recommend deleting all in-text correlation reporting in the 4.1 section, keep the correlation table for interested readers, and then run a quick model (one for each of your literature guided hypothesis, and ignore those for which you have no literature motivation) for your actual inference that accounts for all these dependencies. This will shorten this section, make the results more interpretable and readable, will make full use of your statistical power, and will be cleaner in terms of avoiding violations of independence. At the end of the detailed comments section, I’ll provide additional detail that might be useful.

Response: We thank the Reviewer for their suggestions on the approaches taken in section 4.1 (now section 4.4). In addition to the aggregated sums and proportions of the prevalence and vividness of visual imagery ratings, we have further included a table outlining the sums and proportions of prevalence and vividness across the three excerpts (see Table 5). Further, we have decided to modify the analytical approach of this subsection, as advised by the Reviewer, and run linear mixed models targeting only our hypothesis-driven relationships.

Comment: p. 26-29. ‘How consistent is visual imagery and to what extent does it depend on abstraction’ – As mentioned above, is see a formal issue in analysing and making claims about increases in consistency between Level 2 and Level 3, and suggest deleting these from the results and discussion. […]

Response: We thank the Reviewer for their comment and also helpful example (not included above for conciseness) regarding why the analysis of abstraction is formally inappropriate. We agree and this change has now been implemented in the Results section of the manuscript, with all later related discussion points removed.

Comment: p. 33. Discussion: As mentioned before, I strongly advise removing all level 3 versus level 2 consistency comparisons, instead, you could just focus on discussing the overall findings that hold even in level 3, which is an interesting pattern of shared and idiosyncratic content. Additionally, if you focus on across vs within and remove the between measurement, then the discussion would also require some adjustment but become much more streamlined in the process.

Response: Having adopted much of the Reviewer’s suggested changes for the study’s analytical approach, we agree that rephrasing certain areas of the Discussion is warranted. These changes have been implemented throughout this section.

Comment: p. 35, l. 703 ‘At first glance, our finding show notable differences with a recent investigation by Magulis et al. [6] – Agreed, the overall pattern is the same (higher consistency within person/culture, lower consistency across person/culture, with some degree of overlap remaining). The main difference is not in the results, but the metric, one focusing on continuous similarity on a latent variable (cosine similarity) one focus on more discreet topic overlap (yours). Both have their advantages and disadvantages. Cosine similarity allows continuous weighting of similarity, but doesn’t allow exploring distinct topic prevalence or absence. Your approach allows identifying specific themes, but you lose the ability to weight topic importance. Long story short, Margulis measures similarity, you measure whether (amongst others) two reports cover the same topic, regardless of how intense. As a result, I would recommend being a bit more caution with your statement on p. 36, l..714 ‘It was also a more valuable method with which to approach our dataset’. Your method is great and very useful for your research question, but not necessary inherently ‘more valuable’ for the dataset perse.

Response: We agree with the Reviewer that the phrasing of this paragraph is too strong and overshadows the point that was being made, which was a comparison between our own approach and the approach carried out by Margulis and colleagues in answering very similar research questions. This paragraph has now been rephrased to reflect this more appropriately: ‘Approaches used by other studies to address similar questions have also included frequency-based approaches [7] (i.e., summing occurrences of certain content). Critically, the current research aimed to assess consistency with a focus on the presence or absence of particular themes and topics. By allowing analysis of consistency on the basis of themes and topics, our approach provides a way to estimate consistency levels with a focus on particular aspects of content. However, it would be beneficial to consider the advantages of applying a weighted approach (as Margulis and colleagues have done) when scrutinising visual imagery content in combination with the current study’s methods of assessing discrete overlap. Incorporating both topic overlap as well as weighted similarity might provide insight on the significance of specific types of cross-modal content evidently present (as is seen within our thematic framework) in reports of music-induced visual imagery.’

Comment: p. 39, l. 802 “Musical training did however show a very weak but negative significant relationship with music liking, which could suggest that those with training in music may be mildly uninterested in film music.” – This is a bold conclusion in general, but in particular considering that you’ve only tested three songs and averaged across them. I would recommend just deleting this.

Response: We agree with the Reviewer that this sentence should be excluded from the rationale of this analysis.

Comment: p. 3, l. 805: ‘Our results confirmed our expectation that those who experience a form of synaesthesia are no different in the ways they experience prevalence and vividness of visual imagery across music excerpts and in response to individual listening tracks.’ – I would recommend being more careful with your interpretations on synaesthesia. This is a condition with dramatic differences between individuals, you’ve relied on self-report, without additional information about which senses co-create percepts, and your statistical approach (NHST) by definition cannot support the null-hypothesis, only reject it (only a Bayesian Framework can support the Null). So maybe simply rephrase this along the lines of: […]

Response: We agree with the Reviewer that given that we had not probed deeper into the nature of the synaesthesia experienced by participants, our interpretation of this finding should be rephrased. In fact, in line with our response to an earlier related comment, analyses pertaining to synaesthesia have been removed, due to low uniformity of the type of synaesthesia experienced across our sample.

Comment: p. 40, l ‘820’: “Thus, although the directionality is unclear, increased participation within various art modalities (in this case, visual) may increase imagery engagement with music.” – Also a very exciting finding! However, maybe rephrase is slightly, as it is not only a question of directionality (A -> B vs B->A), but it could also be a question an entirely different sets of variables being responsible that affect both (e.g., C -> A & C ->B). Maybe a phrasing like “Thus, although the general chain of causality is unclear, increased participation within various art modalities (in this case, visual) may increase imagery engagement with music. “

Response: We acknowledge the rationale behind the Reviewer’s comment here and agree. We have rephrased this sentence of the Discussion accordingly: ‘Thus, although the general chain of causality is unclear, increased participation within various art modalities (in this case, visual) may increase imagery engagement with music.’

Comment: p. 40, l. 831: ‘The increase in correlation between across-participant consistency and the majority of behavioural ratings and individual differences could point to shared cultural and societal influences, despite our sample being sourced from a wide range of backgrounds.” – I’m not sure I get this, but it might just be me. Which correlations are increasing? Wouldn’t that require an interaction term? And why does the point towards a shared cultural and societal influence? Maybe you can rephrase and elaborate a bit here.

Response: We agree with the Reviewer that this sentence could have been phrased more clearly. However, in the process of updating certain passages of the Discussion due to the extensive changes in the analysis approach, this sentence has now been deleted as it was no longer relevant.

Comment: p. 40 l.83: “The similar levels between within-participants consistency and ratings that we found may indicate…’’ – Same thing, I struggle parsing the sentence. Is this about narrow distributions of all rating scales and the within-participant consistency? I’m not sure, maybe rephrase and elaborate.

Response: We agree with the reviewer that this sentence could have been phrased more clearly. Similarly to the previous comment, this sentence was modified completely to account for the change in results and we believe now expresses our point more clearly, as follows: ‘The generally positive relationships found between (across- and within-participant) consistency and prevalence and vividness of visual imagery clearly indicate that the qualitative nature of listeners’ visual imagery influences how consistent they are within themselves and with others.’

Online Supplement

Comment: These are very useful, thanks a lot! I think the survey1_ratings and survey2_ratings files still contain participants’ prolific IDs in the anon-id column, in addition to the anonymized ‘ID’ column, you might want to remove those.

Response: We thank the Reviewer for highlighting that the data still contained participants’ Prolific IDs. These have now been removed. The numeric anonymous ID index column was retained though, to clarify which participants in survey 1 had returned to complete survey 2.

Suggested Statistical Approach and Results Structure

We would like to thank the Reviewer for their detailed and thorough suggestions for the analytic approach to our manuscript, especially the lines of R code provided, which were extremely useful. We have chosen to implement the approaches suggested in this section as we agree that running these analyses is more effective in utilising the full range of our dataset, and more efficiently helps us to answer our main research questions. The Results section has also been restructured, following the advice of the Reviewer, along with the relevant introductory and discussion points throughout the manuscript to improve the coherence of the overall narrative given the new changes. 

Comments from Reviewer 2 – Dr Julian Cespedes-Guevara

Comment: (Page 4 Lines 75-82). First, in the introduction, I suggest mentioning Day and Thompson's (2019) work, about how that in many cases, emotional responses to music may occur earlier than visual imagery. This finding suggests that visual imagery is not the necessarily the driving factor behind emotion, but the other way around. The paper does mention this paper in the discussion, but I think it is worth including in the review about the link between imagery and emotion.

Response: We agree with the Reviewer that making this addition will strengthen our point. Day and Thompson's (2019) study has now been mentioned in the Introduction, albeit only briefly: this is in order to not compromise the flow of the Introduction (as follows: ‘Thus, associations between music-induced visual imagery and emotion induction are evident, although the directionality of this relationship is still a widely debated topic supported by contrasting evidence [29,30].’), and because we believe that describing the study in detail is better suited as a later discussion point.

Comment: (Pages 7 and 8: “The current research”). I found it difficult to understand the aims of the study. I had to read them several times to understand the difference between aims 1 and 4. Additionally, it is not immediately clear what the author(s) mean by “formulating a framework” in aim 2. I suggest rewriting these paragraphs so that their meaning is more readily apparent.

Response: We would like to thank the Reviewer for highlighting that our aims (specifically aims 1, 2, and 4) could be expressed more clearly. The goal of Aim 1 (now Aim 4) was to address the relationship between visual imagery prevalence and vividness ratings and ratings of music liking and emotional intensity. Whereas the goal of Aim 4 (now Aim 3) was to investigate potential behavioural and individual factors that may be driving how consistent music-induced visual imagery is across time points and individuals. Additionally, the purpose of Aim 2 (now Aim 1) was to run a thematic analysis on our open-text descriptions of visual imagery content, and to organise the patterns found into a multi-layered framework of music-induced visual imagery content. These aims have now been rephrased (and reordered), as follows:

i. ‘To run a thematic analysis to create a hierarchical framework highlighting the prevalent codes and overarching themes found in descriptions of music-induced visual imagery content,

ii. To ascertain the extent of the consistency of visual imagery within and across individuals during music listening,

iii. To examine potential behavioural factors and individual differences (general visual imagery ability, musical training, and participation in the visual arts) that may be driving the rates of within- and across-participant consistency levels,

iv. To test the extent to which the prevalence and vividness of visual imagery is associated with the emotional intensity and aesthetic appeal of music, as well as an array of individual differences (general visual imagery ability, musical training, and participation in the visual arts).’

Comment: (Pages 8 and 9, Lines 182-204). The author(s) state that they expect to “be able to replicate previous reports of a relationship between visual imagery and both emotion induction and aesthetic appeal” (page 8, lines 182-183). I think this hypothesis needs to be outlined in more detail: what type of relationship did they expect to find? I also suggest labelling the hypotheses stated in this section with numbers and use them throughout the paper to identify them, particularly in the results section.

Response: This hypothesis has now been extended to provide a clearer rationale of what is expected: ‘Further, we predicted that we would be able to replicate previous reports of a relationship between visual imagery and both emotion induction [30,35–37] and aesthetic appeal [26]. We specifically predicted that the prevalence and vividness of visual imagery would both be predicted by ratings of emotional intensity (H3) and music liking (H4), in line with previous reports of positive links between these phenomena [23,25,26,35].’

Further, as suggested by the Reviewer, labels for each hypothesis (i.e., H1, H2, etc.) have been implemented where necessary throughout the manuscript.

Comment: (Page 10, Lines 223-227). Please provide musical details about the musical excerpts that were used: mode, tempo, instrumentation, style, and the movies from where they were taken from.

Response: Additional details pertaining to the musical excerpts have now been added, specifically regarding the movies they are from and their soundtrack number. Unfortunately, upon further search, more specific information regarding the mode, tempo, instrumentation, and style of each piece could not be obtained. As a subjective description of these details did not seem appropriate, no further details have been added.

Thus, this has been amended as follows: ‘Three film music stimuli conveying happy, tender, and fearful emotions were selected from Eerola and Vuoskoski’s [39] database (see, https://www.jyu.fi/hytk/fi/laitokset/mutku/en/research/projects2/past-projects/coe/materials/emotion/soundtracks for access to the original stimuli). These excerpts were obtained from the catalogue of extended 1-min film excerpts (see Appendix of [40], or see, https://www.jyu.fi/hytk/fi/laitokset/mutku/en/research/projects2/past-projects/coe/materials/emotion/soundtracks-1min for access to the original stimuli), validated to be unfamiliar to most listeners and to still convey the intended emotions even in their shorter form. In terms of the films that the tracks were taken from, the excerpt conveying happy emotions was taken from The Untouchables soundtrack (track 6, number 071 from Eerola and Vuoskoski’s set of 110 tracks). The tender excerpt is from the Shine soundtrack (track 10, number 042 from set of 110 tracks). Finally, the fearful excerpt is from the Batman Returns soundtrack (track 5, number 011 from set of 110 tracks). In order to ensure uniformity amongst the musical excerpts, as well as to control the overall length of the survey, all excerpts were edited to last a duration of 45 seconds using Audacity (Version 2.3.2.0). These were also edited to finish with a fade-out to avoid an abrupt ending.’

Comment: (Page 11, Lines 249-250). Did the authors offer any explanation of what synesthesia consists of to the participants?

Response: The participants were provided with a brief general explanation on what synaesthesia consists of, illustrated as follows: “Synaesthesia can be defined as an experience relating to one sense or part of the body by the stimulation of another sense or part of the body”. However, analyses and discussion points pertaining to synaesthesia have now been excluded from the manuscript, given that the data collected for this was not extensive enough to draw convincing conclusions.

Comment: (Page 15, lines 340-347). I suggest expanding the rationale for using both an across-participants and a between-participants measures of consistency. I did not find this explanation offered by the authors to be sufficient to grasp the usefulness of using both measures: “to allow greater comparison of consistency values within and between individuals, a between-participants consistency was calculated by comparing each participant with a random other individual” (lines 345-347).

Response: We would like to thank the Reviewer for highlighting this lack of clarity in the distinction between our across- and between-participant measures. However, it has become clear that the addition of the between-participant measure did not provide a unique contribution to the narrative of the results, over and above the across-participant measure. Thus, the decision was made to exclude this measure from the manuscript entirely.

Comment: (Pages 23-24, lines 841-847; and Pages 26-27, lines 507-510). From my point of view, these sections of the paper are the ones that require revision more urgently, because the author(s) describe the content of the participants’ reported experiences in contradicting terms. For instance, on page 24, lines 485-487, they write: “Whilst not denoting an exact visual image, its presence within reports was considerably prevalent and constituted a prominent attribute of visual imagery experience”. This is confusing, because if the participants did not explicitly report seeing any visual images in their mind, then those experiences do not correspond to “visual experience(s)”. They are indeed, subjective mental experiences while listening to music, but not visual ones. I checked the authors’ original data on the Open Science website, and found that many of the quotes from the participants in the “Feelings & Atmosphere” and “Music” categories correspond to abstract terms such as “Hope and courage", "Sadness", "Haugtiness", "Love, care", "Tranquility peace well-being love", "Fear", "National anthem", "slow music", "A piano like music for study or sleep", "Elevator music", etc. There is nothing in these terms that suggest any visual dimension to them, and therefore, classifying them as “visual imagery” is misleading. From my point of view, these reports correspond to semantic associations about the culturally-shared uses of functions of music.

One important implication of not calling these experiences “visual imagery”, is that the reported percentages of “visual imagery prevalence” should be slightly adjusted.

Response: We appreciate the Reviewer’s comments regarding the consistency and accuracy of how visual imagery is referred to in this section. We did aim to avoid referring to our Associations and References higher order themes as containing pure ‘visual imagery’ and hoped that it would be apparent that these constituted a mixture of ways that participants chose to describe their experience to the music or additional details that participants used to set the scene of their, otherwise, visual experience. This, however, has been clarified in more detail at the beginning of this subsection (as follows, ‘As will be observed below, participant descriptions did not always include experiences that were strictly visual in content and often included descriptions related to affect, musical features, and other forms of mental imagery. This fact has been taken into consideration with regard to analyses and conclusions drawn.’) and the more contradictory statements highlighted by the Reviewer have been either removed or rephrased.

Further, it is not clear whether the percentages that the Reviewer mentions refers to those reported in this subsection, or the percentages of visual imagery prevalence and vividness reported in section 4.1 (now section 4.4). However, assuming that this is in reference to the latter, this percentage is based on the continuous self-report ratings provided by participants on a scale from 1 to 7, and so it is not entirely clear what type of adjustment would be necessary.

Finally, we would like to highlight the additional consistency analyses that we had run using just the Storytelling codes, in order to be certain (in relation to the concerns raised by the Reviewer) that the patterns of differences shown by the models would still stand if the higher order theme containing only visual imagery codes was considered (which they did):

‘Fig D in S2 File presents within- and across-participant distributions using only Storytelling codes, since these, unlike the Associations and References codes, constitute codes that characterise purely visual imagery experiences. We once more show the two distributions to possess distinct shapes and peaks as previously described for all codes. Again, the model comparing Jaccard coefficient values for L2 codes between the within- (Mean = 33.6%, Median = 28.6%) or across-participant (Mean = 21.5%, Median = 21.2%) distributions for Storytelling codes confirmed differences between the two groups (ß = 0.12, SE = 0.01, t = 12.06, p < 0.001). Finally, the model predicting consistency values at the higher granularity L3 codes also demonstrated overall differences between the within- (Mean = 26.1%, Median = 16.6%) and across-participants (Mean = 11.9%, Median = 11.3%) groups (ß = 0.14, SE = 0.01, t = 14.35, p < 0.001).’

Comment: (Pages 25, lines 507-510). Please provide examples of the terms that participants used in this category.

Response: Examples of terms participants used in this subtheme have now been provided immediately after its explanation, as follows: ‘e.g., “As if I were sitting with tea in my hand next to a famous composer like Fryderyk Chopin”, “When the piano started playing I imagined a dark haired male pianist on a black piano…”)’.

Comment: (Page 35, lines 680-688). I suggest that the authors should emphasize that the fact that a significant proportion of the participants’ reports did not have a “visual” quality to them, implies that by narrowing our focus on researching “visual” experiences we may be missing out an important aspect of listeners’ experiences. They could also mention that this has been found in previous research. For instance, Cespedes-Guevara and Dibben (2022) found that this is a common occurrence even when participants are provided with written narratives about the music meaning. Furthermore, those authors suggest that the power of music to evoke semantic associations may be a crucial factor behind listeners’ mind wandering and emotional experiences while listening to music. That interpretation may also help explain the relatively high levels of consistency found on level 2 in the present investigation.

Response: We would like to thank the Reviewer for this suggested addition. We agree that this is a relevant point to make, and it has now been briefly explained in the Discussion, as follows: ‘These findings highlight the considerable prevalence of semantic associations found within listeners’ descriptions in response to the music, even when explicitly instructed to focus their attentions on visual imagery. Such patterns indicate that narrowing one’s research aims on just the visual components of a listeners’ experience could lead to overestimations of its occurrence as well as overshadows its potential unique links to other semantic associations formed in response to the music; one recent study by Cespedes-Guevara and Dibben [50] showed that a considerable portion of listeners’ reports of what went through their minds while listening comprised an array of semantic, personal, as well as visual experiences.’

Comment: (Page 42, lines 861-868). I agree with the authors that using movie soundtracks as stimuli was an important limitation of this study. I suggest also mentioning the fact that they only used 3 stimuli was also a limitation.

Response: We agree with the Reviewer that using only three music stimuli is a key limitation to note. Although this was briefly mentioned, this point has now been emphasised further in the Implications, limitations, and future directions subsection of the Discussion: ‘Our approach may however also suffer from a couple of key limitations; namely that our visual imagery framework was developed in response to only three musical excerpts and were taken from film soundtracks [39]. Our choice of the film genre was partly to ensure that participants were free to provide rich accounts of their visual imagery in response to a programmatic selection of tracks, as well as giving us considerable power with which to analyse consistency, especially when selecting such a low number of listening stimuli. Such decisions mean that our stimuli fall short of being considered entirely comprehensive. It is further a possibility that the predominance of storytelling and media references was an artefact of musical cues present in our chosen tracks that were associated with the development and changes found in film scenes.’

Comments from Reviewer 3 – Prof Marta Olivetti Belardinelli

Comment: Lines 210-218: It would be useful to present the different types of participants in a table, also distinguishing Males and Females. This distinction should also be investigated in the results, all the paper along.

Response: We thank the Reviewer for their suggestion. We have now included a table outlining the sums and proportions of the participants’ countries of residence, rather than describing this in the text (please see Table 1). We would further like to thank the Reviewer for their comment regarding the additional investigation of differences between males and females throughout this manuscript. Unfortunately, though, it remains unclear to us why distinguishing by gender is a theoretically necessary and relevant approach for this study. Given it would greatly lengthen an already very long manuscript and given that we have no evidence that this particular question is a pressing issue for pushing forward visual imagery research, we have opted to not address it in this manuscript. Thank you for your understanding.

Comment: Lines 228-9: the number of the reference should stay after the test name (i.e. after Inventory).

Response: We agree with the Reviewer regarding this change, which has been implemented, as follows: ‘…from Pekala’s Phenomenology of Consciousness Inventory [41].’

Comment: Line 267: a definition of the term “prevalence” is needed.

Response: Agreed. A brief definition of the term “prevalence” (defined as ‘the amount experienced’), as well as the term “vividness” (defined as ‘the clarity with which it was experienced’), has been added.

Comment: Lines 272-3: the order “associated with visual arts,” and synesthesia is the inverse of that indicated in the previous paragraph at lines 250.1.

Response: We thank the Reviewer for highlighting this. However, as a result of the changes made throughout, any analyses and discussion points pertaining to synaesthesia have now been excluded due to low uniformity in the types of synaesthesia present in our dataset.

Comment: Lines 311-3: Please, give more details about the way this hierarchical categorization was performed.

Response: We were a little unclear on what aspects of the hierarchical categorisation the Reviewer thinks require more detail. However, we have included extra information regarding how this process was carried out by the coders, as follows: ‘Commonalities between specific terms were identified from Level 3 (L3) codes, which were sorted and categorised into distinct groups comprising Level 2 (L2) codes. The suitability of the subthemes in this level was reviewed by the two coders, including whether to collapse redundant subthemes or to divide ones that were too diverse. The L2 codes were finally combined to form higher-order Level 1 (L1) themes. The final structure was discussed by the two coders, confirming that it formed an effective hierarchy that offered a parsimonious overview of the content of the free-form descriptions.’

Comment: Line 410: Insert (VA) dopo visual arts and before the coma.

Response: We thank the Reviewer for their suggestion. We agree and have now implemented this change, as follows: ‘In response to the question regarding experience with activities in the visual arts (VA), 20.1% (n = 71) reported that they participate in activities associated with the visual arts, which included activities such as painting, photography, and graphic design. With regard to the averaged excerpt ratings between those who do and do not (NVA) participate in the visual arts, independent samples t-tests showed that those who participate in the visual arts reported significantly more visual imagery prevalence (Mean-VA = 4.63, Mean-NVA = 3.99, t(119.4) = 3.78, p < 0.001) and more vividness (Mean-VA = 4.30, Mean-NVA = 3.74, t(124.1) = 3.37, p < 0.001) than those who do not.’

Comment: Lines 556-7: I think that the indication as first result (as already before in the aims, and afterward line 676) that “visual imagery experience was very prevalent” goes without saying, since according to my comprehension, visual imagery was expressly indicated as the subject’s task. Or were the commitments differently expressed? More interesting is the following indication…

Lines 686-8:…that visual imagery was accompanied by cross-modal aspects, aesthetic evaluation, and emotions, positively related to images vividness.

Response: We would like to thank the Reviewer for their comments regarding the different sections expressing that music-induced visual imagery was highly prevalent. Participants were indeed instructed prior to listening to each excerpt that they should pay attention to any visual imagery that they may be experiencing throughout listening in preparation for the questions that follow. While on one hand it might go without saying that visual imagery was in fact quite prevalent in our results, on the other, and as has been found in our thematic framework as well as in several past studies [1–4], the experience of music-induced visual imagery is often experienced in combination with a number of non-visual thoughts, and participants often form several semantic, emotional, and personal associations with the music. Thus, we aimed to take this into consideration when phrasing our conclusions regarding the results throughout the Discussion.

References

1. Dahl S, Stella A, Bjørner T. Tell me what you see: An exploratory investigation of visual mental imagery evoked by music. Musicae Scientiae. 2022; 10298649221124862. doi:10.1177/10298649221124862

2. Küssner MB, Eerola T. The content and functions of vivid and soothing visual imagery during music listening: Findings from a survey study. Psychomusicology: Music, Mind, and Brain. 2019;29: 90–99. doi:10.1037/pmu0000238

3. Cespedes-Guevara J, Dibben N. The Role of Embodied Simulation and Visual Imagery in Emotional Contagion with Music. Music & Science. 2022;5: 20592043221093836. doi:10.1177/20592043221093836

4. Margulis EH. An Exploratory Study of Narrative Experiences of Music. MUSIC PERCEPT. 2017;35: 235–248. doi:10.1525/mp.2017.35.2.235

Comment: Lines 856-860: I think that the novelty of the methodological approach should be stressed as an important result of the research besides the…:

Lines 863- 964:… very well indicated key limitations.

Response: We thank the Reviewer for their comment. We do touch upon the novelty of our methodological approach at the start of the Implications, limitations, and future directions subsection. However, we have now emphasised this point further in a few more words: ‘With the current research, we have presented a novel methodological approach to probing the content of music-induced visual imagery, a method that we hope will be adopted by future studies seeking to develop the knowledge on the topicality of visual imagery content.’

Comment: Lines 906-918: Given all that is said above the Conclusion should be rewritten in order to put in evidence the bring-home results.

Response: We agree with the Reviewer that the changes proposed thus far merit an amended Conclusion paragraph. The Conclusion has been rephrased to reflect any changes made to the results, as follows: ‘We have presented a detailed investigation into the visual imagery content that listeners experience in response to music. We show that visual imagery is a highly prevalent aspect of individuals’ listening experience, with storytelling being particularly prominent. We also demonstrate the idiosyncrasies of listeners’ content consistency by showing that they were, on average, relatively consistent with themselves across timepoints, in contrast to when compared with other listeners.

The ease with which music appears to elicit visual imagery offers further support for the connection between music and language processing with regard to listeners’ inclination to derive meaning from the music. We anticipate that our research will set a precedence for further studies to develop and hone our understanding of the inherent visual imagery qualities experienced during music listening.’

We look forward to hearing from you soon regarding our submission and we are happy to respond to any further questions and comments that you or the reviewers may have.

With best regards, and on behalf of the co-authors,

Sarah Hashim

Attachment

Submitted filename: Response to Reviewers.docx

Decision Letter 1

Ioanna Markostamou

24 Aug 2023

PONE-D-23-03570R1Music listening evokes story-like visual imagery with both idiosyncratic and shared contentPLOS ONE

Dear Ms Hashim,

Thank you very much for submitting your revised manuscript to PLOS ONE.

I have now received the evaluations of all three of the original reviewers. Please find their comments at the end of this letter. As you will see, the reviewers appreciate very much the changes you have made and point out that their comments and suggestions were well taken into consideration in order to address the issues raised. I also read your revised paper myself and believe that it has now improved substantially.

However, all reviewers have some additional (minor) suggestions for improvement, and I would like to invite you to consider their suggestions in a minor revision of the manuscript.

Please submit your revised manuscript by Oct 08 2023 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

Please include the following items when submitting your revised manuscript:

  • A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'.

  • A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.

  • An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols. Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols.

We look forward to receiving your revised manuscript.

Kind regards,

Ioanna Markostamou, Ph.D.

Academic Editor

PLOS ONE

Journal Requirements:

Please review your reference list to ensure that it is complete and correct. If you have cited papers that have been retracted, please include the rationale for doing so in the manuscript text, or remove these references and replace them with relevant current references. Any changes to the reference list should be mentioned in the rebuttal letter that accompanies your revised manuscript. If you need to cite a retracted article, indicate the article’s retracted status in the References list and also include a citation and full reference for the retraction notice.

[Note: HTML markup is below. Please do not edit.]

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. If the authors have adequately addressed your comments raised in a previous round of review and you feel that this manuscript is now acceptable for publication, you may indicate that here to bypass the “Comments to the Author” section, enter your conflict of interest statement in the “Confidential to Editor” section, and submit your "Accept" recommendation.

Reviewer #1: All comments have been addressed

Reviewer #2: All comments have been addressed

Reviewer #3: (No Response)

**********

2. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: Yes

**********

3. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: (No Response)

**********

4. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: Yes

**********

5. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: Yes

**********

6. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)

Reviewer #1: Dear Dr. Markostamou,

Dear Authors,

I had the great pleasure to review the revised version of the article ‘Music evokes story-like visual imagery with both idiosyncratic and shared content’ submitted for publication to PLOS ONE (PONE-D-23-03570R1; review request received: 01.08.2023, review accepted: 01.08.2023, review submitted: 09.08.2023).

The authors clearly took great care in thoroughly addressing all comments, and I believe the revised version to be an impressive improvement over the last iteration. In particular, all prior comments related to the methodological approach were considered in great detail, including a new statistical approach.

The manuscript is a great contribution to the field and well suited to the audience of PLOS-One.

One final minor suggestion is to include the topic distributions (currently in-text in 4.1.1.-4.13) also in brackets behind the respective topics in Table 2 during the author proof stage. This would provide readers with a very comprehensive understanding of the topic structure when consulting Table 2. But this suggestion is entirely personal preference, and up to the authors to decide.

In summary, I can strongly recommend the manuscript for publication and have no further comments.

Signed,

Steffen Herff

Reviewer #2: I would like to thank the authors for taking into consideration the comments and suggestions I made in the first review to edit the manuscript. I think this new version of the paper represents an improvement to the first one: the paper has now clearer aims, hypotheses and data analyses.

I only have two small suggestions that I think could slightly enhance the readability of the paper:

1) I suggest signalling the hypotheses when reporting each analysis in the results section. For instance, on pages 25, lines 521 onward, the authors could mention that "In order to test hypothesis 2, we asked whether there were differences in consistency when comparing and individual with themselves (...)" etc. Similarly, in the first two paragraph of page 31, they could mention that the models confirm hypotheses 3, 4, 5 and 6, etc.

2) On Page 30, Table 5, I suggest including the aggregated proportions in the table.

These are only suggestions; I leave it to the authors to decide whether incorporating them or not. There’s no need for me to review the paper once more.

Congratulations on your excellent work, I truly think this paper makes an interesting contribution to our understanding of the phenomenon of visual imagery evoked by music.

Reviewer #3: Comments from Reviewer 3 – Prof Marta Olivetti Belardinelli

Comment: Lines 210-218: It would be useful to present the different types of participants in a table, also distinguishing Males and Females. This distinction should also be investigated in the results, all the paper along. Response: We thank the Reviewer for their suggestion. We have now included a table outlining the sums and proportions of the participants’ countries of residence, rather than describing this in the text (please see Table 1). We would further like to thank the Reviewer for their comment regarding the additional investigation of differences between males and females throughout this manuscript. Unfortunately, though, it remains unclear to us why distinguishing by gender is a theoretically necessary and relevant approach for this study. Given it would greatly lengthen an already very long manuscript and given that we have no evidence that this particular question is a pressing issue for pushing forward visual imagery research we have opted to not address it in this manuscript. Thank you for your understanding.

OK for table 1. As regards the gender differences in music and language processing (both implied in the visual imagery task of this research) there is plenty of research assessing it. See for example some of the first studies:

• Koelsch S, Maess B, Grossmann T, Friederici AD. Electric brain responses reveal gender differences in music processing. Neuroreport. 2003 Apr 15;14(5):709-13. doi: 10.1097/00001756-200304150-00010. PMID: 12692468.

• Koelsch S, Grossmann T, Gunter TC, Hahne A, Schröger E, Friederici AD. Children processing music: electric brain responses reveal musical competence and gender differences. J Cogn Neurosci. 2003 Jul 1;15(5):683-93. doi: 10.1162/089892903322307401. PMID: 12965042.

• Koelsch S, Gunter TC, Wittfoth M, Sammler D. Interaction between syntax processing in language and in music: an ERP Study. J Cogn Neurosci. 2005 Oct;17(10):1565-77. doi: 10.1162/089892905774597290. PMID: 16269097.

• Koelsch S. Music-syntactic processing and auditory memory: similarities and differences between ERAN and MMN. Psychophysiology. 2009 Jan;46(1):179-90. doi: 10.1111/j.1469-8986.2008.00752.x. Epub 2008 Nov 21. PMID: 19055508.

• Carrus E, Pearce MT, Bhattacharya J. Melodic pitch expectation interacts with neural responses to syntactic but not semantic violations. Cortex. 2013 Sep;49(8):2186-200. doi: 10.1016/j.cortex.2012.08.024. Epub 2012 Sep 20. PMID: 23141867.

According to my meaning, it is therefore important to cite the topic as a limitation of the study tied to the manuscript length, or as a feature research goal, as you prefer.

Comment: Lines 272-3: the order “associated with visual arts,” and synesthesia is the inverse of that indicated in the previous paragraph at lines 250.1. Response: We thank the Reviewer for highlighting this. However, as a result of the changes made throughout, any analyses and discussion points pertaining to synaesthesia have now been excluded due to low uniformity in the types of synaesthesia present in our dataset. OK

Comment: Lines 311-3: Please, give more details about the way this hierarchical categorization was performed. Response: We were a little unclear on what aspects of the hierarchical categorisation the Reviewer thinks require more detail. However, we have included extra information regarding how this process was carried out by the coders, as follows: ‘Commonalities between specific terms were identified from Level 3 (L3) codes, which were sorted and categorised into distinct groups comprising Level 2 (L2) codes. The suitability of the subthemes in this level was reviewed by the two coders, including whether to collapse redundant subthemes or to divide ones that were too diverse. The L2 codes were finally combined to form higher-order Level 1 (L1) themes. The final structure was discussed by the two coders, confirming that it formed an effective hierarchy that offered a parsimonious overview of the content of the free-form descriptions.’ 27

Comment: Line 410: Insert (VA) dopo visual arts and before the coma. Response: We thank the Reviewer for their suggestion. We agree and have now implemented this change, as follows: ‘In response to the question regarding experience with activities in the visual arts (VA), 20.1% (n = 71) reported that they participate in activities associated with the visual arts, which included activities such as painting, photography, and graphic design. With regard to the averaged excerpt ratings between those who do and do not (NVA) participate in the visual arts, independent samples t-tests showed that those who participate in the visual arts reported significantly more visual imagery prevalence (Mean-VA = 4.63, Mean-NVA = 3.99, t(119.4) = 3.78, p < 0.001) and more vividness (Mean-VA = 4.30, Mean-NVA = 3.74, t(124.1) = 3.37, p < 0.001) than those who do not.’ Good precision.

Comment: Lines 556-7: I think that the indication as first result (as already before in the aims, and afterward line 676) that “visual imagery experience was very prevalent” goes without saying, since according to my comprehension, visual imagery was expressly indicated as the subject’s task. Or were the commitments differently expressed? More interesting is the following indication… Lines 686-8:…that visual imagery was accompanied by cross-modal aspects, aesthetic evaluation, and emotions, positively related to images vividness. Response: We would like to thank the Reviewer for their comments regarding the different sections expressing that music-induced visual imagery was highly prevalent. Participants were indeed instructed prior to listening to each excerpt that they should pay attention to any visual imagery that they may be experiencing throughout listening in preparation for the questions that follow. While on one hand it might go without saying that visual imagery was in fact quite prevalent in our results, on the other, and as has been found in our thematic framework as well as in several past studies [1–4], the experience of music-induced visual imagery is often 28 experienced in combination with a number of non-visual thoughts, and participants often form several semantic, emotional, and personal associations with the music. Thus, we aimed to take this into consideration when phrasing our conclusions regarding the results throughout the Discussion. References 1. Dahl S, Stella A, Bjørner T. Tell me what you see: An exploratory investigation of visual mental imagery evoked by music. Musicae Scientiae. 2022; 10298649221124862. doi:10.1177/10298649221124862 2. Küssner MB, Eerola T. The content and functions of vivid and soothing visual imagery during music listening: Findings from a survey study. Psychomusicology: Music, Mind, and Brain. 2019;29: 90–99. doi:10.1037/pmu0000238 3. Cespedes-Guevara J, Dibben N. The Role of Embodied Simulation and Visual Imagery in Emotional Contagion with Music. Music & Science. 2022;5: 20592043221093836. doi:10.1177/20592043221093836 4. Margulis EH. An Exploratory Study of Narrative Experiences of Music. MUSIC PERCEPT. 2017;35: 235–248. doi:10.1525/mp.2017.35.2.235 It is exactly what I asked to stress.

Comment: Lines 856-860: I think that the novelty of the methodological approach should be stressed as an important result of the research besides the…: Lines 863- 964:… very well indicated key limitations. Response: We thank the Reviewer for their comment. We do touch upon the novelty of our methodological approach at the start of the Implications, limitations, and future directions subsection. However, we have now emphasised this point further in a few more words: ‘With the current research, we have presented a novel methodological approach to probing the content of music-induced visual imagery, a method that we hope will be adopted by future studies seeking to develop the knowledge on the topicality of visual imagery content.’ Great job!

**********

7. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #1: Yes: Steffen A. Herff

Reviewer #2: Yes: Julian Cespedes-Guevara

Reviewer #3: Yes: Marta Olivetti Belardinelli

**********

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at figures@plos.org. Please note that Supporting Information files do not need this step.

PLoS One. 2023 Oct 26;18(10):e0293412. doi: 10.1371/journal.pone.0293412.r004

Author response to Decision Letter 1


18 Sep 2023

Dear Editor,

Thank you for the opportunity to submit a revised version of our manuscript ‘Music listening evokes story-like visual imagery with both idiosyncratic and shared content’ to PLOS ONE. We are grateful to the reviewers for their constructive comments and suggested amendments to our paper. All changes are made on a revised version of the manuscript with tracked changes implemented (entitled ‘Revised Manuscript with Tracked Changes’) as well as in an additional revised unmarked version (entitled ‘Manuscript’), as requested.

Please find below the reviewers’ comments and our responses following each one:

Comments from Reviewer 1 – Dr Steffen A. Herff

Comment: One final minor suggestion is to include the topic distributions (currently in-text in 4.1.1.-4.13) also in brackets behind the respective topics in Table 2 during the author proof stage. This would provide readers with a very comprehensive understanding of the topic structure when consulting Table 2. But this suggestion is entirely personal preference, and up to the authors to decide.

Response: We thank the reviewer for their suggestion. The L1 and L2 topic distributions, originally only presented in the main text, have now also been included in Table 2 next to their corresponding theme label.

Comments from Reviewer 2 – Dr Julian Cespedes-Guevara

I only have two small suggestions that I think could slightly enhance the readability of the paper:

Comment: 1) I suggest signalling the hypotheses when reporting each analysis in the results section. For instance, on pages 25, lines 521 onward, the authors could mention that "In order to test hypothesis 2, we asked whether there were differences in consistency when comparing and individual with themselves (...)" etc. Similarly, in the first two paragraph of page 31, they could mention that the models confirm hypotheses 3, 4, 5 and 6, etc.

Response: We agree with the reviewer that signalling the individual hypotheses across the Results section would provide more clarity. This has been implemented throughout this section as suggested.

Comment: 2) On Page 30, Table 5, I suggest including the aggregated proportions in the table. These are only suggestions; I leave it to the authors to decide whether incorporating them or not. There’s no need for me to review the paper once more.

Response: We agree with this change and have included the aggregated values within the table alongside the individual musical excerpt values (see Table 5). Accordingly, the paragraph preceding this table that originally described the aggregated proportions has been rephrased to describe the table as a whole instead, as follows: ‘Table 5 presents a summary of visual imagery prevalence and vividness ratings divided and aggregated by musical excerpt. These values demonstrate high proportions of visual imagery prevalence and vividness across all musical excerpts, that persist even when one considers the individual excerpt types. Higher prevalence and vividness levels of visual imagery can be seen in response to the Fearful excerpt than the Happy and Tender excerpts, which both exhibit almost equal proportions.’.

Comments from Reviewer 3 – Prof Marta Olivetti Belardinelli

Comment: As regards the gender differences in music and language processing (both implied in the visual imagery task of this research) there is plenty of research assessing it. See for example some of the first studies:

• Koelsch S, Maess B, Grossmann T, Friederici AD. Electric brain responses reveal gender differences in music processing. Neuroreport. 2003 Apr 15;14(5):709-13. doi: 10.1097/00001756-200304150-00010. PMID: 12692468.

• Koelsch S, Grossmann T, Gunter TC, Hahne A, Schröger E, Friederici AD. Children processing music: electric brain responses reveal musical competence and gender differences. J Cogn Neurosci. 2003 Jul 1;15(5):683-93. doi: 10.1162/089892903322307401. PMID: 12965042.

• Koelsch S, Gunter TC, Wittfoth M, Sammler D. Interaction between syntax processing in language and in music: an ERP Study. J Cogn Neurosci. 2005 Oct;17(10):1565-77. doi: 10.1162/089892905774597290. PMID: 16269097.

• Koelsch S. Music-syntactic processing and auditory memory: similarities and differences between ERAN and MMN. Psychophysiology. 2009 Jan;46(1):179-90. doi: 10.1111/j.1469-8986.2008.00752.x. Epub 2008 Nov 21. PMID: 19055508.

• Carrus E, Pearce MT, Bhattacharya J. Melodic pitch expectation interacts with neural responses to syntactic but not semantic violations. Cortex. 2013 Sep;49(8):2186-200. doi: 10.1016/j.cortex.2012.08.024. Epub 2012 Sep 20. PMID: 23141867.

According to my meaning, it is therefore important to cite the topic as a limitation of the study tied to the manuscript length, or as a feature research goal, as you prefer.

Response: We thank the reviewer for their suggestion and useful resources. We have now included a paragraph in the Implications, limitations, and future directions subsection briefly addressing this, including additional references to clarify our points, as follows:

‘The design of the current study required participants to provide unrestricted reports on the content of their visual imagery to music, implicating their ability to verbalise their visual imagery experiences (i.e., constructing their narrative to music [4]). It is generally well evidenced that music and language are reflected by overlapping electrophysiological correlates [16,63–67], but some studies identify differences in this regard (e.g., between males and females [68–71], and as a function of musicianship [72–77]). Neuroscientific studies on music-induced visual imagery are only beginning to emerge (see [61,78] for first evidence of neural signatures). However, we suggest that future studies may seek to combine approaches like those taken in our current paper with emerging insights into neural underpinnings in order to advance knowledge of both the brain and visual imagery during music listening.’

We look forward to hearing from you soon regarding our submission and we are happy to respond to any further questions and comments that you or the reviewers may have.

With best regards, and on behalf of the co-authors,

Sarah Hashim

Attachment

Submitted filename: Response to Reviewers.docx

Decision Letter 2

Ioanna Markostamou

12 Oct 2023

Music listening evokes story-like visual imagery with both idiosyncratic and shared content

PONE-D-23-03570R2

Dear Dr. Hashim,

Thank you for aking into consideration all the comments and suggestions made by the reviewers and submiting your revised manuscript at PLOS ONE.

I have now received the reviews of the three original reviewers, and as you will see from their comments at the end of this letter, all reviewers evaluated very positively your paper in its current form and did not raise any further points of critique.

Therefore, I am happy to inform you that your manuscript has been judged scientifically suitable for publication at PLOS ONE and will be formally accepted for publication once it meets all outstanding technical requirements.

Within one week, you’ll receive an e-mail detailing the required amendments. When these have been addressed, you’ll receive a formal acceptance letter and your manuscript will be scheduled for publication.

An invoice for payment will follow shortly after the formal acceptance. To ensure an efficient process, please log into Editorial Manager at http://www.editorialmanager.com/pone/, click the 'Update My Information' link at the top of the page, and double check that your user information is up-to-date. If you have any billing related questions, please contact our Author Billing department directly at authorbilling@plos.org.

If your institution or institutions have a press office, please notify them about your upcoming paper to help maximize its impact. If they’ll be preparing press materials, please inform our press team as soon as possible -- no later than 48 hours after receiving the formal acceptance. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

Kind regards,

Ioanna Markostamou, Ph.D.

Academic Editor

PLOS ONE

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. If the authors have adequately addressed your comments raised in a previous round of review and you feel that this manuscript is now acceptable for publication, you may indicate that here to bypass the “Comments to the Author” section, enter your conflict of interest statement in the “Confidential to Editor” section, and submit your "Accept" recommendation.

Reviewer #1: All comments have been addressed

Reviewer #2: All comments have been addressed

Reviewer #3: All comments have been addressed

**********

2. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: Yes

**********

3. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: Yes

**********

4. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: Yes

**********

5. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: Yes

**********

6. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)

Reviewer #1: (No Response)

Reviewer #2: I would like to thank the authors for taking into consideration the comments and suggestions I made, and I congratulate them for carrying out a piece of research that advances our understanding of the music-evoked visual imagery phenomenon.

Reviewer #3: My comments were satisfied and the paper sounds clearly ready to be published. I support the publication of this paper

**********

7. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #1: Yes: Dr. Steffen A. Herff

Reviewer #2: Yes: Julian Cespedes-Guevara,PhD

Reviewer #3: Yes: Marta Olivetti Belardinelli

**********

Acceptance letter

Ioanna Markostamou

17 Oct 2023

PONE-D-23-03570R2

Music listening evokes story-like visual imagery with both idiosyncratic and shared content

Dear Dr. Hashim:

I'm pleased to inform you that your manuscript has been deemed suitable for publication in PLOS ONE. Congratulations! Your manuscript is now with our production department.

If your institution or institutions have a press office, please let them know about your upcoming paper now to help maximize its impact. If they'll be preparing press materials, please inform our press team within the next 48 hours. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information please contact onepress@plos.org.

If we can help with anything else, please email us at plosone@plos.org.

Thank you for submitting your work to PLOS ONE and supporting open access.

Kind regards,

PLOS ONE Editorial Office Staff

on behalf of

Dr. Ioanna Markostamou

Academic Editor

PLOS ONE

Associated Data

    This section collects any data citations, data availability statements, or supplementary materials included in this article.

    Supplementary Materials

    S1 File. Appendices.

    (DOCX)

    S2 File. Supplementary.

    (DOCX)

    Attachment

    Submitted filename: PONE-D-23-03570 - REVIEW.pdf

    Attachment

    Submitted filename: PONS review.docx

    Attachment

    Submitted filename: Response to Reviewers.docx

    Attachment

    Submitted filename: Response to Reviewers.docx

    Data Availability Statement

    All data files are available from the OSF database: https://osf.io/nf4x7/?view_only=081602e5aca94788bb959b48ed8b47ef.


    Articles from PLOS ONE are provided here courtesy of PLOS

    RESOURCES