Abstract
In this experiment, we examined pressures that influence linearization decisions when speakers plan and produce spoken descriptions that extend across multiple utterances. More specifically, we aimed to investigate macroplanning in discourse-level production using data from both speech and eye movements. Participants were shown 48 networks that contained two separate branches varying in length (number of nodes) and complexity (number of choice points). Each participant described 24 experimental networks and 24 filler networks while their eye movements and speech were recorded. We found that the location of prespeech fixations did not predict the first described branch. We also found that once speakers began to speak, they described the networks highly incrementally. These results are consistent with extensive planning prior to initiation of a complex description, as well as being indicative of a highly incremental production strategy following the onset of the description. Lastly, we report that speakers prioritize the shorter side of the network, regardless of branch complexity, which may suggest that speakers evaluate perceptual features and prioritize the side that appears to be easier to describe based simply on its length. Overall, our results are consistent with speakers using an apprehension phase to develop a macroplan for multi-utterance descriptions, a plan that is based on perceptual features. These findings offer support for incrementalism in language production, extending this principle to discourse-level production.
Keywords: Eye movements, Language production, Psycholinguistics, Multi-utterance production, Macroplanning
Introduction
Language production allows for internal ideas and concepts to be represented as an external comprehensible message. The flexibility of syntactic forms and the availability of lexical items offer speakers a variety of ways to communicate ideas that originally existed in non-linguistic, multi-dimensional conceptual representations. For example, if a speaker views a scene of a dog running after a man, they may describe this event in the following ways: The dog is chasing the man, the man is being chased by the dog; The animal is pursuing the person, the person is fleeing the animal, and so on. How does a speaker make linguistic decisions, such as assigning the dog to the subject position?
Much of the existing work has focused on decisions involving syntactic assignment or lexical retrieval, which can be considered phrasal or sentence-level planning. In contrast, much less is known about macroplanning, the higher-level processes that guide the overall organization of a sequence of utterances or discourse (Butterworth, 1980; Levelt, 1989). More specifically, macroplanning concerns higher-level aspects of utterance planning. In this stage, a speaker must determine the overall message they wish to convey based on communicative intent or goals. Macroplanning involves selecting the relevant information and sequencing it logically. To maintain the previous example of a man being chased by a dog, macroplanning is responsible for choosing the thematic focus (e.g., Agent or Patient), selecting the discourse starting point (e.g., starting with The man vs. The dog), and how abstract the speaker intends to be (e.g., the dog vs. the animal). Specifically, macroplanning is the conceptual or thematic planning stage, where a speaker determines how to communicate a message, as well as which information is to be included in the resulting discourse. Despite the role macroplanning plays in productive communication, it remains largely understudied.
Macroplanning and multi-utterance language production
Multi-utterance production is defined as language production that extends beyond a single utterance. Some empirical work has identified ordering preferences at the sentence-level, such as identifying how speakers decide which entity will be encoded as the subject, which is often the first item mentioned in a sentence (e.g., Gleitman et al., 2007; Griffin & Bock, 2000). Other factors, like prior knowledge (Bock & Irwin, 1980; Clark & Clark, 1977), context (Bock, 1986), and the availability of other cognitive resources, such as working memory (Levelt, 1989), have also been shown to modulate the order in which items are mentioned. However, when generating a production plan that extends beyond a single sentence, a speaker must determine the overall message they want to convey based on their communicative intention or goals (i.e., macroplanning).
One critical component of macroplanning is linearization (Levelt, 1981), which is the imposition of a linear order on the information that is to be communicated linguistically. In any production task, a speaker must determine the order in which concepts are to be formulated and articulated. As multiple concepts cannot be spoken simultaneously, a speaker must prioritize certain elements of the intended message, settling on an utterance form that places these concepts earlier in time. How speakers approach the problem of linearization may depend on the nature of the information that is to be expressed. For example, an event sequence can be organized temporally, and the logical way to describe such an event would follow the same ordering (Schank & Abelson, 1977). Similarly, human long-term memory is organized according to schemas and scripts, which detail specifications and sequences for certain events and activities (e.g., ordering at a restaurant) (Schank & Abelson, 1977). For example, if asked to describe your morning routine, it is probable that it resembles something of the following: you wake up, you brush your teeth, you make yourself breakfast and coffee, and you depart for work. In descriptions such as the example provided, temporal order provides an inherent linearization strategy (principle of natural order; Levelt, 1989).
When describing something that lacks an inherent temporal order, the solution to the linearization problem takes a different form, as has been shown for descriptions of residences (e.g., Ehrich & Koster, 1983; Linde & Labov, 1975; Shanon, 1984; Ullmer-Ehrich, 1982). Previous work centered on spatial descriptions has suggested that speakers’ movement and spatial information, namely hypothetical movement through the space, are used as linearization strategies (Ehrich & Koster, 1983; Linde & Labov, 1975; Shanon, 1984; Ullmer-Ehrich, 1982). For example, Linde and Labov (1975) had subjects provide verbal descriptions of apartment layouts. Speakers tended to follow a linearization plan that followed a mental tour of the environment, including the use of deictic expressions (e.g., “You walked in the front door. There was a narrow hallway. To the left, the first door you came to was a tiny bedroom…”; Linde & Labov, 1975, p. 927). Though such descriptions are not inherently bound to temporality, a tour structure as a discourse strategy has features of event descriptions, such as the inclusion of temporal expressions (e.g., and then), and may be enacted as an imagined tour of the physical space.
Levelt (1981, 1982) was one of the first to investigate the linearization problem, and to do so, he examined the production of network descriptions. Speakers were presented with simple networks of colored, interconnected nodes, and they were instructed to describe the network in such a way that another subject could correctly reproduce the network from their description alone. Beginning with the starting node, indicated with an arrow in Fig. 1, speakers would then reach a choice point (i.e., white dot; Fig. 1) at which the decision to describe the left side (left branch) or the right side (right branch) of the network must be made. The network structure branches (i.e., dot structures to the left and right of the choice point) varied in length and type (linear, choice, and loops; Fig. 1a–c). Linear figures were ones in which both branches could be described sequentially, with no embedded choice points; choice figures were ones in which either the left or the right branch contained an embedded choice point, necessitating a second linearization decision; and loop figures were ones in which one branch or the other included a series of dots that formed a closed loop.
Fig. 1.

Example of network structures used in Levelt (1981, 1982)
Levelt (1981, 1982) observed three general patterns of linearization: First, speakers described the networks as interconnected structures, where each node had a direct connection to the surrounding nodes (i.e., principle of connectivity). Second, speakers would always return to the nodes that had been previously mentioned in a first-in, last-out fashion (i.e., stack principle). Speakers always returned to the most recent choice point they mentioned, keeping track of these points in a mental stack. Then they would return to each point in reverse order. Last, Levelt (1981, 1982), observed that speakers would begin by describing the branch that incurred the least memory load. Specifically, speakers would describe the shorter branch first (e.g., left side of network in Fig. 1a), or they would describe the branch that did not impose additional choice points (e.g., left branch in Fig. 1b). Levelt (1981) defines this preference as the minimal-load principle: “Order alternative continuations in such a way that the resulting memory load for return addresses is minimal” (Levelt, 1989, p. 95). According to the stack principle, a choice point must be kept in memory until the speaker returns to it. Only after revisiting that point can it be released from memory. Speakers choose to describe the branch that minimizes how long they need to maintain that choice point in memory.
Later work drew similar conclusions: Speakers preferred to describe short branches before long branches and to describe linear branches before choice branches (Ferreira & Henderson, 1998). These results also supported the minimal-load principle. Further, Ferreira and Henderson (1998) found that speakers planned at the branch level rather than at the level of the entire network. They reported that, even when required to begin and end their description at the entry dot (black dot in Fig. 1), speakers still preferred to describe the easy branch before the hard branch. Given this, they concluded that speakers did not plan far enough in advance to identify that, regardless of the branch first described, the overall load would be the same for both possible description strategies, further supporting the incremental nature of discourse-level language planning.
The current study: Investigating macroplanning through network descriptions
The purpose of the current study is to address a limitation of prior work that has used the network description task to study discourse-level production and planning, and that is the reliance on verbal data alone. Information about where speakers look before making decisions would provide valuable information about the plans speakers generate before speaking. For example, if speakers fixate on both branches of the network before beginning to speak, we could conclude that they are evaluating both sides as potential starting points, consistent with extensive utterance preplanning.
In the current study, participants described a series of networks that differed in length and/or structure (Ferreira & Henderson, 1998; Levelt, 1981, 1982) while their eye movements were recorded. First, in accordance with the principle of connectivity, we predicted a tight coupling between fixation location and dot mention (~ 900ms; Griffin & Bock, 2000), with participants describing each dot in spatial order, and with eye movements aligning closely with speech data. Second, we hypothesized that speakers would prioritize the branch that resulted in a lesser memory load and their description and pre-speech fixations would reflect this preference. Ferreira and Henderson (1998) suggest that planning occurs at the branch level. If speakers evaluate the relative cost of describing each branch, we would expect both branches to receive fixations prior to the onset of the description. Furthermore, if these early fixations predict which branch is first described, this may indicate the generation of a specific linguistic plan, rather than macroplanning alone.
Method
Participants
One hundred and eight University of California, Davis undergraduates participated in this experiment in exchange for course credit1. All subjects had normal or corrected-to-normal vision and normal color vision. Participants were naive to the purpose of the experiment and provided informed consent. Eye movement data were inspected for artifacts (e.g., blinks, movement). Thirty-three subjects were excluded from the main analyses due to low eye-tracking accuracy, resulting in 75 subjects that were tracked well (M = 82.5%).
Stimuli
Forty-eight network structures (960 × 720 pixels) were created (Fig. 2). For all stimulus types (experimental and filler), each network consisted of ten uniquely colored dots connected by horizontal and vertical lines. The dots were 54 pixels in diameter and the lines that connected the dots were 112 pixels in length. Each network contained one dot with each of the following colors: red, orange, yellow, green, blue, pink, purple, white, gray, and black. Every color was used once in each network. Networks were labeled using the following convention: left branch type – right branch type. All experimental networks appeared as one of the following types: Linear-Linear, Choice-Choice, Linear-Choice, Choice-Linear. Each network contained (1) a starting dot, as indicated with a teal arrow; (2) a choice dot (the dot directly above the starting dot; (3) a left branch, which was either (a) three or (b) five dots in length and (c) linear or (d) choice in form; a right branch, varying along the same two dimensions (Fig. 2A–D). Subjects described three of each network type for a total of 24 experimental networks.
Fig. 2.

Example of each experimental network type and filler network
Subjects also described 24 filler networks. Filler networks were identical to the experimental networks except that the left and the right branch were the same length. Filler trials were included to ensure that participants would be less likely to identify the hypotheses of the experiment and to inhibit the development of response strategies. Filler networks contained an equal number of dots (4) to either side of the choice point. Specifically, filler networks contained a (1) starting dot; (2) a choice dot, (3) a left branch four dots in length; (4) a right branch, four dots in length. The configuration of the left and right branches for filler items was varied to ensure visual similarity to the experimental networks (Fig. 2).
Apparatus
Eye movements were recorded using an SR Research Eye-Link 1000+ tower mount eyetracker at a sampling rate of 1000 Hz. Viewing was binocular, but eye movements were recorded from the right eye only. To minimize head movements, a chin and forehead rest was used. Experiment presentation was controlled by SR Research Experiment Builder software. Subjects sat 83 cm away from the monitor and the scenes were displayed at a 960 × 720 pixel resolution. Spoken descriptions were recorded using a Roland Rubix USB audio interface and a Shure SM86; however, Experimental Builder downsampled the output files to 24 kHz.
Procedure
The experiment began with a five-point calibration procedure. Successful calibration required an average error below 0.49° and a maximum error below 0.99° (M = 81.5 %). Throughout the duration of the experiment, calibration was maintained using a drift correction procedure. Following successful calibration, subjects were trained on the color names they were to use for the dots in the experimental portion of the task.
First, they were shown a training screen that provided an example of each colored dot along with the corresponding label. This slide provided subjects with a visual display of the colored dots they would encounter in each network, highlighting that each colored dot would only appear once per network description. Then, participants were further instructed to use the provided color label for each colored dot, followed by the word “dot” (e.g., “the green dot”). Color names were selected to be the most basic term available (such as in Berlin & Kay, 1991). Some colored dots had more than one viable color label, such as the purple dot, which could have been referred to as “lilac” or “mauve”. One motivation for pre-training on color terms was to establish consistency across speakers. Additionally, we wanted to ensure that each dot was referred to by a unique color term, which would facilitate the identification of the mention of each colored dot during subsequent analyses of the verbal data. Pre-training on color terms was also done to reduce retrieval time for the appropriate lexical item. Lastly, the stop consonant at the beginning and end of dot affords the opportunity for this or other research teams to perform acoustic analyses in the future.
Following the color label pre-training, participants were instructed as follows: “In this experiment, you will see a series of figures. Your task is to describe the figures so that someone else could draw them from your description alone.” These instructions were selected to match previous iterations of this work (Ferreira & Henderson, 1998; Levelt, 1981, 1982). The final instruction was to begin each network description by beginning with the dot directly above the teal arrow. Subjects were informed that they were allotted unlimited time for their descriptions. Prior to beginning the experimental portion, participants were provided three practice trials during which the experimenter was able to provide corrective feedback (e.g., ensuring subjects referred to the circles as dots and were using the appropriate color terms). Following the three practice trials, subjects were provided with a screen to remind them of their task and that the button box should be used to advance to the next network when they finished their description. Subjects completed three practice network descriptions and 48 experimental network descriptions (24 experimental networks) for a total of 51 network descriptions (see Fig. 3 for experimental schematic).
Fig. 3.

Experimental schematic for network description task, Note. (a) Subjects were provided with task instructions before beginning the experimental portion. (b) A five-point fixation array was used for calibration. (c) Subjects begin the description task and are provided with unlimited time to describe the provided structure and press a button when finished. (d) Subjects were then prompted with a drift correction screen that advanced after experimenter button press
Eye movement data were imported into MATLAB using the Visual EDF2ASC tool in the SR Research DataViewer software. The first fixation and saccadic outliers were excluded (amplitude > 20°). Though successful calibration required an average error below 0.49° and a maximum error below 0.99°, we also determined how well subjects were eye-tracked across all trials (including filler items) by identifying the mean proportion of signal and assigning a percentage for how well subjects were accurately eye-tracked. For data to be included in analysis, subject-level eye-tracking accuracy must have exceeded 70% across all trials (M = 81.5, SD = 6.7). Although the experiment began with a five-point calibration procedure and we required an average error below 0.49° and a maximum error below 0.99° before proceeding with data collection, post-processing revealed that calibration accuracy was often lower than desired. This is why a second threshold of exceeding 70% across all trials was implemented. One possible explanation for this may be the use of a chinrest for a production task. The use of a chinrest while speaking may introduce additional artifacts. Another possible explanation for this may be due to the time during which data for this project were collected. Subjects were run immediately following the return to in-person data collection (i.e., “post” COVID-19), which may have been the source of needing to exclude additional subjects (e.g., dark colored masks interfering with maintaining accurate calibration, as well as subjects getting re-accustomed to the requirements of being tested for an extended period).
Data preprocessing
Verbal data processing
Verbal data preparation
Spoken descriptions provided by all subjects for all scenes were first amplified in Audacity. Then, the .wav files were transcribed using Whisper’s speech-to-text algorithm (Open AI), which converts .wav files to .json files. Using the transcribed spoken descriptions and WhisperX (Bain et al., 2023), the onset and offset for each word were then identified. Following the alignment of speech data, a research assistant naive to the hypotheses manually corrected transcription errors (e.g., deletion of hallucinations). During manual correction, the same research assistant manually identified whether the subject described the left or the right side of the network first. Once corrected, the onset and offset of each word were re-identified using the Montreal Forced Aligner (McAuliffe et al., 2017).
Dot mention
Mentions of color words were identified using the ‘grepl’ function in base R (R Core Team, 2018). In addition to the ten color words, the subjects were instructed to use 200 synonyms for the color words (20 for each color), which were also included (Appendix A) for cases when subjects forgot to follow instructions regarding color terms. Every mention of each dot was coded (Fig. 4).
Fig. 4.

Identification of colored dots for verbal descriptions
Fixation data preparation
Prespeech fixations
Prespeech fixations were identified as any fixations that occurred prior to the onset of the verbal description. Fixation coordinates determined whether the fixation fell within an area of interest (i.e., colored dot) or outside the network structure (i.e., white space). Fixations that fell within the range of (0, y) and (479, y) on the x-axis were coded as a fixation made to the left side of the network. More specifically, if a subject made a fixation that landed anywhere on the y dimension, but did not exceed the midway point of the network (479, y), then the fixation was considered to be a look to the left side of the network. Any fixations that were greater than (481, y) were considered to be a fixation made to the right side of the network. To elaborate, as with the left side of the network fixations, if a subject made a fixation anywhere on the y dimension, and that fixation exceeded 481 on the x dimension, then it was considered to be a look to the right-hand side of the network. The coordinates of the fixations were also inspected to determine which, if any, landed on the midway point (i.e., anywhere on the center of the screen’s y-axis) (480, y) (N = 0).
Colored dot fixation identification
Fixations made to colored dots were identified by using minimal coordinate pairs for each of the network dots. Each colored dot was identified with coordinates for the top left corner (x, y) and bottom right (x, y), creating a boundary to establish the regions (dots) of interest. The resulting boundary box measured 54 × 54 pixels. Any fixations that fell within the bounds of the coordinates (x1, x2, y1, y2) were considered to be a fixation on that dot.
Eye–Voice span estimate (EVS estimate)
To compute the eye–voice span, we operationally defined it as the temporal difference (in ms) between the onset of a fixation and the onset of the corresponding spoken mention. Each mention of a colored dot was paired with a single fixation on a dot of that color. All fixations to that color were treated as candidate matches. For each mention, we calculated the temporal difference between fixation onset and speech onset across all candidate fixations and selected the fixation with the smallest absolute difference. This fixation was used to compute the eye–voice span.
For example, in describing a network, a participant might say, “There is a yellow dot. Next to the yellow dot is a blue dot…”. While describing this network, the participant would have made several separate fixations on the yellow dot. For each fixation, we calculated the difference between the onset of the fixation and the onset of the spoken mention. The fixation with the smallest difference was selected as the one most likely to correspond to the verbal mention and was paired with it for analysis. This procedure was repeated for all mentions of each dot.
Because some participants occasionally used non-standard color labels (e.g., “mustard” for yellow), all deviations were corrected to the pre-specified set of ten colors to ensure consistent pairing of fixations with mentions. These deviations were rare (196 mentions, 0.33% of the total).
Analysis
For the following analyses, mixed-effects models were constructed using the ‘glmer’ function of the ‘lme4’ package in R (Bates et al., 2015); R Core Team, 2018). For all analyses, we started with a maximal random effects structure (including by-subject and by-item intercepts and slopes for all fixed effects), following Barr et al. (2013). In the case of convergence issues, the model was iteratively pruned, dropping the random effects with the least variance, until the structure facilitated convergence (Barr et al., 2013).
Although we had no theoretical grounds for anticipating differences between data collected prior to the onset of the COVID-19 pandemic (pre-pandemic group) and data collected following the return to in-person data collection (post-pandemic group), each of our models also included a fixed effect of Group due to anecdotal reports concerning differences in some subjects’ behavior and compliance with instructions. All models used the default optimizer (bobyqa).
As done previously in Ferreira and Henderson (1998), the design of the experiment was 2 (left side linear vs. left side choice) × 2 (right side linear vs. choice) × 2 (3 vs. 5 circles), with all variables within participants. The labels for the experimental conditions are summarized in Table 1.
Table 1.
Abbreviations for experimental items & expected direction
| Network type | Length | Abbreviation | Expected direction (minimal-load) |
|---|---|---|---|
| Linear-Linear | 3–5 | L3-L5 | Left |
| 5–3 | L5-L3 | Right | |
| Choice-Choice | 3–5 | C3-C5 | Left |
| 5–3 | C5-C3 | Right | |
| Linear-Choice | 3–5 | L3-C5 | Left |
| 5–3 | L5-C3 | Left | |
| Choice-Linear | 3–5 | C3-L5 | Right |
| 5–3 | C5-L3 | Right |
Network description: Apprehension & planning
Pre-speech fixations were discretized such that any look made to the left side of the network (0 < x < 480) was coded as a 0 and any fixation made to the right side of the network (480 < x ≤ 760) was coded as a 1. For this particular analysis, we conducted a logistic mixed-effects regression analysis in which the dependent variable was whether subjects first described the left side of the network (0) or the right side of the network (1), with the proportion of looks made to the left side (fixations to left / all prespeech fixations) prior to beginning each description as the predictor. A second analysis used the direction (left, 0; right, 1) of the final fixation prior to beginning the description as our independent variable.
Network description: Minimal load analysis
To determine which variables affected ordering preferences (i.e., describing the left/right side of the network first), we coded our networks as follows (see Table 1 for information on abbreviations). For the fixed effects, we included fixed effects of Left Branch Type (left side linear vs. left side choice), Right Branch Type (right side linear vs. choice), and Left Side Length (3 vs. 5 circles). Since our networks always contained a branch with either 3 or 5, only the length of the left side was included (as the length of the right side is implicit). To best compare the model results to previous effects (e.g., ANOVAs performed by Ferreira & Henderson, 1998), we used sum-coded contrasts using the function contr.Sum() from the ‘car’ package. The dependent variable, whether the subject first described the left side of the network or the right side of the network, was coded as a 0 (left) or a 1 (right). If a participant first described the left-hand side of the network, this was coded as 0. If a participant began with the right side of the network, this was coded as a 1. Participants’ decisions to describe the left branch first will result in values < 0.5, and to describe the right branch first will result in values > 0.5. Contrast coding followed suggestions from Brehm and Alday (2022).
Network description: Eye–Voice span estimate
To measure the temporal relationship between the look to a colored dot and the mention of the same dot, we calculated the amount of time that elapsed between (1) the onset of each fixation made to every colored dot and (2) the onset of every single mention of the same colored dot. In many cases, subjects mentioned the same colored dot twice in a row, but only the first was the referential mention (e.g., “there is a red dot. From the red dot is a yellow dot…”). To try to isolate the referential look for each mention, we then selected the negative value that was closest to zero. If a subject produced an error in their description, the repair was treated as the mention (e.g., “From the blue, I mean red dot…”, “From the bl..blue dot…”) and the onset of the repair and the offset of the repair were used in measuring the temporal relationship between looks to the dot (e.g., blue dot) and the repaired mention of the same dot.
Results
The average latency to begin the description for all networks was about 1600 ms (M = 1591, SD = 685) and similar latencies were observed for the experimental trials (M = 1582, SD = 694). Prior to beginning the description of the network, subjects made about four fixations (M = 3.6 fixations, SD = 1.9 fixations) (see Table 2 for more detailed statistics). Of the prespeech fixations that fell within the bounds of the colored dots, the majority landed on dots closest to the arrow (Fig. 5).
Table 2.
Average pre-speech (PS) latencies and pre-speech fixations (fixn) made for all network types
| Network type | Avg. PS latency | SD | Avg. PS fixn | SD |
|---|---|---|---|---|
| C3-C5 | 1734 ms | 875.0 ms | 4 fixn | 2.2 fixn |
| C3-L5 | 1569 ms | 668.8 ms | 3 fixn | 1.8 fixn |
| C5-C3 | 1581 ms | 654.2 ms | 4 fixn | 2.0 fixn |
| C5-L3 | 1572 ms | 672.7 ms | 4 fixn | 1.8 fixn |
| L3-C5 | 1505 ms | 613.0 ms | 3 fixn | 1.6 fixn |
| L3-L5 | 1587 ms | 698.7 ms | 4 fixn | 1.9 fixn |
| L5-C3 | 1576 ms | 618.7 ms | 3 fixn | 1.7 fixn |
| L5-L3 | 1608 ms | 628.0 ms | 4 fixn | 1.8 fixn |
Fig. 5.

Prespeech fixation colored dot locations, Note. Coding for dot locations followed Ferreira and Henderson (1998), beginning with coding the expected side (short > long; linear > branching)
Network direction: Proportions
To follow previous analyses (e.g., Levelt, 1981, 1982; Ferreira & Henderson, 1998), we wanted to determine whether subjects would describe the left or the right side of the network first, given each network structure. As with previous reports, many subjects varied their pattern of description (i.e., sometimes began by describing the left, sometimes began by describing the right) while some selected a direction and first described the same side of the network each time (i.e., always started with the left side). Participants who consistently used the same approach across all trials (e.g., always beginning with the left branch) are referred to as single-strategy participants (SSP). Participants who adopted different approaches across trials (e.g., sometimes beginning with the left branch and sometimes with the right branch) are referred to as multi-strategy participants (MSP). Analyses that included all participants will be referred to as all participants.
As with previous work (e.g., F. Ferreira & Henderson, 1998), we will perform our analyses twice: once with all subjects (includes SSP) and once with only MSP (Table 3).
Table 3.
Note. Proportions that were less than 0.5 represent a tendency for participants to describe the left branch first
| Branch type | 3 left, 5 right | 5 left, 3 right | ||
|---|---|---|---|---|
| All participants | MSP | All participants | MSP | |
| Number of circles | ||||
| Linear-Linear | 0.19 | 0.18 | 0.53 | 0.61 |
| Choice-Choice | 0.16 | 0.14 | 0.41 | 0.46 |
| Linear-Choice | 0.22 | 0.22 | 0.50 | 0.57 |
| Choice-Linear | 0.23 | 0.23 | 0.47 | 0.54 |
| 4 left, 4 right | ||||
| All Participants | MSP | |||
| Filler | 0.32 | 0.34 | ||
Proportions that exceed 0.5 represent a tendency for participants to describe the right branch first
Overall, there was a tendency to begin with the left-hand side of the network for both filler and experimental trials. This global tendency may in part be attributable to a left-to-right scanning preference that has been reported for individuals literate in a language that is read from left to right. Additionally, it is possible that cognitive styles, such as the intrinsic versus extrinsic use of left/right may also influence macroplanning and, ultimately, which branch is described first.
For all subjects, including those who described the same side of the network for every single trial, participants tended to describe the shorter branch first (Table 3). More specifically, when the two branches were the same structure (Linear-Linear, Choice-Choice), participants were more likely to begin with the short side, except for C5-C3, where the proportion of responses was nearly split. When the two branches differed in structure, the number of dots, rather than the type, affected the proportion of responses. Specifically, in L5-C3, participants tended to begin with the short side, rather than the longer, but linear, one. The same pattern was observed for C3-L5. After removing the subjects who went to the same side every time (SSP, N = 23), the observed pattern of results was nearly identical.
Network direction: model results
The first-described branch was analyzed through a mixed-effects model (as outlined above under Analysis). The model included fixed effects for length (3/5), branch type (linear/choice), group membership (pre/post), and a three-way interaction between left branch type, right branch type, and length. Random intercepts for both subjects and items were included (Table 4). Across all subjects, including those who described the same side first on very single trial (i.e., SSP), the model reported a significant effect for length (β = −1.313, SE = 0.091, z = −14.419, p < 0.001), left branch type (β = 0.234, SE = 0.078, z = 2.985, p = 0.003) and for right branch type (β = 0.236, SE = 0.078, z = 3.012, p = 0.003). More specifically, when the left side of the network was shorter, speakers were more likely to begin by describing the short side before the long side. When the left side of the network was linear instead of choice, there was an increased probability of describing the right side first. When the right side of the network is choice instead of linear, the probability of describing the choice branch first increases. The model also reported a significant interaction between left branch and right branch type (β = −0.160, SE = 0.078, z = −2.043, p = 0.041). When the left side of the network is linear and the right side of the network is linear, the probability of describing the right branch first decreases. No other interactions were reliably significant, and as expected, there was no reliable effect of group.
Table 4.
Predicted first described branch by network type
| Network type | MSP | Predicted 95% CI | Direction [Feature(s)] |
|---|---|---|---|
| L3-L5 | 0.15 | [0.07, 0.29] | Left [Length] |
| L5-L3 | 0.79 | [0.61, 0.90] | Right [Length] |
| C3-C5 | 0.08 | [0.03, 0.17] | Left [Length] |
| C5-C3 | 0.54 | [0.31, 0.75] | Right [Length] |
| L3-C5 | 0.16 | [0.08, 0.31] | Left [Length + Type] |
| L5-C3 | 0.71 | [0.50, 0.85] | Right [Length] |
| C3-L5 | 0.19 | [0.09, 0.39] | Left [Length] |
| C5-L3 | 0.67 | [0.45, 0.83] | Right [Length + Type] |
Note. The predicted probabilities for the direction mentioned first are presented along with the feature (Type, Length, or Type + Length) most likely responsible for the observed pattern
Eye movements and network planning & descriptions
The location of the prespeech fixations (i.e., the proportion of fixations made prior to the onset of each network description) did not reliably predict the side that would be first described (β = 0.179, z = 0.341, p = 0.733), nor did the side of the final fixation made prior to the onset of the description (β = 0.408, z = 1.170, p = 0.242).
During network descriptions, speakers tended to fixate on the colored dot about 720 ms before mentioning the corresponding dot (M = −720.04 ms, SD = 1110.375). When removing outliers, the average eye-voice span fell to about 600 ms (M = −569.271, SD = 364.924). Our results suggest the following. First, we observed similar apprehension times to other production tasks that differed in scope than the work presented here (e.g., Griffin & Bock, 2000; Gleitman et al., 2007). During this interval, speakers made several fixations before beginning their description of the networks. Of these fixations, the side that was fixated on most prior to description onset did not predict the side that would be first described, nor did the location of the final fixation in the prespeech interval. Our eye-voice span estimate reflects that of similar work, as we observed about a 600-ms difference between the eye (fixation) and voice (mention).
Discussion
The current study examined pressures that influence macroplanning, namely linearization decisions, when speakers plan and produce spoken descriptions that extend across multiple utterances. More specifically, this work had participants describe a series of simple networks consisting of interconnected lines and colored circles while their eye movements were recorded. Previous work suggested that speakers will plan and execute their utterances in a way that will minimize the computational burden that is imposed based on physical properties of the networks (Levelt, 1981, 1982), specifically at the branch level (Ferreira & Henderson, 1998). Overall, our findings were mostly in support of previous work: speakers preferred to describe the easy (based on branch length) branch first.
Apprehension times (i.e., time before onset of description) were similar to previous reports (e.g., Griffin & Bock, 2000, Gleitman et al., 2007), and speakers made about four fixations before beginning to speak. The side on which speakers fixated most often during this interval did not predict the side first mentioned, and the side of the final fixation prior to the onset of the description also did not reliably predict which side of the network would first be described. Lastly, our estimate of the eye–voice span was similar to previous reports (e.g., Griffin & Bock, 2000; Coco & Keller, 2012), but may suggest that speakers describe the networks more incrementally, given that the average onset of a fixation and the average onset of mention was around 600 ms (compared to ~900 ms; Griffin & Bock, 2000). When taken together, these results suggest that speakers prepare, to some extent, a linguistic plan prior to beginning their description, but once they have started, they follow a more strictly incremental production plan. What we do not observe is that subjects simply begin their descriptions with whatever side of the network their eye happened to land on at the beginning of the trial.
To elaborate on our results, our findings are mostly consistent with previous reports (Ferreira & Henderson, 1998; Levelt, 1981, 1982). We observed a strong effect of branch length, in that speakers were more likely to describe the short side before the long side across all network types. Additionally, we observed an effect of left branch type and right branch type, and a significant interaction between the two, suggesting that when both the left branch and right branch are both linear, the probability of describing the left branch first increases. For same-branch-structure networks (L-L, C-C), we found that speakers preferred to begin with the short branch. For mixed-branch-structure networks (L-C, C-L), the length of the branch predicted the side first mentioned. Also consistent with previous analyses, we found that speakers began with the short side of the network when the branch types were mixed (i.e., C5-L3, L3-C5). However, we found that speakers preferred to begin with the short side of the network, even if the branch type of the short side introduced the additional choice point (i.e., C3-L5, L5-C3). This may suggest that our participants prioritized length over branch type.
Previously, the pattern of results suggested that the minimal-load principle is in accordance with easy-first in that the ‘easy’ branch was the side that imposed a computationally less expensive working memory load on the production system based on the size and duration for which choice points must be maintained in memory. In our work, however, it is possible that size and duration were conceptualized at the branch level. Rather than the time for which a choice point is maintained in memory, it may be that speakers are considering the overall time it will take to complete the description of the branch. For example, consider the network structure, C3-L5. The minimal-load principle would suggest that a speaker would begin by describing the right branch (L5) first given that it does not introduce a “return address”, as is imposed by the left branch (C3). However, if a speaker is considering overall duration (i.e., time it takes to complete the description of a branch), the C3 branch would likely take less time given that the description of the left branch only includes the mention of three dots (compared to five dots). This is the pattern we report: speakers describe the shorter branch first, regardless of branch type. This explanation would align with previous reports that suggested that production planning occurred at the branch level (Ferreira & Henderson, 1998). If speakers are generating a production plan based on a comparison of the two branches (rather than a comparison of two network descriptions), then this may explain why we observed an effect of branch length, regardless of branch structure.
This interpretation of our results is also consistent with the concept of perceptual clauses, which are defined as a unit of perceptual information that is used to organize the input to the production system, which is mapped onto a linguistic clause or utterance (Ferreira & Barker, 2024). More specifically, a perceptual clause is used to organize visual information in a manageable way for the language formulator to facilitate fluent and coherent speech. Ferreira and Barker (2024) investigate this idea as it relates to macroplanning in multi-utterance descriptions of real-world scenes. In these scenes, it is up to the speaker to decide where to begin and how to structure the description (i.e., beginning with the room label versus the colors of the walls versus the existence of a couch). They suggest that during the apprehension phase of language planning, focal objects in the scene are identified, which serve as the anchor for smaller, satellite objects. These focal object/satellite object groupings create clusters, which are then able to be mapped onto linguistic constituents. Each of these groupings then serves as a perceptual clause, a unit of visual information taken from the visual scene and to be mapped onto a linguistic utterance.
In the case of network descriptions, speakers may use the pre-speech interval to evaluate the relative perceptual simplicity/complexity of either branch and then use these representations to describe the networks in an orderly way. Subjects were not instructed to describe the networks branch by branch. Their instructions were to begin with a designated dot, then to describe the entirety of the networks. One linearization strategy could have been to describe the dots randomly. Instead, we observed logical groupings of dots, driven by length. During the apprehension phase, the speaker inspects the visual environment and generates significant or meaningful groupings of information, which are then used as input for the formulator, to facilitate an orderly way to describe complex scenes.
As further evidence of this interpretation, the pre-speech latencies and the eye-voice span estimations would also suggest a more incremental approach to describing these networks. Previous production literature has suggested that speakers take about 2 s before beginning to speak when planning single-sentence event descriptions (e.g., Griffin & Bock, 2000; Gleitman et al., 2007) as well as descriptions that span across multiple utterances (e.g., Ferreira & Rehrig, 2019; Henderson et al., 2018). These reported latencies far exceed the interval of time suggested for gist extraction, which occurs within the first few hundred milliseconds of viewing (Castelhano & Henderson, 2007, 2008), which may suggest that some degree of linguistic planning is occurring prior to initiation. The apprehension times reported in this task are consistent with previous work: before beginning a description of a network, speakers took about 1600 ms. During this interval, speakers were generating some form of linearization plan, perhaps the time in which perceptual groupings of information were being chunked as input for the formulator. As with single-sentence and discourse-level planning, speakers make use of an interval of time prior to the onset of their description to generate a linguistic plan for, at least a portion of, the intended utterance.
Additionally, during this time, subjects also made about four fixations, some to the left side of the network, some to the right side of the network, none of which predicted the order of mention. Neither the proportion of fixations made to the left or the right side of the network, or the location of the final fixation prior to speech onset, predicted the first-mentioned direction. Given that participants were provided with pre-training on color terms prior to the experimental portion of the task, as well as three practice trials, it may be that speakers use the pre-speech interval to evaluate the perceptual features of the network (i.e., which side is shorter), and generate a higher-level linearization plan. This may also account for why we observed shorter intervals between fixated locations and subsequent mentions. Given that speakers typically fixate referents shortly before mentioning them (e.g., Coco & Keller, 2012; Griffin & Bock, 2000; Gleitman et al., 2007), the shorter eye-voice spans observed in the current study may reflect the reduced demands of lexical selection: participants were pre-trained on the color labels, thereby eliminating the need to retrieve or select labels at the time of speaking.
Conclusion
In this study, subjects described simple networks with interconnected colored dots while their eye movements were recorded. Our goal was to (1) investigate macroplanning in discourse-level planning and production and to (2) revisit the minimal-load principle and examine which network properties drive linearization using data from both speech and eye movements. And so we return to the question(s): how much of an utterance is prepared prior to beginning to speak and to what extent were planning and execution interleaved. When generating a plan that extends across multiple utterances, speakers must choose where to begin, and determine the items that they wish to include in their intended utterance. When describing spatial networks, the production plan that reduces the memory load for the speaker is driven by length. Our findings suggest that speakers prioritize the ‘easy’ side of the network by beginning with the branch that is shorter in length (or overall duration). The finding that prespeech locations did not predict order of mention is in alignment with an “apprehension” phase (Griffin & Bock, 2000) being used to conceptualize all, or part, of the stimulus. Rather than assigning grammatical roles, speakers instead evaluate some perceptual feature(s) of the network they are to describe and generate a higher-order linearization plan that is driven by these features. This strategy allows a speaker to begin with the easier (shorter) material, requires less time to plan prior to beginning their description, and allows for the interleaving of articulation and description planning for the remaining items. Overall, our findings suggest that speakers relied on perceptually based macroplans that prioritized shorter branches but executed descriptions incrementally, demonstrating both macroplanning and incrementalism in discourse-level production.
Funding
This work was funded by the National Institutes of Health R01HD100516 awarded to Fernanda Ferreira.
Appendix
| white | white | yellow | yellow | orange | orange |
|---|---|---|---|---|---|
| pearl | white | canary | yellow | tangerine | orange |
| alabaster | white | gold | yellow | marigold | orange |
| snow | white | daffodil | yellow | cider | orange |
| ivory | white | flaxen | yellow | rust | orange |
| cream | white | butter | yellow | ginger | orange |
| eggshell | white | lemon | yellow | tiger | orange |
| cotton | white | mustard | yellow | fire | orange |
| chiffon | white | corn | yellow | bronze | orange |
| salt | white | medallion | yellow | cantaloupe | orange |
| lace | white | dandelion | yellow | apricot | orange |
| coconut | white | fire | yellow | clay | orange |
| linen | white | bumblebee | yellow | honey | orange |
| bone | white | banana | yellow | carrot | orange |
| daisy | white | butterscotch | yellow | squash | orange |
| powder | white | dijon | yellow | spice | orange |
| frost | white | honey | yellow | marmalade | orange |
| porcelain | white | blonde | yellow | amber | orange |
| parchment | white | pineapple | yellow | sandstone | orange |
| rice | white | tuscan sun | yellow | yam | orange |
| red | red | pink | pink | purple | purple |
| cherry | red | rose | pink | mauve | purple |
| rose | red | fuchsia | pink | violet | purple |
| jam | red | punch | pink | boysenberry | purple |
| merlot | red | blush | pink | lavender | purple |
| garnet | red | watermelon | pink | plum | purple |
| crimson | red | flamingo | pink | magenta | purple |
| ruby | red | rouge | pink | lilac | purple |
| scarlet | red | salmon | pink | grape | purple |
| wine | red | coral | pink | periwinkle | purple |
| brick | red | peach | pink | sangria | purple |
| apple | red | strawberry | pink | eggplant | purple |
| mahogany | red | rosewood | pink | jam | purple |
| blood | red | lemonade | pink | iris | purple |
| sangria | red | taffy | pink | heather | purple |
| berry | red | bubblegum | pink | amethyst | purple |
| currant | red | ballet slipper | pink | raisin | purple |
| blush | red | crepe | pink | orchid | purple |
| candy | red | magenta | pink | mulberry | purple |
| lipstick | red | hot pink | pink | wine | purple |
| blue | blue | green | green | grey | grey |
| slate | blue | chartreuse | green | shadow | grey |
| sky | blue | juniper | green | graphite | grey |
| navy | blue | sage | green | iron | grey |
| indigo | blue | lime | green | pewter | grey |
| cobalt | blue | fern | green | cloud | grey |
| teal | blue | olive | green | silver | grey |
| ocean | blue | emerald | green | smoke | grey |
| peacock | blue | pear | green | slate | grey |
| azure | blue | moss | green | anchor | grey |
| cerulean | blue | shamrock | green | ash | grey |
| lapis | blue | seafoam | green | porpoise | grey |
| spruce | blue | pine | green | dove | grey |
| stone | blue | parakeet | green | fog | grey |
| aegean | blue | mint | green | flint | grey |
| berry | blue | seaweed | green | charcoal | grey |
| denim | blue | pickle | green | pebble | grey |
| admiral | blue | pistachio | green | lead | grey |
| sapphire | blue | basil | green | coin | grey |
| arctic | blue | crocodile | green | fossil | grey |
| black | black | midnight | black | onyx | black |
| ebony | black | ink | black | pitch | black |
| crow | black | raven | black | soot | black |
| charcoal | black | oil | black | sable | black |
| jet black | black | jade | black | obsidian | black |
| coal | black | spider | black | jade | black |
| metal | black | leather | black |
Footnotes
Conflicts of Interest We have no conflicts of interest to disclose.
Consent to participate Informed consent was obtained from all individual participants included in the study.
Data for 60 of the subjects were collected pre-pandemic. Due to some verbal recording failures and excessive eye-tracking artifacts for a number of subjects, a second group of subjects was run “post” pandemic (i.e., return to in-person data collection). Data were collected until the post-pandemic group matched the post-exclusion pre-pandemic group. Though we do not anticipate a difference between the two groups, membership will be included in our analyses. All data can be found on the OSF: https://osf.io/m7uxd/.
Data Availability
This study was not preregistered. Experimental data, materials, and code are available at https://osf.io/s3j7f/.
Code availability
This study was not preregistered. Experimental data, materials, and code are available at https://osf.io/s3j7f/.
References
- Allum PH, & Wheeldon LR (2007). Planning scope in spoken sentence production: the role of grammatical units. Journal of Experimental Psychology: Learning, Memory, and Cognition, 33(4), 791. [DOI] [PubMed] [Google Scholar]
- Bain M, Huh J, Han T, & Zisserman A (2023). Whisperx: Time-accurate speech transcription of long-form audio. arXiv:2303.00747. [Google Scholar]
- Barr DJ, Levy R, Scheepers C, & Tily HJ (2013). Random effects structure for confirmatory hypothesis testing: Keep it maximal. Journal of Memory and Language, 68(3), 255–278. [Google Scholar]
- Bates D, Mächler M, Bolker BM, & Walker SC (2015). Fitting linear mixed-effects models using lme4. Journal of Statistical Software, 67(1), 1–48. 10.18637/jss.v067.i01 [DOI] [Google Scholar]
- Berlin B, & Kay P (1991). Basic color terms: Their universality and evolution. Oakland: University of California Press. [Google Scholar]
- Brehm L, & Alday PM (2022). Contrast coding choices in a decade of mixed models. Journal of Memory and Language, 125, Article 104334. [Google Scholar]
- Bock JK (1977). The effect of a pragmatic presupposition on syntactic structure in question answering. Journal of Verbal Learning and Verbal Behavior, 16(6), 723–734. [Google Scholar]
- Bock JK, & Irwin DE (1980). Syntactic effects of information availability in sentence production. Journal of verbal learning and verbal behavior, 19(4), 467–484. [Google Scholar]
- Bock JK (1982). Toward a cognitive psychology of syntax: Information processing contributions to sentence formulation. Psychological Review, 89(1), 1. [Google Scholar]
- Bock JK (1986). Meaning, sound, and syntax: lexical priming in sentence production. Journal of Experimental Psychology: Learning, Memory, and Cognition, 12(4), 575. [Google Scholar]
- Bock K, Irwin DE, Davidson DJ, & Levelt WJ (2003). Minding the clock. Journal of Memory and Language, 48(4), 653–685. 10.1016/s0749-596x(03)00007-x [DOI] [Google Scholar]
- Bock K, & Levelt WJ (1994). Language production: Grammatical encoding. In Handbook of Psycholinguistics (pp. 945–984): Academic Press. [Google Scholar]
- Bock K, Loebell H, & Morey R (1992). From conceptual roles to structural relations: bridging the syntactic cleft. Psychological Review, 99(1), 150. [DOI] [PubMed] [Google Scholar]
- Bock JK, & Warren RK (1985). Conceptual accessibility and syntactic structure in sentence formulation. Cognition, 21(1), 47–67. [DOI] [PubMed] [Google Scholar]
- Butterworth B (1980). Evidence from pauses in speech. Language production, 1, 155–176. [Google Scholar]
- Castelhano MS, & Henderson JM (2007). Initial scene representations facilitate eye movement guidance in visual search. Journal of Experimental Psychology: Human Perception and Performance, 33(4), 753. [DOI] [PubMed] [Google Scholar]
- Castelhano MS, & Henderson JM (2008). The influence of color on the perception of scene gist. Journal of Experimental Psychology: Human Perception and Performance, 34(3), 660. [DOI] [PubMed] [Google Scholar]
- Christianson K, & Ferreira F (2005). Conceptual accessibility and sentence production in a free word order language (Odawa). Cognition, 98(2), 105–135. [DOI] [PubMed] [Google Scholar]
- Clark HH (1992). Arenas of language use. Chicago: University of Chicago Press. [Google Scholar]
- Clark HH, & Clark EV (1977). Psychology and language. [Google Scholar]
- Clark HH, & Haviland SE (1974). Psychological processes as linguistic explanation. Explaining Linguistic Phenomena, 91–124. [Google Scholar]
- Coco M, & Keller F (2009). The impact of visual information on reference assignment in sentence production. In Proceedings of the Annual Meeting of the Cognitive Science Society (Vol. 31, No. 31). [Google Scholar]
- Coco MI, & Keller F (2012). Scan patterns predict sentence production in the cross-modal processing of visual scenes. Cognitive Science, 36(7), 1204–1223. [DOI] [PubMed] [Google Scholar]
- Ehrich V, & Koster C (1983). Discourse organization and sentence form: The structure of room descriptions in Dutch. Discourse Processes, 6(2), 169–195. [Google Scholar]
- Ferreira F (1994). Choice of passive voice is affected by verb type and animacy. Journal of Memory and Language, 33(6), 715–736. [Google Scholar]
- Ferreira F, & Barker M (2024). Perceptual clauses as units of production in visual descriptions. Topics in Cognitive Science. [Google Scholar]
- Ferreira F, & Henderson JM (1998). Linearization strategies during language production. Memory & Cognition, 26(1), 88–96. [DOI] [PubMed] [Google Scholar]
- Ferreira F, & Rehrig G (2019). Linearisation during language production: evidence from scene meaning and saliency maps. Language, Cognition and Neuroscience, 34(9), 1129–1139. [Google Scholar]
- Ferreira F, & Swets B (2002). How incremental is language production? Evidence from the production of utterances requiring the computation of arithmetic sums. Journal of Memory and Language, 46(1), 57–84. [Google Scholar]
- Ferreira VS, & Yoshita H (2003). Given-new ordering effects on the production of scrambled sentences in Japanese. Journal of Psycholinguistic Research, 32, 669–692. [DOI] [PubMed] [Google Scholar]
- Garrett MF (1975). The analysis of sentence production. In: Psychology of Learning and Motivation (Vol. 9, pp. 133–177). Academic Press. [Google Scholar]
- Gleitman LR, January D, Nappa R, & Trueswell JC (2007). On the give and take between event apprehension and utterance formulation. Journal of Memory and Language, 57(4), 544–569. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Griffin ZM (2001). Gaze durations during speech reflect word selection and phonological encoding. Cognition, 82(1), B1–B14. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Griffin ZM (2003). A reversed word length effect in coordinating the preparation and articulation of words in speaking. Psychonomic Bulletin & Review, 10(3), 603–609. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Griffin ZM, & Bock K (1998). Constraint, word frequency, and the relationship between lexical processing levels in spoken word production. Journal of Memory and Language, 38(3), 313–338. [Google Scholar]
- Griffin ZM, & Bock K (2000). What the eyes say about speaking. Psychological Science, 11(4), 274–279. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Haviland SE, & Clark HH (1974). What’s new? Acquiring new information as a process in comprehension. Journal of Verbal Learning and Verbal Behavior, 13(5), 512–521. [Google Scholar]
- Henderson JM, Hayes TR, Rehrig G, & Ferreira F (2018). Meaning guides attention during real-world scene description. Scientific Reports, 8(1), 13504. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jescheniak JD, & Levelt WJ (1994). Word frequency effects in speech production: Retrieval of syntactic information and of phonological form. Journal of Experimental Psychology: Learning, Memory, and Cognition, 20(4), 824. [Google Scholar]
- Konopka AE, & Meyer AS (2014). Priming sentence planning. Cognitive Psychology, 73, 1–40. [DOI] [PubMed] [Google Scholar]
- Kuchinsky SE, Bock K, & Irwin DE (2011). Reversing the hands of time: changing the mapping from seeing to saying. Journal of Experimental Psychology: Learning, Memory, and Cognition, 37(3), 748. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Levelt WJ (1981). The speaker’s linearization problem. Philosophical Transactions of the Royal Society of London. B, Biological Sciences, 295(1077), 305–315. [Google Scholar]
- Levelt WJ (1982). Linearization in describing spatial networks. Processes, beliefs, and questions: Essays on formal semantics of natural language and natural language processing, 199–220. [Google Scholar]
- Levelt WJ (1989). Speaking: From Intention to Articulation. Cambridge, MA: MIT Press. [Google Scholar]
- Levelt WJ, & Maasen B (1981). Lexical search and order of mention in sentence production. Crossing the boundaries in linguistics: Studies presented to Manfred Bierwisch (pp. 221–252). Dordrecht: Springer, Netherlands. [Google Scholar]
- Levelt WJ, & Meyer AS (2000). Word for word: Multiple lexical access in speech production. European Journal of Cognitive Psychology, 12(4), 433–452. [Google Scholar]
- Levelt WJ, Roelofs A, & Meyer AS (1999). A theory of lexical access in speech production. Behavioral and Brain Sciences, 22(1), 1–38. [DOI] [PubMed] [Google Scholar]
- Linde C, & Labov W (1975). Spatial networks as a site for the study of language and thought. Language, 924–939. [Google Scholar]
- McAuliffe M, Socolof M, Mihuc S, Wagner M, & Sonderegger M (2017). Montreal forced aligner: Trainable text-speech alignment using kaldi. In: Interspeech, 498–502. [Google Scholar]
- MacDonald MC (2013). How language production shapes language form and comprehension. Frontiers in Psychology, 4, 40296. [Google Scholar]
- McDonald JL, Bock K, & Kelly MH (1993). Word and world order: Semantic, phonological, and metrical determinants of serial position. Cognitive Psychology, 25(2), 188–230. [DOI] [PubMed] [Google Scholar]
- R Core Team. (2018). R: A language and environment for statistical computing. R Foundation for Statistical Computing. Vienna, Austria. Retrieved from https://www.Rproject.org/. [Google Scholar]
- Schank R, & Abelson R (1977). Scripts, plans, goals and understanding. Hillsdale: Erlbaum. [Google Scholar]
- Shanon B (1984). Room descriptions. Discourse Processes, 7(3), 225–255. [Google Scholar]
- Smith M, & Wheeldon L (1999). High level processing scope in spoken sentence production. Cognition, 73(3), 205–246. [DOI] [PubMed] [Google Scholar]
- Swets B, Fuchs S, Krivokapić J, & Petrone C (2021). A cross-linguistic study of individual differences in speech planning. Frontiers in Psychology, 12, Article 655516. [Google Scholar]
- Ullmer-Ehrich V (1982). The structure of living space descriptions. Speech, Place, and Action, 219–249. [Google Scholar]
- Wagner V, Jescheniak JD, & Schriefers H (2010). On the flexibility of grammatical advance planning during sentence production: Effects of cognitive load on multiple lexical access. Journal of Experimental Psychology: Learning, Memory, and Cognition, 36(2), 423. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
This study was not preregistered. Experimental data, materials, and code are available at https://osf.io/s3j7f/.
This study was not preregistered. Experimental data, materials, and code are available at https://osf.io/s3j7f/.
