Abstract
The Parallel Architecture is a conception of the organization of the mental representations involved in language and of the role of language in the mind as a whole. Its basic premise is that linguistic representations draw on three independent generative systems—phonological, syntactic, and semantic structures—plus a system of interface links by which they communicate with each other. In particular, words serve as partial interface links that govern the way they compose into novel sentences.
It is shown that this architecture also provides a natural way to account for our ability to talk about what we see: semantic structure in language has to communicate via interface links with a level of spatial representation that encodes understanding of the physical world. It is suggested that such configurations of independent but linked representations are a widespread feature of cognition.
Keywords: Parallel Architecture, Word, Semantics, Syntax, Phonology, Vision, Music, Modularity
Short abstract
The Parallel Architecture proposes that semantic, syntactic, and phonological structure are independent systems in the language faculty, connected by interface links. This organization can be extended to the relation of language and vision, such that we can talk about what we see.
1.
The Parallel Architecture is a unified theory of the mental representations involved in the language faculty and their interactions with the mind as a whole. I take “mental representations” to be the “data structures” in the brain that collectively create our understanding and experience of the world. Fundamental to this approach is that the brain has to construct the world as we understand and experience it.1 Our repertoire of mental representations creates the dimensions in which this construction takes place. Different organisms presumably have different repertoires of mental representations, affording them differently constructed “worlds.”
The mental representations that support language fall into three basic levels of representation: phonological or sound structure, syntactic or grammatical structure (including morphology, the grammar of words), and semantic or conceptual structure. These levels are related to each other by interfaces, indicated in Fig. 1 by the double‐headed arrows.
Fig. 1.

Outline of the Parallel Architecture.
Under this conception, a sentence is a triple of well‐formed phonological, syntactic, and conceptual structures, with well‐formed links through the interfaces. Three major components have been developed in detail: Simpler Syntax (Culicover, 2013; Culicover & Jackendoff, 2005), Relational Morphology (Booij, 2010; Jackendoff & Audring, 2020), and Conceptual Semantics (Jackendoff, 1983, 1990, 2002, 2007).
Linguistic theory is first and foremost about the structure and organization of the mental representations that support language. But in addition, a theory of the mental representations for language should embed naturally into a psycholinguistic theory, which deals with how language production and language comprehension make use of the mental representations for language; and it should make connections with a theory of language acquisition—how language learners construct the mental representations for language in the context of their environment. The parallel Architecture makes natural connections with theories of processing and acquisition (Jackendoff & Audring, 2020, chapter 7; Huettig, Audring, & Jackendoff, 2022).
Ideally, all of this should be supported by a theory of neural computation: how the neurons actually encode mental representations and the processes that make use of them. By neural computation, I mean not just where in the brain some linguistic process takes place and its time course, but really how the neurons do it. How do the speech sounds /b/ and /p/ differ in the brain? What makes one neural assembly the knowledge of the word dog and another the word cat? How are dog and hot dog partly alike in the brain? And so on. I do not think we know how to answer these questions yet. But I do think linguistic theory is important to neuroscience, because it makes clear what an account of the neural instantiation of language is ultimately responsible for.
To be more concrete, consider the word cat. The word is encoded in all three levels of representation. Every linguistic theory recognizes these three components in some notation or another.
(1) Semantics: CAT1
Syntax: N1
Phonology: /kæt/1
The semantics of the word is the concept of cats—whatever it may be. (We do not understand conceptual structures very well yet, but as a stand‐in, they are customarily notated in capitals.) In syntax, the word cat is encoded as a noun, and in phonology, it is encoded as the pronunciation /kæt/.
What makes (1) the word cat is the fact that the three structures are connected by explicit interface links, notated by the subscript 1 on each structure.2 The subscripts can be thought of as representing the endpoints of association lines between the three structures. They make the word function as part of the interfaces, that is, the double‐headed arrows in Fig. 1. By virtue of these links, hearing the sound /kæt/ activates the phonological structure, and the cosubscript then leads the hearer to activate the syntax and the meaning. And when one wants to express the concept of cats, one activates the semantic structure CAT, and that passes activation through the interfaces to the pronunciation. In other words, the associations between levels are bidirectional.
Central to the Parallel Architecture is that each level of representation has a different repertoire of units. The phonological structure of a word is made up of a sequence of speech sounds (or phonemes), which are combined to form syllables; a word (usually) consists of one or more syllables. Speech sounds in turn have internal structure, coded in terms of distinctive features, illustrated in (2).
-
(2)
/t/ = [+consonant, ‐continuant, ‐voiced, …]
The phoneme /t/ is a stopped sound, contrasting with /s/, which is a continuant; it contrasts with /d/ in being unvoiced. These three consonants contrast with vowels such as /æ/ and /u/, and so on. Sign language gestures show similar organization to some degree (Fischer & Siple, 1990).
The syntactic structure of words encodes parts of speech such as noun and verb, which govern how words combine with each other into phrases and sentences. Syntactic structure also encodes inflection, for instance plural and past tense markers in English and grammatical gender in languages such as French. Inside words, morphosyntax includes compounding, in which a noun like football is composed of two smaller nouns, and derivational morphology, in which a word like baker is composed of the verb bake plus a suffix.
The conceptual structure of words is complex, and as mentioned earlier, there is still a lot to be learned about it. It is fairly clear, however, that its basic units are things like conceptualized objects (inanimate and animate), events, spatial layout, and more abstract concepts, such as numerosity, belief, intention, obligation, and value (Jackendoff, 2007). Conceptual Structure has to form a basis for reference, for instance the ability to label something perceived in the real‐world context as a cat. Conceptual Structure also has to form a basis for inference: if one identifies something as a cat, one can infer that it is an animate physical object of a certain size and shape, that it is a mammal, and that it is a domestic animal. Such inferences cannot be encoded in a feature system parallel to that in phonology. For instance, there surely is no such eccentric feature as [+domestic animal]. Rather, identifying an animal as domestic has to draw on world knowledge of social practice. More generally, it turns out to be very difficult—maybe impossible—to segregate knowledge of word meanings from knowledge of the world (Bolinger, 1965; Jackendoff, 1983, 2002; Lakoff, 1987).
Now, consider the mental representation of whole sentences—the structures constructed online by combining words. The Parallel Architecture again segregates semantic, syntactic, and phonological structures into three independent but linked levels. Fig. 2 is a sketch of the little sentence Jenny bought a bike.
Fig. 2.

The structure of Jenny bought a bike.
The semantics in Fig. 2 says that in the past there was an event of buying, involving two characters, an Agent—the buyer, Jenny, and the Patient—the thing being bought, a bike. Because this is an event of buying, it necessarily has two further characters that this sentence does not mention: the seller and the money being exchanged—where the concept of money is another complicated thing that is embedded in social practice (Searle, 1995).
Turning to the syntax, it is notated in Fig. 2 as a standard simple parse tree, in which the sentence is made up of a noun phrase and a verb phrase, the verb phrase is made up of a verb in the past tense and another noun phrase, which in turn is made up of a determiner and a noun.
The phonology encompasses three sublevels or tiers. Down the middle is the segmental/syllabic tier, which encodes the sequence of speech sounds and their organization into syllables. Above it is the stress grid, which assigns a weight to each syllable. A syllable with one x above it is unstressed. If there are two xs, the syllable is stressed, and the pile of three xs indicates that the syllable in question receives main stress in the sentence.
Below the segmental/syllabic tier is the intonation contour for the sentence, the tune to which it is spoken. The pronunciation notated in Fig. 2 has two prosodic units, set off by brackets, possibly with a slight pause between them. Each of them is pronounced with a sequence of high and low pitches associated with the syllables of the sentence.
But there is more to the structure of the sentence. As in the word cat, the levels of representation are linked, in the manner indicated by the subscripts. For instance, JENNY in the semantics has subscript 2, which links it to the first noun phrase in the syntax and the first two syllables in the phonology. Similarly, the term BIKE in the semantics is linked by subscript 8 to the final noun in the syntax and the final syllable in the phonology. Crucially, there is no single place where the words are fitted into the sentence (so‐called “lexical insertion” or a “lexical level” of processing). Rather, the words are spread across the three levels, connected by interface links; they serve as part of the means by which sounds are related to meanings.
An important feature of the Parallel Architecture is that the levels are not matched one to one. Consider the verb in the syntax. The upper V has subscript 5, which links it to the word bought in the phonology. But morphosyntactically, this upper V actually consists of two parts. One part is the verb stem, which links to the meaning BUY in the semantics, via subscript 4. The other part is the past tense inflection, which links to the time expression PAST in semantics, via subscript 3. In other words, pieces of phonology need not correspond directly to pieces of the semantics, and vice versa. This is what makes buy an irregular verb. This sort of situation is ubiquitous, and it is one of the phenomena that make language complex.
To sum up where we are so far: The overall premise of the Parallel Architecture is that the mental representation of language is made up of independent kinds of information—the semantics, the syntax, and the three tiers of phonology. They align (or misalign, as the case may be) with each other by way of interface links, both in words and in sentences. It is because of this quasi‐independence of the levels of representation that the theory is called the Parallel Architecture.3
The Parallel Architecture contrasts with the outlook of standard generative grammar (Chomsky, 1965, 1981, 1995; Hornstein, 2018), in which the fundamental creative source of language is syntax, and from which semantics and phonology are derived. This view of language may have been attractive early on in modern linguistics. But from the perspective of the Parallel Architecture, it has proven to be a mistake—a mistake that has driven a wedge between linguistics and the rest of cognitive science (see especially Culicover & Jackendoff, 2005).
However, rather than dwelling on the differences between the Parallel Architecture and standard generative grammar, I would like to turn to a very important but (I think) largely neglected issue: How can we talk about what we see (Macnamara, 1978; Miller & Johnson‐Laird, 1975)? For instance, we can point to something in the environment and say That's a cat. How does the meaning of the word cat allow us to do this? And what does the Parallel Architecture offer that helps answer this question?
A preliminary answer is that there has to be some sort of informational conduit between language and the visual system. This sets a challenge for the theory of the visual system as well as for linguistic theory: What is the nature of this conduit, and exactly what does it connect?
Here are some considerations. Language refers to cats, not in terms of collections of pixels, but in terms of object‐centered three‐dimensional descriptions, independent of distance, point of view, and lighting—all the classic visual constancies that go into the understanding of objects (Marr, 1982), including understanding that objects have unseen backs and may even be hollow, like balloons and closed containers. The spatial context may occlude parts of an object or even a whole object, like the picture of a cat behind a bookcase in Fig. 3. If you know it is there, even though you cannot see it, you infer what you would see from a different point of view. And the object is understood as the same whether you are seeing it or not.
Fig. 3.

A bookcase with a cat behind it.
This sort of understanding requires a level of representation that abstracts away from momentary appearance to a more “objective” encoding of the way the world is taken to be. I will call this level Spatial Structure. For a first approximation, it might be thought of as the highest level of representation in the visual system. There are likely other levels of visual representation between Spatial Structure and the rudimentary information coded in V1, perhaps roughly along the lines of Marr's (1982) primal sketch and 2½D sketch, perhaps dividing the work up between complementary streams such as the what‐system and the where‐system (Ungerleider & Mishkin, 1982). Face recognition (Kanwisher, McDermott, & Chun, 1997) might be thought of as an extra tier of Spatial information.
Further consideration suggests that Spatial Structure cannot be just a visual representation. The size and shape of objects and their spatial layout can be determined haptically, that is, through the sense of touch. In addition, information about the spatial configuration of one's body comes from proprioception (Lackner, 1988; Lackner & Dizio, 2000). All three of these—vision, hapsis, and proprioception—have to be correlated with each other in order to understand what is going on in physical space. This job belongs to Spatial Structure. Moreover, Spatial Structure is not just for perceiving: it also has to be used to plan action in the world.
How does Spatial Structure connect with language? Basically, there has to be a partial correspondence between Spatial Structure and Conceptual Structure, mediated by interface links. Hence, for example, the mental representation of the word cat includes not only linked phonological, syntactic, and conceptual structures, but also, linked to them, a piece of Spatial Structure that encodes (to put it roughly) what cats look like and how they behave. This connection is what enables us to identify things in the physical environment that we call cats, and more generally, to talk about what we see.
What sorts of things do we talk about that have to be present in Spatial Structure? We identify parts of objects, such as the cat's head, eyes, ears, legs, and tail. We perceive what might be called “negative objects” such as holes, cracks, dents, mouths, and nostrils. We identify actions like running, jumping, throwing, climbing, falling, and sliding.4 We recognize more abstract sorts of entities such as ends, as in the end of a rope, a table, a road, or even a lecture or movie. In short, there are hundreds of things we have words for that Spatial Representation has to identify.
Fig. 4 sketches the architecture of the whole system. On the top is the Parallel Architecture for language: phonology, syntax, and conceptual structure. In language comprehension, the phonological end is linked to auditory signals; in language production, it is linked to vocal tract instructions. The conceptual end is linked to Spatial Structure, which is in turn connected to all the perceptual systems and to action planning. In other words, as a whole, this is a collection of distinct levels of representation connected by interface links.
Fig. 4.

Language embedded in the mind.
If phonology and syntax are deleted from Fig. 4, the result is a plausible architecture for nonhuman primate minds. Apes likely understand the physical world in much the same way as we do, making use of visual, haptic, and proprioceptive information, and creating action plans. Another conceivable route to spatial representation is echolocation in bats and dolphins. In addition, the social abilities of nonhuman primates suggest that they have a counterpart of conceptual structure, though not as rich as ours—and of course, they cannot talk about it (Cheney & Seyfarth, 2007).
A quite different mental system that lends itself to the Parallel Architecture is music. Fig. 5 is a familiar fragment of music by the Beatles.
Fig. 5.

The Parallel Architecture in music.
Very briefly, following Lerdahl and Jackendoff (1983): Fig. 5 represents four kinds of musical structure. Down the middle is the sequence of notes, with pitch and duration, plus interface links to the words of the song. Below the sequence of notes is grouping structure, the hierarchical articulation of the music in terms of motives and phrases. Above the notes is a configuration of x’s that represents the hierarchical metrical structure of strong and weak beats, rather like the stress grid in phonology (cf. Fig. 2). Above that is a tree structure that represents the relative structural importance of the notes and the resulting patterns of tension and relaxation. In short, music too is organized in terms of independent levels of representation that collaborate through interfaces.
Other systems of cognition also lend themselves to characterization in Parallel Architecture terms. Multimodal expression links linguistic meaning with visual images to form a combined understanding (Cohn & Schilperoord, 2022). Musical performance links visual understanding of the printed notes to musical structure of the sort in Fig. 5, which in turn is linked to motor instructions for singing or playing the piece. Farther afield, games link physical actions (such as kicking a ball into a goal) to an abstract “game tier” that encodes the significance of these actions within the game (such as adding one to the score of the team that the kicker belongs to); learning the game amounts to establishing such links.
Seeing cognition from this perspective sets a major challenge for cognitive neuroscience: How does the brain create and store all these structures and the links between them? Which of them are a product of cognitive evolution, for example, Spatial Structure—and which are cultural innovations, for example, the rules of soccer? And of course, the battles over where language belongs in this dichotomy are legion. The Parallel Architecture provides a perspective for addressing these deeper questions. It is my hope that this perspective will offer fascinating opportunities for collaboration across disciplinary boundaries.
This article is part of the topic “Parallelism in the Architecture of Language,” Giosuè Baggio, Neil Cohn and Eva Wittenberg (Topic Editors).
Notes
This view, taken for granted in contemporary cognitive science, goes back at least to Helmholtz, the gestalt psychologists (e.g., Koffka, 1935; Wertheimer, 1923), and perhaps even Plato and Kant.
The linkage among the three levels is often notated with large square brackets rather than subscripts (e.g., HPSG: Pollard & Sag, 1994).
One might consider each level of representation as an independent module in Fodor's (1963) sense, in that they are strongly domain specific. On the other hand, Fodor offers no way for modules to communicate with one another. In the Parallel Architecture, the interface links serve this essential function.
Action verbs and other concepts that have a temporal dimension pose an interesting challenge to theories of Spatial Structure: how are they encoded in long‐term memory? Although the identification of an act of running—or of a familiar tune—unfolds over time, long‐term memory representation cannot be running the action in a continuous loop: the memory itself basically has to be static (Lashley, 1951).
References
- Bolinger, D. (1965). The atomization of meaning. Language, 41, 555–573. [Google Scholar]
- Booij, G. (2010). Constructional morphology. Oxford: Oxford University Press. [Google Scholar]
- Cheney, D. , & Seyfarth, R. (2007). Baboon metaphysics. Chicago, IL: University of Chicago Press. [Google Scholar]
- Chomsky, N. (1965). Aspects of the theory of syntax. Cambridge, MA: MIT Press. [Google Scholar]
- Chomsky, N. (1981). Lectures on government and binding. Dordrecht: Foris. [Google Scholar]
- Chomsky, N. (1995). The Minimalist Program. Cambridge, MA: MIT Press. [Google Scholar]
- Cohn, N. , & Schilperoord, J. (2022). Reimagining language. Cognitive Science, 46, issue 7. 10.1111/cogs.13174. [DOI] [PubMed] [Google Scholar]
- Culicover, P. W. (2013). Explaining syntax. Oxford: Oxford University Press. [Google Scholar]
- Culicover, P. W. , & Jackendoff, R. (2005). Simpler syntax. Oxford: Oxford University Press. [Google Scholar]
- Fischer, S. D., & Siple, P. (1990). Theoretical issues in sign language research, Volume 1: Linguistics. Chicago, IL: University of Chicago Press. [Google Scholar]
- Fodor, J. A. (1963). The modularity of mind. Cambridge, MA: MIT Press. [Google Scholar]
- Helmholtz, H. v. (1909). Wissenschaftliche Abhandlungen, II .
- Hornstein, N. (2018). The Minimalist Program after 25 years. Annual Review of Linguistics, 4, 49–65. [Google Scholar]
- Huettig, F. , Audring, J. , & Jackendoff, R. (2022). A parallel architecture perspective on pre‐activation and prediction in language processing. Cognition, 224, 1– 15. [DOI] [PubMed] [Google Scholar]
- Jackendoff, R. (1983). Semantics and cognition. Cambridge, MA: MIT Press. [Google Scholar]
- Jackendoff, R. (1990). Semantic structures. Cambridge, MA: MIT Press. [Google Scholar]
- Jackendoff, R. (2002). Foundations of language. Oxford: Oxford University Press. [Google Scholar]
- Jackendoff, R. (2007). Language, consciousness, culture. Cambridge, MA: MIT Press. [Google Scholar]
- Jackendoff, R. , & Audring, J. (2020). The texture of the lexicon: Relational Morphology and the parallel architecture. Oxford: Oxford University Press. [Google Scholar]
- Kanwisher, N. , McDermott, J. , & Chun, M. M. (1997). The fusiform face area: A module in human extrastriate cortex specialized for face perception. Journal of Neuroscience, 17(11), 4302–4311. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Koffka, K. (1935). Principles of gestalt psychology. New York: Harcourt Brace & World. [Google Scholar]
- Lackner, J. (1988). Some proprioceptive influences on the perceptual representation of body shape and orientation. Brain, 111, 281–297. [DOI] [PubMed] [Google Scholar]
- Lackner, J. , & Dizio, P. (2000). Aspects of body self‐calibration. Trends in Cognitive Sciences, 4, 279–288. [DOI] [PubMed] [Google Scholar]
- Lakoff, G. (1987). Women, fire, and dangerous things. Chicago, IL: University of Chicago Press. [Google Scholar]
- Lashley, K. (1951). The problem of serial order in behavior. In Jeffress L. A. (Ed.), Cerebral mechanisms in behavior (pp. 112–136). New York: Wiley. [Google Scholar]
- Lerdahl, F. , & Jackendoff, R. (1983). A generative theory of tonal music. Cambridge, MA: MIT Press. [Google Scholar]
- Macnamara, J. (1978). How do we talk about what we see? Unpublished manuscript, McGill University. [Google Scholar]
- Marr, D. (1982). Vision. San Francisco, CA: Freeman. [Google Scholar]
- Miller, G. , & Johnson‐Laird, P. (1975). Language and perception. Cambridge, MA: Harvard University Press. [Google Scholar]
- Pollard, C. , & Sag, I. (1994). Head‐driven phrase structure grammar. Chicago, IL: University of Chicago Press. [Google Scholar]
- Searle, J. (1995). The construction of social reality. New York: Free Press. [Google Scholar]
- Ungerleider, L. , & Mishkin, M. (1982). Two cortical visual systems. In Ingle D. J., Goodale M. A., & Mansfield R. J. W. (Eds.), Analysis of visual behavior (pp. 549–586). Cambridge, MA: MIT Press. [Google Scholar]
- Wertheimer, M. (1923). Laws of organization in perceptual forms. In Ellis D. (Ed.), A source book of gestalt psychology (pp. 71–88). London: Routledge & Kegan Paul. [Google Scholar]
