Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Sep 4.
Published before final editing as: Curr Biol. 2026 Sep 2:S0960-9822(26)01026-2. doi: 10.1016/j.cub.2026.08.016

Compression as a neural mechanism of mnemonic chunking in prefrontal cortex

Feng-Kuei Chiang 1, Erin L Rich 1,*
PMCID: PMC13540421  NIHMSID: NIHMS2204603  PMID: 42685693

Summary:

Chunking is the mental process of grouping information into meaningful units and plays a central role in many cognitive processes such as working memory, yet the neural mechanisms that form chunks remain unknown. Here we investigated how chunking impacts representations of mnemonic information in the dorsolateral prefrontal cortex (dlPFC). Macaque monkeys performed a self-ordered visuospatial working memory task, in which spontaneous chunking could be behaviorally characterized. We found that chunking improved memory performance overall and produced lower accuracies and longer reaction times at chunk boundaries. In dlPFC, chunking systematically biased neural representations of spatial targets toward common chunk centers. At the same time, chunking reduced the dimensionality of population activity, indicating a compression of mnemonic information. Memory-guided behavior was also biased in the same direction as the neural distortions, suggesting that the compression mechanisms are lossy rather than lossless. Together, these results offer a mechanistic account of chunking as a form of lossy compression in neural coding that sacrifices detail within a chunk but improves overall memory retention.

Graphical Abstract

graphic file with name nihms-2204603-f0001.webp

eTOC Blurb

Chiang and Rich show that macaques spontaneously chunk spatial targets in a self-ordered working memory task. dlPFC neurons compress chunked targets into lower-dimensional representations that are shifted toward chunk centers. Memory-guided behavior reveals similar biases, suggesting a lossy mechanism that sacrifices detail within a chunk to improve overall memory performance.

Introduction

Working memory is the ability to maintain and manipulate information in our mind1 and has a capacity limit of about four separate items25. We overcome this constraint by enlisting mnemonic strategies such as chunking, which involves mentally organizing information into groups. These groups can be learned or spontaneous, and are often established only temporarily to accomplish a task in the moment68. Conceptually, chunking binds together bits of information so they can be remembered and manipulated more efficiently as a single unit. This simple trick allows us to perform cognitive tasks of extraordinary complexity, from language comprehension9 to playing chess10, yet the neural mechanisms that chunk information in working memory have remained unknown.

Some models propose that chunking is a form of data compression that systematically alters the contents of working memory8,1113. In digital media, lossy compression refers to methods that reduce file size by discarding data based on its redundancy or similarity to other information. A similar approach may be employed by resource-limited neural systems, such as those supporting working memory, to improve performance while reducing overall memory demands11,14. To investigate this possibility at the level of neural coding, we focused on the dlPFC. dlPFC is necessary for activities that rely on working memory, such as planning and monitoring self-organized behaviors1517, and dlPFC neurons are known to group items into learned categories1823, suggesting they are naturally suited to establish chunks. Moreover, lesions that include dlPFC produce chunking deficits17,24 and neuroimaging shows dlPFC activation during mnemonic chunking25,26, making this region an important target for understanding chunking mechanisms.

To investigate how neural activity forms chunks, we recorded large populations of dlPFC neurons in macaque monkeys, a species whose prefrontal cortex shares a high degree of neuroanatomical and functional similarity to humans27,28. We first demonstrate that monkeys spontaneously adopt chunking strategies that improve their performance on a visuospatial working memory task. In dlPFC, we found that chunking restructures the substrates of working memory into lower-dimensional representations that are shifted toward other items in the same chunk, consistent with chunk-based compression of information. Probe trials also revealed biases in memory-guided behavior, indicating that a full-fidelity representation could not be reconstituted from its compressed form. Together, these results are consistent with the idea that the brain implements chunking through lossy compression mechanisms11,14.

Results

Monkeys employ chunking strategies

In memory tasks, chunking is typically inferred when items from a list are recalled together6,7, and we used this concept of clustered retrieval from working memory to detect chunking behavior in monkeys. We trained two Rhesus macaques to perform a self-ordered target selection task by making saccades to select each of 8 identical visuospatial targets, returning their gaze to a central fixation point between selections (Figure 1A). Targets could be selected in any order, but reward was only delivered for visiting a new target on each selection. Since selecting a target did not change its appearance, monkeys had to use working memory to remember which targets they had visited and update that information with each new selection. Targets were arranged in pseudorandom spatial configurations around the screen center (Figure S1AB) and appeared in the same configuration for blocks of 40 trials. This allowed us to identify consistent patterns in selection order within each block. The requirement that all target selections begin from the screen center created a center-out pattern that ensured that the order in which the targets were selected did not change the total saccade distance (Figure 1B).

Figure 1. Monkeys demonstrate chunking behavior.

Figure 1.

(A) To initiate a trial, monkeys fixate a central location on the task screen. A configuration of identical targets is presented, and they select one by shifting their gaze to it and holding for 500 ms. If the target has not been selected on that trial, they earn a reward, then return gaze to central fixation to select the next target. Once all targets have been selected, there is an inter-trial interval (ITI) and a new trial starts. Reward contingencies are reset and target color changes to indicate a new trial. (B) Center-out eye trajectories on one example trial. (C) Transition probabilities, defined as the probability of selecting one target following another, across 40 trials of the configuration shown in (B). Line thickness is proportional to probability and arrows indicate direction. Targets B,H,C, and A were assigned to one chunk (red) and targets D, E, F, and G were assigned to the other (blue). (D) Centroids of all target chunks across sessions. Chunk 1 (red) indicates the targets that tend to be selected first. Monkey B chunked targets on the top and bottom of the screen, and Monkey A chunked targets on the upper left and lower right. (E) Both monkeys made fewer revisit errors on blocks with higher MIs. (F) Revisit rates were lower when selecting a target within the same chunk (blue), compared to crossing a chunk boundary (yellow). Revisit rates were normalized by the mean rate across trials with the same number of previously selected targets (Figure S2). Bars = mean of 90 blocks ± SEM. Monkey B: t178df = −6.09, p = 6.7×10−9. Monkey A: t178df = −4.67, p = 5.8×10−6. (G) Reaction times (RTs) when selecting a target within the same chunk (blue) or across the chunk boundary (yellow). Monkey B: t178df = −4.7, p = 4.6×10−6. Monkey A: t178df = −5.85, p = 2.3×10−8. RTs were normalized to the mean of selections with the same number of previously selected targets. See also Figure S1S2.

We found that monkeys frequently selected subsets of targets together, as if they were grouping the targets into 2 or 3 chunks (Figure S1). We used analytic tools from graph theory to quantify these selection patterns on a block-by-block basis and assign each target to a putative chunk (Figure 1C). The resulting patterns showed that the tendency to chunk particular targets was self-organized rather than driven by the specific target arrangements. Both subjects were tested on the same sequence of target configurations, but grouped the targets differently (Figure 1D).

To measure chunking strength, we calculated the modularity index (MI) for each block, which is a normalized measure of transition probabilities within versus across chunks29,30. Higher MI values correspond to more reliable chunking patterns (Figure S1). Both chunk membership and MI were determined only from the order in which targets were selected, but the results predicted other measures of behavior. The probability of revisiting a previously selected target was lower during blocks with higher MI (Pearson correlations: Monkey B: r = −0.43, p = 1.9×10−5, n = 90. Monkey A: r = −0.75, p = 1.6×10−17, n = 90), indicating that more reliable chunking improved working memory performance (Figure 1E, Figure S2). Within a block, monkeys also made slower and less accurate target selections when crossing a chunk boundary, compared to selecting targets within the same chunk (Figure 1FG, Figure S2), similar to patterns in human chunking3134. There was no relationship between MIs for the same configurations across the two subjects (Figure S1E), providing further evidence that specific target configurations were not more likely to be chunked. Taken together, monkeys demonstrated behavior patterns consistent with spontaneous, self-organized chunking.

Chunks have clustered representations

Working memory and cognitive strategies depend on the dlPFC1517,24, where neurons encode items held in working memory3539 and flexibly represent task-relevant rules and categories1823. When remembered items vary in spatial location, dlPFC neurons represent them in a visuospatial reference frame35,36,40. To understand how chunking affects these representations, we recorded neurons from dlPFC with chronically implanted microelectrode arrays (Figure S3AB) and used simultaneous population activity to decode coordinates of each target in visual space (see STAR Methods). We found that the location of the next target the monkey would select could be predicted, with some variability, during the fixation epoch between serial selections (Figure S3C).

We also found that these decoded locations tended to cluster together for targets in the same chunk, consistent with the idea that chunking partially merges neural representations (Figure 2A). To quantify this, we used a clustering algorithm that partitioned the decoded locations into groups, agnostic to the associated target or chunk. The clusters derived from neural representations had a high correspondence with behaviorally-defined chunks (Figure 2BD, Figure S4). Moreover, correspondence increased in higher MI blocks (Figure 2E) with either 2 chunks (Pearson correlations: Monkey B r = 0.34, p = 0.02, n = 43. Monkey A r = 0.38, p = 0.002, n = 67) or 3 chunks (Monkey B r = 0.43, p = 0.003, n = 47. Monkey A r = 0.55, p = 0.007, n = 23.) Therefore, more reliably clustered neural representations occurred with stronger chunking. This suggests that chunking distorts the underlying neural representations, but could also result from spatial clustering of the targets themselves. Therefore, we next assessed the accuracy of neural predictions on a target-by-target basis.

Figure 2. Decoded target locations cluster according to behaviorally-defined chunks.

Figure 2.

(A) Two example blocks with 2 and 3 putative chunks identified from behavior patterns. Points indicate the locations decoded on each target selection, color-coded according to the chunk membership defined behaviorally. Shades of red/blue/green correspond to different targets within a chunk. + indicates the center of the task screen, and ⊗ indicates the centroid of the targets assigned to each chunk. (B) Two example 2-chunk blocks with low (left) or high (right) MIs. Each dot is a decoded target location with face shading indicating group membership defined by k-means clustering of neural predictions of target locations, and edge color indicating chunk membership defined by behavioral patterns of target selections. Matching face and edge colors show correspondence between neurally-defined and behaviorally-defined groups. Solid green circles are actual target locations, + indicates screen center. (C) Same as B, but for two example 3-chunk blocks. (D) Average correspondence between behaviorally-defined and neurally-defined target groups. Neurally-defined groups are based on clustering of decoded target locations. The number of clusters was determined by the number of putative chunks identified from behavior patterns. Correspondence is the proportion in each cluster associated with a target in each behaviorally-defined chunk. (E) Stronger evidence of chunking in behavior correlated with higher correspondence between behaviorally-defined and neurally-defined chunks in both 2- (dark blue) and 3-chunk (light blue) blocks. See also Figures S3S4.

Chunking biases target representations

To determine whether neural representations are distorted by chunking, we tested for systematic biases in the decoded target locations. Specifically, if chunking merges representations together, we expect decoded locations to be shifted toward the centroid of a group of chunked targets. To assess this, we measured Euclidean distances from an actual target or a decoded location to the centroid of a group of chunked targets (Figure S5). Although decoded locations were consistently closer to chunk centers than the actual targets, this could occur if the locations were non-specifically shifted toward the screen center, for instance by decoder noise (Figure S5DF). Therefore, we focused our analyses on the degree to which the decoded locations were displaced in the direction of a chunk center, while controlling for their location relative to the screen center. To do this, we found the displacement of a decoded location from the actual target location, and defined the displacement direction relative to the centroid of a group of chunked targets in polar coordinates (Figure 3A). By defining shifts of a neural representation toward a chunk as a positive angle and away as a negative angle (Figure 3B), we found that representations were consistently displaced toward chunk centers (Circular medians tests compared to 0°, Monkey B n = 720, p = 4.1×10−3. Monkey A n = 720, p = 1.8×10−5. Figure 3C). This suggests a systematic bias in the neural representations of chunked memoranda.

Figure 3. Decoded target locations are shifted toward chunk centers.

Figure 3.

(A) An example block showing the centroids of decoded locations for each target (red/blue dots). Actual target locations are green circles, + = screen center, ⊗ = chunk center. (B) Schematics of the area outlined in (A), showing the angle between vectors from the actual target to screen center and actual target to predicted target. If predictions only shift toward screen center due to noise, this angle should vary symmetrically around 0. Angles toward a chunk center (dashed gray arrows) were labeled positive (first panel) and those away were labeled negative (second panel). (C) Circular histograms of angles shown in (B). Red bars indicate median angles. (D) Median angles increase in higher MI blocks. * = angles significantly greater than 0° (p ≤ 0.01) by circular medians tests. Error bars = SE calculated from circular standard deviation. (E) Schematic showing the angle (θ1) between the screen center (+), actual target (green), and decoded target (blue), and the angle (θ2) between screen center, actual target, and chunk centroid (⊗; gray). These angles were compared with circular correlations. (F) Displacements of a decoded location from the actual target (θ1) were correlated with the angle to the chunk center (θ2). Each point is the location of a chunk center relative to the actual target (green), and colored according to the angle to the predicted target location (θ1). dva = degrees of visual angle. See also Figure S3 and S5S7.

The magnitude of the bias we observed was small but increased in higher MI blocks (Pearson correlation of binned MI versus median angle, r = 0.88, p = 0.02. Figure 3D), showing that representations were more shifted with stronger chunking (Figure S6A). In addition, the relative location of the chunk center determined not only the direction but also the magnitude of the distortion. Small angles were more common when the chunk center was near the actual target location, whereas larger angular shifts occurred when the chunk center was farther away. This pattern was revealed by circular correlations (Figure 3EF). We performed the same analysis on other time epochs surrounding a target selection and found that neural representations were biased toward chunks in a consistent manner throughout the fixation period, when monkeys prepared their next target selection (Figure S6b). However, once the targets were visible and their gaze shifted to a target, the bias disappeared and there was a veridical representation of the target location. This coincided with an apparent change in the neural code approximately 200 ms after the targets appeared (Figure S3C), which likely reflects visuospatial information related to saccade execution.

During our main analysis window, we also confirmed that shifts in neural representations were specifically driven by the centroid of a group of chunked targets and not the location of the previous or next target in a selection sequence, even though these frequently belonged to the same chunk (Figure S6CD). In addition, we created “false chunks” by separating the same targets in each block into groups that were not actually chunked by the monkey. We confirmed that, unlike true chunk centers, the centroid of false chunks did not attract neural representations (Figure S6E). Finally, similar effects were observed across electrode arrays implanted in different subregions of dlPFC (Figure S7). Taken together, there is consistent evidence that chunking shifted neural representations of the selected targets toward the center of a chunk.

Degraded target representations on revisit errors

If the target representations we decoded were related solely to motor preparation for a saccade without a mnemonic component, then we would expect the representation of a target-directed saccade to be similar regardless of whether it was a correct response or a revisit error. To test this, we used the same decoder trained on correct target selections to decode revisits. Two example blocks are shown in Figure 4AB. In each case, decoded locations on correct selections are in the vicinity of the target that will be selected, but locations on revisits were displaced from the corresponding target and often clustered around the screen center. To quantify this across all revisits, we calculated the Euclidean distance between an actual target and the location that was decoded on correct selections and revisits in the same blocks (Figure 4C). There were consistently larger distances on revisits (paired t-tests Monkey B: t29396= −155, p <0.001. Monkey A t18433= −137 p <0.001). In addition, analysis of angular shifts found no consistent deviation toward chunk centers on revisit errors (Figure 4D). Together, this shows that revisits are represented differently than correct selections. Specifically, there are degraded target representations preceding error saccades, which suggests a mnemonic function of these neural representations.

Figure 4. Revisit errors and progression within a block.

Figure 4.

(A-B) Two example blocks showing targets (outlined in black) and decoded locations with matching color codes. Small circles = centroid of decoded points, shading = root mean squared error of the centroid. The first panel of each pair shows data from correct selections (solid circles) and the second shows revisits in the same block (open circles). Targets with only one revisit are shown without shading. Blues = chunk 1, Reds = chunk 2, + = screen center, ⊗ = chunk center. (C) Distances (in degrees of visual angle, dva) between a target and the decoded location on correct selections versus decoded locations when the same target was revisited. Insets show distribution of all distances. (D) Angles from decoded locations on revisit errors did not differ from 0°. (E) Revisits across trials within a block of the same target configuration. Points = error rate per trial across 90 blocks. Statistics = Pearson correlations. (F) Representation shifts on correct selections, quantified by the angle between vectors from the actual target to screen center and actual target to predicted target, as in Figure 3, averaged in sliding windows of 5 trials. Positive angles = shift toward chunk centers. Error bars = SE calculated from the circular standard deviation. Filled circles = angle is significantly greater than 0° (circular means test p ≤ 0.01). Asterisks indicate when high and low MI blocks significantly differ.

Chunking is consistent within a block

Evidence suggests that optimal chunking can be learned from experience11. In our task, monkeys worked with the same target configuration for blocks of 40 consecutive trials, which could provide an opportunity for such learning. However, subjects were also well trained on the task rules, if not the specific configurations, and chunking patterns suggest that they each developed their own strategy, or template, to parse new configurations. Monkey B tended to chunk targets on the top and bottom of the screen and Monkey A tended to chunk targets on the upper left and lower right (Figure 1D). With these templates, targets may be chunked right away, from the first trials in a block. If monkeys become more proficient at chunking with practice on a given configuration, we expect error rates to go down across trials within a block. However, we did not find compelling evidence for this (Figure 4E). Monkey B’s behavior did not change across trials, and Monkey A’s changed modestly from about 1 revisit per 3 trials to 1 revisit per 4 trials. Overall, experience with a configuration didn’t dramatically impact error rates, consistent with the interpretation that monkeys applied a pre-learned chunking strategy.

Next, we assessed whether target representations changed across trials in a block. Here, we used only correct target selections and calculated the angular shifts of decoded locations as we did previously, but grouped them by trial within a block. If the monkeys learned the chunks across repeated trials, we would expect representational shifts to increase as the block proceeds, particularly in high MI blocks. We did find greater shifts (i.e., neural representations moved more toward the chunk center), during high MI blocks compared to low MI blocks, defined by median split (Figure 4F). However, there was no evidence that representations were more shifted later in the trial blocks, and if anything, the magnitude of shifts tended to decrease at the end of the block. The mean shifts differed statistically from 0 less frequently at the end of the block (filled markers), suggesting that chunking-related shifts became less pronounced across trials. In addition, differences between high and low MI blocks (indicated by *) were less common toward the end of the blocks. This change happened earlier in Monkey B and only in the last few trials in Monkey A. Together, these results suggest that practice with a given configuration does not strongly impact neural representations in dlPFC, and the modest effects we found were to decrease rather than increase the influence of chunking. This is consistent with the view that chunking is a strategy that helps the monkeys remember information, and familiarity with a target configuration could reduce reliance on this strategy.

Items in a chunk are not temporally interleaved

Our results suggest that chunking systematically biases neural representations in dlPFC. However, in other paradigms, prefrontal populations briefly and idiosyncratically shift their representations among different items in a display, likely reflecting shifts in the focus of attention4143. If dlPFC temporally interleaves representations of multiple targets in the same chunk44, then averaging over such dynamics could give the appearance of a spatially shifted representation. To test this, we reanalyzed the time epoch as a series of short, overlapping time bins that create trajectories of decoded locations over time prior to each target selection. These trajectories frequently traversed the regions of non-selected targets in the same chunk, which could suggest temporally interleaved representations (Figure 5A). Alternatively, since non- selected targets are clustered around the chunk center, variability around a central representation could also appear in non-target regions by chance.

Figure 5. Chunked representations are shifted, not temporally interleaved.

Figure 5.

(A) Example showing variability of decoded locations over time (gray line) for one target selection. The decoded trajectory is primarily near the selected target (filled circle), but briefly falls in the region of one of the unselected targets in the same chunk (“non-targets”, open circles) and the chunk center (⊗). + = screen center. (B-D) The same block as (A). Heatmaps show all decoded trajectories on selections of each target (filled circle). (E) Gaussian fit of the observations in (B). (F) The same plot as (E) but with the distribution shifted so the peak is on the selected target. (G) Comparisons of decoded trajectories (purple) and Gaussian fits of the decoded trajectories (yellow). Bars show the average (±SEM) number of observations within ± 2° of the selected target, the chunk center, or the average of the non-targets in the same chunk. Pink points indicate the same measures for peak-shifted Gaussians. Because observations in the target region are higher by design, only effects on chunk centers and non-targets were statistically compared. Shifting the distribution to the selected target reduced total observations at the chunk center but not the non-targets (Two-way ANOVAs, chunk center vs. non-targets × unshifted vs. shifted interaction. Monkey B, F1,2879 = 59.12, p = 2.02×10−14; Monkey A, F1,2879 = 3.99, p = 0.046.) * = post-hoc comparisons of chunk centers p ≤ 0.004. Post-hoc comparisons of non-targets p > 0.90 for both subjects.

To determine whether trajectories were disproportionately found in non-target regions, we overlaid all trajectories for a given target to create heatmaps (Figure 5BD), so that tendencies to transiently represent non-targets would appear as higher-density hotspots near the non-targets. Conversely, if non-target regions are decoded because the selected target and chunk center are nearby, then non-target decoding should belong to a smooth spatial distribution associated with the selected target. To test this, we fit each heatmap with a 2D Gaussian (Figure 5E), and compared the density of observations in the non-target regions between the Gaussian fits and the actual spatial distributions. If hotspots were present, there would be more non-target observations in the decoded data compared to the fit, but this was not observed (Figure 5G). In the decoded data, the fewest observations were in the region of non-targets, compared to regions around the selected target or chunk center. Importantly, non-target observations were no more common in the original data than the Gaussian fit of the data, indicating that the decoded trajectories form a smooth distribution without hotspots (two-way ANOVAs of region (target vs. chunk center, vs. non-targets) by real/fit data, main effects of region: Monkey B, F2,4319 = 367.24, p = 5.22 × 10−148; Monkey A, F2,4319 = 367.40, p = 5.22 × 10−151; no main effects of real vs. fit data: Monkey B, F1,4319 = 0.81, p = 0.37; Monkey A, F1,4319 = 2.68×10−5, p = 1.00; no interaction: Monkey B, F2,4319 = 0.49, p = 0.61; Monkey A, F2,4319 = 0.01, p = 0.99). Therefore, neural representations of a selected target are approximated by a single Gaussian spatial distribution, and are not temporally interleaved with representations of non-targets.

To confirm that representations are spatially shifted toward chunk centers, we next moved the peak of the Gaussian fit to the position of the selected target (Figure 5F). If the original representations were biased toward chunk centers, we would expect that shifting the peak to the actual target would reduce observations around the chunk center. As anticipated, the shift increased the density of observations at the target location and decreased the observations at the chunk center (Figure 5G). However, the peak shift did not significantly change observations at the non-target regions, indicating that decoded locations were more consistently displaced in the direction of the chunk center, rather than non-targets.

Chunking reduces neural dimensionality

Our results demonstrate that chunking items in working memory systematically biases their representations in dlPFC, so that items belonging to the same chunk are grouped together. Theoretical accounts have suggested that such distortions are a consequence of information compression, which allows working memory mechanisms to maintain a larger number of items and relationships, but at a lower fidelity11,14. If chunking compresses information, then it should reduce the dimensionality of representations in dlPFC. To test this, we used principal components analysis (PCA) to partition variance in population activity. To equate observations across different sessions, we used the mean firing rates associated with each target in a block on correct selections only (see STAR Methods). We then projected the neural responses from each trial block onto these PCs and calculated the amount of variance in each block accounted for by each dimension. Consistent with the idea that chunking compresses task representations, we found that fewer principal components are needed to account for ≥80% of the total variance in more chunked (higher MI) blocks (Figure 6A). Therefore, the dimensionality of task representations was lower in blocks with stronger chunking. To determine how chunking affected population variance in each dimension, we correlated the amount of variance explained (EV) by each PC with MI, and found that a small number of lower PCs captured more variance as MI increased (positive correlations between EV and MI), while higher PCs captured less variance as MI increased (negative correlations with MI) (Figure 6B). This shows that chunking consolidates information into a smaller number of dominant dimensions, producing a lower-dimensional, more compressed representation of items held in working memory.

Figure 6. Chunking reduces dimensionality of dlPFC activity.

Figure 6.

(A) Scatterplots showing the cumulative number of principal components that account for at least 80% of total variance (PC80), separately for each monkey, quantified with Pearson correlations. (B) Correlations of the variance explained (EV) by each principal component (PC) with MI. Filled points show significant correlations.

Chunking biases memory-guided behavior

So far, we have shown that chunking biases representations of grouped targets away from their location in the physical world and toward a common central position and these biases correspond to a more compact neural code. Information is therefore compressed into chunks. Such compression could be lossy, implying that the original information cannot be reinstated from the compressed code, or lossless, meaning that there is a more efficient repackaging of information with no actual data loss. If information about remembered target locations is lost, then this should be apparent in behavior that relies on these representations. To test this, we trained the same two subjects on a variant of the target selection task that had probe trials interspersed in the standard block structure. During a probe trial, rather than seeing the target configuration, the monkey saw an empty frame positioned in one quadrant of the screen, indicating that he should execute a memory-guided saccade to the location of one hidden target in the frame (Figure 7A). Configurations were created to ensure that each quadrant included exactly two targets, and the subject had to complete two memory-guided saccades, one to each hidden target within the frame, returning gaze to the center in between selections. To encourage subjects to report hidden target locations as accurately as possible, rewards were scaled so they earned more juice for a more accurate report. Moving the frame to different screen quadrants allowed us to test memory for each target within a block of trials. We found that subjects often shifted their gaze within the frame, rather than executing a single saccade, before fixing their eyes for 500 ms to indicate their answer. Therefore, we used the centroid of the eye movement path to test whether their searches were biased toward chunk centers (Figure 7B, Figure S8A). Although both search paths and chunk centers were uniformly distributed around the hidden target locations (Figure S8CD), the search paths were systematically displaced in the direction of the chunk centers (circular test of medians: Monkey B p = 1.65×10−23, Monkey A p = 1.21×10−29, Figure 7CD, Figure S8E). We confirmed that these effects were due to the position of the chunk center relative to the hidden target, and not the other hidden target in the same frame (Figure S8F). Therefore, chunking not only distorts neural representations in dlPFC, but also biases behavioral reports of the remembered information in the same direction. Propagation of distortions from neural representations to behavioral expression suggests that more accurate target information cannot be recovered from the compressed neural code, which is more consistent with a lossy rather than lossless compression mechanism.

Figure 7. Chunking biases memory-guided saccades.

Figure 7.

(A) Schematic of the task with hidden target probe trials (orange shading). Probe trials occurred in the middle of a block of regular trials (top, unshaded). Rather than the target array, an empty frame was shown in one quadrant of the task screen (yellow), indicating the subject had to shift their gaze to the region of one of the targets in the frame. This repeated for each of two targets in the frame, then the regular trials resumed. (B) Example eye trajectories on two hidden target trials, showing a short (left) and long (right) scan path. Trajectories are shown from 200 ms until reward was triggered. The location of the hidden target, fame center, chunk center, and centroid of the eye trajectory are illustrated as labeled in (C), but only the frame was visible to the monkey. (C) Schematic showing the angle that quantifies the bias in gaze paths (θ). Positive angles indicate that the gaze path was biased toward the chunk center, and negative angles indicate a bias away. (D) Circular histograms showing the distribution of θ in each subject. Purple bars show the median angle. * = distribution differs significantly from 0 (circular medians test). See also Figure S8.

Discussion

Many cognitive tasks rely on efficient repackaging of elemental information, or chunking. Behavioral evidence of chunking has been described for decades, but we have so far lacked an understanding of its neural underpinnings. Here, we identified chunking patterns in monkeys’ self-ordered behavior, which allowed us to assess the neural mechanisms that chunk items held in working memory. We found that chunked representations in dlPFC cluster together, inducing concomitant biases in memory reports and reducing the dimensionality of task representations. These results support models of chunking as a form of data compression that sacrifices the fidelity of individual item representations to enhance overall memory retention8,1113.

A striking feature of the compression mechanism we describe is its apparent similarity to human chunking. Previous work has reported chunking-like behavior in monkeys4547, but this study provides the most direct evidence to date that Rhesus monkeys spontaneously adopt mnemonic chunking strategies to support working memory. The behavioral signatures of chunking in macaques parallel those in humans, including reaction time costs at chunk boundaries, reduced error rates within chunks, and systematic biases in recall toward chunk centers6,7,11,3134. This similarity is unlikely to be coincidental. The dlPFC plays a critical role in cognitive control and strategy-based behaviors such as chunking and is most expanded in humans and other primates27,28. Compression-based chunking may be a solution implemented by this shared dlPFC circuitry to maximize the utility of capacity-limited working memory. Future work comparing chunking across species could investigate whether the compression mechanism scales with cognitive capacity, and whether similar solutions have evolved in non-primate species.

Distortions of chunked representations have been predicted by theoretical models that simulate working memory with recurrent neural networks (RNNs)11,14. One type of network maintains information by settling in unique attractor states associated with each stimulus48. When these networks are given a capacity limit that is exceeded, attractors for different stimuli merge as if the items are chunked together11,14. This merging economizes on neural processing resources but systematically biases representations so chunked items appear similar. Supporting this mechanism, we found that memory-guided behavioral reports were biased toward the center of a group of chunked targets, similar to error patterns observed during human chunking11.

Unlike RNN simulations, however, we found that neural representations were biased but not fully merged into a single chunk. This suggests different or additional mechanisms not captured by traditional RNNs, one of which may be transient reweighting of synapses, known as short-term plasticity (STP). STP results in sub-second changes in synaptic efficacy following spiking49 and has been proposed as a mechanism supporting working memory50,51. In chunking, STP could transiently change connectivity among neurons that respond to targets in the same chunk44. Even when the targets are recalled sequentially, such synaptic modifications could bias the population code for the recalled item. More reliable coactivity could also reduce the dimensionality of neural activity, as we observed. Therefore, STP is a promising candidate mechanism for chunking-related compression.

The principle of trading representational fidelity for capacity is well-established in engineering, where lossy compression algorithms such as JPEG image encoding and MP3 audio discard redundant information to reduce storage and transmission costs. Our findings suggest that the brain employs a related but uniquely adaptive strategy to manage the limits of working memory. Monkeys develop chunk structures in the moment that reflect the regularities of their environment and use these to guide ongoing behavior. This approach improved task performance, suggesting that structured compression of active internal representations could be an efficient strategy for capacity-limited systems. Therefore, understanding the neural implementation of this strategy could offer principled inspiration for the design of more efficient cognitive architectures in artificial systems.

In a broader sense, the prefrontal cortex is known to represent task information in compressed form when it is used abstractly52. Abstraction shares features with chunking, as it involves extracting commonalities of individual exemplars, such as visual objects53,54, task rules52,55, or reward predictions56. In this way, abstraction emphasizes a shared concept while individuating particulars are discarded. Similarly, the chunking mechanism we report emphasizes differences between chunks, while reducing differences among items within a chunk. More abstract neural representations are also more compressed in the sense that they have lower dimensionality, paralleling our results in chunking. Therefore, one possibility is that chunking and abstraction lie along a continuum, with high dimensional, veridical representations at one end and fully compressed, abstract representations at the other. From this view, the chunking representations we describe are only partially compressed. Our task requires the monkeys to select targets one at a time, and therefore remember whether each has been visited, which may have stopped the chunked items from collapsing into a fully abstract shorthand that no longer differentiates the constituent items. Such overlearned chunks have been described in language, where words take on abstract semantic representations not directly tied to the letters making them up9.

In summary, we demonstrate a cortical mechanism for chunking that allows the brain to transcend its inherent capacity limits through dynamic internal reorganization. Our results offer a mechanistic entry point for understanding a broad class of cognitive operations that rely on chunking, including learning, language processing, strategic planning, and problem solving9,10,57,58. These same processes are frequently impaired in psychiatric disorders59,60, and understanding their neural underpinnings may also shed light on pathological states of disordered cognition.

Resource availability

Lead contact

Requests for further information and resources should be directed to and will be fulfilled by the lead contact, Erin Rich (elr9746@nyu.edu).

Materials availability

This study did not generate new unique reagents.

Data and code availability

  • Behavioral data, neural data, and metadata have been deposited at figshare.com and are publicly available as of the date of publication at 10.6084/m9.figshare.33087665.

  • All original code has been deposited at figshare.com and is publicly available at 10.6084/m9.figshare.33087665 as of the date of publication.

  • Any additional information required to reanalyze the data reported in this paper is available from the lead contact upon request.

STAR Methods

Experimental model and study participant details

Subjects

Subjects were two male rhesus monkeys (Macaca mulatta), B and A, aged 6 and 8 years, and weighing approximately 8.5–8.8 and 9.1–9.3 kg at the time of data collection. Female subjects were not assessed, which is a limitation of the present study. Monkeys were housed in groups of 2 and 3 individuals, and worked for diluted apple juice rewards. Daily fluid intake was measured and regulated to maintain motivation. All procedures were in accordance with the National Institute of Health guidelines and recommendations and approved by the Icahn School of Medicine at Mount Sinai Animal Care and Use Committee.

Method details

Behavioral task

Subjects sat head-fixed in a primate chair and viewed a computer screen mounted 46cm away. MonkeyLogic software61,62 controlled the behavior interface and subjects’ eye movements were tracked with an infrared camera (ISCAN, Burlington, MA). To begin a trial, monkeys had to fixate a central point on the task screen for 1000 ms, then a configuration of eight targets appeared. Targets were identical filled circles with diameter 1° of visual angle. Target color was green, blue, or white and changed at the beginning of a new trial to indicate that reward contingencies were reset. Subjects had to saccade to shift their gaze to any target and hold fixation (±3° of visual angle) for 500 ms to register a selection. When a selection was made, unselected targets disappeared. If a new target was selected that had not been previously selected on that trial, the monkey earned juice reward (0.33 ml over 500 ms). If they selected a target that had already been visited on that trial, they received a 1000 ms time-out. After the reward or time-out period, the selected target disappeared, the central fixation point reappeared, and monkeys had to fixate centrally again for 1000 ms to make the target configuration reappear so they could make their next selection. The trial was complete when all eight targets had been selected or the subject made 10 consecutive incorrect responses regardless of how many rewards had been collected. Trials were followed by a 1 – 2.5 s ITI before the next trial began. Most trial blocks included 40 completed trials (minimum 32 trials, Table S1) with the same target configuration. Each session included six blocks with different configurations, except for one session (Monkey B) in which a configuration was re-used. Subjects completed 15 sessions, for a total of 90 blocks each. Target presentations in which no target was selected were excluded from further analysis. In total, 31,765 and 30,170 target selections were analyzed for Subjects B and A respectively.

Target locations were drawn from an array of 80 possible locations, arranged on evenly distributed rings around the screen center of 5 different sizes (6°, 8.5 °, 11 °, 13.5 °, 16 ° of visual angle). During pre-training, monkeys saw a random selection of targets with the constraint that there was at least 6° of visual angle between any two targets. During testing, target configurations were pre-determined for each session and the same for both monkeys. Test configurations were created as follows: 8 out of 80 possible targets on horizontal and vertical axes were used for eye-calibration and never shown as targets. The 72 remaining locations were evenly separated into 8 pie-shaped zones converging at the screen center, each covering 45° of a circle and including 9 possible target locations in each (Figure S1A). Target configurations consisted of one of the targets selected randomly from each zone. This ensured that configurations were spatially balanced around the screen during test sessions (Figure S1B). The mean ± SEM distance from centroid of all targets to screen center fixation was 1.2 ± 0.065° of visual angle.

Hidden target task

The hidden target task was similar to the main task, except on every fifth completed trial, the fixation cue was shown as a black square with a yellow outline and one quadrant filled yellow, indicating the upcoming trial would be a probe trial in the corresponding quadrant. After fixation, an empty frame appeared on the cued quarter of the task screen (upper/lower right/left), rather than the target configuration. This indicated that the subject had to fix their gaze at a point within the empty frame to report the remembered location of a target. Targets outside the frame were also not visible and fixing gaze outside the frame had no effect. Target configurations were selected so that the frame always included exactly two targets. Moving the frame on each probe trial ensured that memory for each target was tested. A selection was registered when the subject held their gaze within ± 4 degrees of the target for at least 500 ms. The amount of reward delivered depended on proximity to the nearest target. Four sizes of reward (0.66, 0.33, 0.198, and 0.066 ml) were delivered, depending on the registered gaze distance to selected target (1.3, 2.0, 2.7, and 4.0 degrees). The selected target became visible during the reward epoch and stayed on for 1000 ms on probe trials. Reward amounts were adjusted by changing the duration of fluid delivery. After selecting one hidden target and earning a reward, the subject had to return their eyes to the center fixation point and select the other hidden target in the same frame. After this, they returned their eyes to central fixation and the inter-trial interval started, followed by a new trial with visible targets.

Neural recording

Each subject was implanted with a titanium head positioner (Jerry-Rig, USA) and four 64-channel microelectrode arrays (Utah arrays, Blackrock Microsystems) as shown in Figure S3AB. Array connectors were enclosed in a custom 3d-printed holder affixed to the skull with acrylic. Neurophysiological signals were collected, digitized, and saved with a digital signal processor (Ripple Neural Systems Grapevine). Wideband signal was thresholded, and threshold crossings were captured at 30 kHz and sorted into single units offline, by first auto-sorting63 then manually adjusting the output (Plexon OfflineSorter). Spike times were saved at 1 kHz resolution. Any unit with an overall firing rate <1Hz was excluded from analysis due to difficulty statistically characterizing their responses.

Quantification and statistical analysis

Behavior analyses

Putative chunks were defined separately for each trial block where the same configuration was presented. We used the Louvain community detection algorithm as implemented in Matlab by the Brain Connectivity Toolbox30. Briefly, transition matrices (T) were based on the order of target selections across all saccades in a block (including revisits). Each entry (TX,Y) is the number of times target X was selected on saccade s and Y on s+1, where X and Y are the set of all targets in a block. The Louvain algorithm is an efficient optimization that finds community memberships for each target to maximizes MI64, a measure of overall “chunkiness” of behavior in a block:

MI=12mXY[TX,Y-kXkY2m]δ(cX,cY)

Here, kX=YTX,Y, or the sum of the transition probabilities from target X, cX is the chunk to which X belongs, δ indicates the δ-function that δX,Y=1 if X=Y and zero otherwise, and m=12X,YTX,Y. Therefore, MI is high when TX,Y is high within chunks and low between, and MI is low when transitions are equivalently distributed in T. We used the default resolution parameter defined by the toolbox30, gamma = 1, which should yield a measure of classic modularity. This approach assigned each target to 2–3 non-overlapping clusters and calculated MI, which increased with stronger evidence of reliable clustering.

To relate behavioral measures to chunking tendency, we calculated the total number of revisits per block and compared it to the MI for that block (Figure 1E). We also found the average probability of committing a revisit when a target selection stayed in the same chunk as the previous selection, or crossed a chunk boundary (Figure S2C). To do this, we first found the average rate of revisits for each target selection (1 to 8 for trials with no revisits; 1 to 8 + #revisits otherwise), when that selection was within versus across chunks. To correct for the fact that revisit rates increase across saccades in a trial regardless of chunk boundary (Figure S2A, DE), we subtracted the overall revisit rate across trials for that selection. We then computed the average of these normalized revisit rates within and across chunks, for each of 90 blocks, and compared them with two-sided t-tests (Figure 1F). The same procedure was used to find the average normalized RT for each selection (Figure 1G). RT was defined as the time from targets on until the initiation of the fixation that selected a target. Normalization was necessary here because RTs increase across saccades in a sequence40,65 (Figure S2B, FG). For this analysis, we excluded revisits to remove the possibility that slower RTs were a consequence of the higher rate of revisits across chunk boundaries, and removed any selection that took >1s to initiate.

Target decoding

Target locations were decoded as cartesian coordinates using two separate general linear models, one predicting x-locations and one y-locations, from all isolated neurons in a session (Matlab fitglm.m and predict.m functions). For the main analyses, decoding was performed from average firing rates in a window ± 250 ms around the appearance of the target configuration, unless otherwise specified. This window was chosen a priori to capture the preparatory time before a saccade was executed. Each observation in the decoding matrix was a single target selection, and models were trained using only correct selections (i.e., without revisits) in all blocks of a session, using 10-fold cross validation. Folds were randomized so that training and testing sets were distributed across a session. Prior to decoding, spike trains were smoothed with a 200 ms boxcar, averaged in the 500 ms time window, and z-scored. For cross-temporal and trajectory decoding, an epoch of ± 500 ms around the appearance of the target configuration was used, and the same decoding was performed in 50 ms time windows, iteratively stepped by 5 ms. For cross-temporal decoding (Figure S3C), the same held-out fold was tested on each combination of training and testing time windows, and this was repeated for all 10 folds. For trajectory decoding (Figure 5), only the diagonals of the cross-temporal matrices (training and testing on the same time windows) were included.

To test whether predicted target locations clustered according to behaviorally-defined chunks, we used k-means on predicted target locations (x,y), separately for selections belonging to each trial block, and set the number of clusters equal to the number of behaviorally-identified chunks (Figure 2DE, Figure S4).

To test whether predicted target locations were systematically shifted toward chunk centers, we first found the centroid of all predictions for a given target (xpred,ypred) (8 per block), and the centroid of the actual target locations belonging to each chunk (xchk,ychk). We computed Euclidean distances from each target or decoded target location to the corresponding centroid, and calculated spatial dispersion of chunked points (either true targets or decoded locations) as the sum of the eigenvalues of the covariance matrix, or trace, of x,y coordinates (Figure S5 AD). To assess how these measures would be affected if decoded locations were displaced only toward the center of the screen, we calculated the distance from each target to screen center and reduced that distance by 20% while maintaining the same vector angle to the center and ignoring the chunk center. We then recalculated the distances and dispersion with these shifted targets (Figure S5 DF).

For analyses of angular shifts, we calculated the angle between two vectors: one from the actual target location (xtgt,ytgt) to screen center (0,0), and one from xtgt,ytgt to xpred,ypred. We did this because random noise in the will shift predictions of xtgt,ytgt toward 0,0, so that with no other source of distortion, xpred,ypred should be equally likely to be at positive or negative angles to the vector from xtgt,ytgt to 0,0. To determine if there is a bias toward xchk,ychk, we labeled angles in the same direction as xchk,ychk as positive and angles in the opposite direction as negative (Figure 3AB). Median angles were calculated and compared to 0 degrees (no bias) with circular statistics66.

We compared our main results to the same analyses using “false” chunks, or a designated group of targets that were not actually chunked by the monkey. False chunks were created from the same targets in each block by bisecting the screen in a plane that would rotate the typical chunk centers for each monkey by 90°. Targets for Monkey B were separated into left and right groups, and for Monkey A upper right and lower left (Figure S5 GJ). Only 2-chunk blocks were used to create false chunks. False chunk centers were then computed as the centroid of the grouped targets, identically to true chunk centers.

Dimensionality analysis

To measure dimensionality of task encoding, we used the same matrix of normalized firing rates from the main task as in the decoding analyses (± 250 ms from targets onset). To obtain comparable data sets across sessions with different numbers of trials, we computed the vector of average firing rates for each target in each block (8 targets × 6 blocks), using only correct selections, to create a 48 × n matrix for each session, where n is the number of neurons in the session. We then used PCA to find the 47 PCs and the loadings of each neuron on each component. To understand how dimensionality changes in high and low MI blocks, we projected neural activity in each block onto each PC by multiplying the mean firing rates on that block (8 targets × n neurons) by the PC loadings (n neurons × 47 PCs), and found proportion of total variance in each block explained by each of 47 PCs (Figure 7).

Hidden target behavior analyses

We analyzed the centroids of the monkeys’ gaze path on hidden target trials. All eye positions from 200 ms after the targets appeared until the end of the 500ms hold-target period were included. This time period was chosen to include any search and hold eye positions but exclude the first saccade from central fixation. Prior to analysis, all eye positions were re-centered by subtracting the mean × and y position of the last 500ms of the central fixation before the first hidden target selection of the trial. Any eye positions off the screen were excluded, although this rarely occurred.

Using an analysis analogous to the decoded targets, the centroid of the search path (xeye,yeye) was compared to the actual hidden target location (xtgt,ytgt) and the centroid of the chunk to which the hidden target belonged (xchk,ychk). Since searches tended toward the center of the empty frame, rather than the screen center, we used the center point of the empty frame (x0,y0). Therefore, we found the angle formed by the vector from the xtgt,ytgt to xeye,yeye and from the xtgt,ytgt to x0,y0. If the angle had the same (or different) sign as the angle between vectors from xtgt,ytg to xchk,ychk and from xtgt,ytgt to x0,y0, then the original angle was labeled positive (or negative) (Figure 7C). Chunked targets and chunk centers were determined as described in the main task, using visible target trials only. For the gaze analysis, only hidden target trials were included, and circular statistics determined whether the eye position centroids were significantly shifted toward a chunk center by comparing the median angle to 0 degrees.

To ensure that the results held if we did not account for the frame position, we also found the angle formed by vectors from xeye,yeye to xtgt,ytgt and from xeye,yeye to xchk,ychk (Figure S8B). Since chunk centers (xchk,ychk) were uniformly distributed around the actual targets (Figure S8D), we compared this angle to a uniform distribution with Rayleigh tests for non-uniformity66 (Figure S8E), so that displacement toward 0 degrees would indicate a bias to search in the direction of the chunk center. There was no differential interpretation of positive versus negative angles in this analysis.

Supplementary Material

1

Document S1. Figures S1S8, Table S1.

Key resources table

REAGENT or RESOURCE SOURCE IDENTIFIER
Deposited data
Raw and analyzed data This paper 10.6084/m9.figshare.33087665
Experimental models: Organisms/strains
Rhesus macaques UT MD Anderson N/A
Software and algorithms
Matlab R2021b Mathworks https://www.mathworks.com/
Trellis Ripple Neuro https://rippleneuro.com/support/software-downloads-updates/
Spikesort Issar et al.63 https://github.com/SmithLabNeuro/spikesort
Offline Sorter Plexon https://plexon.com/products/offline-sorter/
NIMH MonkeyLogic Hwang et al.62 https://monkeylogic.nimh.nih.gov/index.html
Circular Statistics Toolbox Berens66 https://www.mathworks.com/matlabcentral/fileexchange/10676-circular-statistics-toolbox-directional-statistics
Brain Connectivity Toolbox Rubinov et al.30 https://sites.google.com/site/bctnet

Highlights.

  • Monkeys spontaneously chunk items in a self-ordered working memory task

  • Chunking biases dlPFC representations toward the center of the chunk

  • Chunking compresses memoranda into lower dimensional representations in dlPFC

  • On probe trials, memory-guided behavior is similarly biased toward chunk centers

Acknowledgments:

The authors thank Peter Rudebeck for surgery assistance and Joey Charbonneau for comments on the manuscript. This work was supported by the National Institutes of Health grant R01MH121480 (ELR), a Whitehall Foundation Research Grant (ELR), and a grant from the Pew Biomedical Scholars Program (ELR).

Footnotes

Declaration of interests: The authors declare no competing interests.

References

  • 1.Baddley AD. (1986). Working Memory. (Oxford: Clarendon Press; ). [Google Scholar]
  • 2.Cowan N. (2001). The magical number 4 in short-term memory: a reconsideration of mental storage capacity. Behav. Brain Sci. 24, 87–114. [DOI] [PubMed] [Google Scholar]
  • 3.Brady TF, Konkle T, Alvarez GA. (2009). Compression in visual working memory: using statistical regularities to form more efficient memory representations. J. Exp. Psychol. Gen. 138, 487–502. [DOI] [PubMed] [Google Scholar]
  • 4.Brady TF, Konkle T, Gill J, Oliva A, Alvarez GA. (2013). Visual long-term memory has the same limit on fidelity as visual working memory. Psychol. Sci. 24, 981–90. [DOI] [PubMed] [Google Scholar]
  • 5.Luck SJ, Vogel EK. (1997). The capacity of visual working memory for features and conjunctions. Nature. 390, 279–81. [DOI] [PubMed] [Google Scholar]
  • 6.Tulving E. (1962). Subjective organization in free recall of “unrelated” words. Psychol. Rev. 69, 344–54. [DOI] [PubMed] [Google Scholar]
  • 7.Tulving E, Sternberg RJ. (1977). The measurement of subjective organization in free recall. Psychological Bulletin. 84, 539–556. [Google Scholar]
  • 8.Chekaf M, Cowan N, Mathy F. (2016). Chunk formation in immediate memory and how it relates to data compression. Cognition. 155, 96–107. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Dehaene S, Meyniel F, Wacongne C, Wang L, Pallier C. (2015). The Neural Representation of Sequences: From Transition Probabilities to Algebraic Patterns and Linguistic Trees. Neuron. 88, 2–19. [DOI] [PubMed] [Google Scholar]
  • 10.Chase WG, Simon HA. (1973). Perception in Chess. Cognitive Psychol. 4, 55–81. [Google Scholar]
  • 11.Nassar MR, Helmers JC, Frank MJ. (2018). Chunking as a rational strategy for lossy data compression in visual working memory. Psychol. Rev. 125, 486–511. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Mathy F, Feldman J. (2012). What’s magic about magic numbers? Chunking and data compression in short-term memory. Cognition. 122, 346–62. [DOI] [PubMed] [Google Scholar]
  • 13.Norris D, Kalm K. (2021). Chunking and data compression in verbal short-term memory. Cognition. 208, 104534. [DOI] [PubMed] [Google Scholar]
  • 14.Wei Z, Wang XJ, Wang DH. (2012). From distributed resources to limited slots in multiple-item working memory: a spiking network model with normalization. J. Neurosci. 32, 11228–40. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Procyk E, Goldman-Rakic PS. (2006). Modulation of dorsolateral prefrontal delay activity during self-organized behavior. J Neurosci. 26, 11313–23. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Petrides M. (1995). Impairments on nonspatial self-ordered and externally ordered working memory tasks after lesions of the mid-dorsal part of the lateral frontal cortex in the monkey. J. Neurosci. 15, 359–75. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Petrides M, Milner B. (1982). Deficits on subject-ordered tasks after frontal- and temporal-lobe lesions in man. Neuropsychologia. 20, 249–62. [DOI] [PubMed] [Google Scholar]
  • 18.Freedman DJ, Riesenhuber M, Poggio T, Miller EK. (2002). Visual categorization and the primate prefrontal cortex: neurophysiology and behavior. J. Neurophysiol. 88, 929–41. [DOI] [PubMed] [Google Scholar]
  • 19.Tsujimoto S, Genovesio A, Wise SP. (2011). Comparison of strategy signals in the dorsolateral and orbital prefrontal cortex. J. Neurosci. 31, 4583–92. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Yamada M, Pita MC, Iijima T, Tsutsui K. (2010). Rule-dependent anticipatory activity in prefrontal neurons. Neurosci. Res. 67, 162–71. [DOI] [PubMed] [Google Scholar]
  • 21.Tsutsui K, Hosokawa T, Yamada M, Iijima T. (2016). Representation of Functional Category in the Monkey Prefrontal Cortex and Its Rule-Dependent Use for Behavioral Selection. J. Neurosci. 36, 3038–48. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Wallis JD, Anderson KC, Miller EK. (2001). Single neurons in prefrontal cortex encode abstract rules. Nature. 411, 953–6. [DOI] [PubMed] [Google Scholar]
  • 23.Ichihara-Takeda S, Funahashi S. (2007). Activity of primate orbitofrontal and dorsolateral prefrontal neurons: task-related activity during an oculomotor delayed-response task. Exp. Brain Res. 181, 409–25. [DOI] [PubMed] [Google Scholar]
  • 24.Gershberg FB, Shimamura AP. (1995). Impaired use of organizational strategies in free recall following frontal lobe damage. Neuropsychologia. 33, 1305–33. [DOI] [PubMed] [Google Scholar]
  • 25.Bor D, Duncan J, Wiseman RJ, Owen AM. (2003). Encoding strategies dissociate prefrontal activity from working memory demand. Neuron. 37, 361–7. [DOI] [PubMed] [Google Scholar]
  • 26.Bor D, Owen AM. (2007). A common prefrontal-parietal network for mnemonic and mathematical recoding strategies within working memory. Cereb. Cortex. 17, 778–86. [DOI] [PubMed] [Google Scholar]
  • 27.Petrides M, Tomaiuolo F, Yeterian EH, Pandya DN. (2012). The prefrontal cortex: comparative architectonic organization in the human and the macaque monkey brains. Cortex. 48, 46–57. [DOI] [PubMed] [Google Scholar]
  • 28.Wise SP. (2008). Forward frontal fields: phylogeny and fundamental function. Trends Neurosci. 31, 599–608. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Leicht EA, Newman ME. (2008). Community structure in directed networks. Phys. Rev. Lett. 100, 118703. [DOI] [PubMed] [Google Scholar]
  • 30.Rubinov M, Sporns O. (2009). Complex network measures of brain connectivity: uses and interpretations. Neuroimage. 52, 1059–69. [DOI] [PubMed] [Google Scholar]
  • 31.Johnson NF. (1966). On the relationship between sentence structure and the latency in generating the sentence. J. Verbal Learn. Verbal Behav. 5, 375–380. [Google Scholar]
  • 32.Johnson NF. (1970). The role of chunking and organization in the process of recall. Psychol. Learn. Motiv. 4, 171–247. [Google Scholar]
  • 33.Anderson JR, Matessa M. (1997). A production system theory of serial memory. Psychol. Rev. 104, 728–748. [Google Scholar]
  • 34.Kennerley SW, Sakai K, Rushworth MF. (2004). Organization of action sequences and the role of the pre-SMA. J. Neurophysiol. 91, 978–93. [DOI] [PubMed] [Google Scholar]
  • 35.Funahashi S, Bruce CJ, Goldman-Rakic PS. (1989). Mnemonic coding of visual space in the monkey’s dorsolateral prefrontal cortex. J. Neurophysiol. 61, 331–49. [DOI] [PubMed] [Google Scholar]
  • 36.Funahashi S, Bruce CJ, Goldman-Rakic PS. (1990). Visuospatial coding in primate prefrontal neurons revealed by oculomotor paradigms. J. Neurophysiol. 63, 814–31. [DOI] [PubMed] [Google Scholar]
  • 37.Goldman-Rakic PS. (1995). Cellular basis of working memory. Neuron. 14, 477–85. [DOI] [PubMed] [Google Scholar]
  • 38.Fuster JM, Alexander GE. (1971). Neuron activity related to short-term memory. Science. 173, 652–4. [DOI] [PubMed] [Google Scholar]
  • 39.Constantinidis C, Funahashi S, Lee D, Murray JD, Qi XL, Wang M, Arnsten AFT. (2018). Persistent Spiking Activity Underlies Working Memory. J. Neurosci. 38, 7020–7028. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Chiang FK, Wallis JD. (2018). Spatiotemporal encoding of search strategies by prefrontal neurons. Proc. Natl. Acad. Sci U.S.A. 115, 5010–5015. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Rich EL, Wallis JD. (2016). Decoding subjective decisions from orbitofrontal cortex. Nat. Neurosci. 19, 973–80. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Buschman TJ, Miller EK. (2009). Serial, covert shifts of attention during visual search are reflected by the frontal eye fields and correlated with population oscillations. Neuron. 63, 386–96. doi: 10.1016/j.neuron.2009.06.020 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Kaufman MT, Churchland MM, Ryu SI, Shenoy KV. (2015). Vacillation, indecision and hesitation in moment-by-moment decoding of monkey motor cortex. Elife. 4, e04677. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Zhong W, Katkov M, Tsodyks M. (2026). Synaptic theory of chunking in working memory. Elife. 15, RP109538. [Google Scholar]
  • 45.Tosatto L, Fagot J, Nemeth D, Rey A. (2024). Chunking as a function of sequence length. Anim. Cogn. 28, 2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Tosatto L, Fagot J, Nemeth D, Rey A. (2022). The Evolution of Chunks in Sequence Learning. Cogn. Sci. 46, e13124. [DOI] [PubMed] [Google Scholar]
  • 47.Holmes CD, Ching S, Snyder LH. (2022). Primates chunk simultaneously-presented memoranda. Front. Behav. Neurosci. 16, 1060193. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Compte A, Brunel N, Goldman-Rakic PS, Wang XJ. (2000). Synaptic mechanisms and network dynamics underlying spatial working memory in a cortical network model. Cereb. Cortex. 10, 910–23. [DOI] [PubMed] [Google Scholar]
  • 49.Mongillo G, Barak O, Tsodyks M. (2008). Synaptic theory of working memory. Science. 319, 1543–6. [DOI] [PubMed] [Google Scholar]
  • 50.Panichello MF, Jonikaitis D, Oh YJ, Zhu S, Trepka EB, Moore T. (2024). Intermittent rate coding and cue-specific ensembles support working memory. Nature. 636, 422–429. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Stokes MG. (2015). ‘Activity-silent’ working memory in prefrontal cortex: a dynamic coding framework. Trends Cogn. Sci. 19, 394–405. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Bernardi S, Benna MK, Rigotti M, Munuera J, Fusi S, Salzman CD. (2020). The Geometry of Abstraction in the Hippocampus and Prefrontal Cortex. Cell. 183, 954–967.e21. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.DiCarlo JJ, Cox DD. (2007). Untangling invariant object recognition. Trends Cogn. Sci. 11, 333–41. [DOI] [PubMed] [Google Scholar]
  • 54.Pagan M, Urban LS, Wohl MP, Rust NC. (2013). Signals in inferotemporal and perirhinal cortex suggest an untangling of visual target information. Nat. Neurosci. 16, 1132–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Courellis HS, Minxha J, Cardenas AR, Kimmel DL, Reed CM, Valiante TA, Salzman CD, Mamelak AM, Fusi S, Rutishauser U. (2024). Abstract representations emerge in human hippocampal neurons during inference. Nature. 632, 841–849. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Chien JM, Wallis JD, Rich EL. (2023). Abstraction of reward context facilitates relative reward coding in neural populations of the macaque anterior cingulate cortex. J. Neurosci. 43, 5944–5962. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Wu L, Knoblich G, Luo J. (2013). The role of chunk tightness and chunk familiarity in problem solving: evidence from ERPs and fMRI. Hum. Brain Mapp. 34, 1173–86. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Schlaghecken F, Stürmer B, Eimer M. (2000). Chunking processes in the learning of event sequences: electrophysiological indicators. Mem. Cognit. 28, 821–31. [DOI] [PubMed] [Google Scholar]
  • 59.Cocchi L, Schenk F, Volken H, Bovet P, Parnas J, Vianin P. (2007). Visuo-spatial processing in a dynamic and a static working memory paradigm in schizophrenia. Psychiatry Res. 152, 129–42. [DOI] [PubMed] [Google Scholar]
  • 60.Huntley J, Bor D, Hampshire A, Owen A, Howard R. (2011). Working memory task performance and chunking in early Alzheimer’s Disease. Br. J. Psychiatry. 198, 398–403. [DOI] [PubMed] [Google Scholar]
  • 61.Asaad WF, Eskandar EN. (2008). A flexible software tool for temporally-precise behavioral control in Matlab. J. Neurosci. Methods. 174, 245–258. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Hwang J, Mitz AR, Murray EA. (2019). NIMH MonkeyLogic: Behavioral control and data acquisition in MATLAB. J. Neurosci. Methods. 323, 13–21. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Issar D, Williamson RC, Khanna SB, Smith MA. (2020). A neural network for online spike classification that improves decoding accuracy. J. Neurophysiol. 123, 1472–1485. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Blondel VD, Guillaume JL, Lamboitte R. (2008). Fast unfolding of communities in large networks. J. Stat. Mech. Theory Exp. E10, P10008. [Google Scholar]
  • 65.Chiang FK, Wallis JD, Rich EL. (2022). Cognitive strategies shift information from single neurons to populations in prefrontal cortex. Neuron. 110, 709–721.e4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.Berens P. (2009). CircStat: A MATLAB Toolbox for Circular Statistics. J. Stat. Softw. 31, 1–21. [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

1

Data Availability Statement

  • Behavioral data, neural data, and metadata have been deposited at figshare.com and are publicly available as of the date of publication at 10.6084/m9.figshare.33087665.

  • All original code has been deposited at figshare.com and is publicly available at 10.6084/m9.figshare.33087665 as of the date of publication.

  • Any additional information required to reanalyze the data reported in this paper is available from the lead contact upon request.

RESOURCES