Significance
A group can often outperform an individual in complex problem solving, even when the group lacks a sophisticated central planner. We develop a simple theory to explain how collective intelligence emerges. The group tries to predict an outcome that depends on many random factors. Even though each individual can observe only a single factor, the collective prediction from averaging across individuals may nonetheless be accurate. We explore two schemes for rewarding individuals to foster collective intelligence. One approach rewards experts: those whose predictions are more accurate. The other approach rewards reformers: those whose predictions may not be accurate, but whose contributions improve the collective prediction. Both reward schemes can foster collective accuracy, but rewarding reformers is more robust and efficient.
Keywords: collective behavior, problem solving, evolutionary dynamics, learning, imitation
Abstract
A collaborative group can often outperform a single individual in complex problem solving, even when information is limited. This phenomenon, called collective intelligence, can be achieved by engineering a central planner who assigns subtasks distributed across the group. But such algorithms cannot explain how natural populations, which often lack sophisticated central control, can nonetheless evolve collective intelligence. In fact, the process of social learning by imitating successful peers will typically reduce diversity and inhibit collective intelligence. Here, we consider a prediction task where the true outcome each round is a continuous quantity that depends linearly on a large number of random causal factors. Each individual can observe only one factor, and the collective prediction is generated by aggregating personal predictions across individuals. We propose two classes of reward structures that guarantee the emergence of collective intelligence through social learning. One scheme provides greater rewards to those individuals (called experts) whose personal predictions are more accurate. The other scheme provides greater rewards to those individuals (called reformers) whose predictions have greater potential to reduce the collective error, even though their personal predictions may be far from the truth. Although both of these payoff structures can provably maintain diversity and establish collective intelligence, we show that rewards based on collective error are more robust to diverse problem settings than rewards based on personal accuracy. Our results show that identifying reformers is more effective than identifying experts in promoting the emergence of collective intelligence.
A group of relatively uninformed individuals can often collaborate to accomplish a task that exceeds the capability of any single individual (1–5). For example, the average estimate of a quantity by a group is often remarkably close to its true value, whether in a single-shot guessing task or in more complex problem solving. This phenomenon in human groups is commonly referred to as collective intelligence.
A body of work in engineering describes how to achieve collective intelligence, at least for distributed computation nodes. Algorithms developed in control theory and machine learning can ensure that a group of nodes will effectively work together to accomplish a complex task. Such algorithms either divide the task into subtasks, each of which can be accomplished in parallel by individual nodes (6–9), or they adaptively optimize hyperparameters to dynamically control the contribution weights from individual nodes (10, 11). These algorithms require sophisticated central control, and they are designed for execution by computers. Such a sophisticated central planner is typically unavailable in natural populations, which nonetheless exhibit collective intelligence (1–5, 12, 13). And so the question remains: How can relatively uninformed individuals, without access to sophisticated central control, nonetheless develop collective intelligence through individual-based processes, such as social learning and peer-to-peer imitation (14–20).
A basic requirement for collective intelligence is a diversity of individual opinions, or strategies, within the population (12, 21–24). However, social influence and learning through imitation permeate the natural world and would seem to winnow the diversity that is required for collective problem solving (21). Consequently, an individual-based mechanism that incentivizes individuals to preserve diversity may be necessary for the emergence of collective intelligence.
Merely maintaining diversity among individuals may be insufficient to achieve collective intelligence. Many real-world problems are inherently complex in ways that diversity alone cannot reconcile (25–29). Moreover, in some settings when numerous factors are observable, which particular factors influence the outcome remains poorly understood. For such problems where solutions depend intricately on numerous factors, it may be impossible for any single individual to aggregate and analyze the requisite data. Even if a diversity of opinions is preserved in a population, and even if a population has access to a simple mechanism for aggregating and averaging individual opinions, it remains unclear what mechanisms will stimulate the correct diversity of individual strategies to produce accurate collective predictions.
A few prior studies have proposed individual incentive structures, such as rewarding accurate individuals (30) or rewarding minority opinions (31, 32), that might facilitate the emergence of collective intelligence in natural populations. These approaches are typically confined to simple value estimation tasks, without correlated factors, or to binary voting scenarios.
Here, we consider prediction tasks where the outcome is continuous and depends on a large number of causal factors, which may be independent or correlated. We consider a setting where each individual has limited information, so that he or she can observe only a single factor when making their individual prediction. A collective prediction is then generated by aggregating individual predictions, through one of several different aggregating rules. Individuals are rewarded with payoffs according to their performance, and they tend to imitate the strategies of others who have higher payoffs, in accordance with the standard assumption of evolutionary game theory that describes the social spread of behaviors (14, 15, 33–35). The key question, then, is what reward structures will incentivize individuals to adopt strategies that establish collective intelligence in the population.
Here, we propose two reward structures that can provably induce the emergence of collective intelligence in these naturalistic settings. We prove that collective intelligence is the asymptotically stable state for both of these structures. One reward structure, called niche expert, is based on the accuracy of each individual’s prediction. Individuals who predict more accurately, and especially those who are observing a factor that is otherwise poorly represented in the population, are rewarded with higher payoffs. The second reward structure, called feedback, is based on the error of the population’s collective prediction. The feedback structure does not reward experts, but rather those individuals, called reformers, whose individual predictions may not be accurate but are nonetheless beneficial for revising the collective prediction toward the truth. Although both of these reward structures provably induce accurate collective predictions, we find that the feedback is more robust to alternative forms of prediction tasks (e.g., task switching or correlated factors) and it exhibits more rapid convergence to collective intelligence. These results imply that rewarding reformers is more effective than rewarding niche experts, in promoting collective intelligence.
Model
We consider a scenario where the environmental outcome changes over discrete time, under the influence of factors denoted by . Let represent the outcome of the environment and denote the value of factor at a given time. Throughout this study we assume a linear relationship between the environmental outcome and the factors (36), expressed as
| [1] |
where () represents the weight of factor ’s contribution to the environmental outcome, and is a constant. The values of the factors vary over discrete time steps, each sampled independently from a distribution with mean and variance , whereas the weights and remain constant (Fig. 1A). We will focus on the case of normally distributed random factors, but our results extend to any distribution with finite variance (see examples in SI Appendix, Fig. S1), as well as some cases in which the outcome is a nonlinear function of factors (SI Appendix, Fig. S2).
Fig. 1.

Individual and collective predictions. (A) In each time step the outcome of the environment, , is a linear combination of random factors () sampled independently from a Gaussian distribution. The linear coefficients () are unknown to individuals. Each individual can observe the value of only a single factor before making their personal prediction for the outcome of the environment. The true environmental outcome is observed only after all players have made their individual predictions. (B) An individual has a strategy composed of two components: which factor, denoted , the individual has chosen to observe, and how the individual believes that factor is correlated with the outcome, denoted and called the individual’s belief. Thus individual makes personal prediction . (C) The collective prediction of the entire population combines the personal predictions of all individuals. There are several ways to combine individual predictions into a collective prediction, expressed by different aggregating functions . We focus on two aggregating functions: averaging across the entire population (Eq. 2) or averaging across clusters of individuals who focus on the same factor (Eq. 5).
We now consider a population of individuals who are tasked with the problem of predicting the outcome of the environment, based on observations of the factors that determine it. Due to cognitive constraints or resource limitations, each individual can observe the value of only a single factor at each time step (Fig. 1B). In particular, when individual chooses to observe factor , they obtain its value . However, , which reflects the true weight of the factor’s influence, remains unknown to the individual. Individual then makes a personal prediction about the environmental outcome based on the observed factor and their belief about the correlation between the observed factor and the outcome. Let represent individual ’s belief: If , individual believes the factor is positively correlated with , whereas if individual believes the factor is negatively correlated with . The magnitude of (i.e., ) reflects the believed strength of the factor’s influence on the outcome. Thus, the personal prediction made by individual is given by . The strategy of individual is denoted by a pair , which prescribes which factor the individual chooses to observe, and how they believe it correlates with the environment.
In addition, we account for the presence of stubborn individuals who do not observe any factors but instead provide constant predictions, even as the values of environmental factors are resampled. To formalize this, we introduce a virtual factor , whose value always equals . Stubborn individuals can be viewed as adopting the strategy , so that their predictions are constant and expressed as . Introducing stubborn individuals is important for obtaining accurate collective predictions, since the combination of stubborn individuals’ beliefs serves to predict the constant term in the environmental outcome.
After all individuals observe one factor and provide their personal predictions, the collective prediction of the group is produced by aggregating the individuals’ opinions (Fig. 1C). Generally, the collective prediction is a function of individual predictions, i.e., . There are many ways to aggregate individual opinions, with different aggregation rules corresponding to different aggregation functions. In this paper, we focus on two kinds of aggregating rules: averaging and clustering. Under the averaging rule, each individual’s prediction is treated as an estimate of the final outcome . The collective prediction is obtained by calculating the average of all individuals’ opinions:
| [2] |
An alternative way to express this aggregate prediction is in terms of the predicted coefficients, , for each factor:
| [3] |
where represents the set of individuals observing factor . The collective prediction can then be expressed as the linear form, .
In contrast to aggregation by equal-weight averaging, the clustering aggregation rule is based on the idea that each individual tries to predict only the contribution of the factor they observe (i.e., ) rather than the final outcome . Under the clustering aggregation rule, therefore, the predicted coefficient is calculated as the average belief among the individuals observing the factor (rather than averaging over all individuals):
| [4] |
Here, denotes the size of cluster , meaning those individuals who observe factor : . The collective prediction is then obtained by summing the predicted effects of all factors,
| [5] |
where any factor with (i.e., no observers) is excluded from the sum.
After a single instance of sampling the environment and forming predictions, individual is rewarded by a payoff according to their performance. In the next section, we introduce several payoff structures that measure the performance of individuals in different ways. After many subsequent instances of prediction (i.e., repeated sampling of the environment), each individual obtains a stationary payoff (i.e., expected payoff ). Individuals then update their strategies by imitating those with larger payoffs. In some parts of our analysis, we also include “mutation” or innovation so that individuals can explore new strategies not presently in the population.
More specifically, strategy imitation occurs as follows. In each discrete generation, two individuals and are selected uniformly at random. With probability , imitates ’s strategy with probability (37, 38)
| [6] |
or adopts a randomly generated strategy with probability . Here, is called the intensity of selection, which quantifies the degree to which payoff affects imitation dynamics. This formulation is standard in evolutionary game theory to describe how agents with bounded rationality update strategies by social learning (39–42), as opposed to perfectly rational agents who would immediately compute Nash equilibria.
Using stationary payoffs to update strategies means that the timescale of repeated environmental sampling and prediction is much faster than that of strategic imitation. In SI Appendix, Fig. S3, we explore alternative relative timescales, including cases where the strategy updating takes place after each instance of prediction. The results demonstrate that our findings remain robust across different timescale configurations. Therefore, our subsequent theoretical analysis continues to assume stationary payoffs (rapid environmental sampling and prediction).
Individuals undergoing the imitation process described above will adjust their strategies over time in pursuit of higher payoffs, which in turn affects their prediction accuracy and subsequent rewards. We study what kind of reward structures can foster collective intelligence. In particular, we propose two reward structures that are based on different forms of information: one reward structure that evaluates an individual’s performance based on their personal success (i.e., the accuracy of their individual prediction) and another payoff that evaluates each individual by the contribution to the collective success (i.e., the error of the group’s prediction).
To evaluate the accuracy of the collective prediction, we define the following metric
| [7] |
Here, the expectation is taken over all possible values of the factors . A larger value of indicates higher prediction accuracy. Specifically, implies that the collective prediction perfectly matches the real outcomes. Importantly, there is no lower bound for , as prediction errors can grow arbitrarily large.
Results
Payoffs Based on Personal Accuracy.
We first examine payoff structures that are based on personal accuracy. A natural reward structure that might perhaps engender collective accuracy is to reward those individuals with larger payoffs when their personal predictions are more accurate. Given individual ’s personal prediction (), such a straightforward payoff structure for individual can be expressed as
| [8] |
Under this payoff structure, those individuals whose predictions are closer to the actual outcome receive higher payoffs, and so we refer to this formulation as the “expert” payoff structure. After a sufficient number of predictions and environmental draws, the expected payoff for individual is given by the expectation of Eq. 8 over all possible values of the factors , namely
| [9] |
where is a constant (the same for all individuals). If (i.e., player is a stubborn individual), the expected payoff still has the same form as Eq. 9 by assuming . Under the expert payoff function, the expected payoff for each individual is independent of the strategies of other individuals. Moreover, among all individuals who choose to observe factor the one whose belief is closest to will obtain the highest payoff, . Thus, within the entire population, the individual who observes the most valuable factor (i.e., for all ) and whose belief is closest to will obtain highest payoff.
Over the course of strategic evolution by imitation, under the expert payoff scheme, all individuals will ultimately converge to adopt this most profitable strategy, . In the stationary state, then, all individuals choose to observe the same factor (factor ) and have the same belief coefficient. As a result, in the absence of mutations (), no matter what aggregation function is employed, the predicted outcome will incorporate no information from factors that are not being observed, so that the collective accuracy can never reach a high level. On the other hand, if the mutation rate is very high then most individuals will employ random strategies with poor estimates of the true coefficients, resulting again in low collective accuracy. For a narrow range of intermediate mutation rates the collective prediction under expert payoffs can reach relatively high accuracy, but perfect prediction is still impossible (SI Appendix, Fig. S4).
The key reason why the expert payoff scheme does not guarantee collective accuracy is that opinion diversity is lost during the process of evolution by social learning. To avoid a population that is entirely focused on a single factor, we propose the “niche expert” payoff structure by setting
| [10] |
Here, is the proportion of individuals in the population who are observing the same factor as individual , . The niche-expert payoff depends not only on the accuracy of an individual’s prediction but also on the frequency of other individuals who are observing the same factor. When a large proportion of individuals observe the same factor , this will decrease the payoffs of those who observe , even though some of these individuals’ personal predictions may be very close to the real outcome, .
We can intuitively understand the equilibrium state for the niche-expert payoff by extending our analysis of the expert payoff case. Given the niche-expert payoff function (Eq. 10), if a factor is observed by a very small proportion of individuals, then choosing to observe factor will produce payoff near , which is the highest payoff possible. As a result of this frequency dependence, more individuals will tend to switch to observe factor . In equilibrium, for , all factors will be observed by a substantial portion of individuals. And among the individuals who observe the same factor, , those whose belief is closer to the true value will obtain a higher payoff, as we saw in our analysis of the expert payoff function. As a result, in the absence of mutation, among the cluster of individuals with strategy observing factor , the strategy whose belief is closest to true coefficient will eventually fix within the cluster, or subpopulation. In other words, in equilibrium all individuals observing factor will have belief . This is the unique stable equilibrium of the evolutionary system, under the expert payoff function.
For the niche-expert payoff, provided the initial composition of the population contains observers for all different factors, and also contains observers whose beliefs are close to the true coefficient for each factor, then the equilibrium collective accuracy under clustering aggregation (Eq. 5) is guaranteed to be close to . By extension, if the initial population contains at least one individual with perfectly accurate belief, for each factor, then the equilibrium collective accuracy will be perfect: . Finally, in the presence of rare mutations, (), the equilibrium collective accuracy is guaranteed to be close to , regardless of initial composition in the population. In SI Appendix, section S2.1, we provide more details about this equilibrium and its provable stability properties. Note that the niche expert does not succeed under the averaging aggregation; and so all subsequent results for the niche expert assume clustering aggregation.
We performed Monte Carlo simulations to verify the effectiveness of the expert and niche-expert payoff structures. All individuals are assigned random initial strategies (observe a random factor with a personal belief sampled from a Gaussian distribution). Fig. 2A shows that expert payoff equipped with the clustering aggregation rule cannot drive the population to produce accurate predictions; but the niche expert equipped with the clustering rule can induce accurate collective predictions. In the stationary state, under the expert payoff nearly all individuals are focused on the same factor (Fig. 2B). But under the niche-expert payoff, each factor is observed by an abundance of individuals and individuals’ beliefs are close to the real coefficients (Fig. 2C, and also see SI Appendix, Fig. S5).
Fig. 2.
The niche-expert payoff structure can facilitate collective intelligence. We consider two kinds of payoff structures that are based on the accuracy of individuals’ personal predictions: the expert (Eq. 8) and niche expert (Eq. 10). Both payoff structures are equipped with the clustering aggregation rule. Under the expert payoff function, an individual receives a higher payoff if their prediction is more accurate (closer to the real outcome); where under niche expert, an individual receives a higher payoff not only if their prediction is accurate but also if few individuals are focused on the same factor as them. (A) The expert payoff function can promote collective accuracy in early stages of behavioral evolution, but collective accuracy subsequently collapses. Whereas the niche-expert structure incentivizes individuals to focus on different factors, which guarantees long-term collective accuracy. (B and C) We show simulation results when the outcome is determined by factors, with coefficients . The stationary strategy distribution in the population for the expert payoff structure (B) is less diverse than for the niche-expert payoff structure (C). Each dot represents an individual’s strategy, including their belief and which factor they observe. (B) For expert payoff, nearly all individuals focus on the factors with the strongest effect on the outcome; but most factors are not observed by any individuals at all. (C) The niche-expert payoff function ensures a diversity of factors being observed, and it results in beliefs that are close to the real coefficient, for each factor. Parameters: , , , (uniform distribution) (A), , (normal distribution), and .
Payoffs Based on Collective Error.
For the expert and niche-expert reward structures, the core idea is that individuals with more accurate personal predictions receive higher payoffs. Thus, these payoffs are based on the accuracy of personal predictions alone. These payoff structures do not incorporate any consideration of the collective prediction produced by the population. Next we consider a different payoff structure that is based on the error of the collective prediction (i.e., ) and how an individual contributes to it:
| [11] |
We call this payoff structure “feedback” because it rewards individuals whose contributions bring the collective prediction closer to the true value. In particular, when collective prediction is lower than the real outcome (i.e., ), individuals whose personal predictions are larger than the collective will receive higher payoff, and vice versa. The core idea of the feedback scheme is not to reward experts (individuals who predict more accurately), but rather to reward individuals whose predictions have more potential to improve the collective prediction, even if their personal predictions may deviate substantially from the outcome . We call such individuals “reformers.”
Unlike a payoff formulation based on personal accuracy, the feedback payoff structure depends on (and requires information about) the collective prediction . As a result, the feedback payoff function depends on the choice of aggregating rule used to form the collective prediction. We consider two versions of the feedback payoff arising from either the average aggregation rule (Eq. 2) or the clustering aggregation rule (Eq. 5).
For a feedback payoff structure (Eq. 11) we can compute the stationary payoff for individual :
| [12] |
Here, is the predicted coefficient of the factor, given by Eq. 3 under the average aggregation rule or by Eq. 4 under the clustering aggregation. It follows from Eq. 12 that all individuals have the same stationary payoff if and only if for all factors . So this is an equilibrium state of the evolutionary process, and the collective prediction will equal the real outcome in equilibrium (). In SI Appendix, section S2.2, we construct a Lyapunov function for this system in the absence of mutation, which decreases monotonically as evolution proceeds until all individuals have the same payoff. Thus, for any initial strategy distribution, provided the real coefficients are in the predictable domain (i.e., given the strategies in the population, there exists at least one strategic composition with equals for all ), the equilibrium state is asymptotically stable. When mutations occur at a low rate, , the collective prediction is still accurate (), and the predictable domain can even increase in size.
Unlike the niche expert, which requires clustering aggregation to produce collective accuracy, the feedback payoff can facilitate collective intelligence under both clustering and averaging aggregation (Fig. 3A). In equilibrium individuals’ beliefs need not be close to the real coefficient, as in the case of the niche-expert payoff, but rather the predicted coefficients must be close to the real coefficients . Thus, the strategy distribution in equilibrium under feedback is not unique, but depends on initial conditions. The equilibrium strategy distribution under feedback payoff is also much more diverse than under the expert or even niche-expert payoff functions (Fig. 3 B and C).
Fig. 3.

Feedback payoff structures can facilitate collective intelligence. The feedback payoff function (Eq. 11) rewards individuals not based on their personal accuracy, but based on a comparison between individual beliefs and collective accuracy. The core idea of this payoff structure is to reward those, called reformers, whose beliefs make the collective prediction more accurate. (A) The feedback payoff has the general form , although the specific function depends upon the aggregation rule used to form the collective prediction (averaging or clustering, Eq. 2 or Eq. 5 respectively). Both averaging feedback and clustering feedback guarantee the long-term accuracy of the collective prediction. (B and C) Unlike payoff functions based on personal accuracy alone (Fig. 2), for feedback structures the stationary beliefs of individuals observing the same factor need not converge to the true coefficient, , but rather are dispersed. Furthermore, the stationary strategy distribution is not unique, but depends on the initial distribution of strategies in the population. Parameters: , , , (B), (C), , and .
Although both clustering feedback and averaging feedback can generate accurate collective prediction, they require different scales of individuals’ beliefs (i.e., the values ). For a fixed population size , with the averaging rule, the predicted coefficient is given by . Thus, if there are more factors, there will be fewer individuals observing each factor, which implies that the magnitude of should be much larger than in order that . For example, when the true coefficients lie in , the initial personal beliefs obey a normal distribution with SD , which is much greater than the true coefficients (Fig. 3). When the environment depends on a larger number of factors ( larger), the magnitudes of initial beliefs should be greater yet. Under clustering feedback, however, no matter how large the problem scale , the magnitude of beliefs can remain the same as the true coefficients, because has the same scale as beliefs (SI Appendix, Fig. S6).
Comparison Between Niche-Expert and Feedback Payoffs.
Although both reward structures have the ability to foster collective intelligence, the niche-expert system has different properties, and outcomes, than the feedback system. To make a fair comparison between the two, we compare niche expert and feedback with the same aggregation rule, clustering aggregation, which requires the same magnitude of beliefs . (Feedback with average aggregation has similar properties as with clustering feedback, except for different requirements on the magnitude of beliefs, see SI Appendix, Fig. S6.)
Niche-expert and feedback payoff schemes respond differently to variation in the nature of the prediction task. We consider two scenarios. In the first scenario, we explore a form of task switching: After the population has already established collective intelligence, we assume that there is a shock so that all coefficients are resampled, requiring that the population find a solution to a different prediction problem. Following the shock, we observe that a population evolving under feedback reward structure quickly recovers and reconverges to a state of perfectly accurate collective prediction. However, under the niche-expert reward structure, it is not guaranteed that the population can reestablish collective intelligence, when mutations are absent (). The reason for this difference arises from the diversity of opinions supported under these two reward structures, shown in Figs. 2C and 3C. When the niche expert reaches equilibrium the belief diversity is lost. All individuals hold beliefs that are equal (or close) to the true coefficients. And so when a shock occurs and the coefficients are resampled from, e.g., a normal distribution, there will be no individuals whose beliefs are close to the new coefficients—so that imitation will not allow the niche-expert system to recover, except in the presence of innovation (Fig. 4A). Under feedback rewards, by contrast, the equilibrium state supports a diversity of beliefs (Fig. 4C). Even if the environment is changed, the new coefficients can still be expressed by individuals’ beliefs after subsequent imitation events.
Fig. 4.
Feedback payoffs are more robust than niche experts to variation in the nature of the prediction task. We consider two variant forms of prediction tasks for an environment that is a linear combination of factors: environmental shocks or correlated factors. We compare the performance of the collective predictions that emerge from feedback payoff versus niche expert, both under the clustering aggregating rule. (A) After reaching the stationary state of perfect collective accuracy, we introduce an environmental shock such that all coefficients of the independent factors are randomly redrawn. Under the clustering feedback payoff, the population rapidly adapts to recover near-perfect collective accuracy. The niche expert can also recover, although more slowly and only in the presence of innovations (mutations) during imitation. Panel (B) illustrates the average variance of beliefs within each cluster, which provides intuition for results in panel (A). For the niche expert, the beliefs within a cluster (individuals focused on the same factor) are all close to the true coefficient in equilibrium, so that the intracluster belief variance decreases toward zero over time. After an environmental shock, beliefs adapt to regain collective accuracy only through the accumulation of many mutations. For the clustering feedback, by contrast, belief diversity within clusters is retained in equilibrium, which guarantees that the new environment can still be expressed by individuals’ beliefs (even without mutation). (C) When the factors are correlated with covariance structure , clustering feedback still produces collective accuracy, but the niche expert fails. (D) To understand the failure of the niche expert for correlated factors we focus on individuals within one cluster, : Each blue dot represents such an individual’s belief. The blue line represents the predicted coefficient (average of beliefs ) within cluster , which converges to a value far from the true coefficient , so that collective prediction cannot be accurate. Note that predicted coefficient within cluster starts near zero, passes through the true value of , and then moves away from the true value, which also explains why the accuracy of niche expert first increases and then decreases (C). Parameters: , , , , , and .
We have also explored prediction tasks when the random factors are correlated with each other. In this case, the niche expert cannot maintain a high level of collective accuracy. However, clustering feedback still provides robust collective intelligence (Fig. 4B). In SI Appendix, section S3.1, we show that for niche-expert payoff, if the factors are correlated, among the individuals observing the same factor the individual with the highest payoff is not the one whose belief equals the coefficient . So in equilibrium, the individuals observing the same factor will have the same belief but not equal to , which means the predicted coefficient (Eq. 4) never equals the real coefficient (Fig. 4D, also see SI Appendix, Fig. S7). Under clustering feedback, by contrast, even though the expected payoff for an individual modified by correlated factors, is still the stable equilibrium state of the system (SI Appendix, section S3.1); and so collective accuracy can still be ensured. Furthermore, if we assume that some factors that influence the outcome can never be observed, the performance of clustering feedback is again far superior to niche expert. In this case, strong correlations between factors will promote collective accuracy under clustering feedback, compared to the case of no correlation for clustering feedback; but correlations of any strength impede accuracy for niche expert (SI Appendix, Fig. S8). This is because feedback payoffs can incentivize individuals to infer partial information about unobservable factors through their correlation with observable factors.
The niche-expert system can be modified, by changing the aggregating rule (SI Appendix, section S3.1), to produce accurate predictions even with correlated factors. However, this form aggregation requires information about the covariance structure of the factors—which may in practice be unknowable—and it works well only if the factors are weakly correlated. Whereas clustering feedback does not require information about covariance structure and it is robust against even strong correlations between factors (SI Appendix, Fig. S9).
Niche-expert and feedback rewards also differ in their response to population size, . In both cases, a large population is always beneficial for collective accuracy (Fig. 5). But the underlying mechanisms are qualitatively distinct. The effectiveness of the niche expert relies on the fact that there are some experts whose beliefs are very close to the real coefficients; and then other individuals will imitate those experts’ strategies. A larger population helps ensure that some individuals’ beliefs are close to the true coefficients, in the initial (randomly generated) composition of the population. Note that even though mutations can introduce new strategies whose beliefs may be close to the true coefficients, small populations remain unable to obtain accurate predictions, because there are very few individuals in each cluster, so that most strategy imitation occurs between different clusters without improving collective accuracy. By contrast, the feedback system does not require that someone’s belief be close to the real coefficient: Accuracy is achieved if the average belief for each factor (i.e., ) equals to the true coefficient. In a large population, the marginal change of after each updating event is relatively small, which helps bring the predicted coefficient close to the real coefficient without overshoot events. And so the effect of population size has to do with the initial population composition for the niche expert; whereas it drives the dynamics of evolution for clustering feedback.
Fig. 5.
Large populations increase the accuracy of collective predictions. We present results for the niche-expert (A) and feedback payoff (B) structures (both with clustering aggregation), with initial beliefs in the population drawn uniformly at random. For a given population size, both payoff schemes perform better when the problem size is smaller (fewer factors, ). However, as the population size increases so does the collective accuracy, eventually reaching perfect accuracy. Evolution under the feedback payoff scheme requires a smaller population size to achieve perfectly accurate collective predictions, compared to the niche-expert scheme. Parameters: , , , and .
Reward structures also have different behavior with respect to convergence speed—which grows with problem complexity under niche expert but is insensitive to problem complexity for feedback (SI Appendix, Fig. S10). Moreover, we have expanded our analysis to include situations when individuals can observe multiple factors at the same time (SI Appendix, section S3.2). In this setting as well, the feedback payoff structure engenders robust collective intelligence, whereas niche expert provides collective accuracy only when each individual observes either a few factors or nearly all factors (SI Appendix, Fig. S11).
Discussion
Collective intelligence, or the “wisdom of crowds,” is a pervasive phenomenon in real-world communities. Whether in decision-making through voting, scientific collaboration, technological innovation, or managing global economic systems, effective collaboration and organization among large numbers of individuals is indispensable for addressing complicated tasks (25–27, 43–45). A large body of evidence has shown that groups of individuals can often outperform single individuals in diverse tasks (1–5, 29). However, large population size alone is not sufficient to ensure the emergence of collective intelligence.
One approach to collective intelligence is to engineer algorithms for efficient distributed computation (6–11, 46, 47). Research in control theory and machine learning has produced remarkable solutions to distributed problem solving. These solutions are intended for execution by computers, though, and they do not describe how natural populations might achieve collective accuracy. Natural populations—even those under common governance—lack the degree of sophisticated central control required to implement solutions from engineering, which involve complex task allocation and dynamic feedback control. Instead, natural populations can access only very simple central authorities that might, for example, aggregate individual opinions and broadcast their average.
In this study, we have developed and analyzed two forms of reward structures (payoff functions) that effectively incentivize accurate collective predictions in a naturalistic population of individuals with limited information. The niche-expert scheme rewards individuals who not only make accurate predictions but also attend to information from underrepresented areas. This ensures a more comprehensive exploration of the problem space. The underlying mechanism of this payoff structure is reminiscent of division of labor in economics (48, 49), where the “invisible hand” balances the supply and demand of labor across various occupations. In equilibrium, switching to another occupation—or attending to another factor—yields no additional benefit.
The second type of reward structure, feedback payoff, differs fundamentally from the niche expert. Rather than rewarding individuals for the accuracy of their own predictions, this structure rewards individuals based on their contribution to reducing the collective prediction error. Reformers who can most effectively reduce the collective error—even if their own predictions are inaccurate—receive higher payoffs. This is roughly analogous to gradient descent optimization, where parameters are adjusted to minimize a global loss function (50, 51).
We have shown that the feedback scheme generates more effective collective accuracy across a wide range of prediction tasks, including environmental shocks or correlated factors. The primary distinction lies in the objective of the individuals under each payoff structure: Niche experts aim to optimize their own predictions, whereas feedback participants focus on improving collective predictions. The alignment of individual goals with collective objectives in the feedback scheme provides more robust performance.
In order to connect our theoretical results to empirical examples of collective intelligence, we must consider what real-world mechanisms could instantiate the niche expert or feedback reward structures we have analyzed. As an illustrative example, the niche-expert reward structure may naturally arise in the realm of scientific collaboration. Scientists are rewarded (by coauthorship on publications and grants) if they have expertise in one aspect of a larger problem (that is, in a single factor) and especially so if their expertise is rare within the scientific community so that they are in demand as coauthors. And so a form of niche-expert reward structure arises naturally in this setting, even without a central authority to measure the frequency of different strategies; it simply requires that individuals with rare expertise be disproportionally rewarded with scientific output.
Likewise, the feedback payoff structure is natural in the setting of a prediction market or a stock market. In this context, individuals buy and sell contracts whose prices reflect the group’s aggregate belief about a future event, such as the outcome of an election. A trader’s profit does not depend on whether his personal forecast is closest to the truth, but on whether his trades move the market price toward the eventual outcome. If a trader pushes the market price upward when the event indeed occurs, he profits; if he pushes the price downward in that case, he loses money. Here, a feedback payoff emerges endogenously from the trading mechanism: Individuals who act as reformers—by nudging the collective prediction closer to the truth—are automatically rewarded, even if their private estimates would not have been accurate in isolation. The same basic feedback operates in stock markets: When a stock is undervalued, traders who buy—even based on overly optimistic predictions—earn profits as the price moves upward.
Although these examples are illustrative of niche-expert and feedback rewards, we emphasize that our model provides only a simplified description of these real-world scenarios. The purpose of the simple model is not to produce a detailed, quantitative description of behavior among scientific collaborators or market traders, but rather to give intuition in a comprehensible framework for how reward structures can stimulate collective accuracy through social learning.
There is an extensive body of empirical literature on collective intelligence in natural populations (1–5, 29, 52–55). Corresponding theoretical studies have proposed various methods to improve collective intelligence in this setting (10, 30–32, 56, 57), and yet the field still lacks a systematic framework for studying how individual-level incentives promote collective accuracy. Most prior theory has focused on estimation tasks where the underlying truth is a constant or a single variable unaffected by external factors (10, 32, 56–58). These models, often grounded in opinion dynamics such as the DeGroot model (56, 58), assume that individuals’ opinions are influenced by their neighbors regardless of their performance. In contrast, our approach uses evolutionary game theory to model social learning through payoff-biased imitation, which describes individuals who are self-motivated and boundedly rational. While a few prior studies have also explored multifactor correlated truths, they are limited to binary voting tasks where both the underlying factors and environmental outcomes are binary (30, 31).
There are several shortcomings of our current analysis, and many directions for future research. We have assumed the population structure is well mixed, yet previous research has demonstrated that population structure can strongly influence information aggregation and behavioral spread by imitation (39, 59–63). There is some promise, then, that the reward structures we have developed may be more effective in structured populations, or they may require different aggregation rules that account for structure. We have also assumed that individuals give predictions based on the (true) value of observed factors. However, individuals might intentionally report false values of factors, in an attempt to achieve higher payoffs. Incentivizing individuals to provide honest predictions may require additional mechanisms, which remain to be explored. Finally, our study is limited to problems of collective prediction, as opposed to more general problem-solving tasks. Prediction of a continuous variable still encompasses a huge range of real-world problems that can be complex; but our analysis still assumes a linear relationship between the outcome and the underlying observable factors, which may be violated in real-world scenarios. Adapting our payoff structures and aggregation rules to account for nonlinear prediction problems is a significant challenge that remains an open problem for future research.
Supplementary Material
Appendix 01 (PDF)
Acknowledgments
G.W. acknowledges support from the China Scholarship Council (No. 202306010132). Q.S. acknowledges support from the National Natural Science Foundation of China (No. 62473252) and from the State Key Laboratory of Autonomous Intelligent Unmanned Systems (No. ZZKF2025-1-4). L.W. acknowledges support from the National Natural Science Foundation of China (No. 62533002 and 62036002). J.B.P. acknowledges support from the US Army Research Office (award W911NF2410393) and the US Office of Naval Research (award N000142412778).
Author contributions
G.W., Q.S., L.W., and J.B.P. designed research; G.W. performed research with J.B.P.; G.W., Q.S., L.W., and J.B.P. analyzed data; and G.W., Q.S., and J.B.P. wrote the paper with input from L.W.
Competing interests
The authors declare no competing interest.
Footnotes
This article is a PNAS Direct Submission.
Contributor Information
Qi Su, Email: qisu@sjtu.edu.cn.
Long Wang, Email: longwang@pku.edu.cn.
Joshua B. Plotkin, Email: jplotkin@sas.upenn.edu.
Data, Materials, and Software Availability
There are no data underlying this work.
Supporting Information
References
- 1.Galton F., Vox populi. Nature 75, 450–451 (1907). [Google Scholar]
- 2.Woolley A. W., Chabris C. F., Pentland A., Hashmi N., Malone T. W., Evidence for a collective intelligence factor in the performance of human groups. Science 330, 686–688 (2010). [DOI] [PubMed] [Google Scholar]
- 3.Yaniv I., Milyavsky M., Using advice from multiple sources to revise and improve judgments. Organ. Behav. Hum. Decis. Process. 103, 104–120 (2007). [Google Scholar]
- 4.Hommes C., Sonnemans J., Tuinstra J., Van De Velden H., Coordination of expectations in asset pricing experiments. Rev. Financ. Stud. 18, 955–980 (2005). [Google Scholar]
- 5.Lorge I., Fox D., Davitz J., Brenner M., A survey of studies contrasting the quality of group performance and individual performance, 1920–1957. Psychol. Bull. 55, 337–372 (1958). [DOI] [PubMed] [Google Scholar]
- 6.Durfee E. H., Lesser V. R., Corkill D. D., Trends in cooperative distributed problem solving. IEEE Trans. Knowl. Data Eng. 1, 63–83 (1989). [Google Scholar]
- 7.Decker K. S., Distributed problem-solving techniques: A survey. IEEE Trans. Syst. Man Cybern. 17, 729–740 (1987). [Google Scholar]
- 8.Smith R. G., Davis R., Frameworks for cooperation in distributed problem solving. IEEE Trans. Syst. Man Cybern. 11, 61–70 (1981). [Google Scholar]
- 9.K. Zhang, Z. Yang, T. Başar, “Multi-agent reinforcement learning: a selective overview of theories and algorithms” in Handbook of Reinforcement Learning and Control, K. G. Vamvoudakis, Y. Wan, F. L. Lewis, D. Cansever, Eds. (Springer International Publishing, Cham, 2021), vol. 325, pp. 321–384.
- 10.Budescu D. V., Chen E., Identifying expertise to extract the wisdom of crowds. Manag. Sci. 61, 267–280 (2015). [Google Scholar]
- 11.De Vincenzo I., Massari G. F., Giannoccaro I., Carbone G., Grigolini P., Mimicking the collective intelligence of human groups as an optimization tool for complex problems. Chaos Solitons Fractals 110, 259–266 (2018). [Google Scholar]
- 12.Aplin L. M., Farine D. R., Mann R. P., Sheldon B. C., Individual-level personality influences social foraging and collective behaviour in wild birds. Proc. R. Soc. B Biol. Sci. 281, 20141016 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.King A. J., Cowlishaw G., When to use social information: The advantage of large group size in individual decision making. Biol. Lett. 3, 137–139 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Taylor P. D., Jonker L. B., Evolutionary stable strategies and game dynamics. Math. Biosci. 40, 145–156 (1978). [Google Scholar]
- 15.Schuster P., Sigmund K., Replicator dynamics. J. Theor. Biol. 100, 533–538 (1983). [Google Scholar]
- 16.Boyd R., Richerson P. J., Culture and the Evolutionary Process (University of Chicago press, 1988). [Google Scholar]
- 17.Szabó G., Tőke C., Evolutionary prisoner’s dilemma game on a square lattice. Phys. Rev. E 58, 69–73 (1998). [Google Scholar]
- 18.Hofbauer J., Sigmund K., Evolutionary game dynamics. Bull. Am. Math. Soc. 40, 479–519 (2003). [Google Scholar]
- 19.Traulsen A., Semmann D., Sommerfeld R. D., Krambeck H. J., Milinski M., Human strategy updating in evolutionary games. Proc. Natl. Acad. Sci. U.S.A. 107, 2962–2966 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Fu F., Rosenbloom D. I., Wang L., Nowak M. A., Imitation dynamics of vaccination behaviour on social networks. Proc. R. Soc. B Biol. Sci. 278, 42–49 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Lorenz J., Rauhut H., Schweitzer F., Helbing D., How social influence can undermine the wisdom of crowd effect. Proc. Natl. Acad. Sci. U.S.A. 108, 9020–9025 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Surowiecki J., The Wisdom of Crowds (Random House, New York, 2005). [Google Scholar]
- 23.Centola D., The network science of collective intelligence. Trends Cogn. Sci. 26, 923–941 (2022). [DOI] [PubMed] [Google Scholar]
- 24.Zafeiris A., Vicsek T., Group performance is maximized by hierarchical competence distribution. Nat. Commun. 4, 2484 (2013). [DOI] [PubMed] [Google Scholar]
- 25.Jegatheesan V., Liow J. L., Shu L., Kim S. H., Visvanathan C., The need for global coordination in sustainable development. J. Clean. Prod. 17, 637–643 (2009). [Google Scholar]
- 26.Bi K., et al. , Accurate medium-range global weather forecasting with 3D neural networks. Nature 619, 533–538 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Helbing D., et al. , Saving human lives: What complexity science and information systems can contribute. J. Stat. Phys. 158, 735–781 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Purves D., et al. , Time to model all life on earth. Nature 493, 295–297 (2013). [DOI] [PubMed] [Google Scholar]
- 29.Kurvers R. H. J. M., et al. , Automating hybrid collective intelligence in open-ended medical diagnostics. Proc. Natl. Acad. Sci. U.S.A. 120, e2221473120 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Hong L., Page S. E., Riolo M., Incentives, information, and emergent collective accuracy. Manag. Dec. Econ. 33, 323–334 (2012). [Google Scholar]
- 31.Mann R. P., Helbing D., Optimal incentives for collective intelligence. Proc. Natl. Acad. Sci. U.S.A. 114, 5077–5082 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Mann R. P., Optimising collective accuracy among rational individuals in sequential decision-making with competition. Collect. Intell. 2, 263391372311764 (2023). [Google Scholar]
- 33.Smith J. M., Evolution and the Theory of Games in Did Darwin Get It Right? (Springer, Boston, MA, 1988), pp. 202–215. [Google Scholar]
- 34.Nowak M. A., Evolutionary Dynamics: Exploring the Equations of Life (Harvard University Press, Cambridge, MA, 2006). [Google Scholar]
- 35.Young H. P., The evolution of social norms. Annu. Rev. Econ. 7, 359–387 (2015). [Google Scholar]
- 36.Wooldridge J. M., Introductory Econometrics: A Modern Approach (Cengage learning, Boston, ed. 6, 2016). [Google Scholar]
- 37.Traulsen A., Pacheco J. M., Nowak M. A., Pairwise comparison and selection temperature in evolutionary game dynamics. J. Theor. Biol. 246, 522–529 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Ohtsuki H., Nowak M. A., The replicator equation on graphs. J. Theor. Biol. 243, 86–97 (2006). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Ohtsuki H., Hauert C., Lieberman E., Nowak M. A., A simple rule for the evolution of cooperation on graphs and social networks. Nature 441, 502–505 (2006). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Harms W., Skyrms B., “Evolution of moral norms” in The Oxford Handbook of Philosophy of Biology, M. Ruse, Ed. (Oxford University Press, 2008).
- 41.Wu B., Altrock P. M., Wang L., Traulsen A., Universality of weak selection. Phys. Rev. E 82, 046106 (2010). [DOI] [PubMed] [Google Scholar]
- 42.Wu B., García J., Hauert C., Traulsen A., Extrapolating weak selection in evolutionary games. PLoS Comput. Biol. 9, e1003381 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Meluso J., Hébert-Dufresne L., Multidisciplinary learning through collective performance favors decentralization. Proc. Natl. Acad. Sci. U.S.A. 120, e2303568120 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Heins C., et al. , Collective behavior from surprise minimization. Proc. Natl. Acad. Sci. U.S.A. 121, e2320239121 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Noothigattu R., et al. , A voting-based system for ethical decision making. Proc. AAAI Conf. Artif. Intell. 32, 1587–1594 (2018). [Google Scholar]
- 46.Wang L., Xiao F., Finite-time consensus problems for networks of dynamic agents. IEEE Trans. Autom. Control 55, 950–955 (2010). [Google Scholar]
- 47.Xiao F., Wang L., Asynchronous consensus in continuous-time multi-agent systems with switching topology and time-varying delays. IEEE Trans. Autom. Control 53, 1804–1816 (2008). [Google Scholar]
- 48.Rodríguez-Clare A., The division of labor and economic development. J. Dev. Econ. 49, 3–32 (1996). [Google Scholar]
- 49.Becker G. S., Human capital, effort, and the sexual division of labor. J. Lab. Econ. 3, S33–S58 (1985). [Google Scholar]
- 50.S. Ruder, An overview of gradient descent optimization algorithms. arXiv [Preprint] (2016). http://arxiv.org/abs/1609.04747 (Accessed 17 May 2025).
- 51.Rumelhart D. E., Hinton G. E., Williams R. J., Learning representations by back-propagating errors. Nature 323, 533–536 (1986). [Google Scholar]
- 52.Tchernichovski O., Frey S., Jacoby N., Conley D., Incentivizing free riders improves collective intelligence in social dilemmas. Proc. Natl. Acad. Sci. U.S.A. 120, e2311497120 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Bernstein E., Shore J., Lazer D., How intermittent breaks in interaction improve collective intelligence. Proc. Natl. Acad. Sci. U.S.A. 115, 8734–8739 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Becker J., Brackbill D., Centola D., Network dynamics of social influence in the wisdom of crowds. Proc. Natl. Acad. Sci. U.S.A. 114, E5070–E5076 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Becker J., Porter E., Centola D., The wisdom of partisan crowds. Proc. Natl. Acad. Sci. U.S.A. 116, 10717–10722 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.Golub B., Jackson M. O., Naïve learning in social networks and the wisdom of crowds. Am. Econ. J. Microecon. 2, 112–149 (2010). [Google Scholar]
- 57.Davis-Stober C. P., Budescu D. V., Dana J., Broomell S. B., When is a crowd wise? Decision 1, 79–101 (2014). [Google Scholar]
- 58.Tian Y., Wang L., Bullo F., How social influence affects the wisdom of crowds in influence networks. SIAM J. Control Optim. 61, 2334–2357 (2023). [Google Scholar]
- 59.Su Q., McAvoy A., Wang L., Nowak M. A., Evolutionary dynamics with game transitions. Proc. Natl. Acad. Sci. U.S.A. 116, 25398–25404 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.McAvoy A., Allen B., Fixation probabilities in evolutionary dynamics under weak selection. J. Math. Biol. 82, 14 (2021). [DOI] [PubMed] [Google Scholar]
- 61.Tarnita C. E., Ohtsuki H., Antal T., Fu F., Nowak M. A., Strategy selection in structured populations. J. Theor. Biol. 259, 570–581 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62.Rand D. G., Arbesman S., Christakis N. A., Dynamic social networks promote cooperation in experiments with humans. Proc. Natl. Acad. Sci. U.S.A. 108, 19193–19198 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Rand D. G., Nowak M. A., Fowler J. H., Christakis N. A., Static network structure can stabilize human cooperation. Proc. Natl. Acad. Sci. U.S.A. 111, 17093–17098 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Appendix 01 (PDF)
Data Availability Statement
There are no data underlying this work.



