Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2022 Sep 1.
Published in final edited form as: Sociol Methods Res. 2019 Nov 20;51(2):566–604. doi: 10.1177/0049124119882479

Meta-Analysis in Sociological Research: Power and Heterogeneity

Guangyu Tong 1, Guang Guo 2
PMCID: PMC9231456  NIHMSID: NIHMS1751944  PMID: 35754525

Abstract

Meta-analysis is a statistical method that combines quantitative findings from previous studies. It has been increasingly used to obtain more credible results in a wide range of scientific fields. Combining the results of relevant studies allows researchers to leverage study similarities while modeling potential sources of between-study heterogeneity. This paper provides a review of the core methodologies of meta-analysis that we consider most relevant to sociological research. After developing the foundation of the fixed-effects and random-effects models of meta-analysis, this paper illustrates the utility of the method with regression coefficients reported from two sets of social science studies. We explain the various steps of the process including constructing the meta-sample from primary studies; estimating the fixed- and random-effects models; analyzing the source of heterogeneity across studies; assessing publication bias. We conclude with a discussion of steps that could be taken to strengthen the development of meta-analysis in sociological research, which will eventually increase the credibility of sociological inquiry via a knowledge-cumulative process.

1. Introduction

Meta-analysis is a method that statistically summarizes evidence from previous studies on the same topic and in similar analytical styles (Cooper and Hedges 2009). Meta-analysis has two advantages. It increases statistical power, often crucially. It also provides a framework within which the heterogeneity of previous studies could be examined (Higgins and Thompson 2002). Heterogeneity refers to the substantial variation in results across studies due to different populations, study contexts, and analytical methods. Sociological research can benefit substantially from increased power and a systematic way of examining heterogeneity.

In many scientific disciplines, meta-analysis has become increasingly, widely, and routinely used to reconcile disagreement across empirical studies that answer the same question. For instance, meta-analysis is popular in psychology, and many psychological researchers view meta-analysis as an indispensable step that “informs us where the field has been”, “cleans up the field”, and “provides suggestions for future research” (Chan and Arvey 2012, p.80). In economics, meta-analytical is also widely used. According to Stanley et al. (2013), more than 200 meta-analysis on economic topics were published per year. It is more and more agreed that however perfect an individual study is conducted, its result is less convincing than a synthesized result.

Sociological researchers have also used meta-analysis to resolve theoretical debates and understand the differences across studies (Amato and Keith 1991; Branigan, McCallum, and Freese 2013; Hedges, Laine, and Greenwald 1994; Shor et al. 2012). By aggregating evidence from multiple studies, meta-analysis increases the statistical power to obtain more conclusive results. For example, Amato and Keith (1991) identify the negative consequence of parental divorce for offspring well-being in adulthood despite inconsistencies in primary results; Hedges et al. (1994) confirm a positive relationship between the input of school resources and student achievement with meta-analysis, though most primary studies suggest a null relationship. In addition, by utilizing study-level features, meta-analysis helps researchers understand the source of between-study heterogeneity. For example, Shor et al. (2012) show that cultural differences and widowhood recency can significantly moderate the effect of widowhood on mortality risk; Branigan et al. (2013) find that heritability of education attainment varies significantly by gender and cohort. These new findings are beyond the scope of a single study but become evident in meta-analysis.

Still, compared to other fields such as psychology and economics, meta-analysis is not used at a large scale in sociology. This is likely due to the high heterogeneity in sociological research. Sociological studies can be heterogeneous due to differences in study populations, sample designs, survey questions, control variables, statistical models, and so on. This heterogeneity is related to uncertainty and research transparency discussed by Young (2009), which are responsible for disagreeing results across sociological studies. Moreover, sociological researchers often motivate their studies by developing “new niches” that have not been studied before, fostering the belief that most sociological studies have not been replicated. Such belief is indeed largely inaccurate, and empirical evidence on many research questions in different subfields has already piled up. For example, large bodies of studies have accumulated on topics of intergeneration transmission of education, income, and wealth inequality (e.g., Mulder et al. 2009; Roksa and Potter 2011), gender pay gap due to motherhood penalty (e.g., Correll, Benard, and Paik 2007), how parents’ incarceration affects children’s outcomes (e.g., Johnson and Easterling 2012), whether strictness benefits the growth of churches (e.g., Thomas and Olson 2010), and so on. However, meta-analysis on these topics has been rarely attempted.

In this article, we provide a review of core methods of meta-analysis for sociological readers and discuss potential changes in research practice that can lead to a broader use of meta-analysis in sociology. We tailor the article to the sociological context and provide practical considerations of meta-analysis for sociological researchers. Our overall objective is to help popularize meta-analysis in sociology where it is appropriate. We provide an article-length text on meta-analysis most relevant to sociological researchers, which is self-contained, conveniently covering a relatively complete body of information on meta-analysis. Our article uses two sociologically realistic examples and focuses on the type of findings (unstandardized coefficients) most often reported in sociological papers. Meta-analysis was first developed and employed primarily by psychological researchers (Glass 1976; Hedges 1983). Correlation coefficients and standardized regression coefficients are typically reported in the primary studies and summarized in meta-analyses. Naturally, much of the meta-analysis literature is centered on the summary of correlation coefficients. In contrast, sociological studies typically report unstandardized regression coefficients. Our worked examples show that meta-analysis can still be used to summarize unstandardized coefficients from linear and logistic regression models, the two statistical models most commonly used in quantitative sociological research. Sociological publications may be more theoretically diversified and empirically heterogeneous than those in some other fields; our examples, nevertheless, demonstrate substantive benefits of meta-analysis in sociological research.

In Section 2, we discuss the common practical issues and corresponding coping strategies during the data preparation for meta-analysis. Moreover, because sociological publications tend to be more heterogeneous and heterogeneity appears to be a major obstacle for the more common use of meta-analysis, in Section 6, our article specifically discusses heterogeneity and provides detailed worked examples of heterogeneity analysis. Heterogeneity analysis is a method that treats the primary results as outcomes and studies their relationship to study-level covariates (Thompson and Higgins 2002). We introduce two methods to analyze heterogeneity, the ANOVA-type approach and meta-regression. Our worked examples particularly illustrate how heterogeneity analysis can be useful for sociological researchers to clarify the source of between-study variation and possibly lead to new insights for a research topic.

In Section 3, we derive the core methods of meta-analysis, fixed-effects and random-effects models, through arguments that are well-known to sociologists wherever we can. Both fixed- and random-effects models are derived in the form of the weighted least squares (WLS) models and can be estimated with WLS in practice. Section 4 illustrates the fixed- and random-effects models with two sociological data examples. In each example, we explain how the estimates can be obtained through hand calculation and using the WLS procedures in Stata. We also illustrate the use of developed Stata functions for meta-analysis “metan” (Harris et al. 2008). Readers interested in the implementation of meta-analysis on other platforms are encouraged to check “meta” and “metafor” packages on R (Viechtbauer 2015) and use “PROC MIXED” and other macros to do meta-analysis with SAS (Arthur Jr, Bennett, and Huffcutt 2001). Section 6 explains the concept of publication bias and how publication bias can be statistically examined and adjusted to yield more robust combined results. A worked example is also used to manifest the substantial benefit from examining publication bias.

In Section 7, we chart the future course of development in meta-analysis, and the discussion is specifically targeted on the changes in research practice of sociology that can lead to an increased use of meta-analysis. In order for meta-analysis to be broadly used, sociological researchers need to be aware of the benefit of meta-analysis and improve the reporting standard in primary studies (e.g., results with different combinations of controls, variance-covariance matrix). Reporting more abundant results in primary studies enables meta-analysis to be conducted more easily in sociology. In the future, researchers working on the same topic can make conscious efforts to build consortiums and collaboration networks that promote the standardized analysis and the sharing of individual-level data. Such practice can reduce the methodological heterogeneity to the minimum and pave the way for more definitive meta-analysis.

Over the past three decades, the methodology of meta-analysis has been gradually and successfully adopted in many scientific fields to summarize existing studies to obtain more credible results. It is high time that the methodology of meta-analysis is made widely accessible for sociological researchers. The methodology can help raise the level of credibility of sociological findings through a knowledge cumulative process.

2. Background

2.1. Literature search and data collection

Meta-analysis begins with collecting all existing evidence on a pre-specified research topic, so that synthesized results are immune from the selection of evidence. A researcher must first clearly define the scope of a research question and set inclusion criteria regarding what outcomes, explanatory variables, and populations are under study. Then, a search over reference databases and retrieval systems (e.g., ProQuest and Google Scholar) is the standard practice to collect relevant primary studies (Reed and Baxter 2009). Such a search should broadly construct search keywords with various combinations because studies on the same topic may be conducted by researchers from different disciplines who subscribe to different vocabularies. Researchers also acknowledge the high heterogeneity in the vocabulary employed in sociology and recommend conducting literature search centering on methodological term—Roelfs et al. (2013) show that searching literature based on measures of concepts can locate a broader range of relevant publications than only using conceptual terms.

Meta-analysts also use auxiliary searching strategies to achieve more comprehensive coverage of literature, such as mining through the references of existing reviews (e.g., papers in Annual Review of Sociology) and consulting with field experts (McManus et al. 1998). Some meta-analysts also spend additional efforts on obtaining less accessible literature (e.g., unpublished manuscripts, working papers, conference proceedings, dissertations). Such efforts can be invaluable for sociological meta-analysis since it has been argued that sociological studies are prone to publication bias (Gerber and Malhotra 2008), which refers to a phenomenon that the results of published studies are more likely to be positive than those of unpublished ones (Chan and Altman 2005; Rosenthal 1979; Rothstein and Hopewell 2009). When the latter is not included in a meta-analysis, the combined result can be biased.

After obtaining relevant primary studies, meta-analysts extract a number of summary statistics and study-level characteristics from them. Typical summary statistics in sociological studies are regression coefficients and standard errors. Test statistics (e.g., t statistics, z scores) and test results (e.g., p-values, ***) are also useful alternatives when standard errors are not reported. Additionally, researchers also code study-level characteristics for potential heterogeneity analysis. When large variation across primary results exists, heterogeneity analysis is a useful tool to identify the reason for such variation. Important study features, such as methodological factors (e.g., study designs, measures, models), contextual features (e.g., study populations, geographical locations, time periods), and extrinsic features (e.g., publication quality) may introduce heterogeneity across studies, and heterogeneity analysis can quantify their influences on primary results and lead to novel findings that are beyond the scope of primary researchers (Lipsey 2009; Valentine 2009).

Two practical guidelines are helpful for collecting high-quality data for meta-analysis. First, when doing a literature search, meta-analysts need to keep detailed records of all the searching strategies and procedures, so that the searched results are replicable. Second, to ensure the reliability of the extracted results from primary studies, at least two coders need to independently extract summary statistics and study characteristics based on predefined, elaborate coding rules (Wilson 2009). A high inter-coder agreement is desirable. Note that meta-analysis is no panacea for correcting potential mistakes in primary studies. Primary results might be affected by measurement error and mistakes in analysis—such as the insufficient adjustment for the impact of complex survey sampling designs on the estimate of standard errors. Solving these problems require information from primary data, which are usually unavailable for meta-analysts. Nonetheless, the combined result in meta-analysis is more trustworthy when such problems potentially exist in one of the primary studies.

2.2. Obtaining an “average” effect from regression coefficients

In this review, we limit our discussion to synthesizing unstandardized regression coefficients because quantitative findings in sociology are typically regression coefficients. Ideally, when all primary studies are essentially the same with respect to the research methodology, such as metrics of predictors and outcomes, model choices, and control variables, meta-analysis is straightforward. In this case, as shown by Becker and Wu (2007), a meta-analysis based on summary statistics of primary studies would yield a result that closely resembles that obtained by jointly analyzing all primary data, and raw regression coefficients can be directly used as effect sizes in meta-analysis. However, sociological studies often have different methodological choices, and their impacts need to be carefully handled in the synthesis of regression coefficients. It is sometimes necessary to convert regression coefficients to a common metric or combine multiple results within a study. Researchers also assess or control for the impact of methodological heteorgeneity with post hoc heteorgeneity analysis, which refers to an analysis that examines whether and how different methodological choices impact the synthesized result. In below, we summarize four common types of methodological heterogeneity and potential solutions.

Differences in predictors.

Primary studies on the same topic may use different sets of predictors. We summarize three commonly-encountered cases here: (1) A predictor is measured in different units. For instance, incomes are measured in dollars or thousands of dollars. In this case, one can rescale regression coefficients accordingly. (2) A predictor is measured in different variable types (e.g., binary, ordinal, continuous). For instance, education attainment may be measured by the number of years of schooling or the multiple-category variable of education levels. In this case, researchers can convert regression coefficients to correlation coefficients, test statistics (Chinn 2000; McCartney and Rosenthal 2000), or semi-partial correlation coefficients (Aloe and Becker 2012). These converted metrics are scale-free and often referred to as the effect size parameter (Glass 1976; Konstantopoulos and Hedges 2004). We explain the conversion to correlation coefficients in Appendix 1.1 with an applied example. (3) A predictor is differently operationalized. For example, depression can be measured by different scales and regression results based on them may reflect the validity and reliability of each scale and therefore cannot be easily converted to be comparable. In such case, researchers can examine the impact of differently operationalized predictors on the synthesized result with a post hoc heterogeneity analysis so that the impact of different predictors can be clarified.

Differences in outcomes.

When the outcome variables of primary studies are in different metrics, the conversions we just discussed in “differences in predictors” can be applied. However, different metrics of outcomes can also lead to the use of different statistical models. For instance, a linear regression is likely to be used when the outcome is the number of years of schooling, whereas a logit model is typically used when educational attainment is measured by a categorical variable. Converting logistic coefficients to linear coefficients has been discussed (Borenstein et al. 2009). We explain them in Appendix 1.2 with an applied example but caution the use of such conversion because two different metrics of outcomes can potentially indicate two qualitatively different research questions. Synthesizing results from the same type of statistical models is always preferred.

Differences in controls.

Primary studies can include different control variables, producing partial regression coefficients conditional on different information. The impact of different control variables depends on the magnitude of multicollinearity, which, if substantial, can lead to significantly different results. This type of heterogeneity cannot be well addressed unless original data or more information from primary data (e.g., the variance-covariance matrix) are known. In practice, meta-analysts can code different sets of controls using a series of dummy variables and examine their impacts in a post hoc heterogeneity analysis (e.g., Doucouliagos and Paldam 2006; Shor et al. 2012, 2013; Stanley 2001).

Multiple results in one study.

A sociological publication often presents results from different models, such as a series of nested models, models based on subsamples, and so on. Deciding on which results to include in a meta-analysis can sometimes be challenging, and a review on the different approaches can be found in Cheung (2014). A general principle is that meta-analysis only needs results from the models that best serve the study goal and are most consistent across studies, and selecting one effect size per study is often followed in practice to ensure the independence across effect sizes (Lipsey and Wilson 2001). In more complicated situations, additional efforts in combining results within one study are necessary before a meta-analysis. One typical example is that a primary study may conduct the same analysis on different subsamples, such as males and females, and results of both sexes should be first aggregated within this study and then meta-analyzed with the other studies unless a gender-specific meta-analysis is needed. For another example, a primary study may use the same statistical model to separately analyze different waves of a panel dataset (collected from the same respondents repeatedly), and findings based upon all waves should be first combined within a study and then used in a meta-analysis. The methods to combine different sets of analysis within a primary study were discussed (Borenstein et al. 2009), and we summarize these methods in Appendix 2.

In summary, when preparing for a meta-analysis, various aspects of differences across primary studies need to be carefully considered. Conversions and post hoc heterogeneity analysis could help alleviate the impact of heterogeneity on the overall estimates of meta-analysis. However, large heterogeneity across primary studies can lead to concerns on the feasibility of a meta-analysis. When a significant level of heterogeneity in the methodologies of primary studies exists, the suitability of a meta-analysis needs to be justified both at the conceptual level (i.e., do different primary studies address the same research question?) and at the practical level (i.e., can a conversion/post hoc heterogeneity analysis lead to desirable results?). Substantive knowledge and established norms for a research topic are always central to inform the suitability of a meta-analysis. The ultimate solution to the methodological heterogeneity in primary studies is non-technical. Future development in consortiums and data harmonization can enable a community of researchers work together prospectively and provide an ultimate solution to the problem of heterogeneity.

In summary, when preparing for a meta-analysis, various aspects of differences across primary studies need to be carefully considered. Conversions and post hoc heterogeneity analysis could help alleviate the impact of heterogeneity in methodologies on the overall estimates of meta-analysis. However, the suitability of meta-analysis can be questionable if only a handful of very different studies are available. When primary studies have apparent heterogeneity in methodologies, the suitability of a meta-analysis needs to be justified both at the conceptual level (e.g., do different primary studies address the same research question?) and at the practical level (e.g., can a conversion/post hoc heterogeneity analysis lead to desirable results?). Substantive knowledge and established norms for a research topic are always central to inform the suitability of a meta-analysis. The ultimate solution to the methodological heterogeneity in primary studies is non-technical. Future development in consortiums and data harmonization can enable a community of researchers work together prospectively and provide an ultimate solution to the problem of heterogeneity.

3. The Core Methods of Meta-analysis

The fixed-effects and random-effects models are the two core methods that serve as the foundation of meta-analysis (Sutton and Higgins 2008). The two models utilize the same summary statistics from primary studies but have different statistical and practical considerations. The fixed-effects meta-analysis assumes that the same effect underlies all the observed effects in primary studies and that the observed differences are due to sampling errors. This assumption is usually referred to as the homogeneity assumption, which, when valid, means that the standard error of each study is merely determined by its sample size and that the pooled result of fixed-effects meta-analysis is generalizable to the population. Some researchers also insist on using fixed-effects model when the homogeneity assumption is violated (Hedges and Vevea 1998). In such case, the fixed-effects meta-analysis calculates a weighted average effect only among the observed studies.

The random-effects meta-analysis extends the fixed-effects meta-analysis by adding the consideration of between-study variation beyond the sampling error, which is the only concern of the fixed-effects model. The sampling error arises from sampling randomness, whereas between-study variation can be introduced by true heterogeneity, such as differences in populations, historical periods, study designs, analytical methods and so on. The random-effects meta-analysis can be viewed as a general model that includes the fixed-effects meta-analysis as a special case.

In practice, the choice between the fixed- and random-effects models depends on the perceived level of heterogeneity (Brockwell and Gordon 2001; Hedges and Vevea 1998). The random-effects model is more appropriate when a substantial amount of between-study variation exists. In practice, a forest plot that aligns 95% confidence intervals of all primary estimates may help determine the level of heterogeneity across studies visually (see the forest plots in Table 1 and Table 2). A forest plot presents quantitative effect sizes and their uncertainty collectively (Lewis and Clarke 2001). It is routinely employed in meta-analytical applications to provide a general impression on the directions and variation of effect sizes across primary studies. Alternatively, determining which model to use by performing a test on the level of heterogeneity (Q-based method) is shown to achieve a low coverage probability of the true effect when the between-study heterogeneity is substantial (Brockwell and Gordon 2001; Hedges and Vevea 1998), which is therefore not recommended. The rest of this section explains the formation of the fixed- and the random-effects models.

Table 1.

Summary statistics and the forest plot of studies on the peer effects on GPA

graphic file with name nihms-1751944-t0007.jpg

Note: All effect sizes are linear regression coefficients. In the forest plot, the whiskers represent the 95ε confidence intervals of the effect sizes. The dashed line is the average effect size in the fixed-effects model, and the sizes of cubes correspond to the weights of each study.

a.

All effect sizes are obtained from linear models that regress one’s own college GPA on peers’ GPA. Foster (2006) and Stinebrickner and Stinebrickner (2006) report results based on males and females separately, and the listed effect sizes are the combined results of subsamples.

Table 2.

Summary statistics and the forest plot of studies on perceived peer norms on adolescent sexual behaviors

graphic file with name nihms-1751944-t0008.jpg

Note: All effect sizes are log odds ratios. In the forest plot, the whiskers represent the 95% confidence intervals of the effect sizes. The dashed line is the average effect size in the random-effects model, and the sizes of cubes correspond to the weights of studies.

a.

Two types of measures of perceived peer norms are employed in primary studies. “Perceived approval” means the perceived peers’ approval of adolescent sexual activity, whereas “friend count” measures the self-reported number of sexually experienced friends.

3.1. The formulation of the fixed-effects model

Assume βj is the observed effect estimate in the jth study, μ is the unknown average effect size we wish to obtain using the fixed-effects model, and εj is the sampling error that follows a normal distribution with mean zero and a study-specific variance σj2. One can write the fixed-effects model based on k independent studies as:

βj=μ+εj,whereεjN(0,σj2),j=1k [Equation 3.1].

The fixed-effects model can be formulated as a weighted least squares (WLS) model with only an intercept, and one can transform it into a linear model with a standard normal error:

βjσj=μσj+εjσj,εjσjN(0,1),j=1k [Equation 3.2].

Because σj2′s are known from primary studies, this model can be estimated as a WLS model, and its solution takes a weighted-average form. Let wj denote 1/σj2. We can write this sum of squares as Q, the weighted sum of squares:

wj=1/σj2 [Equation 3.3],
Q=j=1k(εjσj)2=j=1kwj(μβj)2,j=1k [Equation 3.4].

By setting the first derivative of μ to zero or using the maximum likelihood estimator (MLE), one can obtain the estimate of μ that takes the weighted-average form:

μ^=j=1kwjβjj=1kwj [Equation 3.5].

Given that wj is the inverse variance of each observed effect size, the average effect size, μ^, in Equation 3.6, is simply a weighted average of effect sizes of primary studies, where a primary study with larger uncertainty (larger observed σj2) intuitively carries a smaller weight in the average effect size. The standard error of μ^ is positively related to the observed variance in each study. One can obtain the variance estimator, Var(μ^)=Var(j=1kwjβjj=1kwj)=1j=1kwj, and the standard error of μ^,

se(μ^)=Var(μ^)=(j=1kwj)12,j=1k. [Equation 3.6].

Let v denote (j=1kwj)1, the variance of μ^. The test statistic for H0:μ=0 vs. H1:μ0 is Z=μ^v which follows a standard normal distribution when μ = 0 is true. The two-sided hypothesis test rejects the null hypothesis when Z>z1α2 or Z<z1α2, where α is the test size. Therefore, the 100 (1 − α)% confidence interval (CI) is:

[μ^Z1α2ν,μ^+Z1α2ν] [Equation 3.7].

3.2. The formulation of the random-effects model

A random-effects model of meta-analysis with k independent studies can be expressed as:

βj=μ*+ηj+εj,whereηjN(0,τ2),εjN(0,σj2),j=1k [Equation 3.8].

In this model, βj is the observed effect size of the jth study and, μ* is the unknown average effect size that we want to estimate. Two variance components, ηj and εj, are respectively the between-study variation and the sampling error of each study. The former follows a normal distribution with mean zero and an unknown variance, τ2, and the latter follows a normal distribution with mean zero and a known variance σj2 (σj is often replaced with the observed standard error). The random-effects model generalizes the fixed-effects model in two ways. First, when between-study variation, ηj, degenerates at 0 (τ2=0), the random-effects model is reduced to the fixed-effects model. Second, when τ2 is known, the random-effects model can be estimated as the fixed-effects model (Equation 3.1) because the two variance components can be combined.

Note that an alternative formulation of the between-study variation is through posing a multiplicative parameter ϕ,

βj=μ*+ηj+εj=μ*+θi,whereθjN(0,ϕσj2) [Equation 3.9].

This formulation allows for more flexible account of the heterogeneity because ϕ can be either above or less than one. This formulation is less used in practice and a detailed discussion on its estimation and properties can be found in Stanley and Doucouliagos (2015).

In the traditional formulation, the between-study variation, τ2, can be estimated with various methods, such as the Cochran ANOVA estimator, Paul and Mandel estimator, maximum likelihood estimator, Q- profile method, and DerSimonian and Laird (DL) estimator (Veroniki et al. 2016). The most commonly-used method in practice is the DL estimator, which provides a closed-form, moment-based estimate, and asymptotically unbiased estimate of τ2. The performance of the DL estimator has been intensively studied in simulation studies, and it is shown to be potentially more (upward) biased than some other estimators—such as the Q-profile method, and Paul and Mandel estimator (the restricted maximum likelihood method)—when the number of studies is small, or the true heterogeneity is large (Veroniki et al. 2016). Nevertheless, simulation studies generally conclude the DL estimator performs adequately well and is advantaged in its easiness of implementation (Brockwell and Gordon 2001; Veroniki et al. 2016).

This subsection focuses on the DL estimator. It first requires fitting the fixed-effects model to obtain an average effect size, μ. Then, the total variation across studies is as Q=j=1kwj(μβj)2 (Equation 3.4), where wj=1/σj2. Under the homogeneity assumption, the total variation, Q, conforms to χ2(k1), and its expectation is k − 1 (detailed explanations in Appendix 5.1). The moment-based estimator of τ2 is the difference between the observed Q and expected Q (under the homogeneity assumption) divided by a constant a:

τ^2={Q(k1)awhenQ>k10whenQk1 [Equation 3.10].
wherea=(j=1k1σj2)2j=1k1σj4j=1k1σj2=j=1kwjj=1kwj2j=1kwj [Equation 3.11].

When the observed Q is larger than k − 1, Q − (k − 1) is intuitively the between-study variation that may not be fully accounted for by the sampling error. The detailed procedure of deriving the constant a for this estimator is shown in Appendix 5.2. By contrast, when Qk1, the random-effects model degenerates to the fixed-effects model. After obtaining τ^2, the random-effects model can be estimated as the fixed-effects model with a combined error term ϵj:

βj=μ*+ηj+εj=μ*+ϵj,whereϵjN(0,τ^2+σj2),j=1k [Equation 3.12].

The WLS approach in Section 3.1 can be used to estimate μ*, the average effect size in the random-effects model. By replacing σj2 with τ^2+σj2 in Equation 3.5, one can obtain μ*,

μ*^=j=1kβjσj2+τ2^j=1k1σj2+τ2^=j=1kβjwj*j=1kwj* [Equation 3.13],
wherewj*=1σj2+τ^2 [Equation 3.14].

Intuitively, the weight of each study is determined by the sum of sampling errors and the study-specific variation arising from study-level differences. The variance of μ*^ is Var(μ*^)=Var(j=1kwj*βjj=1kwj*)=1j=1kwj*, and the standard error is

se(μ*^)=Var(μ*^)=(j=1kwj*)12,j=1k [Equation 3.15].

The hypothesis test for H0:μ*=0 v.s. H1:μ*0 is as follows. Let v* denote (j=1kwj*)1, the inverse of the weight sum of all primary estimates. The test statistic for H0 is Z=μ^*v*, which follows a standard normal distribution when μ* = 0 is true. The two-sided test rejects the null hypothesis when |Z|>Z1α2, where α is the test size. The 100 (1 − α)% CI is:

[μ*^Z1α2v*,μ*^+Z1α2v*] [Equation 3.16].

Note that the traditional meta-analysis assumes that τ2 is a constant and does not study its variability (Veroniki et al. 2016). Some recent work (especially in the Bayesian framework) also treats τ2 as a random variable and accounts for its impact on the estimate of wj*. Methods for obtaining the variance estimate of τ2 can be found in Viechtbauer (2007). Moreover, the standard error estimator in Equation 3.15 is based on the assumptions that effect sizes of primary studies follow a normal distribution and are independent from each other. In more complicated cases where these assumptions are violated, robust standard errors can be estimated and interested readers are recommended to read Hedges, Tipton, and Johnson (2010), and Van Den Noortgate and Onghena (2005).

4. Two Illustrative Examples of Meta-analysis

To illustrate the use of the core methods of meta-analysis—the fixed- and random-effects models—we assembled data from two sets of primary studies. We use the fixed-effects model to obtain the average peer effect on academic performance, and the random-effects model to estimate the average effect of perceived peer norms on adolescent sexual behaviors. For each data example, detailed estimation processes are illustrated in hand calculation and Stata (Stata in Appendix 4).

4.1. Obtaining the average peer effect on academic performance with the fixed-effects model

In the 1960s, the landmark Coleman Report concludes that “a pupil’s achievement is strongly related to the educational backgrounds and aspirations of the other students in the school” (Coleman et al. 1966, p.22). Since then, the theory of “peer effect” has become increasingly popular, and studying peer effects on various academic and health outcomes has drawn intense interests among scholars. However, obtaining definitive evidence of peer effects is notoriously difficult because such effects are often confounded by shared environment and self-selection (Kandel 1978; Manski 1993; Moffitt et al. 2001). More recently, this difficulty in studying peer influence has been addressed by using randomized roommate assignment in college dorms (see Sacerdote 2014 for a review).

We obtain a total of five published studies that adopt this design to study peer influence on GPA (see Table 1). The detailed searching procedure is described in Appendix 3.1. Our meta-analysis uses the original regression coefficients without converting them into correlations given the generally identical study design and analysis (i.e., the same randomized assignment design, linear models, both independent and dependent variables using GPA as measures). We want to estimate the average peer effect on GPA with the fixed-effects model. The five estimates in Table 1 exhibit two important patterns. First, the regression coefficients vary between 0.004 and 0.241 and are relatively homogeneous. If the homogeneity assumption of the fixed-effects model holds, the variation arises from sampling randomness rather than study-specific contexts (e.g., different ways data are collected, dorm settings). Second, the widths of the 95ε CIs vary, suggesting a significant amount of sampling uncertainty across these studies.

Figure 1 illustrates the fixed-effects model of meta-analysis with the data example presented in Table 1. Each observed peer effect (βj) equals the true peer effect (μ) plus a study-specific error term (εj). Intuitively, the fixed-effects model assumes that all the five samples are drawn from the same population of college students, but each has a different sampling error. A meta-analysis that combines those five estimates can increase the effective sample size and lead to a more conclusive estimate. We now illustrate the estimating procedure of a fixed-effects model.

Figure 1. The illustration of the fix-effects model of meta-analysis.

Figure 1.

Note: The five normal curves are independent but not identical sampling distributions centered around the mean μ, which is estimated using the fixed-effects model and represented with the vertical line. The sampling variances are estimated in primary studies. Under each bell curve, a corresponding symbol represents the observed effect size in each primary study.

With hand calculation, one can obtain the average peer effect in five steps as shown in Figure 2. (1) Input the columns of effect sizes and standard errors in Table 1. (2) Calculate the weight of each sample according to Equation 3.3. (3) Use Equation 3.5 to obtain the average effect size, μ^ = 0.136. (4) Calculate the standard error, which, according to Equation 3.6, is 0.028. (5) Construct the 95% CI with Equation 3.7: 0.136 ± 1.96 * 0.028, which yields [0.081, 0.192]. Alternatively, one can obtain these results in Stata, and the details are explained in Appendix 4.1.

Figure 2. The five steps of implementing a fixed-effects model manually.

Figure 2.

Note: Step 1 inputs the data.

Step 2 calculates the weight of each study using Equation 3.4: wj=1σj2.

Step 3 calculates the average effect size using Equation 3.6: μ^=j=15wjβjj=15wj.

Step 4 calculates the standard error of the average effect with Equation 3.7: se(μ^)=(j=15wj)12.

Step 5 calculates the confidence interval with Equation 3.8: [μ^z0.975v,μ^+z0.975v], where v=(j=15wj)1.

This meta-analysis confirms a significant, positive peer effect on GPA despite the disagreement across individual studies. One point increase in roommates’ GPA contributes to 0.136 points increment in one’s own GPA. This example shows that fixed-effects meta-analysis can help generate conclusive results and contribute to the verification of the “peer effect” theory. However, our meta-analysis based on only five studies might still be unconvincing for some researchers. For example, it is possible that peer effects only exist in schools with more compulsory settings (e.g., military schools) where roommates are forced to interact more often (Poldin, Valeeva, and Yudkevich 2016). The school settings/culture might be the true explanatory factor for the observed peer effect. To verify this theory, researchers can build a consortium and reach out to different types of schools that assigned roommates randomly. Such a consortium can lead to meta-analysis with a larger sample size and promote the theoretical competition.

4.2. Obtaining the average effect of perceived peer norms on adolescent sexual behaviors with the “random-effects model”.

Our data example of the random-effects model examines the influence of perceived peer norms on adolescent sexual behaviors. Previous studies on this topic yield inconsistent results (see Van de Bongardt et al. (2015) for a review). Although this line of research does not employ the randomization design to obtain a causal effect, identifying the association between peer sociocultural contexts and sexual behaviors of adolescents informs the development of sociological theories, such as social learning theory and social norm theory (Cialdini and Trost 1998). We assembled seventeen studies that specifically focus on whether perceived peer norms predict adolescent sexual behaviors (see Table 2). All the assembled studies use logistic regression models, and we directly synthesize the regression coefficients (log odds) in this example. The detailed searching procedure is described in Appendix 3.2. Note that seventeen studies are considered adequate to achieve an enough coverage probability of the average estimate (close to 95 percent) based on the previous simulation results (Jackson, Bowden, and Baker 2010).

Two features of the seventeen logistic regression coefficients in Table 2 are worth noting. First, effect sizes in primary studies range from −0.05 to 2.51. A meta-analysis can provide a more conclusive result by increasing the sample size from an average of about 800 in the primary studies to more than 13,000 in the pooled result. Even in the presence of between-study heterogeneity, the combined estimate of a random-effects model can be suggestive to the direction of the overall effect. Second, the forest plot in Table 2 suggests ostensible variation in estimates across studies. The sampling error itself may not sufficiently explain this variation, which can also arise from study-specific characteristics, such as the average age of a sample and the different measures of perceived peer norms. A random-effects meta-analysis can therefore quantify the variation by estimating the between-study heterogeneity.

The idea of the random-effects model is illustrated with the data example in Figure 3. Each panel (except for the last panel) is a density plot of the study-specific population for each primary study. Assume that there is an overall population (the last panel) governing all the study-specific populations, and that the mean of the overall population (μ*) is indicated by the dashed line. The random-effects model assumes that the observed effects are drawn from the same overall population in two steps. First, the mean of each study-specific population (μj) is drawn from the overall population with variance as the between-study heterogeneity (τ2). Second, the observed effect of each study (βj) indicated by the solid dot is drawn from a study-specific population with a sampling error that is specific to each study (σj2). This second step is identical to the fixed-effects model. Note that in the random-effects model the true mean of each study-specific population is unknown, and the predicted mean by the best unbiased linear predictor is assumed to be the true mean in Figure 3 for the ease of illustration.

Figure 3. The illustration of the random-effects model of meta-analysis.

Figure 3.

Note: The dashed line represents the estimated average effect of the random-effects model. The density plot in the last panel is based on the estimates of between-study heterogeneity in the random-effects model. The means of the seventeen studies can be viewed as being generated from the overall population with only heterogeneity and study-specific sampling errors. The study-specific true mean is unknown and assumed to be the observed mean for the ease of illustration.

With the ten steps illustrated in Figure 4, one can implement the random-effects meta-analysis manually with the following ten steps. (1) Input the seventeen effect sizes and standard errors in Table 2. (2) Calculate the fixed-effects weight of each primary estimate using Equation 3.3. (3) Calculate the average effect size for the fixed-effects model, μ^, by using Equation 3.5, and we can obtain μ^ = 0.39. (4) Plug μ^ in Equation 3.5 to obtain Q, which yields Q = 41.50. Because Q − (k – 1) = 41.50 − (17 – 1) = 25.50 is larger than zero, we are able to estimate τ2 using Equation 3.10. Otherwise, the random-effects model degenerates to the fixed-effects model. (5) Use Equation 4.5 to obtain a, the constant of the moment-based estimator. We can get a = 432.32. (6) Calculate τ^2 using Equation 3.11, which yields 25.50432.32=0.06. Therefore, the between-study variation follows N (0, 0.06). (7) According to Equation 3.14, we calculate the new weight, wj* for each study. (8) With the new weight, we calculate μ*^, the average effect of the random-effects model according to Equation 3.13, which yields 0.48. (9) Calculate the standard error of μ^ with Equation 3.15, which yields 0.08. (10) The 95% CI is given by Equation 3.16, which is [0.32, 0.65]. One can also implement the random-effects model in Stata (details in Appendix 4.2).

Figure 4. The ten steps of implementing a random-effects model manually.

Figure 4.

Note: Step 1 inputs the data.

Step 2 calculates the fixed-effects weight using Equation 3.3: wj=1σj2.

Step 3 calculates the fixed-effects mean using Equation 3.5: μ^=j=117wjβjj=1kwj.

Step 4 calculates the weighted sum of variance by using Equation 3.4: Q=j=117wj(βjμ^)2,

Step 5 first calculates the square of the fixed-effects weights, wj2 in Column G, and then calculates the constant a in Column H with Equation 3.11: a=j=117wjj=117wj2j=117wj.

Step 6 calculates the between-study variance with Equation 3.10: τ^2=Q(171)a.

Step 7 calculates the updated weights for each primary study with Equation 3.14: wj*=1σj2+τ2^.

Step 8 calculates the average effect of the random-effects model with Equation 3.13: μ*^=j=117βjσj2+τ2^j=1171σj2+τ2^=Σj=117βjwj*j=117wj*.

Step 9 calculates the standard error with Equation 3.15: se(μ*^)=(j=117wj*)12.

Step 10 calculates the confidence interval with Equation 3.16: [μ*^z0.975v*,μ*^+z0.975v*], where v*=(j=117wj*)1.

This example shows that the random-effects meta-analysis can both provide a more conclusive result and quantify between-study heterogeneity. The combined result of this example shows a positive effect of perceived peer norms on adolescent sexual behaviors. The odds ratio e0.48=1.62 indicates that one level increase in the perceived peer norms (higher acceptance to sexual behaviors) increases the odds of sexual initiation by 62 percent. Although this numerical interpretation may be less intuitive given the perceived peer norms is measured differently across the primary studies, this result confirms a positive effect despite the disagreement in primary studies. Moreover, the estimated between-study heterogeneity (τ^2) is equal to 0.06, and the source of this heterogeneity might be explained by certain study-level features. We introduce the methods to study the source of between-study heterogeneity in the next Section.

5. Heterogeneity Analysis

After obtaining an average effect size, meta-analysts can employ heterogeneity analysis to understand how methodological factors (e.g., differences in study design, measurement, and data analysis) and study-level characteristics (e.g., study populations, geographical locations) may contribute to differences in effect sizes across primary studies. Heterogeneity analysis can be particularly useful for sociologists for two reasons. First, it can clarify the impact of different methodologies on the pooled result. Methodological differences across studies are common in sociology. Heterogeneity analysis codes them as study-level covariates and examines their impact. Second, heterogeneity analysis can identify study-level characteristics that are responsible for the variation in primary results. This aspect is particularly beneficial to sociological researchers because they can employ heterogeneity analysis to obtain novel findings that are beyond the scope of primary researchers. Note that the heterogeneity analysis in our discussion addresses between-study variation, which is different from the framework of “modeling uncertainty” developed by Young and colleagues (Young 2009; Young and Holsteen 2017) who consider within-study variation caused by modeling choices.

Some sociological researchers have already taken advantage of heterogeneity analysis to obtain important findings, such as the work by Shor et al. (2012) and Branigan et al. (2013) mentioned earlier in the introduction. Sometimes, even insignificant results in heterogeneity analysis can offer fresh insights. For example, Shor, Roelfs, and Yogev (2013) show that the effect of social support on mortality hazard is invariant across geographical and cultural contexts, indicating a ubiquitous mechanism of social integration; Amato and Gilbreth (1999) show that the beneficial effect of quality interactions with nonresident fathers on children’s well-being generally does not vary across age groups, race, and gender of children, confirming the universally important role of father in socialization. In all these studies, the study-level factors examined are either not a concern or impossible to be evaluated by primary researchers.

In the following, we describe the three steps of a heterogeneity analysis: (1) testing for between-study heterogeneity, (2) measuring heterogeneity, and (3) analyzing the source of heterogeneity. In the third step, we introduce two methods, the ANOVA-type approach and meta-regression. We use the two aforementioned data examples to illustrate how these steps are implemented and how this analysis can lead to substantive findings.

5.1. Testing for heterogeneity

To conduct a heterogeneity analysis, we first perform a test on whether effect sizes are homogeneous across studies. If the homogeneity hypothesis cannot be rejected, sampling errors are likely to explain the variation across studies, and a further assessment of heterogeneity may be unnecessary. Otherwise, rejecting the homogeneity hypothesis may indicate the existence of between-study heterogeneity, which is potentially caused by diverse study-level characteristics that need further examination.

Classic literature utilizes the Q statistic (Equation 3.4) to test for heterogeneity (Hedge and Vevea 1998). The Q statistic is calculated as the weighted sum of squared deviations across all effect sizes. Under the null hypothesis that all k effect sizes are equal (H0:β1=β2==βk) the Q statistic conforms to χ2(k1), and the alternative hypothesis is that at least one effect size is different from the others. The derivation for the distribution of this statistic is in Appendix 5.1, and a more detailed discussion on the properties of the Q statistic can be found in Hoaglin (2016).

To calculate the Q statistic, we plug the estimated average effect from the fixed-effects meta-analysis into Equation 3.4. We illustrate this test for heterogeneity with the two data examples. In the example of peer effects on GPA, we have already calculated μ^ = 0.136 in the fixed-effects model. With Equation 3.4, we obtain Q = 6.20. Using a Chi-square table, we can find that 6.20 is smaller than 9.49, the 95 percent critical value of χ2 (5 – 1). Then, the null hypothesis that all effect sizes are equal cannot be rejected at α = 0.05. This result indicates that the between-study heterogeneity is insubstantial, and a further assessment of its source may be unnecessary.

In the example of perceived peer norms on adolescent sexual behaviors, the Q statistic is already calculated (Step 4 in the manual way in Section 4.2, Q = 41.50). The Q statistic conforms to χ2 (17 – 1) under the null hypothesis that all effect sizes are equal. In a Chi-square table, the 95 percent critical value of χ2 (16) is 26.30, which is smaller than 41.50. The null hypothesis is rejected at α = 0.05, suggesting substantial between-study heterogeneity.

5.2. Measuring heterogeneity

After testing for heterogeneity, meta-analysts hope to quantify and interpret heterogeneity meaningfully. With the simple transformation shown in Equation 5.1, we obtain the heterogeneity index, I2, which measures the proportion of between-study variation out of the total variation when Q > (k − 1) (Higgins and Thompson 2002). This index is widely used to facilitate the interpretation of the between-study heterogeneity (Huedo-Medina et al. 2006).

I2=(Q(k1)Q)*100% [Equation 5.1]

Intuitively, I2 calculates the proportion of the excessive variation compared to the expected amount of variation under the homogeneous condition. Similar to the proportion of residuals in a linear model (1 − R2), I2 is the weighted proportion of the unexplained sum of squares out of the total sum of squares bounded between 0% to 100%. A larger I2 means larger heterogeneity. Generally speaking, an I2 less than 50% indicates small heterogeneity, between 50% and 75% indicates medium heterogeneity, and above 75% indicates high heterogeneity (Higgins et al. 2003).

Back to the example of peer effects on GPA, we can calculate I2=6.20(51)6.20*100%=35.5% It means that 35.5% of the total weighted sum squares cannot be explained by the sampling errors, which suggests small heterogeneity. Similarly, we can obtain I2 for the data example of perceived peer norms on adolescent sexual behaviors: I2=41.50(171)41.50*100%=61.4%. It suggests that 61.4% of the total weighted sum squares is not explained by the sampling errors, which indicates medium heterogeneity. These results are consistent with the results of testing for heterogeneity.

5.3. Analyzing the source of heterogeneity

Meta-analysts often explain heterogeneity with study-level factors when the level of heterogeneity is substantial. The study-level factors should be pre-registered before the heterogeneity analysis in an effort to avoid data dredging (Thompson and Higgins 2002). The general idea of this analysis is to use study-level characteristics as covariates to explain effect sizes. Note that when the number of studies is small (e.g., five studies), the heterogeneity analysis is not particularly meaningful.

Two types of methods have often been employed to analyze heterogeneity: the ANOVA-type approach and meta-regression. The ANOVA-type approach identifies factors that can partition a heterogeneous group of studies into (more) homogeneous subgroups of studies (Deeks, Altman, and Bradburn 2008). It applies when a study-level factor of interest is categorical (or when the linearity assumption between an ordinal covariate and an effect size does not hold). The ANOVA-type approach is a simple and effective way of analyzing heterogeneity, and we explain the details of the ANOVA-type approach with a worked example in Appendix 7.

However, when a meta-analyst encounters continuous study-level covariates or need to analyze multiple study-level factors concurrently, meta-regression is more useful as a general framework for assessing between-study heterogeneity. The idea of meta-regression is to incorporate the linear regression scheme into meta-analysis models, enabling the analysis of multiple study-level covariates concurrently (Thompson and Higgins 2002). It also provides a regression-type interpretation over the results. In this subsection, we explain the meta-regression model and illustrate the utility of the model with the example of perceived peer norms on adolescent sexual behaviors. With this example, we hope to clarify whether the two study features presented in Table 2 can explain the between-study variation and lead to additional insights for the research on this topic.

The meta-regression model can be expressed as an extended form of the random-effects model presented in Equation 4.1. Assume we have p study-level covariates, x1, x2xp, and their corresponding regression coefficients are γ1,γ2γp, (with the intercept γ0). The meta-regression model can be expressed as:

βj=γ0+γ1xj1+γ2xj2++γpxjp+ηj+εj,whereηjN(0,τ2),εjN(0,σj2),j=1k [Equation 6.2]

The inference of this model is similar to those in the fixed- and random-effects meta-analysis. Specific discussions on how to fit a meta-regression model with an unknown τ2 are lengthy (e.g., Berkey et al. 1995; Thompson and Higgins 2002; Van Houwelingen, Arends, and Stijnen 2002). In practice, a commonly-used solution is an extended form of DL estimator, which provides a closed-form solution to all parameters in two steps: (1) estimate τ2 with the moment-based method and (2) estimate γ0,γ1γp in a WLS model (Konstantopoulos and Hedges 2004). The details of how γ0,γ0γp are estimated in WLS are shown in Appendix 8 for interested readers. Sometimes, τ2 may be assumed to be zero when researchers believe the between-study variation is trivial after adjusting for covariates. This model then degenerates to a WLS model.

The rest of this subsection illustrates the utility of the meta-regression model with the example of perceived peer norms on adolescent sexual behaviors. Two covariates presented in Table 2 are analyzed: the perceived peer norms measured as the perceived number of sexually experienced peers or the perceived peers’ approval of adolescent sex; the average age of each sample. For the latter, the strength of peer effects may vary across different stages of adolescence. This example examines whether these two covariates can explain the between-study variation.

The implementation of meta-regression with Stata is explained in Appendix 8. We present the results of meta-regression in Table 3, in which model 1 and model 2 include one of the two study-level covariates and model 3 includes both covariates. Consistent with the results obtained using the ANOVA-type approach, the average age of sample does not significantly affect the effect sizes (model 1), whereas using the measure of the perceived number of sexually experienced friends yields a significantly larger effect than the perceived approval measure (model 2). The difference is 0.41. When both covariates are included (model 3), the average age of the sample is positively related to the effect size, indicating a significant age effect (14 percent per year) conditional on the different measures of perceived norms.

Table 3.

Results of the heterogeneity analysis using meta-regression

Model 1 Model 2 Model 3
Average Age 0.11 0.14*
(0.07) (0.06)
Peer Norm Measure 0.41** 0.51**
(ref=Approval) (0.15) (0.17)
Intercept −1.07 0.74*** −1.12
(0.99) (0.13) (0.88)
τ 2 0.07*** 0.04* 0.04*

Note:

*

p value <0.05,

**

p value <0.01,

***

p value <0.001.

This meta-regression analysis provides two additional insights. First, it clarifies the impact of different measures. The manner in which the perceived peer norms are measured can impact the magnitude of peer effects. Second, the meta-regression uncovers a novel age trend in peer effects conditional on the different measures; the perceived peer norms have a stronger impact on mid-to-late adolescence. Both findings are beyond the scope of primary researchers and become evident through meta-regression. Nevertheless, we highlight two cautionary notes regarding its implementation and the interpretation of the results of meta-regression. First, what covariates a heterogeneity analysis needs to examine should be pre-registered before its implementation. Conducting heterogeneity analysis in a “data dredging” manner is plagued by the multiple testing problem (Thompson and Higgins 2002). It is recommended that researchers consider a limited number of study-level factors in a heterogeneity analysis and pre-justify the necessity of examining them (Deeks et al. 2008) Second, the evidence drawn from the heterogeneity analysis should only be interpreted at the study level. Such consideration stems from the concern of “ecological fallacy”, which means that results of a group-level analysis cannot be interpreted reductively as an individual-level effect (Firebaugh 2008). The results of heterogeneity analysis that utilizes study-level characteristics can be weaker or even disappear at the individual-level when confounders exist (Thompson and Higgins 2002). Nonetheless, heterogeneity analysis can at least inform provide researchers with potential mechanisms of interest and contribute to the improved design of primary studies in the future.

6. Sensitivity Analysis: Evaluating Publication Bias

Evaluating publication bias as a sensitivity analysis has become a routine in meta-analysis to improve the validity of results. The combined results in meta-analysis need to be interpreted with caution because the collected primary results could be affected by publication bias, namely, statistically insignificant results are less likely to be published. An ideal data collection for meta-analysis should locate all the published and unpublished studies. However, even the best data collection practice can only locate part of the unpublished studies, and results that are insignificant or contradictory to hypotheses are very often not documented and unavailable for meta-analysts. Two classes of methods have been developed in the literature of publication bias, the selection models and the funnel plot-based approaches (Lin and Chu 2018). Selection models postulate certain selection mechanisms that lead to the missing observations (in unpublished studies). This method is often complicated and has limited applicability in practice (see Sutton et al. (2000) for a review), and the focus of this section is the funnel plot-based approaches.

The funnel plot-based approaches assume that, with no publication bias, primary results should appear symmetry in a funnel plot, which is usually a scatterplot of standard errors against effect sizes in all primary studies (Copas and Shi 2000; Sterne and Egger 2005). An asymmetric funnel plot can indicate the potential existence of publication bias, and different procedures have been developed to test and correct for the asymmetry of a funnel plot, such as, the Egger’s regression test (Egger et al. 1997) and the trim-and-fill method (Duval and Tweedie 2000). In meta-analytical applications, it has become a routine practice to perform a visual check on publication bias by using the funnel plot, the funnel plot assumes that the results of studies with lower precision (i.e., studies with larger variance and smaller sample sizes) are likely to be more varied than those with higher precision (i.e., studies with smaller variance and larger sample sizes). An asymmetric funnel plot often misses results with large variance and small effect sizes—they are likely to be the unobserved, insignificant results that are more difficult to be published.

Figure 5 presents the funnel plot of the seventeen primary estimates of perceived peer norms on adolescent sexual behaviors. Results with small standard errors are more centered around the average effect indicated by the vertical dashed line (estimated with the fixed-effects model), and results with large standard errors deviate more from the average effect. The empty space on the left corner suggests that some studies with large standard errors and small effect sizes are missing, indicating potential publication bias. The outliers on the right side indicate the potential existence of heterogeneity. Note that the asymmetry of a funnel plot can be caused by reasons other than publication bias, such as data irregularities, true heterogeneity in primary samples, language bias (selective inclusion in English), and familiarity bias (i.e., selective inclusion based on an meta-analyst’s field) (Rothstein, Sutton, and Borenstein 2006). It can also reflect whether a true theory competition exists regarding a research topic, and insignificant results are more likely to be published when some researchers believe the true effect is null (Doucouliagos and Stanley 2013). As such, the funnel plot always needs to be interpreted with cautions and only serves as a sensitivity analysis in meta-analytical applications.

Figure 5. Funnel plot of the data example of perceived peer norms on adolescent sexual behaviors.

Figure 5.

Note: Each dot represents a study. The vertical dashed line indicates the average effect size estimated by a fixed-effects model, which approximates the true average effect size. The dashed lines that mark the two sides of the isosceles triangle show the 95% confidence region around the fixed-effect estimate.

For interested readers, we illustrate two statistical methods—the Egger’s regression test (Egger et al. 1997) and the trim-and-fill method (Duval and Tweedie 2000)—that test for the symmetry of a funnel plot and potentially correct for publication bias in the synthesized result in Appendix 6. Note that the validity of these methods is based on the assumption that the asymmetry of a funnel plot is caused by the publication bias. The Egger’s regression test is shown to more frequently detect publication bias in various settings and outperform the other competitors (Lin et al. 2018), while the “trim and fill” method can provide an estimate that potentially reduce the publication bias and is recommended as a sensitivity analysis (Peters et al. 2007). In practice, combining different statistical tests and procedures in the assessment of publication bias is recommended (Lin et al. 2018). It is also important to know that statistical testing and correction methods often only address the publication bias in an ad hoc manner rather than providing a fundamental solution. Non-statistical approaches should still be the primary focus of addressing publication bias. In sociological research practice, a more fundamental solution for publication bias is to encourage publishing replication studies regardless of sample sizes and significance. With specified authorship rules, interested researchers on a same topic and build consortiums and collaboration networks that attract and provide guidelines for researchers interested in conducting small or replication studies. Publication bias can be better addressed in such way.

7. Conclusion: A Proposal for Developing Sociological Meta-analysis

Over the past three decades, meta-analysis has played an increasingly important role in many scientific fields. The expanding use of meta-analysis fully reflects the ultimate pursuit of scientific research, which continuously attains cumulative, synthetic, and comprehensive knowledge that informs empirical consensus and theoretical development. Meta-analysis is how such knowledge is produced. Although meta-analysis has been attempted in sociological studies (Branigan et al. 2013; Hedges et al. 1994; Shor et al. 2012), the benefit of meta-analysis is yet fully realized. We propose that an increased use of meta-analysis can benefit the entire discipline of sociology by improving the robustness of the knowledge-cumulating process (e.g., the sense that pooled results are more trusted) and the rigidness of sociological research practice (e.g., reducing the tendency of selective review of existing literature, critical evaluation of publication bias). However, the broad use of meta-analysis depends on the accumulation of an adequate number of primary studies with comparable study designs, measurements, models, and so on. In this section, we conclude our article with a proposal for developing sociological meta-analysis with four stages.

Stage 1. Increased awareness of the importance of meta-analysis: researchers should be made aware of the benefit of meta-analysis, which increases statistical power and provides a systematic way to examine heterogeneity. Researchers should understand that differences in methodological choices across primary studies can lead to difficulties in the future meta-analysis.

Stage 2. Improved reporting standards in primary studies: Reporting abundant results in primary studies is the key to the broad application of meta-analysis. In “The Handbook of Research Synthesis”, Cooper and Hedge (2009, p564–565) suggest the importance of establishing reporting standards in scientific journals for primary studies and give examples of how such reporting standards in educational and psychological research have benefited the broad use of meta-analysis. Establishing reporting standards can also benefit the development of sociological meta-analysis in the future. For example, primary results are encouraged to provide results of standardized analysis so that the impact of measurement in different scales is reduced. They can provide results of a series of models from basic to more sophisticated to ensure the inclusion of a compatible model that might be included in a future meta-analysis. As Young and Holsteen (2017) suggest, primary studies can quantify and report the within-study variation of results across models when possible. Moreover, additional information of primary analysis can further promote the use of advanced meta-analytical techniques that improve the efficiency of estimates and produce results less affected by measurement error or mistakes in primary studies. For example, the variance-covariance matrix or the full correlation matrix is crucial for the use of more advanced methods, such as the multivariate meta-analysis that jointly synthesizes two or more correlated effect sizes in each primary study (Jackson, Riley, and White 2011), and the meta-analytical structural equation modeling that synthesizes primary studies in a more integrative way (Cheung 2015). Reporting the covariance information of multivariate analysis could greatly facilitate the use of those more advanced meta-analytical techniques.

Stage 3. Conducting meta-analysis on existing studies: Researchers assemble summary statistics from available primary studies to conduct a meta-analysis in an effort to increase the statistical power and clarify and explain heterogeneity across results. The procedures of conducting meta-analysis at this stage is the focus of this paper, which includes collecting primary results, using statistical models to combine primary results (i.e., fixed- and random-effects models), evaluating publication bias, and performing heterogeneity analysis.

Stage 4. Developing collaboration networks and consortiums for meta-analysis: A formal collaborative network among sociological researchers may be developed when enough resources can be mobilized to address a common research question. Such a formal collaboration network can promote data harmonization and integration across studies. The between-study heterogeneity arising from differences in methodologies can be reduced to the minimum by performing a coordinated meta-analysis in two steps (Debray et al. 2015). First, researchers develop a research protocol to coordinate analytical methods with standardized concepts, measures, and analytical methods, and perform analysis separately to obtain study-specific estimates. Second, they combine study-specific estimates and conduct heterogeneity analysis. This two-stage approach is especially valuable for situations when only some primary data are available but the others are not. Moreover, when all the primary data can be jointly analyzed in a research synthesis project, it can lead to the most accurate synthesized results and heterogeneity analysis (e.g., not prone to ecological fallacy). Hierarchical or multilevel models can be directly applied in this case (Debray et al. 2015). In long-term, with developed authorship rules, consortiums can also attract interested researchers to share insignificant or unpublished results, alleviating publication bias and promoting theoretical competition. Existing examples of consortiums and collaborative networks include Consortium on Health and Aging: Network of Cohorts in Europe and the United States (CHANCES) (Boffetta et al. 2014) and non-profit organizations like Campbell Collaboration and Cochrane Collaboration that facilitate the evaluation over social policies and interventions across the world (Borenstein 2009; Boruch 2005).

These four stages of developing sociological meta-analysis signify significant changes in the research practice of sociology. When carefully and routinely applied, meta-analysis will provide a systematic way of summarizing available evidence and notably increase the credibility of sociological inquiry.

Appendix 1. Converting among linear regression coefficients, logistic regression coefficients, and correlation coefficients.

In this appendix, we introduce the conversions among the linear regression coefficient, logistic regression coefficient, and correlation coefficient. The aim is to help combine different types of results in a meta-analysis. In this set of conversions, the correlation coefficient is often used as a common metric when different types of summary statistics need to be synthesized (McCartney and Rosenthal 2000). One can also refer to Borenstein et al. (2009), Chinn (2000), and McCartney and Rosenthal (2000) for more discussions on these conversions.

Appendix 1.1. Converting linear regression coefficients to partial correlation coefficients.

To convert a linear regression coefficient to a partial correlation coefficient, the first step is to calculate the t statistic of a regression coefficient. Assume β is a linear regression coefficient, and σ is its standard error, then

t=βσ [Equation A1.1]

, and the partial correlation coefficient is:

r=tt2+np1 [Equation A1.2]

, where n is the sample size of a study, p is the number of predictors in a linear regression model other than the intercept, and therefore np − 1 is the degree of freedom of the t statistic. Conversely, the equation that converts a partial correlation coefficient to a t statistic with the degree of freedom of np − 1 is:

t=rnp11r2 [Equation A1.3].

In practice, because a correlation coefficient is strictly bounded between −1 and 1, meta-analysts often transform it to a Fished-Z transformed coefficient that ranges across the entire real line. Researcher often use Fisher-Z transformed correlation coefficients (z) in meta-analysis. To transform a correlation coefficient into a Fisher-Z transformed correlation coefficient, one can use,

Z=12log(1+r1r) [Equation A1.4].

The variance of a Fisher-Z transformed correlation coefficient can be approximated by,

Var(z)=1n3 [Equation A1.5]

, where n is the total number of observations of a primary study.

After obtaining a synthesized z-transformed correlation z* and its standard error σZ*, one can convert it back to a correlation coefficient to obtain a more sensible interpretation. The equations for this transformation are,

r*=e2z*1e2z*+1 [Equation A1.6],
σr*=2(e2(z*+σz*)e2z*)(e2(z*+σz*+1))*(e2z*+1) [Equation A1.7].

Now, suppose we want to combine two primary studies that studied the impact of education on income. Both studies used the same measure of income (in 10k, annually), the same set of six control variables (p=6), and linear regression models. The only difference is that the first study measured education as years of education and obtained the result as β1=0.300 (σ1=0.200), whereas the second study measured education as whether having a college degree and obtained the result as β2=2.100 (σ2=0.700). The first study had a sample size of 1,000, and the second study had a sample size of 3,000. To combine these two results, one may first want to convert them to partial correlation coefficients in the following steps. (1) With Equation A 1.1, one can obtain t1=0.3000.200=1.500, and t2=2.1000.700=3.000 (2) With Equation A 1.2, one can get r1=1.5001.5002+100061=0.048, and r2=3.0003.0002+300061=0.055. (3) To transform them to Fisher-Z transformed correlation coefficients, one can use Equation A 1.4 and get z1=12log(1+0.04810.048)=0.048, and z2=12log(1+0.05510.055)=0.055. (3) With Equation 1.5, one can obtain the approximate standard errors of z1 and z2, which are se(z1) = 0.032 and se(z2) = 0.018.

Now, suppose that the fixed-effects model has been employed to combine these two effect sizes and obtained the synthesized z-transformed correlation coefficient as z* = 0.053 and its standard error as σz* = 0.016. We can then transform it back to a correlation coefficient. With Equation A 1.6, one has r*=e2*0.0531e2*0.053+1=0.053, and with Equation 1.7, one has σr*=2(e2(0.053+0.016)e2*0.053)(e2(0.053+0.016)+1)*(e2*0.053+1)=0.016. This result suggests that the strength of correlation is 0.053 (se=0.016) between education and income given the information of those two primary studies, though they have different metrics of education.

Appendix 1.2. Converting logistic regression coefficients to linear regression coefficients.

To convert a logistic regression coefficient (log odds ratio) to a standardized linear regression coefficient (or a standardized mean difference), one can use the below formulas proposed by Hasselblad and Hedges (1995),

d=Log(OR)*3π [Equation A1.8],
σd=σLog(OR)*3π2 [Equation A1.9].

Now, suppose two studies on the racial achievement gap (whites v.s. blacks) measure the outcome of education as the years of education and whether having a college degree. These two different metrics lead to the use of linear and logistic models respectively. Assume that the logistic model yielded the result as odds ratio = 1.570, and standard error σLog(OR)=0.090, and one wants to convert this result to a standardized linear regression coefficient. With Equation A1.8, one can obtain d=log(1.570)*3π=0.249, and with Equation A1.9, one can get, σd=0.090*3π2=0.027. If the linear model in the other study is standardized, one can combine its result with d directly. However, if it is not standardized, one needs to convert both of them to correlation coefficients by following the steps in Appendix 1.1.

Appendix 2. Combining multiple results within a study.

Appendix 2.1. Combining results from independent subsamples

Assume β1 and β2 are regression coefficients of two independent subsamples, σ12 and σ22 are their variances, and n1 and n2 are the corresponding sample sizes. The combined effect size β¯ is in the weighted average form as follows,

β¯=n1*β1+n2*β2n1+n2 [Equation A2.1].

The variance of β¯ can be obtained with the below equation,

σ¯2=(n11)σ12+(n21)σ22+n1n2n1+n2(β1β2)2n1+n21 [Equation A2.2].

In this case, the pooled average effect size, β¯, has an effective sample size of n1 + n2. In a more general case, the formulas to combine 1 independent results β1βn within a study, each with a sample size of n1 and variance of σi2, are as follows,

β¯=i=1pni*βii=1pni [Equation A2.3],
σ¯2=i=1pσi2(ni1)+i=1pni(βiβ¯)2i=1pni1 [Equation A2.4].

This synthesis increases the effective sample size to i=1pni.

Researchers also point out that the above method may be susceptible to the “Simpson Paradox” problem (Borenstein et al. 2009), which in this context means that the trends observed in subgroups might disappear if the primary data of them were jointly analyzed. Therefore, the use of this combination method needs to be treated with caution.

Appendix 2.2. Combining correlated results

We start with a simple case that combines two regression coefficients. Assume β1 and β2 are the two regression coefficients obtained with the same statistical analysis on two different waves of a panel study, and β¯ is the average effect size within a study. Also, assume σ12 and σ22 are variance of β1 and β2, and σ¯2 is the variance of β¯, then,

β¯=12(β1+β2) [Equation A2.5]
σ¯2=14(σ12+σ22+2rσ1σ2) [Equation A2.6]

The quantity, r, in the Equation A2.6 is the Pearson correlation between the two effect sizes, a quantity that is unlikely reported in a primary study. In practice, researchers can give a convenient estimate of r based on previous literature or their field-specific expertise. Conducting a sensitive analysis by choosing different r’s within a range of values is sometimes employed (Borenstein et al. 2009).

Similarly, the formulas to combine 1 estimates from a primary study are as follows,

β¯=1pi=1pβi [Equation A2.7]
σ¯2=1p2(i=1pσi2+2i<jrijσiσj) [Equation A2.8]

, where rij is the Pearson correlation between the ith and jth effect sizes. This combination does not increase the statistical power, and the effective sample size remains the study sample size.

Appendix 3. The data collection process of the two data examples.

Appendix 3.1. Peer influence on academic performance

In this data example, we collect studies that adopted the random assignment design to study whether peers influence academic performance among college students. We only include studies that measure the academic performance (both egos and peers) as the grade point average (GPA), and exclude studies that measure peer academic performance alternatively as roommates’ college admission exam scores (i.e., SATs, ACTs). Studies with the latter measure usually explains the peer influence poorly and is not the focus of this data example.

We set the time range of our literature retrieval from Jan. 2000 to Feb. 2017, given the first published study with randomized roommate assignment was at Dartmouth College by Sacerdote (2001). The primary method of retrieval was by searching through a number of major electronic databases including Google Scholar, ProQuest, Web of Science, and EBSCO. Various combination of the following key terms are used: peer influence, peer effect*, random assign*, roommate stud*, randomized design, randomized roommate stud*, GPA, academic performance, achievement, and academic outcome*. We also mine through the citations of a critical review paper by Sacerdote (2014). This literature search generates 740 unique results, and we ultimately locate fourteen qualified studies after examining the abstracts of these papers.

In the next step, by intensively reading through the contents of studies, we exclude five studies that have irrelevant contexts of peer influence. We also find an unpublished dissertation by uses the same dataset as in a later published paper by Levy (2000), Kremer and Levy (2008), and another earlier, unpublished version of Sacedote (2001). We thus only include the published ones. We finally exclude another two publications (Carrell, Fullerton, and West 2009; Zimmerman 2003) for using SAT scores as predictors. The number of qualified studies ends up being five (i.e., Foster 2006; Kremer and Levy 2008; Lyle 2007; Sacerdote 2001; Stinebrickner and Stinebrickner 2006).

Appendix 3.2. Perceived peer social norms on adolescent sexual behaviors

We first restrict the scope of this meta-analysis to studies on how perceived peer norms impact sexual behaviors among adolescents (aged 10–19). The outcome is measured as sexual initiation. We then retrieve studies from four electronic databases and searching engine (i.e., GoogleScholar, PubMed, Web of Science and ProQuest) with various combinations of the keywords, including peer influence, peer effect, peer norm, social norm, social integration, peer conform*, social pressure, peer value, culture, teenager, adolescent sexuality, sexual health, adolescen* sexual behavior*, sexual activit*, adolescen* health, youth health, first sex, sexual intercourse, hav* sex, sex debut, timing, and risky behavior*. The time frame of our search is until May 2017. We also mine through the citations of a previously published meta-analysis on similar topics (Van de Bongardt et al. 2015). The first literature search yields a total of about 11000 results, and among which 66 studies potentially qualify and are included for further examination.

After carefully reading through the content of these studies, we first exclude 43 studies that are not in good alignment with the scope of our meta-analysis. Another 6 closely-related studies are also excluded for using different study designs that are inconsistent with the other studies. These studies include Carvajal et al. (1999), Gillmore et al. (2002), Doornwaard et al. (2015), Reitz et al. (2015), Van De Bongardt et al. (2014) and Pai and Lee (2012).

We finally locate 17 studies on the association between perceived peer norms and adolescent sexual behaviors. Listed in Table 2, these studies are (Akers 2011; Bersamin et al. 2005; Cabral et al. 2017; DiIorio et al. 2001, 2004; Kabiru et al. 2010; Kawai et al. 2008; Kinsman et al. 1998; Laflin, Wang, and Barry 2008; L’engle and Jackson 2008; Little and Rankin 2001; Maguen and Armistead 2006; Potard, Courtois, and Rusch 2008; Rosenthal, Smith, and De Visser 1999; Sieving et al. 2006; Van De Bongardt et al. 2014; Villarruel et al. 2004). Note that Dilorio et al. (2004) only indicate the effect size is insignificant but do not report the standard error, we assume the effect size is one standard deviation from zero for convenience. Sensitivity analysis shows that different values of standard errors (0.27–0.54) does not affect the synthesized result. We are aware of the existence of more advanced methods to impute missing variance (e.g., Chowdhry, Dworkin, and McDermott 2016).

Appendix 4. STATA code for the data examples of the fixed-effects and random-effects meta-analysis.

Appendix 4.1. Implementing the fixed-effects model in Stata

This appendix provides the Stata code for implementing the fixed-effects meta-analysis. The data example in Section 4.1 is used here. With the following five steps, one can obtain the fixed-effects estimates manually,

*Step (1)

input beta se

0.141 0.200

0.115 0.097

end

*Step (2)

gen w = 1/se^2

*Step (3)

egen w_sum = sum(beta*w)

egen total_weight = sum(w)

gen miu=w_sum/total_weight

*Step (4)

gen se_miu= total_weight ^−1/2

*Step (5)

gen conf_left=miu − 1.96*se_miu

gen conf_right=miu+1.96*se_miu

Alternatively, one can use the built-in function, vwls, designed for fitting WLS models. The Stata code for the implementation of the fixed-effects model is as follows. One can first generate a non-zero constant vector indicating the intercept and then use vwls to obtain the average effect size, standard error and CI in one step.

*The alternative approach using vwls function

gen x = 1

vwls beta x, sd(se)

Moreover, the build-in package “metan” can also provide a one-step solution.

*The alternative approach using metan function

metan beta se, fixed

Appendix 4.2. Implementing the random-effects model in Stata

This appendix provides the Stata code for implementing the random-effects meta-analysis. The data example in Section 4.2 is used here. With the following six steps, one can obtain the random-effects estimates manually.

*Step (1)

input beta se

0.67 0.34

0.08 0.40

end

*Step (2)

gen x = 1

vwls beta x, sd(se)

*Step (3)

gen Q = 1/se^2

gen dev=(beta−0.39)^2

egen q=sum(dev*w)

*Step (4)

egen sum_weight= sum(w)

egen sum_weight_sq=sum(w^2)

gen a=sum_weight-sum_weight_sq/sum_weight

gen tao_sq = (q-(17-1))/a

*Step (5)

gen se_new = (tao_sq+se^2)^0.5

*Step (6)

vwls beta x sd(se_new)

Alternatively, the random-effects model can be easily implemented with the package “metan”,

*The alternative approach using metan function

metan beta se, random

Appendix 5. The moment-based estimator of the random-effects model.

Appendix 5.1. The property of the Q statistic

This appendix shows that the Q statistic under the homogeneity assumption follows a χ2 distribution. The original discussion on this property can be found in Cochran (1954). Define β=(β1βk), ω=(w1wk), W=diag(ω), 1=(1)k×1, I=diag(1), b=j=1kwi, and μ¯=1ωβω1=b11ωβ. The Q statistic can be expressed in a quadratic form as,

Q=β(Ib1ω1)W(Ib11ω)β [Equation A5.1].

For convenience, if we denote A=(Ib1ω1)W(Ib11ω), we can have Q in the following quadratic form: Q=βAβ. Under the homogeneity assumption of βN(0,W1), we can easily show that AW1 is idempotent. Therefore, we have by the classical result of quadratic form (Schott 2016),

Qχ2(rank(AW1))=χ2(trace(AW1))=χ2(k1) [Equation A5.2].

Although a Chi-square test for heterogeneity with the Q statistic is widely used in practice, we caution that the above deviation is only asymptotically valid and relies on other assumptions on the primary studies. For instance, when ni, the sample size of each primary study is small or the variance estimates in primary studies do not approximate the true variance of studies, the exact distribution of Q is complicated, and the direct use of Q as a test statistic is problematic. Moreover, if the normality assumption on the primary effect sizes is violated, the χ2(k1) is no longer a valid large-sample approximation. An exact but more complicated first- and second-moment of Q was derived by (Kulinskaya, Dollinger, and Bjørkestøl 2011). Furthermore, when the number of primary studies (k) is small, the testing for heterogeneity with Q could be replaced with different approximate distributions of Q, such as Gamma approximation, saddle-point approximation, and Pearson approximation, to obtain more conservative testing results (Biggerstaff and Jackson 2008). Studying heterogeneity has always been a major challenge in the field of meta-analysis, and as noted by Hoaglin (2016) in a recent review on this topic, finding more valid testing procedures for heterogeneity is still an ongoing research subject.

Appendix 5.2. The calculation of τ2

This appendix illustrates how the moment-based estimator for τ2 is obtained. This moment-based estimator was originally provided by DerSimonian and Laird (1986). We know from Appendix 5.1 that Qχ2(k1). Define μ=(1)k×1, under the null hypothesis (τ2=0),

E(Qτ2=0)=(k1)+μAμ

Under the alternative hypothesis (τ2>0):

E(Qτ2>0)=βAβ=μAμ+trace(AΣ),whereΣ=diag(τ2+w11,,τ2+wk1).

We can calculate the differences between the expectation of Q under the null and alternative hypotheses:

E(Qτ2>0)E(Qτ2=0)=trace(AΣ)(k1)
=trace(diag(1+τ2wi))b1wwdiag(τ2+wi1)(k1)
=k+τ2i=1kwib1i=1kwi2(τ2+wi1)(k1)=
=k1+τ2(i=1kwii=1kwi2i=1kwi)(k1)=τ2(i=1kwii=1kwi2i=1kwi)

Hence, E(Qτ2>0)E(Qτ2=0) can be expressed as τ2 multiply by a constant a:

a=E(Qτ2>0)E(Qτ2=0)τ2=i=1kwii=1kwi2i=1kwi

Because τ2>0 is needed to ensure E(Qτ2>0)E(Qτ2=0)>0, only when Q>k1 can we calculate τ^2 as Q(k1)a. Otherwise, the random-effects model cannot be fitted.

Appendix 6. The use of Egger’s regression test and the “trim and fill” method.

The symmetry of a funnel plot can be tested using Egger’s regression test (Egger et al. 1997). The procedure of this test is to first regress the Z-scores (i.e., effect sizes divided by standard errors) on precision (i.e., the inverse of standard errors) and then test whether the intercept of this regression model is significantly above zero. A significant non-zero intercept usually indicates publication bias. Improved versions of the Egger’s regression test, including Begg-Mazumdar’s test, Harbord-Egger’s test, and modified Macaskill’s test are designed to achieve smaller type-I errors, especially for odds ratios (Peters et al. 2006; Sutton and Higgins 2008).

We illustrate the Egger’s regression test with the data example of perceived peer norms on adolescent sexual behaviors. This test can be easily implemented manually. We first calculate the Z scores of effect sizes in Table 2 and then regress them against the inverse of standard errors. The t statistic of the intercept equals to 3.01 (β0 = 1.89, se = 0.62, df = 16), which is significantly above zero (p <0.01). The test result suggests that the funnel plot is asymmetric. This result is consistent with the visual judgment using the funnel plot.

To correct for the publication bias in the synthesized result, the non-parametric “trim and fill” method (Duval and Tweedie 2000) is commonly used in practice. The idea of the “trim and fill” method is to impute a few hypothetical studies through an iterative procedure so that the symmetry of the funnel plot can be improved. Details of the algorithm are explained by Duval and Tweedie (2000), and readers can implement it in STATA with the metatrim function (Steichen 2010). Because of its assumption-free nature and good performance in practice, the “trim and fill” method is much more often used than its parametric alternatives, such as a priori weigh functions (Vevea and Woods 2005) and pseudo-data simulation (Bowden, Thompson, and Burton 2006). Note that the non-parametric nature of the “trim and fill” method limits its utility in developing hypothesis-testing procedures for the symmetry of funnel plots (Duval and Tweedie 2000).

Figure 6A. The funnel plot with four imputed studies.

Figure 6A.

Note: The 17 solid dots represent the original studies, and the 4 empty dots are the imputed studies. The dashed lines that mark the two sides of the isosceles triangle show the 95% confidence intervals around the fixed-effect estimate on the original 17 studies. The vertical dashed line indicates the new average effect estimated by the random-effects model with all the 21 studies.

We use the example of peer norms on adolescent sexual behaviors to illustrate the “trim and fill” method. Figure 6A is a new funnel plot with four imputed effect sizes (the empty dots) in addition to the original seventeen estimates (the solid dots, which is the same as those in Figure 4). These four imputed effect sizes and standard errors are β18 = −0.41, σ18 = 0.32, β19 = −0.79, σ19 = 0.48, β20 = −1.16, σ20 = 0.68, β21 = −1.81, σ21 = −0.80. Following the procedures in Section 4.4, we re-do the random-effects meta-analysis with the total of 21 estimates, and obtain the new between–study variance component (τ2^=0.10), average effect (μ*^=0.38), and standard error (se(μ*^)=0.10). The new average effect is indicated by the vertical dashed line in Figure 6A. This result still indicates a significantly positive effect of perceived peer norms (higher acceptance) on adolescent sexual behaviors (p < 0.01).

Appendix 7. The ANOVA-type approach of heterogeneity analysis

The ANOVA-type approach identifies factors that can partition a heterogeneous group of studies into homogeneous subgroups of studies (Deeks, Altman, and Bradburn 2008). It applies when a study-level factor of interest is categorical (or when the linearity assumption between an ordinal covariate and an effect size is immaterial). Now, assume a categorical variable Z that can partition studies into g subgroups. By using either the fixed-effects or random-effects meta-analysis, we can obtain a pooled average, μi, a standard error of this pooled average, σi, and a heterogeneity statistic, Qi, for each subgroup. See more discussion on whether fixed- or random-effects model should be chosen by Hedges and Pigott (2004) and Overton (1998). However, the fixed-effects models are more likely to be employed because it can provide a more robust estimate of Qi when sample sizes are small in subgroups. Given the overall heterogeneity for all studies, QT (which is obtained using the same method as calculating Qi ), the between-subgroup heterogeneity QB can be written as,

QB=QTi=1gQi [Equation A7.1].

QB can be interpreted as the heterogeneity explained away by Z, and we can perform a one-sided test on QB. The null hypothesis is H0:QB=0 v.s. H1:QB>0, where QB under the null hypothesis conforms to χ2(g1). Rejecting the null hypothesis indicates that the variable Z can substantively explain the variation across subgroups. The quantity QBQT can also be interpreted as the proportion of variance explained by Z, which is similar to R2 for a linear model.

We illustrate the ANOVA-type approach with the example of perceived peer norms on adolescent sexual behaviors. We separately test two study-level factors—the average ages of primary analytical samples and the two measures of perceived peer norms (shown in Table 2)—that can potentially introduce heterogeneity. The average age is dichotomized by a binary variable of whether the average age of respondents is above 14 years old, which is commonly recognized as the threshold between early and middle adolescence (e.g., Clark-Lempers, Lempers, and Ho 1991; Liu 2005). For each subgroup, we obtain Qi by implementing the fixed-effects meta-analysis manually or in STATA following the steps in Section 3.4 and calculate the between-study heterogeneity, QB, with Equation A7.1.

We present in Table 7A the results of this analysis. The total heterogeneity QT is 41.50. The group with the average age above 14 years old (9 studies) has the heterogeneity of 26.78, whereas the group 14 years old or below (8 studies) has the heterogeneity of 13.43. The heterogeneity between the two age groups is therefore 1.29 (using Equation A 7.2). An χ2 test on H0:QB=0 v.s. H1:QB>0 shows that the variance explained away by the two age groups is not significantly larger than 0 at α = 0.05 level, given the 95% critical value of χ2 (1) is 3.84. This result indicates that the effect of perceived peer norms on adolescent sexual behaviors is potentially consistent across the two stages of adolescence.

Moreover, the studies that adopt the measure of the perceived peers’ approval of adolescent sex have the within-group heterogeneity of 13.24 (10 studies), whereas the within-group heterogeneity for studies using the measure of the perceived number of sexually experienced friends is 15.85 (7 studies). The between-group heterogeneity is therefore 41.50–13.24–15.85=12.41. Given the 99.9% critical value of χ2(1) is 10.83, the χ2 test on H0:QB=0 (and Ha:QB>0) shows that this between-group variation is significantly greater than 0 at α = 0.001 level. The two different measures of perceived peer norms also explain 30ε of the total heterogeneity (QBQT*100%=12.4141.50*100%=30%). This result helps clarify the source of heterogeneity from measurement choices and inform future researchers of the properties of these measures: measuring peer norms with the perceived number of sexually experienced friends can lead to significantly larger estimates than the perceived approval measure.

Table 7A.

Results of the heterogeneity analysis using the ANOVA-type approach

Number of Studies Sample Size Average Effect Size S.E. Q QB
Average Age
Above 14 9 6163 0.32*** 0.07 26.78 1.29
   14 or Below 8 6790 0.43*** 0.06 13.43
Peer Norm Measure
Approval 10 5988 0.29*** 0.05 13.24 12.41***
Fri. Count 7 6965 0.65*** 0.09 15.85
Total 17 12953 0.39*** 0.04 41.50

Note:

***

p value <0.001

Appendix 8. Estimating the parameters in meta-regression with the WLS estimator.

This appendix explains how a meta-regression model can be estimated in two steps: estimating τ2 and obtaining the WLS estimates of meta-regression coefficients. The original discussion on this method can be found in Berkey et al. (1995). We estimate τ2 using an extended version of the DerSimonian and Laird estimator. The procedure is the similar to the one that obtains a random-effects model of meta-analysis shown in Section 4.2 and 4.3. Assume a meta-regression with p covariates on a collection of k studies. Q is the heterogeneity statistic defined in Equation 3.5 β=(β1βk), γ=(γ0γ1γp), X=[1x11x12x1p1x21x22x2p11xk1xk2xkp], ω=(w1wk), where wj=1σj2, and W=diag(ω).

(1) Estimating τ2:

τ^2=Q(kp)c [Equation A8.1]
wherec=tr(W)tr((XWX)1XW2X) [Equation A8.2].

Here we postulate that Q is a larger quantity than k-p, because otherwise the analysis of heterogeneity is not well motivated.

(2) Estimating γ:

After the τ2 is estimated, the original model is reduced to a WLS model:

βj=γ0+γ1xj1+γ2xj2++γpxjp+ξj,
whereξjN(0,τ^2+σj2),j=1k. [Equation A8.3]

Redefine wj*=1σj2+τ^2, ω*=(w1*wk*), and W*=diag(ω*), then

γ^=(XW*X)1XW*β [Equation A8.4],

and the covariance matrix of γ^ is:

Σ^=(XW*X)1 [Equation A8.5].

The meta-regression model can be implemented using Stata in two steps: (1) estimate τ2 following Equation A8.1 and A8.2 in Appendix 8, and (2) use vwls function in Stata or Equation A8.4 and A8.5 to obtain estimates and standard errors of γ0,γ1,γ2γp. For the developed commands for meta-regression in Stata, please refer to the metareg package developed by Harbord and Higgins (2008). The metareg function in this package can be used to produce the results in Table 3,

metaregbetax,wsse(se)mm

, where x indicates the covariates and “mm” indicates the moment-based estimator for the between-study heterogeneity.

Contributor Information

Guangyu Tong, Duke University.

Guang Guo, University of North Carolina at Chapel Hill.

References:

  1. Aloe Ariel M. and Becker Betsy Jane. 2012. “An Effect Size for Regression Predictors in Meta-analysis.” Journal of Educational and Behavioral Statistics 37(2):278–297. [Google Scholar]
  2. Amato Paul R. and Keith Bruce. 1991. “Consequences of Parental Divorce for Adult Well–Being: A Meta-Analysis.” Journal of Marriage and the Family 53:43–58. [Google Scholar]
  3. Arthur Winfred Jr, Bennett Winston, and Huffcutt Allen I. 2001. Conducting Meta-Analysis Using SAS. Psychology Press. [Google Scholar]
  4. Becker Betsy Jane and Wu Meng-Jia. 2007. “The Synthesis of Regression Slopes in Meta-analysis.” Statistical Science 414–429.
  5. Berkey Catherine S., Hoaglin David C., Mosteller Frederick, and Colditz Graham A. 1995. “A Random-Effects Regression Model for Meta-Analysis.” Statistics in Medicine 14(4):395–411. [DOI] [PubMed] [Google Scholar]
  6. Boffetta Paolo, Bobak Martin, Axel Borsch-Supan Hermann Brenner, Eriksson Sture, Grodstein Fran, Jansen Eugene, Jenab Mazda, Juerges Hendrik, Kampman Ellen, and others. 2014. “The Consortium on Health and Ageing: Network of Cohorts in Europe and the United States (CHANCES) Project—Design, Population and Data Harmonization of a Large-Scale, International Study.” European Journal of Epidemiology 29(12):929–936. [DOI] [PubMed] [Google Scholar]
  7. Borenstein Michael, Hedges Larry V., Higgins Julian PT, and Rothstein Hannah R. 2009. Introduction to Meta-Analysis. Chichester, UK: John Wiley σ Sons, Ltd. [Google Scholar]
  8. Boruch Robert F. 2005. “Place Randomized Trials in Education, Criminology, Welfare and Health.” The Annals of the American Academy of Political and Social Science 599:6–18. [Google Scholar]
  9. Branigan Amelia R., McCallum Kenneth J., and Freese Jeremy. 2013. “Variation in the Heritability of Educational Attainment: An International Meta-Analysis.” Social Forces 92(1):109–140. [Google Scholar]
  10. Brockwell Sarah E. and Gordon Ian R. 2001. “A Comparison of Statistical Methods for Meta-analysis.” Statistics in Medicine 20(6):825–840. [DOI] [PubMed] [Google Scholar]
  11. Chan An-Wen and Altman Douglas G. 2005. “Identifying Outcome Reporting Bias in Randomized Trials on PubMed: Review of Publications and Survey of Authors.” British Medical Journal 330(7494):753. [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Chan MeowLan Evelyn and Arvey Richard D. 2012. “Meta-Analysis and the Development of Knowledge.” Perspectives on Psychological Science 7(1):79–92. [DOI] [PubMed] [Google Scholar]
  13. Cheung Mike W. L. 2014. “Modeling Dependent Effect Sizes with Three-Level Meta-Analyses: A Structural Equation Modeling Approach.” Psychological Methods 19(2):211. [DOI] [PubMed] [Google Scholar]
  14. Cheung Mike W. L. 2015. Meta-Analysis: A Structural Equation Modeling Approach. John Wiley & Sons. [Google Scholar]
  15. Chinn Susan. 2000. “A Simple Method for Converting an Odds Ratio to Effect Size for Use in Meta-Analysis.” Statistics in Medicine 19(22):3127–3131. [DOI] [PubMed] [Google Scholar]
  16. Cialdini Robert B. and Trost Melanie R. 1998. “Social Influence: Social Norms, Conformity and Compliance.” Pp. 151–92 in The handbook of social psychology. Vol. 2, edited by Gilbert DT, Fiske ST, and Lindzey G Boston: McGraw-Hill. [Google Scholar]
  17. Coleman James S., Campell E, Hubson C, McPartland J, Mood A, Weinfeld F, and York R 1966. Equality of Educational Opportunity. Washington D.C.: U.S. Government Printing Office. [Google Scholar]
  18. Cooper Harris H. and Hedges Larry V. 2009. “Research Synthesis as a Scientific Process.” Pp. 3–18 in, edited by Cooper HH, Hedges LV, and Valentine JC New York: Russell Sage Foundation. [Google Scholar]
  19. Copas John and Shi Jian Qing. 2000. “Meta-Analysis, Funnel Plots and Sensitivity Analysis.” Biostatistics 1(3):247–262. [DOI] [PubMed] [Google Scholar]
  20. Correll Shelley J., Benard Stephen, and Paik In. 2007. “Getting a Job: Is There a Motherhood Penalty?” American Journal of Sociology 112(5):1297–1338. [Google Scholar]
  21. Debray Thomas PA, Moons Karel GM, van Valkenhoef Gert, Efthimiou Orestis, Hummel Noemi, Groenwold Rolf HH, Reitsma Johannes B., and GetReal Methods Review Group. 2015. “Get Real in Individual Participant Data (IPD) Meta-Analysis: A Review of the Methodology.” Research Synthesis Methods 6(4):293–309. [DOI] [PMC free article] [PubMed] [Google Scholar]
  22. Deeks Jonathan J., Altman Douglas G., and Bradburn Michael J. 2008. “Statistical Methods for Examining Heterogeneity and Combining Results from Several Studies in Meta-Analysis.” Systematic Reviews in Health Care: Meta-Analysis in Context, Second Edition 285–312.
  23. Doucouliagos Chris and Stanley Tom D. 2013. “Are All Economic Facts Greatly Exaggerated? Theory Competition and Selectivity.” Journal of Economic Surveys 27(2):316–339. [Google Scholar]
  24. Doucouliagos Hristos and Paldam Martin. 2006. “Aid Effectiveness on Accumulation: A Meta Study.” Kyklos 59(2):227–254. [Google Scholar]
  25. Duval Sue and Tweedie Richard. 2000. “Trim and Fill: A Simple Funnel-Plot–Based Method of Testing and Adjusting for Publication Bias in Meta-Analysis.” Biometrics 56(2):455–463. [DOI] [PubMed] [Google Scholar]
  26. Egger Matthias, George Davey Smith Martin Schneider, and Minder Christoph. 1997. “Bias in Meta-Analysis Detected by a Simple, Graphical Test.” British Medical Journal 315(7109):629–634. [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Firebaugh Glenn. 2008. Seven Rules for Social Research. NJ: Princeton University Press. [Google Scholar]
  28. Gerber Alan S. and Malhotra Neil. 2008. “Publication Bias in Empirical Sociological Research: Do Arbitrary Significance Levels Distort Published Results?” Sociological Methods σ Research 37(1):3–30. [Google Scholar]
  29. Glass Gene V. 1976. “Primary, Secondary, and Meta-Analysis of Research.” Educational Researcher 5(10):3–8. [Google Scholar]
  30. Harris R, Bradburn M, Deeks J, Roger Harbord D Altman, and Sterne J 2008. “Metan: Fixed-and Random-Effects Meta-Analysis.” Stata Journal 8(1):3. [Google Scholar]
  31. Hedges Larry V. 1983. “A Random Effects Model for Effect Sizes.” Psychological Bulletin 93(2):388. [Google Scholar]
  32. Hedges Larry V., Laine Richard D., and Greenwald Rob. 1994. “An Exchange: Part I: Does Money Matter? A Meta-Analysis of Studies of the Effects of Differential School Inputs on Student Outcomes.” Educational Researcher 23(3):5–14. [Google Scholar]
  33. Hedges Larry V., Tipton Elizabeth, and Johnson Matthew C. 2010. “Robust Variance Estimation in Meta-Regression with Dependent Effect Size Estimates.” Research Synthesis Methods 1(1):39–65. [DOI] [PubMed] [Google Scholar]
  34. Hedges Larry V. and Vevea Jack L. 1998. “Fixed-and Random-Effects Models in Meta-analysis.” Psychological Methods 3(4):486. [Google Scholar]
  35. Higgins Julian PT, Thompson Simon G., Deeks Jonathan J., and Altman Douglas G. 2003. “Measuring Inconsistency in Meta-Analyses.” BMJ: British Medical Journal 327(7414):557. [DOI] [PMC free article] [PubMed] [Google Scholar]
  36. Higgins Julian and Thompson Simon G. 2002. “Quantifying Heterogeneity in a Meta-analysis.” Statistics in Medicine 21(11):1539–1558. [DOI] [PubMed] [Google Scholar]
  37. Hoaglin David C. 2016. “Misunderstandings about Q and ‘Cochran’s Q Test’in Meta-Analysis.” Statistics in Medicine 35(4):485–495. [DOI] [PubMed] [Google Scholar]
  38. Huedo-Medina Tania B., Julio Sánchez-Meca Fulgencio Marín-Martínez, and Botella Juan. 2006. “Assessing Heterogeneity in Meta-Analysis: Q Statistic or I2 Index?” Psychological Methods 11(2):193. [DOI] [PubMed] [Google Scholar]
  39. Jackson Dan, Bowden Jack, and Baker Rose. 2010. “How Does the DerSimonian and Laird Procedure for Random Effects Meta-Analysis Compare with Its More Efficient but Harder to Compute Counterparts?” Journal of Statistical Planning and Inference 140(4):961–970. [Google Scholar]
  40. Jackson Dan, Riley Richard, and White Ian R. 2011. “Multivariate Meta-Analysis: Potential and Promise.” Statistics in Medicine 30(20):2481–2498. [DOI] [PMC free article] [PubMed] [Google Scholar]
  41. Johnson Elizabeth I. and Easterling Beth. 2012. “Understanding Unique Effects of Parental Incarceration on Children: Challenges, Progress, and Recommendations.” Journal of Marriage and Family 74(2):342–356. [Google Scholar]
  42. Kandel Denise B. 1978. “Homophily, Selection, and Socialization in Adolescent Friendships.” American Journal of Sociology 84(2):427–436. [Google Scholar]
  43. Konstantopoulos Spyros and Hedges Larry V. 2004. “Meta-Analysis.” Pp. 281–297 in The Sage handbook of quantitative methodology for the social sciences, edited by Kaplan D Thousand Oaks, CA: Sage publications. [Google Scholar]
  44. Lewis Steff and Clarke Mike. 2001. “Forest Plots: Trying to See the Wood and the Trees.” Bmj 322(7300):1479–1480. [DOI] [PMC free article] [PubMed] [Google Scholar]
  45. Lin Lifeng and Chu Haitao. 2018. “Quantifying Publication Bias in Meta-Analysis.” Biometrics 74(3):785–794. [DOI] [PMC free article] [PubMed] [Google Scholar]
  46. Lin Lifeng, Chu Haitao, Mohammad Hassan Murad Chuan Hong, Qu Zhiyong, Cole Stephen R., and Chen Yong. 2018. “Empirical Comparison of Publication Bias Tests in Meta-analysis.” Journal of General Internal Medicine 33(8):1260–67. [DOI] [PMC free article] [PubMed] [Google Scholar]
  47. Lipsey Mark W. 2009. “Identifying Interesting Variables and Analysis Opportunities.” Pp. 147–158 in The handbook of research synthesis and meta-analysis, edited by Cooper HH, Hedges LV, and Valentine JC New York: Russell Sage Foundation. [Google Scholar]
  48. Lipsey Mark W. and Wilson David B. 2001. Practical Meta-Analysis. Thousand Oaks, CA: Sage Publications, Inc. [Google Scholar]
  49. Manski Charles F. 1993. “Identification of Endogenous Social Effects: The Reflection Problem.” The Review of Economic Studies 60(3):531–542. [Google Scholar]
  50. McCartney Kathleen and Rosenthal Robert. 2000. “Effect Size, Practical Importance, and Social Policy for Children.” Child Development 71(1):173–180. [DOI] [PubMed] [Google Scholar]
  51. McManus RJ, Sue Wilson BC Delaney, Fitzmaurice DA, Hyde CJ, Tobias RS, Jowett Sue, and Hobbs FDR 1998. “Review of the Usefulness of Contacting Other Experts When Conducting a Literature Search for Systematic Reviews.” British Medical Journal 317(7172):1562–1563. [DOI] [PMC free article] [PubMed] [Google Scholar]
  52. Moffitt Terrie E., Caspi A, Rutter M, and Silva PA 2001. Sex Differences in Antisocial Behaviour: Conduct Disorder, Delinquency, and Violence in the Dunedin Longitudinal Study. Cambridge, UK: Cambridge university press. [Google Scholar]
  53. Mulder Monique Borgerhoff, Bowles Samuel, Hertz Tom, Bell Adrian, Beise Jan, Clark Greg, Fazzio Ila, Gurven Michael, Hill Kim, and Hooper Paul L. 2009. “Intergenerational Wealth Transmission and the Dynamics of Inequality in Small-Scale Societies.” Science 326(5953):682–688. [DOI] [PMC free article] [PubMed] [Google Scholar]
  54. Peters Jaime L., Sutton Alex J., Jones David R., Abrams Keith R., and Rushton Lesley. 2007. “Performance of the Trim and Fill Method in the Presence of Publication Bias and Between-Study Heterogeneity.” Statistics in Medicine 26(25):4544–62. [DOI] [PubMed] [Google Scholar]
  55. Poldin Oleg, Valeeva Diliara, and Yudkevich Maria. 2016. “Which Peers Matter: How Social Ties Affect Peer-Group Effects.” Research in Higher Education 57(4):448–468. [Google Scholar]
  56. Reed Jeffrey and Baxter Pam. 2009. “Using Reference Databases.” Pp. 73–102 in HANDBOOK OF RESEARCH SYNTHESIS AND META-ANALYSIS, edited by Cooper HH, Hedges LV, and Valentine JC New York: Russell Sage Foundation. [Google Scholar]
  57. Roelfs David J., Shor Eran, Falzon Louise, Davidson Karina W., and Schwartz Joseph E. 2013. “Meta-Analysis for Sociology–a Measure-Driven Approach.” Bulletin of Sociological Methodology 117(1):75–92. [DOI] [PMC free article] [PubMed] [Google Scholar]
  58. Roksa Josipa and Potter Daniel. 2011. “Parenting and Academic Achievement: Intergenerational Transmission of Educational Advantage.” Sociology of Education 84(4):299–321. [Google Scholar]
  59. Rosenthal Robert. 1979. “The File Drawer Problem and Tolerance for Null Results.” Psychological Bulletin 86(3):638. [Google Scholar]
  60. Rothstein Hannah R. and Hopewell Sally. 2009. “Grey Literature.” in The handbook of research synthesis and meta-analysis, edited by Cooper H, Hedges LV, and Valentine JC New York: Russell Sage Foundation. [Google Scholar]
  61. Rothstein Hannah R., Sutton Alexander J., and Borenstein Michael. 2006. Publication Bias in Meta-Analysis: Prevention, Assessment and Adjustments. John Wiley σ Sons. [Google Scholar]
  62. Sacerdote Bruce. 2014. “Experimental and Quasi-Experimental Analysis of Peer Effects: Two Steps Forward?” Annual Review of Economics 6(1):253–272. [Google Scholar]
  63. Shor Eran, Roelfs David J., Curreli Misty, Clemow Lynn, Burg Matthew M., and Schwartz Joseph E. 2012. “Widowhood and Mortality: A Meta-Analysis and Meta-Regression.” Demography 49(2):575–606. [DOI] [PMC free article] [PubMed] [Google Scholar]
  64. Shor Eran, Roelfs David J., and Yogev Tamar. 2013. “The Strength of Family Ties: A Meta-analysis and Meta-Regression of Self-Reported Social Support and Mortality.” Social Networks 35(4):626–638. [Google Scholar]
  65. Stanley Tom D. 2001. “Wheat from Chaff: Meta-Analysis as Quantitative Literature Review.” The Journal of Economic Perspectives 15(3):131–150. [Google Scholar]
  66. Stanley Tom D. and Doucouliagos Hristos. 2015. “Neither Fixed nor Random: Weighted Least Squares Meta-Analysis.” Statistics in Medicine 34(13):2116–2127. [DOI] [PubMed] [Google Scholar]
  67. Stanley Tom D., Doucouliagos Hristos, Giles Margaret, Heckemeyer Jost H., Johnston Robert J., Laroche Patrice, Nelson Jon P., Paldam Martin, Poot Jacques, and Pugh Geoff. 2013. “Meta-Analysis of Economics Research Reporting Guidelines.” Journal of Economic Surveys 27(2):390–394. [Google Scholar]
  68. Sterne Jonathan AC and Egger Matthias. 2005. “Regression Methods to Detect Publication and Other Bias in Meta-Analysis.” Publication Bias in Meta-Analysis: Prevention, Assessment and Adjustments 99–110.
  69. Sutton Alex J., Duval SJ, Tweedie RL, Abrams Keith R., and Jones David R. 2000. “Empirical Assessment of Effect of Publication Bias on Meta-Analyses.” Bmj 320(7249):1574–1577. [DOI] [PMC free article] [PubMed] [Google Scholar]
  70. Sutton Alexander J. and Higgins Julian. 2008. “Recent Developments in Meta-Analysis.” Statistics in Medicine 27(5):625–650. [DOI] [PubMed] [Google Scholar]
  71. Thomas Jeremy N. and Olson Daniel VA. 2010. “Testing the Strictness Thesis and Competing Theories of Congregational Growth.” Journal for the Scientific Study of Religion 49(4):619–639. [Google Scholar]
  72. Thompson Simon G. and Higgins Julian. 2002. “How Should Meta-Regression Analyses Be Undertaken and Interpreted?” Statistics in Medicine 21(11):1559–1573. [DOI] [PubMed] [Google Scholar]
  73. Valentine Jeffrey C. 2009. “Judging the Quality of Primary Research.” Pp. 129–46 in The Handbook of Research Synthesis and Meta-analysis, edited by Cooper HH, Hedges LV, and Valentine JC New York: Russell Sage Foundation. [Google Scholar]
  74. Van de Bongardt Daphne, Reitz Ellen, Sandfort Theo, and Deković Maja. 2015. “A Meta-analysis of the Relations between Three Types of Peer Norms and Adolescent Sexual Behavior.” Personality and Social Psychology Review 19(3):203–234. [DOI] [PMC free article] [PubMed] [Google Scholar]
  75. Van Den Noortgate Wim and Onghena Patrick. 2005. “Parametric and Nonparametric Bootstrap Methods for Meta-Analysis.” Behavior Research Methods 37(1):11–22. [DOI] [PubMed] [Google Scholar]
  76. Van Houwelingen, Hans C, Arends Lidia R., and Stijnen Theo. 2002. “Advanced Methods in Meta-Analysis: Multivariate Approach and Meta-Regression.” Statistics in Medicine 21(4):589–624. [DOI] [PubMed] [Google Scholar]
  77. Veroniki Areti Angeliki, Jackson Dan, Viechtbauer Wolfgang, Bender Ralf, Bowden Jack, Knapp Guido, Kuss Oliver, Julian PT Higgins Dean Langan, and Salanti Georgia. 2016. “Methods to Estimate the Between-Study Variance and Its Uncertainty in Meta-analysis.” Research Synthesis Methods 7(1):55–79. [DOI] [PMC free article] [PubMed] [Google Scholar]
  78. Viechtbauer Wolfgang. 2007. “Confidence Intervals for the Amount of Heterogeneity in Meta-analysis.” Statistics in Medicine 26(1):37–52. [DOI] [PubMed] [Google Scholar]
  79. Viechtbauer Wolfgang. 2015. “Package ‘Metafor.’” The Comprehensive R Archive Network.
  80. Wilson David B. 2009. “Systematic Coding.” Pp. 159–76 in The Handbook of Research Synthesis and Meta-analysis, edited by Cooper HH, Hedges LV, and Valentine JC New York: Russell Sage Foundation. [Google Scholar]
  81. Young Cristobal. 2009. “Model Uncertainty in Sociological Research: An Application to Religion and Economic Growth.” American Sociological Review 74(3):380–397. [Google Scholar]
  82. Young Cristobal and Holsteen Katherine. 2017. “Model Uncertainty and Robustness: A Computational Framework for Multimodel Analysis.” Sociological Methods σ Research 46(1):3–40. [Google Scholar]

Reference:

  1. Akers Ronald L. 2011. Social Learning and Social Structure: A General Theory of Crime and Deviance. New Brunswick, NJ: Transaction Publishers. [Google Scholar]
  2. Berkey Catherine S., Hoaglin David C., Mosteller Frederick, and Colditz Graham A. 1995. “A Random-Effects Regression Model for Meta-Analysis.” Statistics in Medicine 14(4):395–411. [DOI] [PubMed] [Google Scholar]
  3. Bersamin Melina M., Walker Samantha, Waiters Elizabeth D., Fisher Deborah A., and Grube Joel W. 2005. “Promising to Wait: Virginity Pledges and Adolescent Sexual Behavior.” Journal of Adolescent Health 36(5):428–436. [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Biggerstaff Brad J. and Jackson Dan. 2008. “The Exact Distribution of Cochran’s Heterogeneity Statistic in One-Way Random Effects Meta-Analysis.” Statistics in Medicine 27(29):6093–6110. [DOI] [PubMed] [Google Scholar]
  5. Borenstein Michael, Hedges Larry V., Higgins Julian PT, and Rothstein Hannah R. 2009. Introduction to Meta-Analysis. Chichester, UK: John Wiley σ Sons, Ltd. [Google Scholar]
  6. Bowden Jack, Thompson John R., and Burton Paul. 2006. “Using Pseudo-Data to Correct for Publication Bias in Meta-Analysis.” Statistics in Medicine 25(22):3798–3813. [DOI] [PubMed] [Google Scholar]
  7. Cabral Patricia, Wallander Jan L., Song Anna V., Elliott Marc N., Tortolero Susan R., Reisner Sari L., and Schuster Mark A. 2017. “Generational Status and Social Factors Predicting Initiation of Partnered Sexual Activity among Latino/a Youth.” Health Psychology 36(2):169. [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Carrell Scott E., Fullerton Richard L., and West James E. 2009. “Does Your Cohort Matter? Measuring Peer Effects in College Achievement.” Journal of Labor Economics 27(3):439–464. [Google Scholar]
  9. Carvajal Scott C., Parcel Guy S., Karen Basen-Engquist Stephen W. Banspach, Coyle Karin K., Kirby Douglas, and Chan Wenyaw. 1999. “Psychosocial Predictors of Delay of First Sexual Intercourse by Adolescents.” Health Psychology 18(5):443. [DOI] [PubMed] [Google Scholar]
  10. Chinn Susan. 2000. “A Simple Method for Converting an Odds Ratio to Effect Size for Use in Meta-Analysis.” Statistics in Medicine 19(22):3127–3131. [DOI] [PubMed] [Google Scholar]
  11. Chowdhry Amit K., Dworkin Robert H., and McDermott Michael P. 2016. “Meta-Analysis with Missing Study-Level Sample Variance Data.” Statistics in Medicine 35(17):3021–3032. [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Clark-Lempers Dania S., Lempers Jacques D., and Ho Camilla. 1991. “Early, Middle, and Late Adolescents’ Perceptions of Their Relationships with Significant Others.” Journal of Adolescent Research 6(3):296–315. [Google Scholar]
  13. Cochran William G. 1954. “The Combination of Estimates from Different Experiments.” Biometrics 10(1):101–129. [Google Scholar]
  14. Deeks Jonathan J., Altman Douglas G., and Bradburn Michael J. 2008. “Statistical Methods for Examining Heterogeneity and Combining Results from Several Studies in Meta-Analysis.” Systematic Reviews in Health Care: Meta-Analysis in Context, Second Edition 285–312.
  15. DerSimonian Rebecca and Laird Nan. 1986. “Meta-Analysis in Clinical Trials.” Controlled Clinical Trials 7(3):177–188. [DOI] [PubMed] [Google Scholar]
  16. DiIorio Colleen, Dudley William N., Kelly Maureen, Soet Johanna E., Mbwara Joyce, and Potter Jennifer Sharpe. 2001. “Social Cognitive Correlates of Sexual Experience and Condom Use among 13-through 15-Year-Old Adolescents.” Journal of Adolescent Health 29(3):208–216. [DOI] [PubMed] [Google Scholar]
  17. DiIorio Colleen, Dudley William N., Soet Johanna E., and McCarty Frances. 2004. “Sexual Possibility Situations and Sexual Behaviors among Young Adolescents: The Moderating Role of Protective Factors.” Journal of Adolescent Health 35(6):528–e11. [DOI] [PubMed] [Google Scholar]
  18. Doornwaard Suzan M., Ter Bogt Tom FM, Reitz Ellen, and Van Den Eijnden Regina JJM. 2015. “Sex-Related Online Behaviors, Perceived Peer Norms and Adolescents’ Experience with Sexual Behavior: Testing an Integrative Model.” PloS One 10(6):e0127787. [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Egger Matthias, George Davey Smith Martin Schneider, and Minder Christoph. 1997. “Bias in Meta-Analysis Detected by a Simple, Graphical Test.” British Medical Journal 315(7109):629–634. [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Foster Gigi. 2006. “It’s Not Your Peers, and It’s Not Your Friends: Some Progress toward Understanding the Educational Peer Effect Mechanism.” Journal of Public Economics 90(8):1455–1475. [Google Scholar]
  21. Gillmore Mary Rogers, Archibald Matthew E., Morrison Diane M., Wilsdon Anthony, Wells Elizabeth A., Hoppe Marilyn J., Nahom Deborah, and Murowchick Elise. 2002. “Teen Sexual Behavior: Applicability of the Theory of Reasoned Action.” Journal of Marriage and Family 64(4):885–897. [Google Scholar]
  22. Harbord Roger M. and Higgins Julian PT. 2008. “Meta-Regression in Stata.” The Stata Journal 8(4):493–519. [Google Scholar]
  23. Hasselblad Vic and Hedges Larry V. 1995. “Meta-Analysis of Screening and Diagnostic Tests.” Psychological Bulletin 117(1):167. [DOI] [PubMed] [Google Scholar]
  24. Hedges Larry V. and Pigott Therese D. 2004. “The Power of Statistical Tests for Moderators in Meta-Analysis.” Psychological Methods 9(4):426. [DOI] [PubMed] [Google Scholar]
  25. Kabiru Caroline W., Beguy Donatien, Undie Chi-Chi, Zulu Eliya Msiyaphazi, and Ezeh Alex C. 2010. “Transition into First Sex among Adolescents in Slum and Non-Slum Communities in Nairobi, Kenya.” Journal of Youth Studies 13(4):453–471. [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. Kawai Kosuke, Kaaya Sylvia F., Kajula Lusajo, Mbwambo Jessie, Kilonzo Gad P., and Fawzi Wafaie W. 2008. “Parents’ and Teachers’ Communication about HIV and Sex in Relation to the Timing of Sexual Initiation among Young Adolescents in Tanzania.” Scandinavian Journal of Social Medicine 36(8):879–888. [DOI] [PubMed] [Google Scholar]
  27. Kinsman Sara B., Romer Daniel, Furstenberg Frank F., and Schwarz Donald F. 1998. “Early Sexual Initiation: The Role of Peer Norms.” Pediatrics 102(5):1185–1192. [DOI] [PubMed] [Google Scholar]
  28. Kremer Michael and Levy Dan. 2008. “Peer Effects and Alcohol Use among College Students.” The Journal of Economic Perspectives 22(3):189–189. [Google Scholar]
  29. Kulinskaya Elena, Dollinger Michael B., and Bjørkestøl Kirsten. 2011. “On the Moments of Cochran’s Q Statistic under the Null Hypothesis, with Application to the Meta-Analysis of Risk Difference.” Research Synthesis Methods 2(4):254–270. [DOI] [PubMed] [Google Scholar]
  30. Laflin Molly T., Wang Jing, and Barry Maxine. 2008. “A Longitudinal Study of Adolescent Transition from Virgin to Nonvirgin Status.” Journal of Adolescent Health 42(3):228–236. [DOI] [PubMed] [Google Scholar]
  31. L’Engle Kelly Ladin and Jackson Christine. 2008. “Socialization Influences on Early Adolescents’ Cognitive Susceptibility and Transition to Sexual Intercourse.” Journal of Research on Adolescence 18(2):353–378. [Google Scholar]
  32. Levy Dan Maurice. 2000. “Family Income and Peer Effects as Determinants of Educational Outcomes.” Evanston, IL.
  33. Little Craig B. and Rankin Andrea. 2001. “Why Do They Start It? Explaining Reported Early-Teen Sexual Activity.” Sociological Forum 16:703–729. [Google Scholar]
  34. Liu Ruth X. 2005. “Parent-Youth Closeness and Youth’s Suicidal Ideation: The Moderating Effects of Gender, Stages of Adolescence, and Race or Ethnicity.” Youth σ Society 37(2):145–175. [Google Scholar]
  35. Lyle David S. 2007. “Estimating and Interpreting Peer and Role Model Effects from Randomly Assigned Social Groups at West Point.” The Review of Economics and Statistics 89(2):289–299. [Google Scholar]
  36. Maguen Shira and Armistead Lisa. 2006. “Abstinence among Female Adolescents: Do Parents Matter above and beyond the Influence of Peers?” American Journal of Orthopsychiatry 76(2):260. [DOI] [PubMed] [Google Scholar]
  37. McCartney Kathleen and Rosenthal Robert. 2000. “Effect Size, Practical Importance, and Social Policy for Children.” Child Development 71(1):173–180. [DOI] [PubMed] [Google Scholar]
  38. Overton Randall C. 1998. “A Comparison of Fixed-Effects and Mixed (Random-Effects) Models for Meta-Analysis Tests of Moderator Variable Effects.” Psychological Methods 3(3):354. [Google Scholar]
  39. Pai Hsiang-Chu and Lee Sheuan. 2012. “Sexual Self-Concept as Influencing Intended Sexual Health Behaviour of Young Adolescent Taiwanese Girls.” Journal of Clinical Nursing 21(13–14):1988–97. [DOI] [PubMed] [Google Scholar]
  40. Peters Jaime L., Sutton Alex J., Jones David R., Abrams Keith R., and Rushton Lesley. 2006. “Comparison of Two Methods to Detect Publication Bias in Meta-Analysis.” JAMA 295(6):676–680. [DOI] [PubMed] [Google Scholar]
  41. Potard Courtois, Courtois R, and Rusch E 2008. “The Influence of Peers on Risky Sexual Behaviour during Adolescence.” The European Journal of Contraception σ Reproductive Health Care 13(3):264–270. [DOI] [PubMed] [Google Scholar]
  42. Reitz Ellen, van de Bongardt Daphne, Baams Laura, Doornwaard Suzan, Dalenberg Wieke, Dubas Judith, van Aken Marcel, Overbeek Geertjan, Bogt Tom ter, van der Eijnden Regina, and others. 2015. “Project STARS (Studies on Trajectories of Adolescent Relationships and Sexuality): A Longitudinal, Multi-Domain Study on Sexual Development of Dutch Adolescents.” European Journal of Developmental Psychology 12(5):613–626. [Google Scholar]
  43. Rosenthal Doreen A., Smith Anthony MA, and De Visser Richard. 1999. “Personal and Social Factors Influencing Age at First Sexual Intercourse.” Archives of Sexual Behavior 28(4):319–333. [DOI] [PubMed] [Google Scholar]
  44. Sacerdote Bruce. 2001. “Peer Effects with Random Assignment: Results for Dartmouth Roommates.” The Quarterly Journal of Economics 116(2):681–704. [Google Scholar]
  45. Sacerdote Bruce. 2014. “Experimental and Quasi-Experimental Analysis of Peer Effects: Two Steps Forward?” Annual Review of Economics 6(1):253–272. [Google Scholar]
  46. Schott James R. 2016. Wiley: Matrix Analysis for Statistics. 3rd ed. Hoboken, NJ: Wiley. [Google Scholar]
  47. Sieving Renee E., Eisenberg Marla E., Pettingell Sandra, and Skay Carol. 2006. “Friends’ Influence on Adolescents’ First Sexual Intercourse.” Perspectives on Sexual and Reproductive Health 38(1):13–19. [DOI] [PubMed] [Google Scholar]
  48. Steichen Thomas. 2010. “METATRIM: Stata Module to Perform Nonparametric Analysis of Publication Bias.” Statistical Software Components.
  49. Stinebrickner Ralph and Stinebrickner Todd R. 2006. “What Can Be Learned about Peer Effects Using College Roommates? Evidence from New Survey Data and Students from Disadvantaged Backgrounds.” Journal of Public Economics 90(8):1435–1454. [Google Scholar]
  50. Sutton Alexander J. and Higgins Julian. 2008. “Recent Developments in Meta-Analysis.” Statistics in Medicine 27(5):625–650. [DOI] [PubMed] [Google Scholar]
  51. Van de Bongardt Daphne, Hanneke De Graaf Ellen Reitz, and Deković Maja. 2014. “Parents as Moderators of Longitudinal Associations between Sexual Peer Norms and Dutch Adolescents’ Sexual Initiation and Intention.” Journal of Adolescent Health 55(3):388–393. [DOI] [PubMed] [Google Scholar]
  52. Van de Bongardt Daphne, Reitz Ellen, Sandfort Theo, and Deković Maja. 2015. “A Meta-Analysis of the Relations between Three Types of Peer Norms and Adolescent Sexual Behavior.” Personality and Social Psychology Review 19(3):203–234. [DOI] [PMC free article] [PubMed] [Google Scholar]
  53. Vevea Jack L. and Woods Carol M. 2005. “Publication Bias in Research Synthesis: Sensitivity Analysis Using a Priori Weight Functions.” Psychological Methods 10(4):428. [DOI] [PubMed] [Google Scholar]
  54. Villarruel Antonia M., Jemmott John B. III, Jemmott Loretta S., and Ronis David L. 2004.
  55. “Predictors of Sexual Intercourse and Condom Use Intentions among Spanish-Dominant Latino Youth: A Test of the Planned Behavior Theory.” Nursing Research 53(3):172–181. [DOI] [PubMed] [Google Scholar]
  56. Zimmerman David J. 2003. “Peer Effects in Academic Outcomes: Evidence from a Natural Experiment.” The Review of Economics and Statistics 85(1):9–23. [Google Scholar]

RESOURCES