Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2018 May 1.
Published in final edited form as: Can Psychol. 2016 Oct 6;58(2):140–147. doi: 10.1037/cap0000074

Reporting Practices and Use of Quantitative Methods in Canadian Journal Articles in Psychology

Alyssa Counsell 1,*, Lisa L Harlow 2
PMCID: PMC5494980  NIHMSID: NIHMS817038  PMID: 28684887

Abstract

With recent focus on the state of research in psychology, it is essential to assess the nature of the statistical methods and analyses used and reported by psychological researchers. To that end, we investigated the prevalence of different statistical procedures and the nature of statistical reporting practices in recent articles from the four major Canadian psychology journals. The majority of authors evaluated their research hypotheses through the use of analysis of variance (ANOVA), t-tests, and multiple regression. Multivariate approaches were less common. Null hypothesis significance testing remains a popular strategy, but the majority of authors reported a standardized or unstandardized effect size measure alongside their significance test results. Confidence intervals on effect sizes were infrequently employed. Many authors provided minimal details about their statistical analyses and less than a third of the articles presented on data complications such as missing data and violations of statistical assumptions. Strengths of and areas needing improvement for reporting quantitative results are highlighted. The paper concludes with recommendations for how researchers and reviewers can improve comprehension and transparency in statistical reporting.

Keywords: Canadian psychology, quantitative methods, statistics, review, reporting practices


Quantitative methods are widely used in psychology, but not without controversy and debate. For example, a number of articles and books are devoted to discussing the use of null hypothesis significance testing (NHST: e.g., Chow, 1996; Cumming, 2012; Harlow, Mulaik, & Steiger, 2016; Kline, 2013). The NHST debate led to the publication of guidelines from the American Psychological Association (APA) through the creation of a task force on statistical inference (Wilkinson & the APA Task Force on Statistical Inference, 1999). These guidelines are described in detail in the current edition of the publication manual (APA, 2010). Given the wide number of individuals arguing against NHST or at the very least for better supplementing of NHST information, we believe that the following paper contributes by providing information about the extent to which recent empirical articles in Canadian journals have incorporated these guidelines into their reporting practices.

Reporting Quantitative Results

Significance tests

The majority of empirical articles in psychology use NHST (Rodgers, 2010) despite considerable opposition to an exclusive focus on dichotomous significance tests (e.g., Cohen, 1994; Cumming, 2012; Kline, 2013; Rozeboom, 1997; Schmidt & Hunter, 1997; Wilkinson et al., 1999). Amidst these opposing perspectives, a number of researchers endorse the use of significance tests in some circumstances, particularly if accompanied by relevant effect sizes and confidence intervals (CIs) (e.g., Abelson, 1997; Denis, 2003; Hagen, 1997; Harlow, 2010; Harris, 1997; Mulaik, Raju, & Harshman, 1997). The publication manual (APA, 2010) recommends reporting full results from hypothesis tests (including the test statistic, degrees of freedom, and exact p value), but also recommends including information about measures of magnitude and CIs. Given that most journals follow the APA publication manual for reporting practices, it is unsurprising that applied researchers continue to rely on reporting NHST results.

Effect sizes and confidence intervals

Many of those opposed to NHST consider effect sizes the viable alternative. Cumming (2008; 2012) and Thompson (2007) have strongly argued against using NHST and advocate reporting effect sizes and their associated CIs without tests of statistical significance. Their reasoning is that effect sizes provide information about the magnitude or importance of an effect, which is really what researchers want, rather than whether a nil hypothesis has been rejected. Effect sizes are particularly useful when accompanied by information about variability around the effect (e.g., CI); and if the CI contains the null hypothesis value (e.g., a mean difference of 0), lack of statistical significance can be inferred. Wilkinson and the Task Force (1999) explicitly stated that effect sizes should always be presented, while a measure of variability such as a CI should be included on any effect size reported. It is becoming more common for journals to require effect sizes but CIs are not typically required, so the extent to which individuals are reporting effect sizes and CIs together is not clear.

Visual representations of data

Figures allow researchers to present a large amount of data in an efficient manner so that readers may examine the results in a more comprehensive manner than solely providing results of significance tests. The APA publication manual devotes a large section to tables and figures in order to discuss the many purposes for visual displays of data such as exploration, communication, calculation, storage or decoration. A number of books and articles on graphical expressions of data have been written (e.g., Cleveland, 1993; Friendly, 2000; Friendly & Meyer, 2015, Margolis & Pauwels, 2011, Tukey, 1977). Further, Wilkinson and the Task Force (1999) recommend researchers include high quality figures with indicators of variability to convey statistical findings.

Data complications

Researchers rarely collect data without complications. Issues such as missing data (due to attrition, nonresponses, etc.) or violations of statistical assumptions (e.g., nonnormally distributed variables) should be considered when reporting the results of statistical tests. These complications are described below.

Missing Data

Numerous resources discuss missing data in psychology (Allison, 2002; Baraldi, & Enders, 2013; Enders, 2010; Little & Rubin, 2002). Several strategies for dealing with missing data have been suggested in the past (e.g., listwise deletion, pairwise deletion, mean substitution), however, these methods have been called into question and newer methods have been advocated, including multiple imputation and full information maximum likelihood (Graham, 2009). Multiple imputation is a multistage process whereby the missing data points are replaced by a score predicted from a regression line calculated by including other relevant variables. With full information maximum likelihood, one does not replace missing data points, but instead produces model estimates using all of the available information from the data.

Unfortunately, the most common methods used by applied researchers tend to be those that rely on software defaults (e.g., listwise deletion in SPSS) rather than recommendations by methodologists (Bodner, 2006; Wood, White, & Thompson, 2004). The APA task force on statistical inference stated that excluding cases with missing data is “among the worst methods available for practical applications” (Wilkinson et al., 1999, p. 598). Aside from issues with overreliance on simple methods for dealing with missing data, both Kline (2013) and Reinhart (2015) discuss how many articles do not explicitly state how the research dealt with missing data problems at all.

Statistical Assumptions

Parametric statistical tests must satisfy a number of statistical assumptions in order for valid interpretation of the results. Unfortunately, articles include little information about statistical assumptions (Kline, 2013). A lack of information about statistical assumptions could stem from assumptions not being tested, mistaken information about the robustness of statistical tests, (e.g., Bradley, 1978; Glass, Peckham, & Sanders, 1972), being unaware of the importance of attending to statistical assumptions, or that the data have met the statistical assumptions but the researchers have simply not reported it. The decision for a researcher to use a parametric test should depend on which of these scenarios occur.

Improving Psychological Science through Reporting Practices

All of the issues described thus far contribute to the larger problem in psychology of lack of transparency and issues with replication. Special issues of journals are devoted to this topic (e.g., Perspectives on Psychological Science), along with a number of general psychology articles discussing research practice (e.g., Anderson & Maxwell, 2016; Funder et al., 2014; Kline, 2013, Nosek, Spies, & Motyl, 2012; Wilkinson et al., 1999). While we believe that these articles are of paramount importance, it remains worthwhile to examine the impact of such papers on articles in practice. We believe that it is important not only to discuss areas for improvement when applied researchers do not follow some of these recommended guidelines, but also discuss what researchers are doing well. As such, this paper contributes to the literature by providing concrete information about what and how recent Canadian journal articles in psychology are reporting.

Purpose of the Current Study

The purpose of the current paper is to investigate the statistical practices used in recent psychology articles in Canadian journals. Specifically, we aim to examine the frequency of specific statistical procedures (e.g., correlation, t-test, etc.) and the types of information reported in conjunction with a quantitative analysis (e.g., figures, effect sizes, confidence intervals). We further sought to assess whether articles were using statistical procedures appropriate for their research design. The paper will discuss both strengths and limitations of the current reporting practices of the articles in Canadian journals and conclude with recommendations for reporting quantitative results.

Method

The current study examined all issues of the four major Canadian psychology journals published in 2013. The journals included: Canadian Psychology (CP), the Canadian Journal of Experimental Psychology (CJEP), the Canadian Journal of Behavioural Science (CJBS), and the Canadian Journal of School Psychology (CJSP). The first three journals are both Canadian Psychological Association (CPA) and American Psychological Association (APA) journals. After excluding articles that were not empirical studies (e.g., editorials, book reviews, commentaries, and theory/review papers), our first task was to classify the articles based on whether they used qualitative or quantitative analyses. For articles that included quantitative methods, we examined specific types of quantitative information and sought to assess the appropriateness of the quantitative methods used. As an example for assessing appropriateness, if a researcher had data with a dependency structure (e.g., individuals nested within couples and both were included in the study), it would be inappropriate to use a traditional linear regression model instead of one that takes the dependency into account (e.g., a multilevel model). To collect information regarding whether specific types of information were included from a quantitative analysis, the first author coded whether the information was included and the second author reviewed the coding and made suggestions when needed.

Results

Descriptive Information About the Articles

There were 126 articles in all of the issues published during 2013 of the four Canadian journals. Articles that were excluded were 25 editorials, introductions to a special issue, book or test reviews, commentaries (on other papers or conference activities), 31 theory or review papers, and two papers that contained a simulation study (whereby the majority of the coding did not apply). This left 68 articles with an empirical study. Of these, 63 (92.7%) included at least one quantitative analysis, whereas 5 (7.3%) included a qualitative analysis. Most of these articles were written in English (N = 59, 86.7%) and 9 were written in French (13.3%). As the focus was on quantitative reporting practices, the rest of the paper will not include information from the studies that employed qualitative methods.

Prevalence of Statistical Analyses and Inferential Procedures

Although there were 63 studies with quantitative information, most articles had several unique statistical analyses, such that there were 151 analyses investigated in the current study. An analysis was considered unique if it was used to answer a question of substantive interest, was not used as a manipulation check or to equate groups based on demographic information, and was not used to supplement another analysis (e.g., presenting a correlation matrix when the main analysis is a multiple regression model). Table 1 presents a breakdown of the types of statistical analyses of which the 151 procedures were comprised.

Table 1.

Prevalence of Statistical Analyses

Analysis N %
ANOVA 38 25.1
Z or t test on means 23 15.2
Multiple regression 21 13.9
Correlation 15 9.9
Chi square 15 9.9
Structural Equation Models 8 5.3
Logistic Regression 4 2.6
Factor Analysis (or Principle Components Analysis) 4 2.6
Descriptive Only 4 2.6
ANCOVA 3 2.0
Multilevel/Mixed Effects Models 3 2.0
Generalized linear models 3 2.0
MANOVA 2 1.3
Mann-Whitney U test 2 1.3
Z test on dependent correlations 2 1.3
Meta-analysis 1 0.7
Discriminant function analysis 1 0.7
Robust canonical correlation 1 0.7
MANCOVA 1 0.7
Total 151 100.0

Note. ANOVA = analysis of variance; ANCOVA = analysis of covariance; MANOVA = multivariate analysis of variance; MACOVA = multivariate analysis of covariance.

From Table 1, one can see that the most popular methods were ANOVA and z or t-tests. In fact, 40% of the analyses included a univariate mean comparison. Tests of univariate mean comparisons were highly representative of the articles in the CJEP. Analyses that examined associations amongst variables (multiple regression, correlation, and chi square) were also frequently used in the four journals (34% of the analyses used one of these three techniques). Multivariate and modeling techniques tended to be used less frequently and with a wide range of techniques employed (e.g., structural equation modeling, logistic regression, mixed effects models, and generalized linear models). Four of the analyses (2.6%) included only descriptive statistics (means, odds ratios, etc.) to answer their research question.

Types of Quantitative Information Reported

Whereas the prevalence of statistical methods is informative, investigating the types of statistical inference information presented from such analyses will provide information about reporting practices and areas for improvement. This information is presented in Table 2.

Table 2.

Content Reported from Statistical Analyses

Inference Information N %
Reported significance test 138 91.4
    Includes effect size 128 92.7
    Reported dichotomous p values 73 *52.9
    Reported exact p values 71 *51.4
    Includes standard error 34 24.6
Effect size reported with no sig. test 12 7.9
Reported confidence interval on effect 16 10.6
Includes figure with data 47 31.1
    Figure has error bar 25 53.2
Includes information on missing data 46 30.5
Includes information on at least one statistical assumption 44 29.1
Total Analyses 151

Note: If information is indented, the percentage refers to the parent category and not the total number of analyses. For example, of the 47 analyses that included a figure, 25 included an error bar such that 53.2% included it, whereas only 16.6% of all analyses included a figure with an error bar.

*

These do not add up to 100% because sometimes researchers reported both dichotomous and exact p values within the same analysis.

Significance tests and effect sizes

Almost all of the articles presented significance tests with their analysis (91.4%), and these included inconsistencies with their reporting of p values. Dichotomous and exact p values were reported with almost equal frequency across the articles surveyed, although in many instances a researcher would report both dichotomous and exact p values within the same analysis. For example, when reporting the results of a multiple regression analysis, a researcher may have presented the exact p value for the overall model's significance test (e.g., p = .023), but then reported p < .05 from a predictor variable's hypothesis test. For this reason the percentage of significance tests that included dichotomous or exact p values does not sum to 100% in Table 2. Few analyses (24.6%) reported the standard error associated with their test statistic and p value. One author did not report any statistical information from their analysis because it was not statistically significant.

The majority of articles presented an effect size. In fact, of the 138 analyses that reported a significance test, 128 (92.7%) included an effect size. Twelve analyses included an effect size without a hypothesis test, but only one of these included a CI on the effect. In fact, CIs were rarely employed regardless of whether a significance test was used since only 16 (10.6%) of all analyses included them. Of the 140 analyses that included an effect size (i.e., regardless of whether a significance test was used), 40% reported only unstandardized effects (e.g., raw means, medians, unstandardized regression weights, odds ratios), 27% reported only standardized effects (e.g., standardized regression weights or factor loadings, η2, R2, RMSEA), and 33% reported both unstandardized and standardized effect sizes.

Figures

Forty-seven (31.1%) analyses included a visual representation of the data alongside their statistics. Of these 47, only 25 (53.2%) of them included an indication of variability such as error bars on the CI or standard error. In general the plots tended to be simple bar charts presenting a small number of group means.

Missing data and statistical assumptions

Only 30.5% of the analyses included explicit information about how much missing data was present and how the researcher dealt with this issue. That being said, examining the degrees of freedom from the analyses often allowed us to determine whether there was any missing data and if so, whether pairwise or listwise deletion was used. In fact, it appeared as though many of the articles in the CJEP had complete cases. If information was presented about missing data, few articles reported a missing data strategy other than a simple deletion technique or mean substitution.

The number of analyses that included information about statistical assumptions was similar to those reporting on missing data (29.1%). While just under one third of analyses included some information about statistical assumptions, only two analyses included information about whether all of their statistical assumptions were met. The other 42 included limited information and typically only addressed one of their statistical assumptions (e.g., were the data normally distributed when conducting an ANOVA?). In some cases, authors attempted to address a statistical assumption, but did so incorrectly. One example of this is where a researcher failed to examine the normality of the regression residuals and instead examined the distributions of the independent and dependent variables.

Appropriateness Ratings of Statistical Procedures

We initially sought to provide information about whether authors implemented the most appropriate statistical analysis based on their research design and sample. The challenge was that few articles presented enough information in the results section to adequately assess whether their statistical choice was appropriate. The lack of transparency around statistical assumptions was one of the biggest issues for assessing whether authors used appropriate statistical methods. Almost all of the researchers chose statistical tools that adequately complemented their research design, but without information on statistical assumptions, it is impossible to provide reliable validation for an author's choice of statistical test. Thus, we decided against presenting our appropriateness ratings here; if interested, readers may request this information from the first author.

Discussion

The current research examined the methodological trends of psychology articles in the four major Canadian journals in psychology during 2013. After examining the prevalence of qualitative and quantitative methods in empirical articles, we investigated in detail the statistical methods used in the articles and information researchers reported alongside their analyses. Getting a view of the landscape in these articles offered greater awareness of which analyses are currently being used in the Canadian literature and how this information is reported. Investigating this type of information provided insight into what researchers are doing well and what could be done to improve the nature of inference in future studies.

Before discussing the specific statistical tools used in the articles it is worth discussing how authors overwhelming use quantitative approaches in empirical studies. Few articles included any qualitative data or information. While Gergen, Josselson, and Freeman (2015) argue that qualitative information allows for a more pluralistic or holistic view of individuals, we believe that both quantitative and qualitative methods have their own merits. Including a mixed methods approach with both quantitative and qualitative tools may provide a richer account of a particular phenomenon.

Quantitative Methods Used in Empirical Articles

Of the articles under investigation, the types of statistical methods used were largely univariate in nature. In fact, the majority of the papers included methods taught at an undergraduate level (i.e., t-tests, ANOVA, chi square, correlation and multiple regression). However, they also represented popularities in certain fields. For example, experimental articles (mostly in the CJEP) overwhelming used ANOVA. Observational studies typically included correlation or multiple regression analyses. These findings contrast with those from Harlow, Korendijk, Hamaker, Hox, and Duerr (2013) who examined the extent of multivariate methods and statistical inference procedures used in eight European psychology journals. Their study found that 57% of the articles used multivariate methods. Whereas parsimony is important and few articles in the current study used a more complicated model when a simpler one would suffice, some articles would have benefited from incorporating their research hypotheses into a larger multivariate model. In general, researchers were much more likely to conduct several univariate models than to include a multivariate analysis. For example, instead of running several multiple regression models, a researcher could use a path analysis model or structural equation model. Multivariate models also allow for a focus on a more cohesive, integrated understanding of the nature of the data with respect to the research questions asked. The challenge is that multivariate models tend to require larger sample sizes and the median sample size in the empirical articles surveyed was only 89 (but ranged from N = 5 to 44, 560). The choice to include several univariate models instead of one larger multivariate model, however, remains dependent on the nature of one's research questions.

Hypothesis Testing and Effect Sizes

Despite calls for reducing reliance on NHST, the majority of the articles surveyed used significance tests. However, the constant calls for reporting effect sizes appears to have had an effect on the Canadian psychology articles as just over 90% of the analyses that used a significance test also included a standardized or unstandardized effect size. Few articles presented an effect size without hypothesis testing, and few of the analyses’ results included a confidence interval. In fact, CIs were not typically reported as a supplement for NHST, nor were they included on effect sizes without statistical significance tests. In general, the articles that did not use significance tests tended to be descriptive studies or present survey results.

In coding whether an analysis included an effect size or not, we adopted a broad framework for effect sizes such that both unstandardized and standardized measures were included. Information about the prevalence of each was included at the request of an anonymous reviewer. We adopted this broader definition because the goal is to provide readers with the magnitude of an effect, such that readers can see the practical significance of the findings. This can be achieved in a number of different ways. As stated by the Task Force on Statistical Inference, “if the units of measurement are meaningful on a practical level (e.g., number of cigarettes smoked per day), then we usually prefer an unstandardized measure (regression coefficient or mean difference) to a standardized measure (r or d)” (Wilkinson et al., 1999 p. 599). However, they go on to describe how it is important to include comments that place these effect sizes within a relevant theoretical context. We noticed that this is an area requiring improvement. Researchers are becoming more aware of the importance of presenting effect sizes, but they are not discussing them further or situating them within the larger body of literature.

Visual Displays of Data

Data visualization can be an incredibly useful tool for presenting statistical information. This study demonstrated that high quality informative graphics were not being utilized in the majority of these research articles. Specifically, less than one third of the analyses included a graphical representation of the data and only half of these included a measure of variability such as an error bar on the CI or standard error. For the figures that were included, many of them were unnecessary, presenting simple bar charts plotting means from t-tests or one-way ANOVAs. Given that in this particular investigation of psychology articles, the majority of analyses used univariate mean comparisons or simple correlation, presenting complex figures may not be as crucial. However, better visualization methods could be used. For example, it may be helpful to include boxplots instead of bar charts, since boxplots include information about distribution shape, central tendency, variability, and outliers.

General Transparency and Detailing Important Statistical Information

In general, the articles included a great amount of detail in the methods section, which allows other researchers to attempt replication, but the information provided in the results section could be more comprehensive. The majority of articles in the study included the information required by the publication manual such as test statistic, df, p value, effect size, but few articles presented their data and analyses in sufficient detail so that a reader could justify the authors’ conclusions. For example, it was common for researchers to say that they conducted an ANOVA or F-test, without specifying which type. This term could refer to between subjects, within subjects, mixed effects, factorial, etc. If not explicitly stated, we used the model degrees of freedom and information from the design in the methods section to identify the type of ANOVA used. Having to identify an ANOVA by degrees of freedom is particularly problematic when they may have been adjusted due to missing data or robust alternatives (e.g., Greenhouse-Geisser epsilon), or may involve a typographical error. This is a simple detail that would highly improve the clarity of one's statistical test for readers.

A related issue was the lack of information about how missing data and statistical assumptions were addressed. This poses a real problem for validating statistical decisions and was the biggest issue in trying to assess the appropriateness of a researcher's statistical test. If researchers are not testing for statistical assumptions, their choice of statistical test is likely problematic since statistical assumptions in psychological research are frequently violated in practice (Blanca, Arnau, Lopez-Montiel, Bono, & Bendayan, 2011; Keselman et al., 1998, Micceri, 1989). Using traditional parametric statistical tests with assumption violation has implications such as higher Type I or Type II error rates depending on the nature of the violations (Coombs, Algina, & Oltman, 1996; Cribbie, Fiksenbaum, Wilcox, & Keselman, 2012; Glass et al., 1972; Lix, Keselman, & Keselman, 1996). As researchers’ conclusions, implications, and suggestions for future directions are all based on the results assuming valid parametric procedures, when statistical assumptions are violated but not addressed, one runs the risk of presenting useless, misleading, or potentially harmful results. Missing data poses a similar problem. We acknowledge, however, that strategies for dealing with missing data are numerous and may be complicated, as the method adopted should depend on a number of factors such as the research design and reason for the missing data. While this is a topic beyond the scope of this discussion, we simply advocate for explicating the amount (or lack) of missing cases and strategy used for clarity and justification.

Why Is This Important Information Missing?

We do not think that researchers are hiding data issues, but instead that applied researchers, reviewers, and editors do not immediately realize the importance of such information for critical evaluation of the work. Recent research by Hoekstra, Kiers, and Johnson (2012) suggests that researchers have limited knowledge about the robustness of parametric tests and how and why they should examine statistical assumptions. With limited space in journals, specific details of statistical tests may be lost; researchers focus their page space on the discussion and conclusions — what their results actually mean and why others should care. Further, researchers report what is required by the journal. If editors and reviewers do not require certain types of statistical information, it is unlikely to be reported since that page space can be used elsewhere.

Another issue is that the importance of quantitative skills and training tends not to be emphasized enough in both undergraduate and graduate level training. In fact, on average, graduate students in Canada are only required to take two statistics courses by the end of their doctoral degree (Counsell, Cribbie, & Harlow, 2016). Given this limited training, many applied researchers rely on the information presented from software. For example, if software does not report a CI on Cohen's d, it is unlikely that a researcher will calculate one his or herself. Recommendations by the publication manual and by journals may not be strong enough. Instead researchers should be required to discuss these types of information in their articles so that others have the necessary material to validate the quantitative decisions in publications.

Reporting Practices for a Better Science

The current study contributes to the field by providing a recent snapshot of the state of Canadian psychology articles with regards to statistical methods and inferential procedures used. It is important to monitor and report on practices and trends of a discipline to capitalize on strengths and address limitations. Some of the issues that arose in this study belong to a larger group of issues that need to be addressed. Transparency in reporting and research practices (e.g., Nosek, et al., 2012), replication (e.g., Anderson & Maxwell, 2016; Open Science Collaboration, 2015), and using valid and reliable methods and instruments are a few issues. Along with the APA publication manual and Task Force guidelines and recommendations, other researchers have published recommendations. For example, Funder and colleagues (2014) have outlined a number of recommendations put forth by the Society of Personality and Social Psychology Task force on Publication and Research Practices. Nuijten and colleagues (2015) examined articles from eight major journals from 1985 to 2013 and found a large percentage of reporting errors and include some recommendations for researchers in an effort to improve the dependability of psychological research. Cousineau (2014) discusses the importance and need for replication studies so that the field can work towards building a body of scientific knowledge rather than simply publishing articles with an acceptable p value. Recommendations are helpful but in order for the reporting practices of researchers to improve, journals must insist on reporting certain types of information. As Funder et al. (2014) note, “to make our field more amenable to these [recommended] practices, it is important for all of us, including editors, reviewers, and those who make hiring/promotion decisions, to educate ourselves about their value” (p. 9).

Limitations and Future Directions

Before making recommendations, it is important to note that the current paper has a few limitations. One limitation was the narrow scope of the journals examined. Articles from 2013 may not necessarily be representative of the publications in Canadian journals in psychology. That being said, we sought to have a narrow scope to allow for in depth investigation of the quantitative information included. Furthermore, our results are consistent with previous literature on reporting practices (e.g., Hoekstra et al., 2012; Reinhart, 2015). As previously noted, ongoing research on reporting practices and methods remains important to monitor and improve upon our discipline. A second limitation of our paper and future direction of this research is to examine another potential data complication of outlier detection and removal. Although we did not investigate the prevalence of reporting on outliers, it is both relevant and important for researchers to discuss, just as we recommend doing for missing data and statistical assumptions. Future studies investigating quantitative reporting should be encouraged to examine whether and how articles report outliers in their data.

Recommendations and Conclusion

Reporting practices have come a long way in psychology. The changes from each edition of the APA publication manual highlight this progress, and journals now require more information from authors than any previous year. That being said, we believe more can be done. For this reason, we include a list of recommendations driven by the current study's results that will benefit both applied researchers and the reviewers and editors of journals. They are as follows:

  1. Think about conceptualizing a larger model or using a multivariate method instead of running several univariate analyses. Larger models are not always necessary or feasible, but at very least, researchers should consider whether their hypotheses could be answered by one model instead of running multiple smaller analyses.

  2. Be explicit about which statistical test you have conducted. This can be as simple as stating that you conducted a paired samples t-test as opposed to reporting conducting a “t-test” or “mean comparison.” Specifying the type of ANOVA (e.g., between groups factorial) or type of regression analysis (linear multiple regression) would improve the reader's comprehension and higher potential for reproducing results.

  3. Report on the amount of missing data present and how you dealt with it in your analyses. In cases with a lot of missing data, or where data are not missing completely at random, consider strategies other than deletion methods or mean substitution (see Graham, 2009) to prevent biasing results. It is important to consider the research design and possible reasons for missing data as well.

  4. Present information about whether statistical assumptions were met. If they were not, how were the violations addressed and were there other unanticipated data complications?

  5. Present the data graphically if visualization allows for readers to better see trends and patterns, but do not include graphics that are redundant or unhelpful (e.g., simple bar charts with two or three groups). We would always recommend presenting a figure for factorial ANOVA models that include an interaction; this allows readers to see the nature of the interaction, as this cannot be easily evaluated from information provided by significance tests alone.

  6. Always include some type of effect size and its associated confidence interval. This can be in the form of unstandardized units such as mean differences, or standardized units such as Cohen's d. Point estimates of effect size provide readers with important information, but including the variability on the effect (e.g., CI) provides more information. Effect sizes should also be explored in the paper's discussion section, as it is more informative than simply reporting whether a finding was statistically significant or not.

Overall, we found it encouraging that most of the articles we examined from these four Canadian journals in psychology reported effect sizes, along with information about statistical tests and associated p values. Researchers are encouraged to also include confidence intervals to highlight the degree of uncertainty around their effect sizes. Although not all computer programs provide CI information, online sources are available (e.g., Soper, 2006-2016). It would also be helpful to provide more information on missing data and assumptions to allow for more accurate assessment on the adequacy of the study and its findings. Whereas we hope that our suggested recommendations can help researchers incorporate better reporting practices into their papers, we also call on statistical educators and quantitative methodologists to provide support and guidance regarding these issues. We are further aware that these implementations will take time and depend on reinforcement from journal editors and research associations such as the CPA and APA.

References

  1. Abelson RP. The surprising longevity of flogged horses: Why there is a case for the significance test? Psychological Science. 1997;8:12–15. doi: 10.1111/j.1467-9280.1997.tb00536.x. [Google Scholar]
  2. Allison PD. Missing data. Sage Publications; Thousand Oaks, CA: 2002. [Google Scholar]
  3. American Psychological Association [APA] Publication manual of the American Psychological Association. 6th ed. APA.; Washington, DC: 2010. [Google Scholar]
  4. Anderson S, Maxwell S. There's more than one way to conduct a replication study: Beyond statistical significance. Psychological Methods. 2016;21:1–12. doi: 10.1037/met0000051. doi: 10.1037/met0000051. [DOI] [PubMed] [Google Scholar]
  5. Baraldi AN, Enders C. Missing data methods. In: Little's TD, editor. The Oxford Handbook of Quantitative Methods in Psychology: Vol. 2: Statistical Analysis. Oxford; New York, NY: 2013. pp. 635–664. [Google Scholar]
  6. Blanca MJ, Arnau J, Lopez-Montiel D, Bono R, Bendayan R. Skewness and kurtosis in real data samples. Methodology. 2011;92:78–84. doi: 10.1027/1614-2241/a000057. [Google Scholar]
  7. Bodner TE. Missing data: Prevalence and reporting practices. Psychological Reports. 2006;99:675–680. doi: 10.2466/PR0.99.3.675-680. doi: 10.2466/PR0.99.3.675-680. [DOI] [PubMed] [Google Scholar]
  8. Bradley JV. Robustness? British Journal of Mathematical and Statistical Psychology. 1978;31:144–152. doi: 10.1111/j.2044-8317.1978.tb00581.x. [Google Scholar]
  9. Chow SL. Statistical significance: Rationale, validity and utility. Sage; Beverly Hills, CA: 1996. [DOI] [PubMed] [Google Scholar]
  10. Cleveland WS. Visualizing data. Hobart Press; Summit, NJ: 1993. [Google Scholar]
  11. Cohen J. The earth is round (p < .05). American Psychologist. 1994;49:997–1003. doi: 10.1037/0003-066X.49.12.997. [Google Scholar]
  12. Coombs WT, Algina J, Oltman DO. Univariate and multivariate omnibus hypothesis tests selected to control type I error rates when population variances are not necessarily equal. Review of Educational Research. 1996;66:137–179. doi: 10.3102/00346543066002137. [Google Scholar]
  13. Counsell A, Cribbie RA, Harlow LL. Increasing literacy in quantitative methods: The key to the future of Canadian psychology. Canadian Psychology. 2016;57:193–201. doi: 10.1037/cap0000056. [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Cousineau D. Restoring confidence in psychological findings: A call for direct replication studies. The Quantitative Methods for Psychology. 2014;10:77–79. [Google Scholar]
  15. Cribbie RA, Fiksenbaum L, Wilcox RR, Keselman HJ. Effects of nonnormality on test statistics for one-way independent groups designs. British Journal of Mathematical and Statistical Psychology. 2012;65:56–73. doi: 10.1111/j.2044-8317.2011.02014.x. doi: 10.1111/j.2044-8317.2011.02014.x. [DOI] [PubMed] [Google Scholar]
  16. Cumming G. Replication and p intervals: p values predict the future only vaguely, but confidence intervals do much better. Perspectives on Psychological Science. 2008;3:286–300. doi: 10.1111/j.1745-6924.2008.00079.x. doi: 10.1111/j.1745-6924.2008.00079.x. [DOI] [PubMed] [Google Scholar]
  17. Cumming G. Understanding the new statistics: Effect sizes, confidence intervals, and meta-analysis. Routledge; New York, NY: 2012. [Google Scholar]
  18. Denis D. Alternatives to null hypothesis significance testing. Theory & Science. 2003;4:1–21. Available at: http://theoryandscience.icaap.org/content/vol4.1/02_denis.html. [Google Scholar]
  19. Enders CK. Applied missing data analysis. Guilford Press; New York: 2010. [Google Scholar]
  20. Friendly M. Visualizing categorical data. SAS Institute, Inc.; Cary, NC: 2000. doi: 10.1002/jhbs.20078. [Google Scholar]
  21. Friendly M, Meyer D. Discrete data analysis with R: Visualization and modeling techniques for categorical and count data. CRC Press; Boca Raton, FL: 2015. [Google Scholar]
  22. Funder DC, Levine JM, Mackie DM, Morf CC, Sansone C, Vazire S, West SG. Improving the dependability of research in personality and social psychology: Recommendations for research and educational practice. Personality and Social Psychology Review. 2014;18:3–12. doi: 10.1177/1088868313507536. doi:10.1177/1088868313507536. [DOI] [PubMed] [Google Scholar]
  23. Gergen KJ, Josselson R, Freeman M. The promises of qualitative inquiry. American Psychologist. 2015;70:1–9. doi: 10.1037/a0038597. doi: 10.1037/a0038597. [DOI] [PubMed] [Google Scholar]
  24. Glass GV, Peckham PD, Sanders JR. Consequences of failure to meet assumptions underlying the fixed-effects analysis of variance and covariance. Review of Educational Research. 1972;42:237–288. [Google Scholar]
  25. Graham JW. Missing data analysis: Making it work in the real world. Annual Review of Psychology. 2009;60:549–576. doi: 10.1146/annurev.psych.58.110405.085530. doi: 10.1146/annurev.psych.58.110405.085530. [DOI] [PubMed] [Google Scholar]
  26. Hagen RL. In praise of the null hypothesis statistical test. American Psychologist. 1997;52:15–24. doi: 10.1037/0003-066X.52.1.15. [Google Scholar]
  27. Harlow LL. On scientific research: The role of statistical modeling and hypothesis testing. Journal of Modern Applied Statistical Methods. 2010;9:348–358. Available at: http://digitalcommons.wayne.edu/jmasm/vol9/iss2/4. [Google Scholar]
  28. Harlow LL, Korendijk E, Hamaker EL, Hox J, Duerr SR. A meta-view of multivariate statistical inference methods in European psychology journals. Multivariate Behavioral Research. 2013;48:749–774. doi: 10.1080/00273171.2013.822784. doi: 10.1080/00273171.2013.822784. [DOI] [PubMed] [Google Scholar]
  29. Harlow LL, Mulaik SA, Steiger JH, editors. What if there were no significance tests? Classic Edition Routledge; New York, NY: 2016. [Google Scholar]
  30. Harris RJ. Significance tests have their place. Psychological Science. 1997;8:8–11. doi: 10.1111/j.1467-9280.1997.tb00535.x. [Google Scholar]
  31. Hoekstra R, Kiers HAL, Johnson A. Are assumptions of well-known statistical techniques checked, and why (not)? Frontiers in Psychology. 2012;3:1–9. doi: 10.3389/fpsyg.2012.00137. doi: 10.3389/fpsyg.2012.00137. [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Keselman HJ, Huberty CJ, Lix LM, Olejnik S, Cribbie R, Donahue B, Levin JR. Statistical practices of educational researchers: An analysis of their ANOVA, MANOVA, and ANCOVA analyses. Review of Educational Research. 1998;68:350–386. doi: 10.3102/00346543068003350. [Google Scholar]
  33. Kline R. Beyond significance testing: Reforming data analysis methods in behavioral research. 2nd ed. American Psychological Association; Washington, DC: 2013. [Google Scholar]
  34. Little RJA, Rubin DB. Statistical analysis with missing data. 2nd ed. Wiley; New York, NY: 2002. [Google Scholar]
  35. Lix LM, Keselman JC, Keselman HJ. Consequences of assumption violations revisited: A quantitative review of alternatives to the one-way analysis of variance F test. Review of Educational Research. 1996;66:579–619. doi: 10.3102/00346543066004579. [Google Scholar]
  36. Margolis E, Pauwels L, editors. The SAGE handbook of visual research methods. SAGE Publications; Thousand Oaks, CA: 2011. [Google Scholar]
  37. Micceri T. The unicorn, the normal curve, and other improbable creatures. Psychological Bulletin. 1989;105:156–166. doi: 10.1037/0033-2909.105.1.156. [Google Scholar]
  38. Mulaik SA, Raju NS, Harshman RA. There is a time and place for significance testing. In: Harlow LL, Mulaik SA, Steiger JH, editors. What if there were no significance tests. Erlbaum; Mahwah, NJ: 1997. pp. 65–115. [Google Scholar]
  39. Nosek BA, Spies JR, Motyl M. Scientific utopia II. Restructuring incentives and practices to promote truth over publishability. Perspectives on Psychological Science. 2012;7:615–631. doi: 10.1177/1745691612459058. doi: 10.1177/1745691612459058. [DOI] [PMC free article] [PubMed] [Google Scholar]
  40. Nuijten MB, Hartgerink CHJ, Marcel ALM, van Assen MALM, Epskamp S, Wicherts JM. The prevalence of statistical reporting errors in psychology (1985–2013). Behavior Research Methods. 2015 Oct; doi: 10.3758/s13428-015-0664-2. online. doi: 10.3758/s13428-015-0664-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  41. Open Science Collaboration Estimating the reproducibility of psychological science. Science. 2015;349(6251) doi: 10.1126/science.aac4716. http://dx.doi.org/10.1126/science.aac4716. [DOI] [PubMed] [Google Scholar]
  42. Reinhart A. Statistics Done Wrong: The Woefully Complete Guide. No Starch Press; 2015. [Google Scholar]
  43. Rodgers JL. The epistemology of mathematical and statistical modeling: A quiet methodological revolution. American Psychologist. 2010;65:1–12. doi: 10.1037/a0018326. doi: 10.1037/a0018326. [DOI] [PubMed] [Google Scholar]
  44. Rozeboom WW. Good science is abductive, not hypothetico-deductive. In: Harlow LL, Mulaik SA, Steiger JH, editors. What if there were no significance testing. Erlbaum; Mahwah, NJ: 1997. pp. 335–391. [Google Scholar]
  45. Schmidt FL, Hunter JE. Eight common but false objections to the discontinuation of significance testing in the analysis of research data. In: Harlow LL, Mulaik SA, Steiger JH, editors. What if there were no significance testing? Erlbaum; Mahwah, NJ: 1997. pp. 37–64. [Google Scholar]
  46. Soper D. [March 28, 2016];Confidence interval calculators. 2006-2016 ( http://www.danielsoper.com/statcalc/category.aspx?id=4)
  47. Thompson B. Effect sizes, confidence intervals, and confidence intervals for effect sizes. Psychology in the Schools. 2007;44:423–432. doi: 10.1002/pits.20234. [Google Scholar]
  48. Tukey JW. Exploratory data analysis. Addison-Wesley; Reading, MA: 1977. [Google Scholar]
  49. Wilkinson L, The APA Task Force on Statistical Inference Statistical methods in psychology journals guidelines and explanations. American Psychologist. 1999;54:594–604. doi: 10.1037/0003-066X.54.8.594. [Google Scholar]
  50. Wood AM, White IR, Thompson SG. Are missing outcome data adequately handled? A review of published randomized controlled trials in major medical journals. Clinical Trials Review. 2004;1:368–376. doi: 10.1191/1740774504cn032oa. doi: 10.1191/1740774504cn032oa. [DOI] [PubMed] [Google Scholar]

RESOURCES