Abstract
Objectives
Most pharmaceutical investigations have relied on p values to infer conclusions from their study findings. Central to this paradigm is the concept of null hypothesis significance testing. This approach is however fraught with overuse and misinterpretations. Several alternatives have already been proposed, yet uptake remains low. In this study, we aimed to discuss the pitfalls of p value-based testing and to provide readers with the basics to apply Bayesian statistics.
Methods
Jeffreys’s Amazing Statistical Package (JASP) was used to evaluate the effect of a clinical pharmacy (CP) intervention (opposed to usual care) on the number of emergency department (ED) visits without hospital admission. Basic Bayesian terminology was explained and compared with classical p value-based testing. In the study example, a Cauchy prior distribution was used to determine the effect size with a scale parameter r=0.707 at location=0 and Bayes factors (BF) were subsequently estimated. A robustness analysis was then performed to visualise the impact of different r values on the BF value.
Results
A BF of 4.082 was determined, indicating that the observed data were about four times more likely to occur under the alternative hypothesis that the CP intervention was effective. The median effect size of the CP intervention on ED visits was found to be 0.337 with a 95% credible interval of 0.074 to 0.635. A robustness check was performed and all BF values were in favour of the CP intervention.
Conclusion
Bayesian inference can be an important addition to the statistical armamentarium of pharmacists, who should become more acquainted with the basic terminology and rationale of such testing. To prove our point, Jeffreys’ approach was applied to a CP study example, using an easy-to-use software program JASP.
Keywords: Jeffreys, JASP, statistical analysis, clinical pharmacy, older inpatients, Bayes factor
Introduction
Generating evidence with null hypothesis significance testing
A majority of controlled investigations has relied on null hypothesis significance testing (NHST) to determine whether a researched intervention was effective in improving a predefined primary outcome measure. Owing to the large uptake and success of NHST, it is now accepted that clinical pharmacists can play an important role in reducing morbidity by identifying and managing drug-related problems.1–3 A certain level of evidence in support of clinical pharmacy (CP) services can be gathered from studies in which multifaceted interventions were performed in older medical inpatients.4 In that specific setting, several investigations have already shown a reduction of unplanned readmissions.4
The conceptual framework of NHST uses a numerical value, the p value, to reject a predefined null hypothesis (H0). In this view, the null hypothesis is considered to be the default position of there not being any association between several phenomena (eg, that a certain intervention will not impact the number of readmissions). This conceptual framework draws from both Fisher’s use of the p value and Neyman’s and Pearson’s concept of hypothesis testing.5
Central to the p value-driven paradigm is that a statement regarding H0 is provided indirectly, as the p value is the probability of finding data at least as extreme as those observed, assuming that H0 is true.6 The p value therefore does not allow to infer the probability of H0 itself being true; it does not allow to ‘accept’ the null (or an alternative) hypothesis. It merely captures the probability of finding the observed data or extremer given a multitude of assumptions, including but not limited to H0 being true. Importantly, these assumptions also include correct patient enrolment, follow-up and analysis.6 A low p value might then be explained by a deviation of any of these assumptions.5 7–9
Limitations of NHST and p value-based testing
The use of p values has come under scrutiny the past decades, mainly due to overuse and misinterpretations, even leading to a replication crisis in some instances.5 6 10 For example, this includes the frequently heard statement that the actual p value is equal to the probability that the intervention is true; it is not, this probability is better captured by other measures, such as the posterior odds (reconverted to a probability) provided by Bayesian analysis.5 7 Also, the arbitrary use of a significance level of 0.05 (or 5%) has been increasingly met with criticism.6 10
Moreover, readers might interpret that another metric, frequently used in NHST, the 95% CI conveys a 95% probability that this specific interval contains the real population value; again, this is not the case. This would rather imply that the credibility—and not the confidence—interval be determined.6 Also here, a Bayesian approach should be used to estimate the credibility interval. Over-reliance on p values has clouded the interpretation of study findings the past decades.9
In sum, the p value (as part of NHST) can be viewed as p(data at least as extreme as those observed|null hypothesis), meaning the probability of data given a certain hypothesis. Importantly, NHST does not include any direct information on the alternative hypothesis (H+). Also, neither prior nor posterior plausibility of H0 and H+ are included in the NHST paradigm.5 7–9
Bayesian alternative
Other approaches have already been proposed to fix the way researchers should deal with statistics and also how we could define the concept of probability.10 In this paper, we will focus on Bayesian analysis, which is one of the proposed alternatives to NHST.
Bayesian analysis can be summarised by the following equation, also known as the Bayes’ rule (in its odds form): ‘prior odds times Bayes factor (BF)=posterior odds’, for example, ‘prior odds of H+times BF=posterior odds of H+’.7 9 11 The prior odds are determined by the researcher’s baseline beliefs or knowledge. The BF captures the relative change (or update) of the baseline evidence about hypotheses (H0 or H+), based on the observed study results.7 In other words, it enables researchers to infer how probable (or plausible) H0 or H+ both are after having observed the study data, which is a more intuitive approach pf interpreting study results.7 9 12 Bayesian inference aims to provide information on the plausibility (or probabilities) of research hypotheses. For additional background information on Bayesian statistics, we wish to refer to the comprehensive and introductory review of Etz et al.12 They provided an annotated reading list for interested readers.
Several packages have been developed in the open-source statistical package R. Yet, not all healthcare providers have the time needed to become acquainted, let alone become proficient in the use of R or the Bayesian approach. Hence, we promote the use of easy-to-use software such as Jeffreys’s Amazing Statistical Package (JASP).11 13
To the best of our knowledge, only two relevant CP reports have been published that have applied Bayesian statistics. First, Hemming et al used a Bayesian analysis in their re-analysis of the results of the PINCER trial.14 Second, Kawaguchi et al used a Bayesian approach to take into account prior preferences in their discrete choice experiment.15 The basic principles of Bayesian statistical analysis were only very briefly explained in both publications.
Objective
In this report, we hence aim to provide a pragmatic and easy-to-understand framework, worked out for CP researchers to evaluate the level of evidence that can be gathered from observed study results. To this end, a previously reported dataset was selected and JASP was used to perform a Bayesian analysis.9 11 16
Methods
Setting and design
This report concerns a practical guide on how to use Bayesian statistics and in that view, a post hoc statistical analysis was performed, using selected data retrieved from a previously published report on a CP intervention which took place on geriatric wards of University Hospitals of Leuven, Belgium.16
The study was designed as a monocentric, prospective, controlled study (ClinicalTrials.gov identifier: NCT01513265). Geriatric inpatients (n=172) were enrolled with few exclusion criteria and followed up for 3 months after discharge from the teaching hospital. The CP intervention consisted of a medication reconciliation and inpatient medication review.
Outcome measures
The study was powered to see an effect on the number of discontinued or tapered drugs from the pre-admission medication list. Several secondary outcome measure were evaluated. For the Bayesian example, we chose the outcome measure of the number of emergency department (ED) visits without a subsequent hospital admission, owing to its clinical relevance and use in comparable clinical investigations.2 3
Statistical analysis: G*Power and JASP
In the original study, sample size estimations were performed a priori with G*Power 3.1.17 In this post hoc analysis, G*power 3.1 was also used to visualise the original two-sided p value, being the tail-end integral of the H0 distribution based on the observed test statistic. In addition, the frequentist effect size according to Cohen’s d as well as the a posteriori power were determined to contrast these results to the Bayesian approach.
Jeffreys’s Amazing Statistical Package
JASP was used to evaluate the effect of the intervention (opposed to usual care) on the number of ED visits without hospital admission.11 13 JASP is a freely available software package (https://jasp-stats.org/) with an intuitive graphical user interface, not unlike other commercial and well-known statistical software packages such as the Statistical Package for the Social Sciences.11 JASP provides frequentist (classical) as well as Bayesian statistical inferential tests and is based on the R package Bayes Factor developed by Morey and Rouder.11 In this report, JASP was used to perform both classical NHST (p value based) and Bayesian analyses.
In Bayesian statistics, a prior probability distribution has to be defined which follows from the Bayes’ rule. Briefly, priors can be considered to be subjective or informative and are then based on the researcher’s knowledge or personal beliefs. On the other hand, objective priors can also be preferred, which reflect as few as possible assumptions about the prior distribution (of H0 or H+). In this analysis, a Bayesian t-test was selected in JASP. Based on the previous work of Morey et al, an objective Cauchy prior distribution was used for the effect size with a scale parameter r=0.707 at location=0, meaning that there was a 50% prior chance that the real CP effect size would fall between −0.707 and 0.707.11 This type of distribution favours H0 and hence renders more robust positive results (in terms of the BF).9 11 The Bayesian median effect size including the 95% credibility interval was estimated as well.9 11 BF values were determined for the alternative hypothesis that the CP intervention had an effect on ED visits and the null hypothesis of equipoise. BF were determined with H0 defined as equipoise between usual care and the CP intervention (ie, both have equal odds of being true or both have a 50% probability), and for H+ that the intervention was better than the usual care group, meaning that we expected more ED visits in usual care patients during follow-up. A one-side test was hence selected. The BF was presented as the change in probability that the CP intervention was indeed better than usual care in reducing ED visits in geriatric patients (ie, H+) and visualised using a pizza or pie chart. In other words, a higher BF value indicated more evidence in support of the CP intervention.18 In case of prior expected equipoise (ie, 50% probabilities allocated to H+ and H0), the BF equals the posterior odds.
Several r values have been proposed to account for the prior knowledge (or beliefs), concerning the intervention. A sensitivity, or a robustness, analysis, was performed to visualise the impact of different r values on the BF value. The user prior was the aforementioned Cauchy prior of r=0.707, the wide prior (=Jeffreys’ default prior) was defined as r=1.00 and the ultrawide prior as r=√2.11 The r value associated with the highest BF was determined as well.
Results
In total, follow-up data on ED visits were available for 166 out of 172 patients. Nearly all patients were octogenarian, multimorbid and had been admitted through the ED.
Fewer ED visits without admission occurred in the intervention group compared with the control group, with a mean difference of 0.077 (95% CI 0.009 to 0.145), which coincided with a p value of 0.0201. This result corresponded to a frequentist Cohen’s d of 0.36 with an a posteriori power (1–β) of approximately 60%, as depicted in figure 1.
Figure 1.
Probability density plots of the null and alternative hypotheses based on the observed study data and using a t distribution. A t distribution is shown with t values and the X-axis and P(X) on the Y-axis. The vertical lines signify critical t values. The shaded area is the β value, defined as the type II error, which is the probability of incorrectly not rejecting the null hypothesis. The two-sided α (ie, α/2) characterises the rejection region or the type I error, also the probability of incorrectly rejecting the null hypothesis.
Using JASP, descriptive parameters were determined and have been summarised in table 1. A BF of 4.082 was determined in support of H+, with an error% of ~1.818e−4; a reciprocal BF in favour of H0 was found to be 0.245. This was visualised, using a pie or pizza chart, depicting the proportional weighing of the evidence of both the H+ and H0 (figure 2). The posterior odds were converted to probabilities: 4.082/(1+4.082), signifying 80% of the evidence was in favour of H+ and 20% in favour of H0. Departing from a position of equipoise (ie, p(H0)=p(H+)=50%), the probability that CP intervention was effective in reducing ED visits hence increased from 50% to approximately 80%.
Table 1.
The number of emergency department (ED) visits without a subsequent hospital admission
| Group | N | Mean number of ED visits | SD | SE | 95% credible interval | |
| Lower | Upper | |||||
| Control | 79 | 0.089 | 0.286 | 0.032 | 0.025 | 0.153 |
| Intervention | 87 | 0.011 | 0.107 | 0.011 | 0.000 | 0.034 |
Figure 2.

The visualisation of the Bayesian output, based on the observed study results. BF+0 is the Bayes factor in support of H+ and BF0+ in favour of H0. The wheel contains the proportional weighing of the evidence for H+ and H0. The probability density plot below shows the prior and posterior probabilities and a visual representation of the Savage-Dickey density ratio (ie, the grey dots); the ratio of these heights equates to the BF for H+ vs H0.11
The Bayesian approach was visualised by showing how the prior distribution was updated by observing the data found in the study. Figure 2 shows this change from prior to posterior knowledge based on the observed study results, depicted by plotting the two probability density plots as a function of the effect size. The Bayesian median effect size of the CP intervention was found to be 0.337 with a 95% credible interval of 0.074 to 0.635.
A robustness check was performed and several BF values were derived, as a function of the different predefined priors. These results have been summarised in figure 3. All BF values pointed towards the found evidence being in favour of H+.
Figure 3.

Bayes factor (BF) robustness check. BF+0 is the BF in support of the alternative hypothesis. H+, the alternative hypothesis; H0, null hypothesis; r, scale parameter of the Cauchy distribution, ie, the Cauchy prior width, of the effect size which is used to describe the prior belief.
Discussion
This report is to the best of our knowledge the first in which a Bayesian framework is proposed for pharmacist researchers. A Bayesian example was worked out and JASP was used to evaluate the impact of a CP intervention in older inpatients. JASP is an attractive alternative to other well-known statistical software programs, which enables researchers to intuitively apply Bayesian inferential testing to their study findings.13
Study results from a previously published study were used as a statistical example. Using JASP, the impact on ED visits without subsequent hospital admission was assessed. Importantly, this study was neither designed nor powered to observe these findings.16 Furthermore, no correction (eg, Bonferroni correction) was made for multiple testing at the time, precluding any straightforward interpretation of the originally reported p values, which was also considered to be a limitation in the original manuscript. Conversely, multiple testing is not a major issue in Bayesian testing.9
In this one-sided Bayesian analysis, we found that the observed study results generated a BF of 4.082 in favour of the CP intervention. This should be interpreted as follows: the observed study results increased the probability that the CP intervention was indeed effective in terms of reducing ED visits in older adults, by a factor of 4.082. Bearing in mind the assumed equipoise between usual care and the CP intervention, it can equally be stated that about 80% of the posterior evidence was in support of the CP intervention and 20% in support of usual care. This value of 20% is not negligible however and reinforces the need to collect more data to increase the level of evidence in support of the CP intervention. An important additional benefit of the Bayesian approach is that previous data can be incorporated seamlessly in a follow-up analysis.9
When using different predefined priors in a robustness analysis, it was concluded that the evidence remained positive (and robust). Bayesian analysis has the advantage that priors can be defined, hence allowing for taking into account previous evidence. Conversely, sufficient effort should be invested in defining priors, as they also need to be defined. The BF values provided more relevant information on the CP intervention than the previously reported p value, both of which should preferentially be interpreted in a continuous (and not a dichotomous) manner.6 A median Cohen’s d of 0.36 was found, which should be viewed as a small to moderate effect size, yet for an important economic and clinical outcome measure. Importantly, the several components of the original intervention were completely restricted to the actual hospital stay and did hence not include systematic patient education or the provision of transitional or postdischarge care.16 This effect coincided with a p value of 0.0201, which was below the predefined yet arbitrary threshold for significance of 0.05. This tail-end integral (as depicted in figure 1) conveys a 2.01% probability of observing the observed difference or extremer (ie, larger than the results we found), assuming that the entire statistical model was correct, which includes the validity of H0. It could perhaps be viewed as (weak to moderate) evidence against H0 of there being no impact of CP services on ED visits. For completeness’ sake, this does not mean that there was 2.01% probability that our results were due to chance or that we inferred a 2.01% probability in favour of H0.6
The use of p value-based testing has seen a very wide uptake, in nearly all medical disciplines, and nearly all investigations have reported p values in order to reject a predefined (null) hypothesis. The large ‘success’ of p values can be explained by the current academic curricula that rely heavily on frequentist statistics. Another factor is its availability, as most statistical software programs with easily understandable graphical user interfaces have a certain proclivity towards p value-based testing.9 11 The major pitfall of NHST however is that it does not allow researchers to determine the amount of evidence in favour of their intervention. We hence applied Jeffreys’ approach to our study results and wish to promote a ‘double’ and easy-to-interpret approach, doing frequentist and Bayesian analyses side by side. This allows for a more clear interpretation of study results, and offers the advantage of re-imagining what it means to infer conclusions based on observed study results, which also can be seen in figure 2.
Other approaches than Bayesian analyses have already been proposed. Several experts have propagated among others the a priori publication of the statistical plan, the reporting of the false positive rate and lowering the thresholds of significance (eg, to p<0.001).10 This does not mean that all p values should be discarded. They remain unmistakably useful for showing statistical significance and decision making, but only when used appropriately.6 10 Even in that best-case scenario, it would still be very informative however to provide information on the probability of a research hypothesis given observed (study) data, that is, p(hypothesis|data). There currently remains no reason to not include Bayesian approaches in data analysis.13
The evidence base up to now for CP interventions in older inpatients has been low to moderate. Large multicentre randomised controlled trials have been rather the exception than the rule.4 Recently, several large-scale investigations have however been initiated, the first of which has recently been published by Ravn-Nielsen et al.2 Our findings on ED visits co-align with those from Ravn-Nielsen et al, and those from Gillespie et al, further supporting the evidence base, from a Bayesian viewpoint, that clinical pharmacists can safely improve patient trajectories after hospital discharge.2 3
In sum, it was inferred that the observed data increased the plausibility that a CP intervention was effective by a factor of 4.082, which equals a moderate evidence base. Future investigations should ascertain whether such interventions in older inpatients are robustly able to reduce all-cause ED visits and by extension all-cause unplanned readmissions. In conclusion, we promote the uptake of Bayesian approaches to draw conclusions from observed study results.11 12 Jeffreys’ approach was applied to previously published data, using the easy-to-use software JASP.13
Conclusion
Bayesian inference is an important addition to the statistical armamentarium of pharmacists, who should become more acquainted with the basic terminology and rationale of such testing. Jeffreys’ approach was applied to a CP study example, using the easy-to-use software JASP.
Key messages.
What is already known on this subject
P value-based testing is frequently overused to discard a null hypothesis (H0) as part of null hypothesis significance testing (NHST).
Importantly, in NHST there is no statement on the validity of the alternative hypothesis (H+), only whether or not H0 should be discarded.
What this study adds
Other valid approaches should be explored, such as Bayesian statistical analysis. Bayesian analysis has the potential drawback of being more labour-intensive, but has the benefit of providing more intuitive results.
There are currently easy-to-use software programs that can be readily used to perform Bayesian analysis (eg, Jeffreys’s Amazing Statistical Package).
Acknowledgments
The authors would like to thank EJ Wagenmakers for his crucial input on the final manuscript regarding Bayesian analysis and the Jeffreys’s Amazing Statistical Package output.
Footnotes
Contributors: LRVdL performed the analysis and provided the first draft. JH, KW, IS, JT and JF participated in the study concept and design and preparation of the final manuscript. All authors read and approved the final version of the manuscript.
Funding: This investigation (design and execution) was not funded. LRVdL received a partial research scholarship of the University Hospitals Leuven.
Competing interests: None declared.
Provenance and peer review: Not commissioned; internally peer reviewed.
Data availability statement
All data relevant to the study are included in the article or uploaded as supplementary information.
Ethics statements
Patient consent for publication
Not required.
Ethics approval
This study was approved by the hospital’s Ethics Committee.
References
- 1. Kaboli PJ, Hoth AB, McClimon BJ, et al. Clinical pharmacists and inpatient medical care: a systematic review. Arch Intern Med 2006;166:955–64. 10.1001/archinte.166.9.955 [DOI] [PubMed] [Google Scholar]
- 2. Ravn-Nielsen LV, Duckert ML, Lund ML, et al. Effect of an in-hospital multifaceted clinical pharmacist intervention on the risk of readmission: a randomized clinical trial. JAMA Intern Med 2018. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3. Gillespie U, Alassaad A, Henrohn D, et al. A comprehensive pharmacist intervention to reduce morbidity in patients 80 years or older: a randomized controlled trial. Arch Intern Med 2009;169:894–900. 10.1001/archinternmed.2009.71 [DOI] [PubMed] [Google Scholar]
- 4. Skjøt-Arkil H, Lundby C, Kjeldsen LJ, et al. Multifaceted pharmacist-led interventions in the hospital setting: a systematic review. Basic Clin Pharmacol Toxicol 2018;123:363–79. 10.1111/bcpt.13030 [DOI] [PubMed] [Google Scholar]
- 5. Goodman SN. Toward evidence-based medical statistics. 1: the P value fallacy. Ann Intern Med 1999;130:995–1004. 10.7326/0003-4819-130-12-199906150-00008 [DOI] [PubMed] [Google Scholar]
- 6. Greenland S, Senn SJ, Rothman KJ, et al. Statistical tests, P values, confidence intervals, and power: a guide to misinterpretations. Eur J Epidemiol 2016;31:337–50. 10.1007/s10654-016-0149-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. Goodman SN. Toward evidence-based medical statistics. 2: the Bayes factor. Ann Intern Med 1999;130:1005–13. 10.7326/0003-4819-130-12-199906150-00019 [DOI] [PubMed] [Google Scholar]
- 8. Gelman A, Shalizi CR. Philosophy and the practice of Bayesian statistics. Br J Math Stat Psychol 2013;66:8–38. 10.1111/j.2044-8317.2011.02037.x [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9. Wagenmakers E-J, Marsman M, Jamil T, et al. Bayesian inference for psychology. Part I: theoretical advantages and practical ramifications. Psychon Bull Rev 2018;25:35–57. 10.3758/s13423-017-1343-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10. Leek J, McShane BB, Gelman A, et al. Goodman Sn: five ways to fix statistics. Nature 2017;551:557–9. [DOI] [PubMed] [Google Scholar]
- 11. Wagenmakers E-J, Love J, Marsman M, et al. Bayesian inference for psychology. Part II: example applications with JASP. Psychon Bull Rev 2018;25:58–76. 10.3758/s13423-017-1323-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Etz A, Gronau QF, Dablander F, et al. How to become a Bayesian in eight easy steps: an annotated reading list. Psychon Bull Rev 2018;25:219–34. 10.3758/s13423-017-1317-5 [DOI] [PubMed] [Google Scholar]
- 13. Perezgonzalez JD, Frias-Navarro MD. Retract p < 0.005 and propose using JASP,365 instead. F1000Res. 6, 2017. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14. Hemming K, Chilton PJ, Lilford RJ, et al. Bayesian cohort and cross-sectional analyses of the pincer trial: a pharmacist-led intervention to reduce medication errors in primary care. PLoS One 2012;7:e38306. 10.1371/journal.pone.0038306 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15. Kawaguchi T, Azuma K, Yamaguchi T, et al. Preferences for pharmacist counselling in patients with breast cancer: a discrete choice experiment. Biol Pharm Bull 2014;37:1795–802. 10.1248/bpb.b14-00452 [DOI] [PubMed] [Google Scholar]
- 16. Van der Linden L, Decoutere L, Walgraeve K, et al. Combined use of the rationalization of home medication by an adjusted STOPP in older patients (RASP) list and a pharmacist-led medication review in very old inpatients: impact on quality of prescribing and clinical outcome. Drugs Aging 2017;34:123–33. 10.1007/s40266-016-0424-8 [DOI] [PubMed] [Google Scholar]
- 17. Faul F, Erdfelder E, Lang A-G, et al. G*Power 3: a flexible statistical power analysis program for the social, behavioral, and biomedical sciences. Behav Res Methods 2007;39:175–91. 10.3758/BF03193146 [DOI] [PubMed] [Google Scholar]
- 18. Fagan TJ. Letter: nomogram for Bayes theorem. N Engl J Med 1975;293:257. 10.1056/NEJM197507312930513 [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
All data relevant to the study are included in the article or uploaded as supplementary information.

