Skip to main content
PLOS One logoLink to PLOS One
. 2024 Jun 5;19(6):e0300001. doi: 10.1371/journal.pone.0300001

Audit and feedback to change diagnostic image ordering practices: A systematic review and meta-analysis

Oluwatosin Badejo 1, Maria Saleeb 1, Amanda Hall 1,2, Bradley Furlong 1, Gabrielle S Logan 1, Zhiwei Gao 2, Brendan Barrett 2,3, Lindsay Alcock 4, Kris Aubrey-Bassler 1,2,*
Editor: Joshua Robert Zadro5
PMCID: PMC11152319  PMID: 38837994

Abstract

Background

Up to 30% of diagnostic imaging (DI) tests may be unnecessary, leading to increased healthcare costs and the possibility of patient harm. The primary objective of this systematic review was to assess the effect of audit and feedback (AF) interventions directed at healthcare providers on reducing image ordering. The secondary objective was to examine the effect of AF on the appropriateness of DI ordering.

Methods

Studies were identified using MEDLINE, EMBASE, CINAHL, Cochrane Central Register of Controlled Trials and ClinicalTrials.gov registry on December 22nd, 2022. Studies were included if they were randomized control trials (RCTs), targeted healthcare professionals, and studied AF as the sole intervention or as the core component of a multi-faceted intervention. Risk of bias for each study was evaluated using the Cochrane risk of bias tool. Meta-analyses were completed using RevMan software and results were displayed in forest plots.

Results

Eleven RCTs enrolling 4311 clinicians or practices were included. AF interventions resulted in 1.5 fewer image test orders per 1000 patients seen than control interventions (95% confidence interval (CI) for the difference -2.6 to -0.4, p-value = 0.009). The effect of AF on appropriateness was not statistically significant, with a 3.2% (95% CI -1.5 to 7.7%, p-value = 0.18) greater likelihood of test orders being considered appropriate with AF vs control interventions. The strength of evidence was rated as moderate for the primary objective but was very low for the appropriateness outcome because of risk of bias, inconsistency in findings, indirectness, and imprecision.

Conclusion

AF interventions are associated with a modest reduction in total DI ordering with moderate certainty, suggesting some benefit of AF. Individual studies document effects of AF on image order appropriateness ranging from a non-significant trend toward worsening to a highly significant improvement, but the weighted average effect size from the meta-analysis is not statistically significant with very low certainty.

Introduction

Up to thirty percent of diagnostic imaging (DI) tests may be unnecessary [1, 2] and this excess use increases healthcare costs, wait times, and the likelihood of patient harm [3]. Unwarranted DI testing often leads to incidental findings which can in turn lead to a cascade of further unnecessary tests and treatments [4, 5]. For example, more liberal use of imaging for back pain has been associated with higher rates of surgery and other procedures and higher healthcare costs, as well as longer absence from work [6]. Incidental findings can lead to increased patient anxiety, financial burden, and ultimately delays in necessary treatment [5, 7], while also exacerbating long wait times for patients who do require these tests [8]. Physical harm to patients is also important to consider, as some types of imaging such as computed tomography (CT) involve exposure to high doses of ionizing radiation, which may lead to an increased risk of iatrogenic cancers [9, 10].

Audit and feedback (AF) has been implemented in healthcare settings as a strategy to modify behaviours in delivering health care services, including DI test ordering [11]. AF provides summaries of clinical performance over a specified period to health care providers, with the aim of motivating behaviour change. A Cochrane review of 70 randomized trials in healthcare settings revealed moderate quality evidence that AF has a moderate effect (dichotomous outcome: median adjusted risk difference of 4.3%, IQR 0.5% to 16% (49 studies); continuous outcome: median adjusted percent change of 1.3% (interquartile range (IQR) 1.3% to 28.9% (26 studies)) on increasing health professional compliance with desired behaviour when compared to usual practice [11]. This review included 4363 providers or provider groups from 49 trials that examined dichotomous outcomes and 1266 providers or provider groups from 21 trials that examined continuous outcomes. However, this review examined AF that targeted multiple issues, including the management of diabetes mellitus, blood pressure control, inappropriate antibiotic prescribing, X-ray utilization rates and more. The large range of topics addressed in this review and the heterogeneity in outcomes make it difficult to draw conclusions about the effect of AF on DI ordering, and the effect estimate for this area was not provided [11].

The objective of the current review was to determine the effect of AF interventions on DI ordering rates and DI ordering appropriateness. We also completed a comprehensive description of DI AF interventions using the template for intervention description and replication (TIDieR) checklist [12].

Materials and methods

Our review protocol was developed in line with recommendations from the Cochrane Effective Practice and Organization of Care (EPOC) group [13] and was prospectively registered with Open Science Framework (https://osf.io/5dczr) [14]. Although this group is no longer active, the resources are still published online [13] and additional information can be found in the Cochrane handbook [15]. Initially, we proposed to include both randomized controlled trials (RCTs) and some observational designs; however, our literature search discovered a sufficient number of RCTs and a discrepancy in results between the RCTs and the observational studies. The primary analyses were therefore limited to RCTs due to higher quality evidence, and the findings of the observational studies are reported in the S1 Appendix. We also proposed to compare AF interventions to a different active intervention, but elected to remove this comparison to simplify interpretation of the findings. Finally, the original protocol included only the appropriateness of image orders as the sole outcome, but we elected to add a total DI orders outcome because it directly aligned with the purpose of this review.

Data sources

We identified studies using a systematic search of MEDLINE (PubMed), EMBASE, CINAHL, the Cochrane Central Register of Controlled Trials and the ClinicalTrials.gov registry. Our search strategy was modelled after that of the Cochrane review [11], but it was adapted by an information specialist to ensure it included sufficient terms related to diagnostic imaging (S1 Appendix). These search strategies also underwent peer review using the Peer Review of Electronic Search Strategy (PRESS) guidelines [16]. We searched for full-text articles available up to December 20th, 2022 with no earlier date restriction. These database searches were supplemented with electronic and manual searches, including forward tracking to identify papers that cited the studies already included in the review.

Inclusion and exclusion criteria

Study design

RCTs with no restriction on language, geographic setting or year of publication were included in the primary analyses. We planned to translate articles published in a language other than English using Google Translate (https://translate.google.com/) which is as accurate as human translators for the languages commonly used for science [17]. The results of controlled before-after, non-randomized controlled studies and interrupted time series analyses are included in the S1 Appendix. Uncontrolled studies, case series and case reports were excluded.

Population

We included studies targeting health-care professionals who order DI in the routine management of their patients. Studies were excluded if the target population was healthcare professionals who do not normally order imaging tests such as pharmacists, radiologists, technicians, and medical students.

Intervention and comparator

Studies that provided feedback on individual clinician or clinician group ordering compared to a target recommended by local, regional or national guidelines, or a benchmark such as test ordering among peer clinicians were included. Studies that examined AF as the sole strategy or AF as the integral part of a multi-faceted intervention were included. Similar to the Cochrane review of AF [11], we considered AF to be “integral” if the other features of the intervention were unlikely to be offered without AF or if other components of the intervention were optional and therefore not necessarily received or used by subjects in the intervention group. For example, we included studies of AF combined with an educational session, but we excluded a comparison of AF combined with an electronic, point-of-care decision support tool vs usual practice from this sub-analysis.

We were primarily interested in the comparison of AF to a usual practice control group. We did not consider the provision of paper or digital clinical practice guidelines to be an active intervention as provision of guidelines alone is rarely associated with measurable behaviour change [18]. Thus, groups receiving guidelines together with AF were categorized as “AF alone,” and those receiving guidelines alone were categorized as “usual practice.” If comparison groups were not explicitly defined, they were assumed to be usual practice. Some papers studied AF combined with another intervention verses the other intervention alone, which we included in a separate subgroup of the meta-analyses as the effect of adding AF to another intervention may be different than the effect of AF alone.

Studies were excluded if audits occurred without feedback, if they occurred during a patient visit, or if feedback was given in real time during or shortly after a patient encounter. If feedback was given for hypothetical situations or was a reminder without reference to specific ordering behaviour, it was also excluded. We were also exclusively interested in the effect of AF on diagnostic imaging and we therefore excluded studies that focussed on screening tests such as mammograms.

Outcome

The primary outcome of this review was the total number of DI tests ordered and the secondary outcome was the appropriateness of test orders. Appropriateness was measured as the total number or proportion of imaging tests that were classified as concordant with a standard of care, such as clinical practice guidelines, according to the individual study authors. Some papers examined AF of non-DI tests in addition to DI orders, but only the DI-specific outcome data were included.

When studies reported more than one measure of the same outcome, we extracted (in order of preference): post-intervention continuous measure adjusted for baseline values, change from baseline continuous measure, post-intervention continuous measure (no adjustment for baseline values), then post-intervention dichotomous measure. Odds ratios for dichotomous outcomes were converted to continuous outcomes as described in the Cochrane Handbook [15, Section 10.6] to be included in the continuous meta-analyses.

Study selection

We uploaded the identified citations to the web-based systematic review software platform, Covidence [19]. Duplicates were identified automatically by Covidence or manually during screening. The titles and abstracts of all articles were screened independently by two review authors and screening conflicts were resolved by a third reviewer (OB, MS, AH, KAB). Pilot screening of 10 studies was undertaken to ensure uniformity in screening procedure. Full-text screening followed the same process of review and conflict resolution.

Data extraction and quality assessment

We extracted information on study characteristics, population, intervention and outcome from each of the included studies according to the TIDieR recommendations [12]. We also extracted all data relating to the outcomes described above including raw numbers, proportions and effect estimates where provided using a modified version of the Cochrane Effective Practice and Organization of Care (EPOC) data collection checklist [20]. Data were extracted independently by two reviewers and discrepancies in the extracted data were resolved by discussion or involvement of a third reviewer (OB, MS, AH, KAB).

The risk of bias for each included study was independently reviewed by two authors as high, low, or unclear using the Cochrane risk of bias tool version 1 (OB, BF, MS, KAB) [21], which has since been updated [22, 23] but was the version in use at the time this study was originally conceptualized. Studies with high risk of bias in at least one of the domains for assessment received an overall judgment of high risk. Studies with baseline imbalances in study group characteristics that were greater than would be expected due to chance were classified as high risk. Blinding study participants to AF interventions is not possible and the primary outcome was objective, so studies were not classified as high risk in this domain if investigators treated all study groups equally (e.g. both intervention and comparator groups were aware they were participating in a study). In addition, because our primary outcome was objective, the unavailability of a pre-published protocol was not considered an indicator of selective reporting. Discrepancies in ratings were resolved through discussion or by involvement of a third reviewer.

Data synthesis and analysis

All studies identified the individual clinician or clinical team as the unit of study participation. While some studies reported outcomes at the clinician or team level, some studies only reported results at the patient level (e.g., proportion of patient visits at which a DI test was ordered) or at the study group level (e.g., total DI orders in the group), without mentioning any adjustment for clustering of observations. Others reported results at the study participant level (e.g., mean number of DI orders per clinician), and other studies reported both. Although effect estimates measured at the study group or individual patient level are representative, variance is likely underestimated unless the analyses adjust for correlated observations. Therefore, we preferentially extracted data at the participating clinician level. We included data that did not appear to be adjusted for clustering but noted this when evaluating risk of bias and in our results.

Where possible, the means of multiple outcomes from the same paper (e.g. DI ordering for different imaging types) were included in the meta-analyses as recommended in the Cochrane Handbook [15, Section 6.5.2.10]. When it was not possible to determine the mean of outcomes (e.g., only odds ratios and 95% CIs reported), we included data for the more frequent outcome. Data were compiled into meta-analyses and forest plots, and heterogeneity was estimated using I2 and Chi2 statistics using Review Manager (RevMan) software [24]. Potential sources of heterogeneity are explored qualitatively in Results and Discussion. We considered presenting data in subgroups by imaging modality or by target organ, but there are few studies and heterogeneity exists primarily within potential subgroups rather than between subgroups so we elected to present data without such grouping. Because of variability in the continuous outcome measures used between studies, we combined results using the standardized mean difference (SMD). We then rescaled the summary SMD from the total imaging orders meta-analysis into units of the mean difference between intervention and control in the number of DI tests ordered per 1000 patient consultations, which was the outcome used across a plurality of studies in the meta-analysis, including the largest. This conversion of SMD to natural units is recommended by the Cochrane Collaboration to enhance interpretability of the SMD [25, 26]. SMD measures outcomes in units of the standard deviation; therefore, to convert the summary SMD and its confidence interval (CI) into natural units, we chose the weighted (by trial sample size) average of the SDs from each of the studies that used the DI tests per 1000 patients outcome. We also expressed this outcome as a percentage of the weighted average of the baseline, pre-intervention DI tests per 1000 patients from each of these same studies. The SMDs from the secondary outcome meta-analysis were similarly converted into the difference between intervention and control in the percentage of image test orders that were considered to be appropriate.

Summary of findings and GRADE strength of evidence

Two authors (OB, MS), with resolution of disagreement by a third author (KAB), applied the Grading of Recommendations Assessment, Development and Evaluation (GRADE) approach to summarize our findings and rate the strength of evidence [15, Chapter 14]. The GRADE approach uses risk of bias, inconsistency, indirectness, imprecision, and evidence of publication bias to assign a level of certainty to the body of evidence regarding each outcome or comparison. Because this analysis was restricted to RCTs, we began with an assumption of high certainty evidence. Evidence was downgraded if a majority of studies in a given comparison were considered to be at unclear or high risk of bias. Inconsistency was assessed using I2 values, with downgrading of one level for comparisons with an I2 greater than or equal to 60%. To determine indirectness, factors such as population, the interventions and co-interventions, as well as the DI modality were considered. As the primary objective was to study the effect of AF on all diagnostic imaging utilization (X-ray, ultrasound, echocardiogram, CT, MRI), evidence was downgraded for indirectness if a comparison included two or fewer DI modalities. There is relatively little guidance in the literature on how to assess imprecision in reviews that use SMD as an outcome measure and when there is no clear consensus on what is considered a minimally important difference. We elected to follow the convention that a standardized mean difference (SMD) between 0.5 and 0.8 indicates a moderately effective intervention. If a CI was greater than the midpoint of that range, 0.65 units, we downgraded the certainty of evidence by one level [15]. For context, an SMD of 0.65 units in our primary analysis translates into a 14.3% reduction in DI ordering when converted into natural units as described above. Finally, to assess publication bias, we subjectively assessed the symmetry of the funnel plots (S3 and S4 Figs in S1 Appendix) and elected to downgrade evidence if there was clear asymmetry.

Results

Eleven RCTs met the inclusion criteria from an initial literature search that identified 4493 papers (Fig 1) [2737]. One non-randomized controlled trial (NRCT) [38] and 5 observational studies were included in the Appendix (S1-S4 Tables and S1, S2 Figs in S1 Appendix) [3944]. All of these studies included at least one comparison that met our inclusion criteria. Some of the studies examined additional interventions which did not meet our inclusion criteria, and comparisons involving these other interventions were excluded. While the certainty of evidence for most comparisons was judged to be very low to low based on risk of bias, indirectness and imprecision (Table 1), we considered the evidence for the effect of AF on our primary outcome of all DI orders (both subgroups) to be of moderate certainty. The rationale for these certainty of evidence ratings is provided in the footnotes of Table 1 and the included studies are described in Table 2.

Fig 1. PRISMA diagram for study identification, screening, and exclusions.

Fig 1

Table 1. Summary of findings.

Effect of audit and feedback on diagnostic imaging requests
Population: Healthcare providers
Setting: Healthcare
Outcome: Number of diagnostic imaging requests
Comparison Participants (studies) Anticipated effects a Quality Comments
Comparator (tests/1000 patients) Intervention vs comparator MD (95% CI)
AF alone vs usual practice or guideline provision 4190 31.5 tests 1.4 fewer ⨁⨁◯◯ Negative MD favors AF
(6 RCTs) (-2.8 to 0.0) Lowc,d
AF+ other intervention vs other intervention alone 107 20.3 tests 1.7 fewer ⨁◯◯◯ 1 RCT examined CTPA and the other TTE
(2 RCTs) (-4.3 to 0.9) Very low c,e,f
AF vs comparator (both subgroups) 4297 31.3 tests 1.5 fewer ⨁⨁⨁◯
(8 RCTs) (-2.6 to -0.4) Moderatec
Outcome: Appropriateness of diagnostic imaging requests
Comparison Participants (studies) Anticipated effects a Quality Comments
Comparator (percent appropriate) b Intervention vs comparator MD (95% CI)
AF alone vs usual practice or guideline provision 317 82.9% 0.9% greater ⨁◯◯◯ Positive MD favors AF
(4 RCTS) (-4.5 to 6.3%) Very Lowc,d,e,g
AF+ other intervention vs other intervention alone 107 - 7.0% greater ⨁◯◯◯ 1 RCT examined CTPA and the other TTE
(2 RCTs) (2.6 to 11.5%) Very lowc,e,f
AF vs comparator (both subgroups) 424 82.9% 3.1% greater ⨁◯◯◯
(6 RCTs) (-1.5 to 7.7%) Very lowc,d,e

Abbreviations: AF: audit and feedback; CTPA: computed tomography pulmonary angiogram; MD: mean difference; RCT: randomized controlled trial; TTE: trans-thoracic echocardiogram

a. Note that a reduction in the number of orders and an increase in the appropriateness of orders were considered favorable outcomes. Thus, the sign of a favorable result is reversed for each outcome.

b. No papers in the second subgroup used the % appropriate outcome

GRADE rating explanations

c. Downgraded due to high or unclear risk of bias in a majority of studies. See Table 4. Most studies were unclear on their procedure for allocation concealment.

d. Downgraded due to inconsistency (I2≥60%)

e. Downgraded due to imprecision. (95% CI for the SMD > 0.65)

f. Downgraded due to indirectness as studies only evaluated the effect of AF on CT scans for pulmonary embolism and transthoracic echocardiography.

g. Downgraded due to indirectness as 2 of the 4 studies only evaluated lumbar and knee radiographs, and the remaining 2 studies analyzed echocardiogram ordering

Table 2. Summary characteristics of included RCTs.

REFERENCE / COUNTRY TARGET PROVIDER TARGET BEHAVIOUR AF DESCRIPTION COMPARISON DESCRIPTION SAMPLE SIZE (A) PHYSICIAN/PATIENT GENDER AND AGE PRIMARY OUTCOME(S) OUTCOME (B)
WINKENS, 1995
NL
Primary care physicians (c) Decrease in various x-rays (d) and US orders AF with critical feedback comments on individual test orders Same intervention on different, non-imaging tests (No intervention) 79 physicians (C:39, I:40) /~187 000 patients Physicians:
Proportion female (C: 10.2%, I: 10%)
Patient: not reported
Mean number of tests ordered
Mean percentage of tests which were guideline-appropriate
Total imaging
Appropriateness
KERRY, 2000
UK
Primary care physicians Decrease in all x-rays AF alone Paper guidelines 69 practices (C:36, I:33)/ 175 physicians Not reported Mean percent reduction in total number tests ordered compared to baseline Total imaging
ECCLES, 2001
UK
Primary care physicians Decrease in lumbar spine and knee x-rays AF alone Paper guidelines 121 practices (C:61, I:60) Not reported Mean number of tests/1000 patients
Odds Ratio for an appropriate test in the intervention vs control groups
Total imaging
Appropriateness
ROBLING, 2002
UK
Primary care physicians Increase in guideline appropriate lumbar spine and knee MRIs AF alone Paper guidelines 19 practices (C:10, I:9)
95 requests (C:53, I:42)
Not reported Percentage of requests that were guideline appropriate Appropriateness
VERSTAPPEN, 2003
NL
Primary care physicians Decrease in x-rays of shoulder, spine, hip and knee AF + clinician education Same intervention for other (non-imaging) tests 25 groups (C:12, I:13)/ 163 physicians (C:75, I:88) Physicians: Proportion female (C: 16% control, I: 17% intervention).
Mean age (C: 46%, I: 45.8%)
Patients: Percent of patients over 65:
(C: 15%, I: 13%)
Mean number of tests/physician
Mean number of guideline appropriate tests/physician
Total imaging
Appropriateness
BHATIA, 2014
USA
Cardiology fellows Decrease in rarely appropriate TTE AF + clinician education + pocket card No intervention 1 hospital/24 physicians (C:12, I:12)/ 1213 patients (C:600, I:613) Physicians: not reported
Patients: Average age (C:65, I: 64)
Proportion female (C: 32%, I: 33%)
Mean number of tests/physician
Proportion of tests that were guideline appropriate
Total imaging
Appropriateness
RAJA, 2015
USA
Emergency physicians Increase in guideline- appropriate and decrease in overall use of CTPA AF + computerized decision support Computerized decision support 1 ED/43 physicians (C:21, I:22)/2167 patients (C:1149, I:1018) Physicians:
Mean age (C: 41.2, I: 39.4)
Proportion female:
(C: 29%, I:32)
Patients: not reported.
Number of tests/patient seen
Proportion of tests that were guideline appropriate
Total imaging
Appropriateness
DUDZINSKI, 2016
USA
Academic cardiologists Decrease in rarely appropriate TTE AF + clinician education Clinician education 1 hospital/ 66 physicians (C:33, I:33)/16075 tests (C:8166, I:7909) Physicians: Age and gender not reported.
Patients: Mean age (C: 66, I: 65)
Proportion female (C: 39%, I: 41.8%)
Total tests ordered
Proportion of tests that were guideline appropriate
Total imaging
Appropriateness
BHATIA, 2017
CA, USA
Primary care physicians and cardiologists Decrease in rarely appropriate TTE AF + education No intervention 8 hospitals/ 153 physicians (C:79, I:74)/ 14697 tests (C:7798, I:6899) Not reported Mean number of tests/physician
Mean percentage of tests that were guideline- appropriate
Total imaging
Appropriateness
ZAFAR, 2019
USA
Primary Care Physicians and Nurse practitioners Decrease in LS MRI orders AF + clinical decision support Clinical decision support 8 Practices/ 52 providers (C:26, I:26))/ 5142 visits (C:2021, I:3121) Number not reported Proportion of visits for LBP on which a test was ordered Total imaging
O’CONNOR, 2022
AU
Primary Care physicians Decrease in CT, MRI, X-Ray, US orders AF alone No intervention 2271 practices/ 3660 physicians (C:727, I:2933) Physicians:
Proportion 60 or over (C: 42%, I: 40%)
Proportion female (C: 37%, I: 39.5%)
Patients: Not provided.
Rate of request/1000 patients Total imaging

(a)Data from this column that are missing were not reported. (I) refers to intervention; (C) refers to control. The highest level that includes numbers for (I) and (C) is the level at which experimental groups were randomized.

(b)Outcomes included total imaging or appropriate test orders as described in Methods

(c)Primary care physicians may include family, general practice and general internal medicine physicians

(d)Includes chest, cervical spine, thoracic spine, lumbar spine, pelvis/hip, knee, ankle and sinus x-rays

Abbreviations: AF, audit and feedback; AU, Australia; CA, Canada; CT, Computed Tomography; CTPA, CT Pulmonary Angiogram; LBP, low back pain; MRI, Magnetic Resonance Imaging; NL, Netherlands; TTE, Transthoracic Echocardiogram; US, Ultrasound; UK, United Kingdom; USA, United States of America.

Intervention fidelity, bias, and certainty of evidence

The AF interventions are described in Table 3. One study directly assessed if AF reports had been opened by tracking logins to an online system and found that 61% of participants logged in at least once [28]. Verstappen et al. reported that 100% of study participants attended an in-person education session at which AF reports were discussed [35] and O’Connor et al. reported the percentage of AF reports sent by post that were returned unopened was 4.9%-14.7%, dependent on which intervention they received [32]. However, the lack of a returned envelope was not considered sufficient proof of an AF receipt so we indicated “Not reported” for this variable. No other studies clearly reported AF receipt or other measures of intervention fidelity (Table 3).

Table 3. Description of AF interventions according to TiDIER recommendations.

Winkens et al., 1995 Kerry et al., 2000 Eccles et al., 2001 Robling et al., 2002 Verstappen, et al. 2003 Bhatia et al., 2014 Raja et al., 2015 Dudzinski et al., 2016 Bhatia et al., 2017 Zafar et al., 2019 O’Connor et al, 2020
Who and where?
Provider type PCPs (a) PCPs PCPs PCPs PCPs Card, GIM res. EP Cardio PCPs, Card PCPs PCPs
AF provided to Individuals or group Individual Individual Individual Individual Individual Individual Individual Individual Individual Individual Individual
AF delivered directly to provider? Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes
Inpatient/Outpatient setting? Unclear Unclear Outpatient Unclear Outpatient Outpatient Inpatient Outpatient Outpatient Outpatient Outpatient
Content of AF reports
Imaging modality X-Ray, US X-Ray X-Ray MRI X-Ray Echo CTPA Echo Echo MRI CT, MRI, X-Ray, US
Desired change in ordering Decrease Decrease Decrease Decrease Decrease Decrease Decrease Decrease Decrease Decrease Decrease
Patient outcomes (findings on imaging test) Yes No No No No Yes No No No No No
Other info (e.g. costs, guidelines, doses) Yes Yes No No Yes Yes No No Yes No Yes
AF of Individual or group Individual Group Individual Group Individual Individual Individual Individual Both Individual Individual
Feedback about Individual cases or aggregate cases Individual Aggregate Aggregate Aggregate Aggregate Aggregate Both Aggregate Aggregate Aggregate Aggregate
Comparison provided(b) Own, Peers Own Peers Peers Peers None Peers None Own, peers Own, peers Peers
Explicit target provided No No No No Yes No No No No No No
Action plan provided (c) Yes No No No Yes No No No No No No
Graphical elements No No No No Yes No Yes No Yes No Yes
When and how much AF?
Time period of audit data 1 Month 6 Months 6 Months Unclear 6 Months 1 Month 3 Months 1 Month 1 Month Unclear 12 months
Lag between audit and feedback Weeks Months Months Unclear Unclear Brief Brief Brief Weeks Months Months
Frequency (number of times given) 5 1 2 1 3 9 4 6 17 Unclear 1 or 2
Time in between reports 6–7 Months N/A 6 Months N/A Unclear 1 Month 3 Months 1 Month 1 Month 4–6 Months 12 months
Duration of intervention (months) 32 12 12 Unclear 6 9 12 6.4 17 11 12 or 24
How was AF delivered?
Verbal or written Written Written Written Both Both Written Written Written Written Written Written
Delivery mode Unclear Unclear Post Post, In Person Post, In Person E-mail E-mail E-mail E-mail Unclear Post
Source of in-person delivery N/A N/A N/A Expert Leader N/A N/A N/A N/A N/A N/A
Asked to reflect on AF Yes No No Yes Yes No No No No No No
Who provided AF?
Who conducted the audit? Authority(d) Unclear Researcher Unclear Unclear Researcher Researcher Researcher Researcher Unclear Researcher
Who developed the feedback report? Authority(d) Unclear Researcher Unclear Unclear Researcher Researcher Researcher Researcher Unclear Unclear
Fidelity
Planned assessment of AF receipt? Not Reported Not Reported Not Reported Not Reported Yes Not Reported Not Reported Not Reported Yes Not Reported Not reported
AF receipt by providers Not Reported Not Reported Not Reported Not Reported 100% Not Reported Not Reported Not Reported 61% Not Reported Not reported

Abbreviations: AF, Audit and Feedback; Card, Cardiologists; CTPA, Computed Tomography Pulmonary Angiogram; Echo, Echocardiography; EP, Emergency Physicians; GIM, General physicians; MRI, Magnetic Resonance Imaging; N/A, not applicable; PCP, Primary care; Res, residents; US, ultrasound; XR, X-ray

(a) Primary care physicians may include family, general practice and general internal medicine physicians

(b) Comparison provided Includes own/ peers’ previous performance, national benchmark

(c) Action plan refers to an explicit verbal or written plan to achieve desired targets. We did not consider the provision of appropriate use criteria separate from the AF reports to be an action plan.

(d) Refers to hospital, health maintenance organization (HMO) or health authority

Note: For multifaceted interventions, we assessed the characteristics of the audit and feedback component

The risk of bias was judged to be unclear in a majority of studies and only two studies were thought to be at low risk of bias (Table 4) [32, 35]. Nine of the eleven studies did not report on the concealment of study group allocation, but almost all studies were downgraded based on more than just that item. Zafar et al. [37] used hierarchical regression to adjust for clustering in some of their analyses; however, those results were not suitable for this meta-analysis. The data that are suitable for synthesis from this study do not appear to be adjusted for correlated observations and it was therefore reported as high risk for this reason. A risk of bias table was not included for the secondary outcome as evaluations were similar.

Table 4. Risk of bias for each domain and overall judgement for risk of bias for each included RCT.

Bias Winkens 1995 Kerry 2000 Eccles 2001 Robling 2002 Verstappen 2003 Bhatia 2014 Raja 2015 Dudzinski 2016 Bhatia 2017 Zafar 2019 O’ Connor 2022
Random Sequence Gen.
Allocation concealment
Blinding of participants
Blinding of outcome assessment
Incomplete outcome data
Selective reporting
Other (baseline imbalance)
Other (no clustering adjustment)
OVERALL JUDGMENT

Low risk; Unclear Risk; High risk

The effect of AF on the total number of DI requests (Fig 2)

Fig 2. Effect of audit and feedback on the number of diagnostic imaging requests.

Fig 2

The AF groups in this figure include audit and feedback alone and audit and feedback as the main component of a multi-faceted intervention. The control group includes usual care or the provision of paper guidelines only (subgroup 1) or an active control group that was compared against the same intervention with the addition of AF (subgroup 2). Note that the results from Raja et al and Zafar et al may not be adjusted for correlated observations.

Ten trials examined the effect of AF on the number of diagnostic imaging requests, nine of which are presented in the forest plots (Fig 2). Six trials used a control group that did not receive any intervention or only received practice guidelines [27, 28, 30, 31, 32, 35] and three additional studies measured the effect of adding AF to another intervention [29, 33, 37]. The remaining study describes baseline imbalances between the control and intervention groups which results in very similar post-intervention outcomes (not shown) [36]. However, the authors of this study report a 4% reduction in test ordering over the study period in the intervention group (p-value = 0.11), but they don’t report control group data and we were therefore not able to include these results in the meta-analysis. The post-intervention values were not included in the meta-analysis because of bias due to the baseline imbalance [36]. Although there was a fair degree of heterogeneity in the types of imaging that were addressed by the AF interventions (Table 3), results for the pooled analyses of the primary outcome were only moderately heterogeneous (I2 = 45%, p = 0.08) and heterogeneity existed mostly within potential subgroups (e.g. Bhatia 2014, 2017 and Dudzinski, 2016, all of which examined echocardiography), rather than between subgroups. We therefore decided not to pool our results by imaging modality and/or target organ.

The meta-analysis demonstrates a statistically significant reduction in total DI test ordering (SMD = -0.22, 95% CI = -0.38 to -0.06, p-value = 0.009), which translates into 1.5 fewer image test orders per 1000 patients seen (95% CI -2.6 to -0.4) in the intervention vs the control groups. The GRADE quality of evidence for this summary effect was rated as moderate, but the rating for each subgroup was very low to low (Table 1). The weighted mean average number of DI tests ordered during the pre-intervention period of the three studies that used this outcome was 31.3 orders per 1000 patients seen [30, 32, 33]. Thus, audit and feedback was associated with a 4.9% (95% CI 1.3 to 8.4) greater reduction in test ordering than control. This finding is driven primarily by a single study which includes almost 70% of the participants in the meta-analysis [32]; however, the results of most other studies were similar (I2 = 45%, p-value = 0.08). Only one study, which examined the effect of AF on echocardiogram ordering practices showed a higher rate of ordering in the AF vs the usual practice group, though this difference was not significant [27]. Interestingly, two other studies on echocardiogram ordering from the same research group showed the opposite, non-significant trend towards reduced ordering in the AF group [28, 29].

In the first subgroup of Fig 2A, Kerry et al., Eccles et al., and O’Connor et al. examined AF alone [3032]. The remaining studies in this subgroup examined AF as the core part of a multi-faceted intervention, including a discussion or education session [27, 35], or an education session together with the provision of a mobile application to assist with decision-making [28]. In the second subgroup included in Fig 2A, the studies examined AF added to electronic clinical decision support [33] or an educational session [29]. The results of the two subgroups in this analysis were similar (I2 = 0%, p-value = 0.83) suggesting that the effect of AF is similar when implemented on its own or when added to another intervention, although only 2 studies were included in the second subgroup. The additional study included in Fig 2B (dichotomous outcome), which investigated the effect of AF added to real-time alerts implemented at the point of electronic ordering [37], reports similar findings.

The effect of AF on the appropriateness of diagnostic imaging requests (Fig 3)

Fig 3. Effect of audit and feedback on the appropriateness of diagnostic imaging requests.

Fig 3

The AF groups in this figure include audit and feedback alone and audit and feedback as the main component of a multi-faceted intervention. The control group includes usual care or the provision of paper guidelines only (subgroup 1) or an active control group that was compared against the same intervention with the addition of AF (subgroup 2). Although Dudzinski et al and Bhatia et al (2017) papers found no significant difference in the appropriateness outcome analyzed in our meta-analysis, both papers found a significant reduction in “rarely appropriate” echocardiograms in their AF intervention group (Odds Ratio (OR) = 0.59, 95% CI 0.39–0.88, p = 0.01 and OR = 0.75, 95% CI 0.57–0.99, p = 0.039, respectively). Note that favors AF is on the right side of the axis.

Whereas a decrease in total imaging was considered favorable, an increase in appropriateness was considered favorable. Thus, studies favoring AF are presented on opposite sides of the vertical axis in the forest plots for each of these outcomes (Figs 2 and 3). Four studies evaluated the effect of AF on the appropriateness of DI requests compared to usual practice, [27, 28, 30, 34] and two additional studies evaluated AF added to electronic clinical decision support [33] or an educational session [29]. All studies included in this section were also included in the primary outcome analyses (Fig 2A), with the exception of Robling et al. [34]. Results for the appropriateness outcome were mixed compared to the total imaging outcome, but overall, AF had no significant effect on appropriateness (SMD = 0.27, 95% CI = -0.13, 0.66, p-value = 0.18), with a high degree of heterogeneity (I2 = 70%, p-value = 0.005). This SMD translates into a 3.1% (95% CI -1.5 to 7.7%) higher proportion of image orders that were considered to be appropriate in the AF vs the comparator groups. Three of the six studies that examined appropriateness were deemed to be at high risk and the remainder were deemed to be at unclear risk of bias.

The two studies that examined AF added to another intervention were consistent (I2 = 0%, p-value = 0.48) in finding that AF improved appropriateness (SMD = 0.60, 95% CI = 0.22, 0.99, p-value = 0.002), despite substantial differences in the co-interventions examined in those two studies.

The effect of AF alone vs the effect of AF added to another intervention (subgroup 1 vs subgroup 2)

The effect of AF alone is presented in subgroup 1 and the effect of AF added to another intervention is presented in subgroup 2 of Figs 2 and 3. Although one might expect that co-interventions would “dilute” the effectiveness of AF, our findings suggest that may not be the case, albeit on the basis of a limited number of studies. The SMD for AF added to another intervention for both outcomes is higher than the SMD for AF alone, although this is only significant for the appropriateness outcome.

Results from observational studies (S1 Appendix)

The meta-analyses of the observational study data were similar to those included in the main text, with a higher degree of variability contributing to non-significant summary results. Almost all studies were considered to be at high risk of bias (S2 Table in S1 Appendix). The single study judged to be at low risk of bias was an interrupted time series analysis of clinical data related to a national intervention to improve the management of back pain, including a reduction in the use of imaging tests [43]. This study is notable because of its strong design that mitigates many of the limitations of observational analyses, its low risk of bias and the dramatic 10.9% reduction in imaging (albeit with a high degree of imprecision: 95% posterior interval = 0.85–20.9%) after the introduction of their AF intervention, resulting in substantial cost savings [43].

Discussion

This review includes 11 RCTs that assessed the effect of audit and feedback on diagnostic image test ordering. Our meta-analyses demonstrated a significant, 4.9% reduction in total number of DI orders but variable and non-significant results on the appropriateness of orders. The evidence for the primary, total DI orders outcome was judged to be of moderate certainty but the evidence for all other comparisons was found to be very low to low certainty and these final results should therefore be interpreted with caution. For context, the Cochrane review on all uses of AF for healthcare found a roughly 1.3% improvement in practice associated with AF interventions [11]; thus, the effectiveness of AF appears to be larger when used on DI ordering. The two studies that were deemed to be at low risk of bias in our review contrasted in their findings, with one study finding a significant, modest reduction in DI ordering after AF [32], while the other found no significant effect [35]. Neither of these low-risk studies examined the appropriateness of DI requests; thus, the results for this outcome must be interpreted with greater caution.

All studies that reported appropriateness expressed this outcome as a proportion of total image orders. Our finding that AF interventions result in a decrease in total image ordering but no statistically significant change in the proportion of appropriate orders, suggests that appropriate and inappropriate tests may therefore be reduced at a similar rate following AF. This may increase the risk of delayed or missed diagnoses due to a reduction in appropriate testing and disproportionately harm people who generally receive lower imaging rates, particularly minority groups and people of color[45, 46]. However, DI appropriateness criteria are relatively crude measures that often do not address a substantial grey area in clinical decision-making; thus, we cannot infer that reductions in “appropriate” imaging automatically result in patient harm [27, 29].

Although our meta-analyses demonstrated no significant effect on the appropriateness outcome, several papers found a significant benefit on a related outcome. While the effect of AF on appropriateness in Dudzinski et al. and Bhatia et al. 2017 [28, 29] was non-significant (Fig 3), these authors found significant effects on “rarely appropriate” (i.e., inappropriate) imaging requests. Bhatia et al. 2014 [27] found significant effects on both appropriate and rarely appropriate imaging requests. This discordance in the statistical significance of two related outcomes (appropriateness and inappropriateness) is not unexpected, especially when there is a substantial difference in the frequency of these outcomes. The work of the Cochrane collaboration demonstrates that statistical significance is more likely for less frequent outcomes [15, Section 6.4.1.5]. We chose to analyze “appropriateness” rather than “inappropriateness,” as not all papers reported both outcomes and this allowed a greater number of studies to be included in our meta-analyses.

Recommendations to enhance the effectiveness of AF

A meta-regression completed as part of the Cochrane AF review found that low baseline performance, repeated delivery of AF reports, a supervisor or colleague as the source of feedback, both verbal and written delivery of feedback and the provision of explicit targets and an action plan were all associated with improved effectiveness of AF [11]. While most of the studies included in our review did not comment on baseline performance, the single study with the most dramatic effect on DI ordering selectively enrolled high test-ordering clinicians [32]. This study also found that receiving two instances of AF reduced ordering to a greater degree than one report [32]. The frequency of AF provision amongst the other studies included in our review range from one to seventeen. Comparing across these studies, we did not observe an association between the numbers of reports received and reduced ordering; in fact, the largest effect sizes were observed in the studies that provided one to two reports. Our review does not support the recommendations that AF is provided by a supervisor or colleague, that AF should be provided both verbally and written, or that specific targets or action plans be provided with AF, albeit on the basis of a limited number of studies that examined these aspects. Thus, our results should be considered inconclusive regarding the effectiveness of these features in AF for DI requests.

Although we did not find support for the recommendation that AF reports be delivered by a supervisor or colleague, presumably the value of this method is the perception of reliability and importance of the information. This factor is often considered critical when pursuing clinician behaviour change [47, 48]. Additionally, having reports delivered by a supervisor may motivate clinicians to change their behavior to maintain their professional reputation with their supervisors and peers [49]. In the three studies focusing on echocardiogram ordering, the greatest benefit came in the study targeting cardiology and general internal medicine residents compared to other studies that enrolled independently practicing physicians [2729]. While these 3 studies did not include delivery of AF reports by an individual, it may be that the clinicians in training were more likely to perceive the information as trustworthy or they were more motivated by a desire to achieve professional norms [48].

Another factor that was not examined in the Cochrane meta-regression [11] was the effect of visual appearance on the effectiveness of AF. While the four studies in our review that included graphical elements in their AF reports do not appear to be associated with improved AF performance, O’Connor et al. compared an enhanced to a standard visual display of their AF data in their factorial design trial, and found that the enhanced display outperformed the standard version [32]. In this study, both standard and enhanced versions of the report included graphical information, but the enhanced version added highlighting to draw attention to indicators of higher utilization. Enhanced visual displays such as this could increase the effectiveness of AF, without substantially increasing costs and resource utilization.

A final consideration is the effectiveness of AF alone verses the effect when AF is added to another intervention. The evidence is very low certainty, but our findings suggest the possibility that AF may be more effective when added to another intervention than it is when implemented alone.

Limitations

Most of the studies included in this review were determined to be either at high or unclear risk of bias, and the quality of evidence for most comparisons was assessed as very low to low. Although the papers included in this review examined a range of imaging modalities, six of the studies exclusively examined a single, less commonly used imaging modality, sometimes just for a specific indication such as pulmonary embolism, or knee and back pain. The effectiveness of AF may vary across different modalities or indications and therefore these findings may be hard to generalize across different modalities and indications. While all the imaging modalities are used for diagnosis, some of the tests such as echocardiogram are more commonly used to monitor for progression of previously diagnosed conditions such as valvular heart disease, congestive heart failure and aortic dilation than they are for the initial diagnosis of those conditions, which may also affect the results of an AF intervention. Because the effectiveness of AF may vary dependant on indication or imaging modality, future studies could restrict the analyses to further investigate the effect of AF on these specific indications or modalities.

Conclusions

This review reports moderate quality evidence that AF and AF added to other interventions likely has a small but variable effect on the total number of DI requests, but results for improvements in the appropriateness of those requests are equivocal and of very low quality. The observation that AF may reduce total imaging requests with no change in appropriateness, suggests that both clinically indicated and inappropriate tests are reduced at a similar rate, raising the possibility of adverse clinical outcomes. Future studies of AF interventions should pay careful attention to study design and reporting standards to improve the quality and reliability of evidence, and they should consider studying harm outcomes.

Supporting information

S1 Appendix

S1 Fig. a. Effect of audit and feedback in observational studies on the number of diagnostic imaging requests (continuous outcome) (4–6). b. Effect of audit and feedback in observational studies on the number of diagnostic imaging requests (dichotomous outcome) (7, 8). S2 Fig. Effect of audit and feedback in observational studies on image order appropriateness (dichotomous outcome) (7). S3 Fig. Funnel plot of RCTs analyzing the total image order outcome. We did not consider this figure to be indicative of publication bias. The study in the bottom right favored the control intervention, not AF. S4 Fig. Funnel plot of RCTS analyzing the appropriateness of image orders outcome.We did not consider this figure to be indicative of publication bias. S1 Table. Description of AF interventions using TiDIER recommendations (1). Abbreviations: AF, Audit and Feedback; CT, Computed Tomography; Echo, Echocardiography; GIM, General physicians; Res, residents; Gov., Government; Mm; MRI, Magnetic Resonance Imaging; N/A, not applicable; PCP, Primary care physicians (e) PCPs refers to primary care physicians and may include family, general practice and general internal medicine physicians, (f) The term residents also refers to registrars (g) Comparison provided Includes own/ peers’ previous performance, national benchmark. Note: For multifaceted interventions, we assessed the characteristics of the audit and feedback component. S2 Table. a. Risk of Bias for NRCTs using the Risk Of Bias In Non-randomized Studies—of Interventions (ROBINS-I) tool (2). b. Risk of Bias for observational studies using Effective Practice and Organisation of Care (EPOC) recommendations (3). c. Risk of Bias for interrupted time series studies using Effective Practice and Organisation of Care (EPOC) recommendations (3). Legend: Low risk; Indeterminate Risk; High risk. S3 Table. Effect of audit and feedback in a non-randomized, crossover design study on the number of diagnostic imaging request 9).*no p-values, standard deviations, or confidence intervals provided. S4 Table. Effect of audit and feedback with another intervention vs usual care in an interrupted time-series analysis (10). S1 File. EMBASE search strategy. S2 File. CINAHL search strategy. S3 File. PubMED search strategy. S4 File. References for supporting information.

(ZIP)

pone.0300001.s001.zip (5.6MB, zip)
S1 Checklist. PRISMA checklist.

(DOCX)

pone.0300001.s002.docx (30KB, docx)

Acknowledgments

We would like to acknowledge Bethan Copsey from the University of Leeds for her advice on the statistical considerations for the meta-analyses.

Data Availability

All relevant data are within the manuscript and its Supporting information files.

Funding Statement

O.B received the Memorial University of Newfoundland’s NL SUPPORT – Support for People and Patient Oriented Research and Trials Grant (https://www.mun.ca/nlcahr/nlcahr-funding-programs/nl-support/nl-support-patient-oriented-research-grant/). There is no grant number associated with this award. The funder played no role in the study design, data collection and analysis, decision to publish, or preparation of the manuscript.

References

  • 1.Canadian Institute for Health Information. Unnecessary Care in Canada. Ottawa, ON: CIHI; 2017. [Google Scholar]
  • 2.Smith-Bindman R, Kwan ML, Marlow EC, Theis MK, Bolch W, Cheng SY, et al. Trends in Use of Medical Imaging in US Health Care Systems and in Ontario, Canada, 2000–2016. JAMA. 2019;322(9):843–56. doi: 10.1001/jama.2019.11456 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Oren O, Kebebew E, Ioannidis JPA. Curbing Unnecessary and Wasted Diagnostic Imaging. JAMA. 2019;321(3):245–6. doi: 10.1001/jama.2018.20295 [DOI] [PubMed] [Google Scholar]
  • 4.Lumbreras B, Donat L, Hernández-Aguado I. Incidental findings in imaging diagnostic tests: a systematic review. Br J Radiol. 2010;83(988):276–89. doi: 10.1259/bjr/98067945 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Ganguli I, Simpkin AL, Lupo C, Weissman A, Mainor AJ, Orav EJ, et al. Cascades of Care After Incidental Findings in a US National Survey of Physicians. JAMA Network Open. 2019;2(10):e1913325-e. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Lurie JD, Birkmeyer NJ, Weinstein JN. Rates of Advanced Spinal Imaging and Spine Surgery. Spine. 2003;28(6):616–20. doi: 10.1097/01.BRS.0000049927.37696.DC [DOI] [PubMed] [Google Scholar]
  • 7.Lemmers GPG, van Lankveld W, Westert GP, van der Wees PJ, Staal JB. Imaging versus no imaging for low back pain: a systematic review, measuring costs, healthcare utilization and absence from work. Eur Spine J. 2019;28(5):937–50. doi: 10.1007/s00586-019-05918-1 [DOI] [PubMed] [Google Scholar]
  • 8.Vogel L. Nearly a third of tests and treatments are unnecessary: CIHI. Canadian Medical Association Journal. 2017;189(16):E620–E1. doi: 10.1503/cmaj.1095417 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Smith-Bindman R, Lipson J, Marcus R, al e. Radiation dose associated with common computed tomography examinations and the associated lifetime attributable risk of cancer. JAMA Internal Medicine. 2009;169(22):2078–86. doi: 10.1001/archinternmed.2009.427 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Bora A, Açıkgöz G, Yavuz A, Bulut MD. Computed tomography: Are we aware of radiation risks in computed tomography? Eastern Journal Of Medicine. 2014;19(4):164–8. [Google Scholar]
  • 11.Ivers N, Jamtvedt G, Flottorp S, Young JM, Odgaard-Jensen J, French SD, et al. Audit and feedback: effects on professional practice and healthcare outcomes. Cochrane Database Syst Rev. 2012(6):Cd000259. doi: 10.1002/14651858.CD000259.pub3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Hoffmann TC, Glasziou PP, Boutron I, Milne R, Perera R, Moher D, et al. Better reporting of interventions: template for intervention description and replication (TIDieR) checklist and guide. BMJ: British Medical Journal. 2014;348:g1687. doi: 10.1136/bmj.g1687 [DOI] [PubMed] [Google Scholar]
  • 13.Cochrane Effective Practice and Organisation of Care Working Group. EPOC resources for review authors Oslo, Norway: Norwegian Institute of Public Health; 2021 [updated January 2022. https://epoc.cochrane.org/resources/epoc-resources-review-authors.
  • 14.Foster ED, Deardorff A. Open Science Framework (OSF). J Med Libr Assoc. 2017;105(2):203–6. [Google Scholar]
  • 15.Cochrane Handbook for Systematic Reviews of Interventions version 6.3 (updated February 2022). 6.2 ed. Higgins JPT, Thomas J, Chandler J, Cumpston M, Li T, Page MJ, et al., editors. Chichester (UK): John Wiley & Sons; 2022. [Google Scholar]
  • 16.McGowan J, Sampson M, Salzwedel DM, Cogo E, Foerster V, Lefebvre C. PRESS Peer Review of Electronic Search Strategies: 2015 Guideline Statement. J Clin Epidemiol. 2016;75:40–6. [DOI] [PubMed] [Google Scholar]
  • 17.Aiken M. An Updated Evaluation of Google Translate Accuracy. Studies in Linguistics and Literature. 2019;3:p253. [Google Scholar]
  • 18.Cabana MD, Rand CS, Powe NR, Wu AW, Wilson MH, Abboud PA, et al. Why don’t physicians follow clinical practice guidelines? A framework for improvement. Jama. 1999;282(15):1458–65. doi: 10.1001/jama.282.15.1458 [DOI] [PubMed] [Google Scholar]
  • 19.Covidence systematic review software Veritas Health Innovation, Melbourne, Australia [www.covidence.org.
  • 20.Cochrane Effective Practice and Organisation of Care Review Group. Data Collection Checklist Ottawa, ON: Institute of Population Health, University of Ottawa; 2002 [updated June 2002. https://epoc.cochrane.org/sites/epoc.cochrane.org/files/public/uploads/datacollectionchecklist.pdf.
  • 21.Higgins JPT, Altman DG, Gøtzsche PC, Jüni P, Moher D, Oxman AD, et al. The Cochrane Collaboration’s tool for assessing risk of bias in randomised trials. BMJ. 2011;343:d5928. doi: 10.1136/bmj.d5928 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.RoB 2: A revised Cochrane risk-of-bias tool for randomized trials. https://methods.cochrane.org/bias/resources/rob-2-revised-cochrane-risk-bias-tool-randomized-trials. [DOI] [PubMed]
  • 23.Sterne JAC, Savović J, Page MJ, Elbers RG, Blencowe NS, Boutron I, et al. RoB 2: a revised tool for assessing risk of bias in randomised trials. BMJ. 2019;366:l4898. doi: 10.1136/bmj.l4898 [DOI] [PubMed] [Google Scholar]
  • 24.Review Manager (RevMan) [Computer program]. Version 5.4. 5.4 ed: The Cochrane Collaboration; 2020.
  • 25.Guyatt GH, Thorlund K, Oxman AD, Walter SD, Patrick D, Furukawa TA, et al. GRADE guidelines: 13. Preparing Summary of Findings tables and evidence profiles-continuous outcomes. Journal of Clinical Epidemiology. 2013;66(2):173–83. doi: 10.1016/j.jclinepi.2012.08.001 [DOI] [PubMed] [Google Scholar]
  • 26.Thorlund K, Walter SD, Johnston BC, Furukawa TA, Guyatt GH. Pooling health-related quality of life outcomes in meta-analysis-a tutorial and review of methods for enhancing interpretability. Res Synth Methods. 2011;2(3):188–203. doi: 10.1002/jrsm.46 [DOI] [PubMed] [Google Scholar]
  • 27.Bhatia RS, Dudzinski DM, Malhotra R, Milford CE, Yoerger Sanborn DM, Picard MH, et al. Educational intervention to reduce outpatient inappropriate echocardiograms: a randomized control trial. JACC Cardiovasc Imaging. 2014;7(9):857–66. doi: 10.1016/j.jcmg.2014.04.014 [DOI] [PubMed] [Google Scholar]
  • 28.Bhatia RS, Ivers NM, Yin XC, Myers D, Nesbitt GC, Edwards J, et al. Improving the Appropriate Use of Transthoracic Echocardiography: The Echo WISELY Trial. J Am Coll Cardiol. 2017;70(9):1135–44. doi: 10.1016/j.jacc.2017.06.065 [DOI] [PubMed] [Google Scholar]
  • 29.Dudzinski DM, Bhatia RS, Mi MY, Isselbacher EM, Picard MH, Weiner RB. Effect of Educational Intervention on the Rate of Rarely Appropriate Outpatient Echocardiograms Ordered by Attending Academic Cardiologists: A Randomized Clinical Trial. JAMA Cardiology. 2016;1(7):805–12. doi: 10.1001/jamacardio.2016.2232 [DOI] [PubMed] [Google Scholar]
  • 30.Eccles M, Steen N, Grimshaw J, Thomas L, McNamee P, Soutter J, et al. Effect of audit and feedback, and reminder messages on primary-care radiology referrals: a randomised trial. Lancet. 2001;357(9266):1406–9. doi: 10.1016/S0140-6736(00)04564-5 [DOI] [PubMed] [Google Scholar]
  • 31.Kerry S, Oakeshott P, Dundas D, Williams J. Influence of postal distribution of the Royal College of Radiologists’ guidelines, together with feedback on radiological referral rates, on X-ray referrals from general practice: a randomized controlled trial. Fam Pract. 2000;17(1):46–52. doi: 10.1093/fampra/17.1.46 [DOI] [PubMed] [Google Scholar]
  • 32.O’Connor DA, Glasziou P, Maher CG, McCaffery KJ, Schram D, Maguire B, et al. Effect of an Individualized Audit and Feedback Intervention on Rates of Musculoskeletal Diagnostic Imaging Requests by Australian General Practitioners: A Randomized Clinical Trial. Jama. 2022;328(9):850–60. doi: 10.1001/jama.2022.14587 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Raja AS, Ip IK, Dunne RM, Schuur JD, Mills AM, Khorasani R. Effects of Performance Feedback Reports on Adherence to Evidence-Based Guidelines in Use of CT for Evaluation of Pulmonary Embolism in the Emergency Department: A Randomized Trial. AJR Am J Roentgenol. 2015;205(5):936–40. doi: 10.2214/AJR.15.14677 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Robling MR, Houston HL, Kinnersley P, Hourihan MD, Cohen DR, Hale J, et al. General practitioners’ use of magnetic resonance imaging: an open randomized trial comparing telephone and written requests and an open randomized controlled trial of different methods of local guideline dissemination. Clin Radiol. 2002;57(5):402–7. doi: 10.1053/crad.2001.0864 [DOI] [PubMed] [Google Scholar]
  • 35.Verstappen WH, van der Weijden T, Sijbrandij J, Smeele I, Hermsen J, Grimshaw J, et al. Effect of a practice-based strategy on test ordering performance of primary care physicians: a randomized trial. Jama. 2003;289(18):2407–12. doi: 10.1001/jama.289.18.2407 [DOI] [PubMed] [Google Scholar]
  • 36.Winkens RA, Pop P, Bugter-Maessen AM, Grol RP, Kester AD, Beusmans GH, et al. Randomised controlled trial of routine individual feedback to improve rationality and reduce numbers of test requests. Lancet. 1995;345(8948):498–502. doi: 10.1016/s0140-6736(95)90588-x [DOI] [PubMed] [Google Scholar]
  • 37.Zafar HM, Ip IK, Mills AM, Raja AS, Langlotz CP, Khorasani R. Effect of Clinical Decision Support-Generated Report Cards Versus Real-Time Alerts on Primary Care Provider Guideline Adherence for Low Back Pain Outpatient Lumbar Spine MRI Orders. AJR Am J Roentgenol. 2019;212(2):386–94. doi: 10.2214/AJR.18.19780 [DOI] [PubMed] [Google Scholar]
  • 38.Freeborn DK, Shye D, Mullooly JP, Eraker S, Romeo J. Primary care physicians’ use of lumbar spine imaging tests: effects of guidelines and practice pattern feedback. Journal of general internal medicine. 1997;12(10):619,Äê25. doi: 10.1046/j.1525-1497.1997.07122.x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Berwick DM, Coltin KL. Feedback reduces test use in a health maintenance organization. Jama. 1986;255(11):1450–4. [PubMed] [Google Scholar]
  • 40.Bhatia RS, Milford CE, Picard MH, Weiner RB. An educational intervention reduces the rate of inappropriate echocardiograms on an inpatient medical service. JACC Cardiovasc Imaging. 2013;6(5):545–55. doi: 10.1016/j.jcmg.2013.01.010 [DOI] [PubMed] [Google Scholar]
  • 41.Cammisa C, Partridge G, Ardans C, Buehrer K, Chapman B, Beckman H. Engaging physicians in change: results of a safety net quality improvement program to reduce overuse. Am J Med Qual. 2011;26(1):26–33. doi: 10.1177/1062860610373380 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Halpern DJ, Clark-Randall A, Woodall J, Anderson J, Shah K. Reducing Imaging Utilization in Primary Care Through Implementation of a Peer Comparison Dashboard. J Gen Intern Med. 2021;36(1):108–13. doi: 10.1007/s11606-020-06164-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Morgan T, Wu J, Ovchinikova L, Lindner R, Blogg S, Moorin R. A national intervention to reduce imaging for low back pain by general practitioners: a retrospective economic program evaluation using Medicare Benefits Schedule data. BMC Health Services Research. 2019;19(1):983. doi: 10.1186/s12913-019-4773-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Salehi L, Jaskolka J, Yu H, Ossip M, Phalpher P, Valani R, et al. The impact of performance feedback reports on physician ordering behavior in the use of computed tomography pulmonary angiography (CTPA). Emerg Radiol. 2023;30(1):63–9. doi: 10.1007/s10140-022-02100-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Ross AB, Kalia V, Chan BY, Li G. The influence of patient race on the use of diagnostic imaging in United States emergency departments: data from the National Hospital Ambulatory Medical Care survey. BMC Health Serv Res. 2020;20(1):840. doi: 10.1186/s12913-020-05698-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Ross AB, Rother MDM, Miles RC, Flores EJ, Boakye-Ansa NK, Brown C, et al. Racial and/or Ethnic Disparities in the Use of Imaging: Results from the 2015 National Health Interview Survey. Radiology. 2022;302(1):140–2. doi: 10.1148/radiol.2021211449 [DOI] [PubMed] [Google Scholar]
  • 47.Brehaut JC, Colquhoun HL, Eva KW, Carroll K, Sales A, Michie S, et al. Practice Feedback Interventions: 15 Suggestions for Optimizing Effectiveness. Ann Intern Med. 2016;164(6):435–41. doi: 10.7326/M15-2248 [DOI] [PubMed] [Google Scholar]
  • 48.Michie S, Johnston M, Abraham C, Lawton R, Parker D, Walker A. Making psychological theory useful for implementing evidence based practice: a consensus approach. Qual Saf Health Care. 2005;14(1):26–33. doi: 10.1136/qshc.2004.011155 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Korenstein D, Gillespie EF. Audit and Feedback—Optimizing a Strategy to Reduce Low-Value Care. JAMA. 2022;328(9):833–5. doi: 10.1001/jama.2022.14173 [DOI] [PubMed] [Google Scholar]

Decision Letter 0

Joshua Robert Zadro

28 Nov 2023

PONE-D-23-34965Audit and Feedback to change diagnostic image ordering practices: A systematic review and meta-analysis.PLOS ONE

Dear Dr. Aubrey-Bassler,

Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process. Both reviewers have provided a comprehensive assessment of the paper and raised important points to consider.

Please submit your revised manuscript by Jan 12 2024 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

Please include the following items when submitting your revised manuscript:

  • A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'.

  • A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.

  • An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols. Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols.

We look forward to receiving your revised manuscript.

Kind regards,

Joshua Robert Zadro, PhD

Academic Editor

PLOS ONE

Journal Requirements:

When submitting your revision, we need you to address these additional requirements.

1. Please ensure that your manuscript meets PLOS ONE's style requirements, including those for file naming. The PLOS ONE style templates can be found at 

https://journals.plos.org/plosone/s/file?id=wjVg/PLOSOne_formatting_sample_main_body.pdf and 

https://journals.plos.org/plosone/s/file?id=ba62/PLOSOne_formatting_sample_title_authors_affiliations.pdf

2. Please review your reference list to ensure that it is complete and correct. If you have cited papers that have been retracted, please include the rationale for doing so in the manuscript text, or remove these references and replace them with relevant current references. Any changes to the reference list should be mentioned in the rebuttal letter that accompanies your revised manuscript. If you need to cite a retracted article, indicate the article’s retracted status in the References list and also include a citation and full reference for the retraction notice.

[Note: HTML markup is below. Please do not edit.]

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #1: Yes

Reviewer #2: Yes

**********

2. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #1: No

Reviewer #2: I Don't Know

**********

3. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #1: Yes

Reviewer #2: Yes

**********

4. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #1: Yes

Reviewer #2: Yes

**********

5. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)

Reviewer #1: PONE-D-23-34965

Audit and Feedback to change diagnostic image ordering practices: A systematic review and meta-analysis

Comments to editor

Thank you for the opportunity to review this manuscript.

Comments to author

I thank the authors for providing the opportunity to review this manuscript. The authors aimed to assess the effect of auditing and feedback interventions directed at healthcare workers to reduce imaging orders. The authors ran a comprehensive search and included 11 RCT’s with 4311 participants in total. This review provides an interesting insight into the effects of audit and feedback interventions on the ordering of images in primary care, especially given the high rates of unnecessary imaging currently. There are some matters that will need to be addressed before this review could be considered suitable for publication.

Primarily, the concerns are regarding the way in which the studies have been pooled together. Different conditions will have a lower or higher threshold for further testing and pooling these studies together may not appropriate. For example, studies done on cardiologist will yield very different results to studies in a musculoskeletal setting as one condition (chest pain) would be considered life threatening compared to the other (i.e., knee pain)

Major issues/comments

OVERALL

- My impression is that the authors have looked at the effect of AF intervention to reduce imaging. My concern is that the pooling of studies with different conditions (i.e., musculoskeletal based studies vs cardiovascular studies) includes too much heterogeneity and means the results are not entirely interpretable. This is because the threshold for serious pathology will be different e.g., chest pain will likely get imaged and have a low threshold for further testing, whereas knee pain may have a higher threshold for further testing as it may not be considered life threatening.

ABSTRACT

METHODS

- One overall comment is that it may not be appropriate to pool studies that used diagnostic imaging for different body parts. For example, the threshold for suspicion of serious pathology is going to be different for the lumbar spine verses a pulmonary condition. The setting of which the imaging is ordered is also something that should be pooled separately. Serious pathology is more prevalent in emergency departments settings and clinicians may have a lower threshold to image compared to primary care. Therefore, pooling all the studies together does not seem appropriate due to the heterogeneity between studies. My suggestion would be pool studies based on musculoskeletal conditions, respiratory, cardiovascular etc. Happy for this to be discussed with the Editor, however a meta-analysis inclusive of all body parts seems inappropriate in my view.

RESULTS

Study Characteristics

- No description of participant characteristics e.g., age, gender

Overall comment

- This section is written well on a complicated issue. However as per the comments in the methods section, I do not feel that the way the data is presented accurately highlights the impact of AF interventions, as the effect will differ for different body parts. This is evident for example in Figure 2A Eccles 2001 (musculoskeletal) and Bhatia 2014 (cardiology) and the varying results of these two studies.

DISCUSSION

Strengths and limitations

- If the Editor decides the current meta-analysis is ok, then I think the heterogeneity between the studies need to be clearly stated. It needs to be clear that the pooling of different conditions adds to the heterogeneity in the data and means the interpretation of these results should be taken very cautiously. Again, my recommendation is to reconsider the methods.

Reviewer #2: Thank you for the opportunity to review this manuscript titled “Audit and Feedback to change diagnostic image ordering practices: A systematic review and meta-analysis.”. This is an important topic and will be of interest to administrators and policy makers looking to reduce overuse of imaging which wastes resources and can worsen patient outcomes. This is a high quality systematic review that follows EPOC recommendations, uses the TIDieR statement describe interventions and was prospectively registered.

**********

6. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #1: No

Reviewer #2: Yes: Gemma Altinger

**********

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at figures@plos.org. Please note that Supporting Information files do not need this step.

Attachment

Submitted filename: PONE-D-23-34965_Review_CH.docx

pone.0300001.s003.docx (19.4KB, docx)
Attachment

Submitted filename: PONE-D-23-34965 Reviewer Comments.docx

pone.0300001.s004.docx (16.3KB, docx)
PLoS One. 2024 Jun 5;19(6):e0300001. doi: 10.1371/journal.pone.0300001.r002

Author response to Decision Letter 0


8 Feb 2024

Please see our response attached to this portal. We have also provided it here for convenience.

February 7, 2024

Dear Editors,

We thank the reviewers for providing positive feedback and the opportunity to revise our manuscript. In response to the comments, suggestions, and guidance received from the reviewers, we have modified our manuscript and addressed the feedback. We hope our revisions will justify publication. Please note that we have grouped reviewer comments that are related and require a similar response, and we have not listed positive reviewer comments that did not require a response from us. Additionally, please note that references 5&6 have been added to our manuscript as they are more recent than the previous ones, and references 22 & 23 have been included to notify of the change from the Cochrane risk bias tool version 1 to the updated version 2.

RESPONSE TO REVIEWER 1:

Comments to author

-Primarily, the concerns are regarding the way in which the studies have been pooled together. Different conditions will have a lower or higher threshold for further testing and pooling these studies together may not appropriate. For example, studies done on cardiologist will yield very different results to studies in a musculoskeletal setting as one condition (chest pain) would be considered life threatening compared to the other (i.e., knee pain)

Comment 2:

OVERALL

- My impression is that the authors have looked at the effect of AF intervention to reduce imaging. My concern is that the pooling of studies with different conditions (i.e., musculoskeletal based studies vs cardiovascular studies) includes too much heterogeneity and means the results are not entirely interpretable. This is because the threshold for serious pathology will be different e.g., chest pain will likely get imaged and have a low threshold for further testing, whereas knee pain may have a higher threshold for further testing as it may not be considered life threatening.

Comment 3:

METHODS

- One overall comment is that it may not be appropriate to pool studies that used diagnostic imaging for different body parts. For example, the threshold for suspicion of serious pathology is going to be different for the lumbar spine verses a pulmonary condition. The setting of which the imaging is ordered is also something that should be pooled separately. Serious pathology is more prevalent in emergency departments settings and clinicians may have a lower threshold to image compared to primary care. Therefore, pooling all the studies together does not seem appropriate due to the heterogeneity between studies. My suggestion would be pool studies based on musculoskeletal conditions, respiratory, cardiovascular etc. Happy for this to be discussed with the Editor, however a meta-analysis inclusive of all body parts seems inappropriate in my view.

Comment 5:

Overall comment

- However as per the comments in the methods section, I do not feel that the way the data is presented accurately highlights the impact of AF interventions, as the effect will differ for different body parts. This is evident for example in Figure 2A Eccles 2001 (musculoskeletal) and Bhatia 2014 (cardiology) and the varying results of these two studies.

Comment 6:

DISCUSSION

Strengths and limitations

- If the Editor decides the current meta-analysis is ok, then I think the heterogeneity between the studies need to be clearly stated. It needs to be clear that the pooling of different conditions adds to the heterogeneity in the data and means the interpretation of these results should be taken very cautiously. Again, my recommendation is to reconsider the methods.

Response 1:

We sincerely thank the reviewer for their time and effort reviewing our manuscript. We acknowledge the point raised about appropriate grouping of studies and we have put a lot of thought into this issue. Our results for the primary outcome of total DI utilization are relatively homogeneous (I2=45%, Figure 2A). The primary source of heterogeneity for the pooled analyses (Bhatia 2014) examines the use of echocardiography in cardiology learners and is very similar to two other papers that examine the use of echo, mostly in cardiologists (Dudzinski 2016 and Bhatia 2017), both of which had outcomes that were similar to the pooled result. i.e. the main heterogeneity lies within one of the proposed pools rather than between pools. Given this issue and the small number of studies included in this review, we worry that adding additional layers of pooling will complicate data interpretation and data presentation may suffer; we therefore propose to maintain the pooling as it currently stands. In support of this approach, the 2012 Ivers Cochrane review (1), the largest systematic review on the subject of audit and feedback (AF), grouped studies of healthcare AF interventions with substantially more variability than our own, ranging from chronic disease management to antibiotic prescribing and diagnostic imaging utilization.

We have added several lines summarizing the rationale provided here to our paper (Lines 213-217)

1. Ivers N, Jamtvedt G, Flottorp S, Young JM, Odgaard-Jensen J, French SD, et al. Audit and feedback: effects on professional practice and healthcare outcomes. Cochrane Database Syst Rev. 2012(6):Cd000259.

RESULTS

Comment 4:

Study Characteristics

- No description of participant characteristics e.g., age, gender

Response 2:

Thank you for bringing this to our attention. We have now included participant and patient age and gender when available in Table 2.

RESPONSE TO REVIEWER 2:

Reviewer #2

Comment 1:

Introduction:

Well written, clear, and relevant.

Line 5 – I think it would be more powerful to mention the types of unnecessary treatments, e.g. surgery and opioid use in the context of back pain/ The treatment cascade can have significant risk of harm, particularly if it arose from unnecessary testing. This could be emphasised.

Line 12 – Could you add the GRADE rating/certainty of evidence for the Cochrane review here?

Response 3:

Thank you for giving us a chance to strengthen our introduction. From lines 4-9, the impact of unnecessary treatments has been further emphasized. The GRADE rating of moderate has been included for the Cochrane review on line 16.

Comment 4:

Methods:

Lines 84 to 87 – Could you please expand further on how you judged appropriateness, e.g. which guidelines you compared the practice to, to decide if it was appropriate. Or, did the review authors reply on what the trial authors deemed appropriate? This could be more clear.

Response 4:

This has been clarified in the Methods on line 90 to explain that appropriateness was judged by the individual study authors. Thank you for bringing this to our attention. We did not have access to individual DI test ordering data that would have been necessary to gauge appropriateness ourselves.

Comment 4:

Line 109 – Please specify if you used the Cochrane ROB tool 1 or 2

Response 4:

This has been specified on line 113.

Comment 5:

Lines 110 to 112 – Significant differences in base characteristics would be by nature due to random chance, assuming randomisation procedure was appropriate. Issues with randomisation would be picked up with the Cochrane ROB tool, therefore I don’t think this extra step is necessary.

Response 5:

Thank you for giving us the opportunity to explain how this downgrade was decided on. We downgraded the evidence in studies because of baseline imbalance that was “greater than expected due to chance.” The Cochrane collaboration allows for results to be downgraded for this reason (Cochrane Handbook, Section 8.3.2).

One of the papers in our review reported a significant benefit of AF in reducing DI ordering in terms of a change from baseline measure (comment in text only - data not shown), but the post-intervention ordering data which was presented suggested that the effect of AF was essentially neutral (ordering in fact slightly increased in the AF group). We not only downgraded the risk of bias assessment for this paper, but we also decided to exclude this result from out meta-analysis, both of which were justified in our opinion. Excluding the result from the meta-analysis would have been harder to justify had we not downgraded the evidence for baseline imbalance.

Comment 6:

Line 115 – In similar reviews, lack of a pre-published protocol is considered a risk for selective reporting. Could you please elaborate on why you did not consider this to be a risk?

Response 6:

Because our primary outcome was objective, we decided not to consider this a risk. We have added a comment on line 120 to clarify this.

Comment 7:

Results:

Line 212 – Given the studies ranges from very low-moderate quality, consider clarifying that it is a low to moderate certainty evidence that AF can reduce imaging.

Response 7:

This has been added on lines 221-223, thank you for this suggestion.

Comment 8:

Line 233 – word “to” not required

Response 8:

This has been fixed, thank you for alerting us.

Comment 9:

Discussion:

Line 280-287 - Another implication that could be considered is how AF may reduce imaging for those who are typically underserviced, increasing harm to these patients. For instance, interventions and recommendations to reduce antibiotics for all children with otitis media can harm Indigenous children who are underserviced and under treated. Are there certain populations that could experience fewer DI than is clinically indicated?

Response 9:

This is a very important point to consider, and thank you for bringing this to our attention. This has been added to the discussion (line 296), as it has been documented that people of colour receive less diagnostic imaging than their white counterparts, and could be disproportionately affected by a reduction of appropriate imaging tests.

Comment 10:

Line 291 – It is interesting that “rarely appropriate” imaging requests had significant effects. You have provided the explanation of frequency, but could there also be clinician agreement and endorsement of the inappropriateness of imaging in those cases which increased the effect?

Response 10

Thank you for allowing us to explain our rationale. Any reduction in inappropriateness should result in an increase in appropriateness (when expressed as proportions), unless providers only changed their documentation without any change to their ordering practices. However, if this was the case, we would expect total ordering (our primary outcome - figure 2) to be unchanged, which was not the case.

Alternatively the reviewer may have understood that appropriateness was determined in part by the ordering clinicians; however, that is not the case - in all studies, appropriateness was determined retrospectively by (usually) blinded study personnel reviewing ordering materials. Please feel free to have the reviewer clarify this comment so that we might address it better.

Comment 11:

Line 315 – Other reasons a supervisor of colleague may be used is because this can threaten professional status and peer opinion. Clinicians may be motivated by wanting to maintain their reputation.

Response 11:

Thank you for this suggestion as this likely another motivating factor for physicians when being engaged in AF through their supervisor. This has been added to the discussion from lines 330-331.

Comment 13:

Tables and Figures:

Figure 1 and 2 have AF on the left, but figure 3 has AF on the right. At a glance, this looks like the effects were in the opposite direction. If possible, keep consistent across figures.

Otherwise tables and figures are excellent.

Response 13:

This is a great point as this may confuse readers. A favourable primary outcome was considered to be negative (reduction in total image ordering) whereas a favourable secondary outcome is positive (increase in appropriateness). We debated presenting the data as the reviewer suggests, but elected not to as that may cause confusion for other readers. This explanation has been added to the legend of Figure 3 and bolded to clear up any confusion for readers.

Comment 14:

Please submit PRISMA checklist.

Response 14:

An older version of the PRISMA checklist was attached with our previous submission, but we have now included the most up to date version. Several changes were made to the abstract to meet the PRISMA abstract checklist and to keep the length less than 300 words.

Sincerely,

Kris

Kris Aubrey-Bassler

Associate Professor and Director

On behalf of the co-authors

Attachment

Submitted filename: Response to Reviewers_v3.docx

pone.0300001.s005.docx (56.8KB, docx)

Decision Letter 1

Joshua Robert Zadro

20 Feb 2024

Audit and Feedback to change diagnostic image ordering practices: A systematic review and meta-analysis.

PONE-D-23-34965R1

Dear Dr. Aubrey-Bassler,

We’re pleased to inform you that your manuscript has been judged scientifically suitable for publication and will be formally accepted for publication once it meets all outstanding technical requirements.

Within one week, you’ll receive an e-mail detailing the required amendments. When these have been addressed, you’ll receive a formal acceptance letter and your manuscript will be scheduled for publication.

An invoice for payment will follow shortly after the formal acceptance. To ensure an efficient process, please log into Editorial Manager at http://www.editorialmanager.com/pone/, click the 'Update My Information' link at the top of the page, and double check that your user information is up-to-date. If you have any billing related questions, please contact our Author Billing department directly at authorbilling@plos.org.

If your institution or institutions have a press office, please notify them about your upcoming paper to help maximize its impact. If they’ll be preparing press materials, please inform our press team as soon as possible -- no later than 48 hours after receiving the formal acceptance. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

Kind regards,

Joshua Robert Zadro, PhD

Academic Editor

PLOS ONE

Additional Editor Comments (optional):

I thank the authors for adequately addressing the reviewers comments and congratulate them on putting together a very important paper.

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. If the authors have adequately addressed your comments raised in a previous round of review and you feel that this manuscript is now acceptable for publication, you may indicate that here to bypass the “Comments to the Author” section, enter your conflict of interest statement in the “Confidential to Editor” section, and submit your "Accept" recommendation.

Reviewer #2: All comments have been addressed

**********

2. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #2: Yes

**********

3. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #2: I Don't Know

**********

4. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #2: Yes

**********

5. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #2: Yes

**********

6. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)

Reviewer #2: (No Response)

**********

7. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #2: Yes: Gemma Altinger

**********

Acceptance letter

Joshua Robert Zadro

4 Mar 2024

PONE-D-23-34965R1

PLOS ONE

Dear Dr. Aubrey-Bassler,

I'm pleased to inform you that your manuscript has been deemed suitable for publication in PLOS ONE. Congratulations! Your manuscript is now being handed over to our production team.

At this stage, our production department will prepare your paper for publication. This includes ensuring the following:

* All references, tables, and figures are properly cited

* All relevant supporting information is included in the manuscript submission,

* There are no issues that prevent the paper from being properly typeset

If revisions are needed, the production department will contact you directly to resolve them. If no revisions are needed, you will receive an email when the publication date has been set. At this time, we do not offer pre-publication proofs to authors during production of the accepted work. Please keep in mind that we are working through a large volume of accepted articles, so please give us a few weeks to review your paper and let you know the next and final steps.

Lastly, if your institution or institutions have a press office, please let them know about your upcoming paper now to help maximize its impact. If they'll be preparing press materials, please inform our press team within the next 48 hours. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

If we can help with anything else, please email us at customercare@plos.org.

Thank you for submitting your work to PLOS ONE and supporting open access.

Kind regards,

PLOS ONE Editorial Office Staff

on behalf of

Dr. Joshua Robert Zadro

Academic Editor

PLOS ONE

Associated Data

    This section collects any data citations, data availability statements, or supplementary materials included in this article.

    Supplementary Materials

    S1 Appendix

    S1 Fig. a. Effect of audit and feedback in observational studies on the number of diagnostic imaging requests (continuous outcome) (4–6). b. Effect of audit and feedback in observational studies on the number of diagnostic imaging requests (dichotomous outcome) (7, 8). S2 Fig. Effect of audit and feedback in observational studies on image order appropriateness (dichotomous outcome) (7). S3 Fig. Funnel plot of RCTs analyzing the total image order outcome. We did not consider this figure to be indicative of publication bias. The study in the bottom right favored the control intervention, not AF. S4 Fig. Funnel plot of RCTS analyzing the appropriateness of image orders outcome.We did not consider this figure to be indicative of publication bias. S1 Table. Description of AF interventions using TiDIER recommendations (1). Abbreviations: AF, Audit and Feedback; CT, Computed Tomography; Echo, Echocardiography; GIM, General physicians; Res, residents; Gov., Government; Mm; MRI, Magnetic Resonance Imaging; N/A, not applicable; PCP, Primary care physicians (e) PCPs refers to primary care physicians and may include family, general practice and general internal medicine physicians, (f) The term residents also refers to registrars (g) Comparison provided Includes own/ peers’ previous performance, national benchmark. Note: For multifaceted interventions, we assessed the characteristics of the audit and feedback component. S2 Table. a. Risk of Bias for NRCTs using the Risk Of Bias In Non-randomized Studies—of Interventions (ROBINS-I) tool (2). b. Risk of Bias for observational studies using Effective Practice and Organisation of Care (EPOC) recommendations (3). c. Risk of Bias for interrupted time series studies using Effective Practice and Organisation of Care (EPOC) recommendations (3). Legend: Low risk; Indeterminate Risk; High risk. S3 Table. Effect of audit and feedback in a non-randomized, crossover design study on the number of diagnostic imaging request 9).*no p-values, standard deviations, or confidence intervals provided. S4 Table. Effect of audit and feedback with another intervention vs usual care in an interrupted time-series analysis (10). S1 File. EMBASE search strategy. S2 File. CINAHL search strategy. S3 File. PubMED search strategy. S4 File. References for supporting information.

    (ZIP)

    pone.0300001.s001.zip (5.6MB, zip)
    S1 Checklist. PRISMA checklist.

    (DOCX)

    pone.0300001.s002.docx (30KB, docx)
    Attachment

    Submitted filename: PONE-D-23-34965_Review_CH.docx

    pone.0300001.s003.docx (19.4KB, docx)
    Attachment

    Submitted filename: PONE-D-23-34965 Reviewer Comments.docx

    pone.0300001.s004.docx (16.3KB, docx)
    Attachment

    Submitted filename: Response to Reviewers_v3.docx

    pone.0300001.s005.docx (56.8KB, docx)

    Data Availability Statement

    All relevant data are within the manuscript and its Supporting information files.


    Articles from PLOS ONE are provided here courtesy of PLOS

    RESOURCES