Skip to main content
PLOS One logoLink to PLOS One
. 2025 Jun 25;20(6):e0324684. doi: 10.1371/journal.pone.0324684

Prevalence of symptom exaggeration among North American independent medical evaluation examinees: A systematic review of observational studies

Andrea J Darzi 1,2,3, Li Wang 2,3,4, John J Riva 1, Rami Z Morsi 5, Rana Charide 1, Rachel J Couban 2, Samer G Karam 1, Kian Torabiardakani 2,3, Annie Lok 3, Shanil Ebrahim 1, Sheena Bance 6, Regina Kunz 7, Gordon H Guyatt 1, Jason W Busse 1,2,3,4,*
Editor: Thiago P Fernandes8
PMCID: PMC12193048  PMID: 40561085

Abstract

Background

Independent medical evaluations (IMEs) are commonly acquired to provide an assessment of impairment; however, these assessments show poor inter-rater reliability. One potential contributor is symptom exaggeration by patients, who may feel pressure to emphasize their level of impairment to qualify for incentives. This study explored the prevalence of symptom exaggeration among IME examinees in North America, which if common may represent an important consideration for improving the reliability of IMEs.

Methods

We searched CINAHL, EMBASE, MEDLINE and PsycINFO from inception to July 08, 2024. We included observational studies that used a known-group design or multi-modal determination method. Paired reviewers independently assessed risk of bias and extracted data. We performed a random-effects model meta-analysis to estimate the overall prevalence of symptom exaggeration and explored potential subgroup effects for sex, age, education, clinical condition, and confidence in the reference standard. We used the GRADE approach to assess the certainty of evidence.

Results

We included 44 studies with 46 cohorts and 9,794 patients. The median of the mean age was 40 (interquartile range [IQR] 38–42). Most cohorts included patients with traumatic brain injuries (n = 31, 67%) or chronic pain (n = 11, 24%). Prevalence of symptom exaggeration across studies ranged from 17% to 67%. We found low certainty evidence suggesting that studies with a greater proportion of women (≥40%) may be associated with higher rates of exaggeration (47%, 95%CI 36–58) vs. studies with a lower proportion of women (<40%) (31%, 95%CI 28–35; test of interaction p = 0.02). Possible explanations include biological differences, greater bodily awareness, or higher rates of negative affectivity. We found no significant subgroup effects for type of clinical condition, confidence in the reference standard, age, or education.

Conclusion

Symptom exaggeration may occur in almost 50% of women and in approximately a third of men undergoing IMEs. The high prevalence of symptom exaggeration among IME attendees provides a compelling rationale for clinical evaluators to formally explore this issue. Future research should establish the reliability and validity of evaluation criteria for symptom exaggeration and develop a structured IME assessment approach.

Background

In 2022, Statistics Canada found that 8.0 million Canadian adults reported a disability [1] and in 2020, 64.4 million Americans reported living with disability [2]. Individuals suffering from a disabling injury or illness may be eligible to receive financial compensation and services based on their level of impairment. Determinations of impairment often rely on independent medical evaluations (IMEs), which are requested by a third party, such as an insurance company or employer, and conducted by a clinician who is not part of the patient’s regular medical team [3]. Underlying this process is the concern that treating clinicians may have difficulty providing impartial assessments of their patients [4,5]. Such concerns are supported by a trial that randomized 5,888 individuals in Norway to an independent assessment or usual care and found 29% of IMEs recommended less sick leave than the treating physician (68% the same, and 3% a longer duration) [6].

Despite their widespread use and far-reaching consequences, the consistency and reliability of IMEs has been challenged. The most recent systematic review found that clinical experts assessing the same patients often dissented on whether they were disabled from working (median inter-rater reliability 0.45) [7]. Although this review suggested that standardization of the assessment process may improve the reliability of IMEs, [7] two subsequent studies failed to support this hypothesis [8]. Another potential source of variability in IME assessments is symptom exaggeration [3]. IME assessors may focus too narrowly on a biomedical model to explain symptoms, without giving sufficient attention to psychosocial and work-related factors that may influence how individuals present their symptoms [3,9].

Patients referred for IMEs often present with subjective complaints (e.g., mental illness, chronic pain) and may feel pressure to emphasize their level of impairment to qualify for wage replacement benefits, receiving time off work, or other incentives [3,10,11]. Patients’ presentation may also be affected if they perceive the assessor as representing the referring agency rather than their interests. Whether or not IME assessors consider symptom exaggeration has the potential to lead to very different conclusions; however, the prevalence of exaggeration among IME attendees is uncertain and individual studies report rates as low as 17% [12] or as high as 67% [13]. Also, terminology such as exaggeration, malingering, or over-reporting are defined inconsistently across studies, making it difficult to distinguish intentional deception from psychological amplification of distress [4,14]. We undertook the first systematic review of observational studies to explore the prevalence of symptom exaggeration among IME examinees in North America.

Methods

We conducted our systematic review in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) and Meta-analysis of Observational Studies in Epidemiology (MOOSE) checklists [15,16]. (See S1 and S2 Checklists in the supplemental material) We registered our protocol on the Open Science Framework (Registration DOI: https://doi.org/10.17605/OSF.IO/64V2B) [17]. After registration but prior to data analysis, we included five meta-regressions/subgroup analyses to explore variability among studies reporting the prevalence of symptom exaggeration: (1) proportion of female participants, (2) older age, (3) level of formal education, (4) clinical condition, and (5) level of confidence in the reference standard used in the approach for evaluating symptom exaggeration.

Data sources and searches

An experienced medical librarian (RJC) developed database-specific search strategies (S1 Table) and conducted a systematic search in CINAHL, EMBASE, MEDLINE and PsycINFO, from inception through July 08, 2024. We included English, French or Spanish studies to reduce language bias. The search strategies were developed using a validation set of known relevant articles and included a combination of MeSH headings and free text key words, such as malinger* or litigation or litigant or “insufficient effort” and “independent medical examination” or “independent medical evaluation” or “disability” or “classification accuracy”. We did not use any filters for our searches to maximize sensitivity. We screened the reference lists of all included studies for additional eligible articles.

Study selection

Six reviewers screened the titles and abstracts of all retrieved citations, independently and in duplicate, and subsequently the full texts of potentially eligible studies, using standardized and pre-tested forms [18]. A third senior reviewer resolved disagreements when necessary.

Eligible studies: (i) enrolled individuals presenting for an IME in North America, (ii) in the presence of external incentive (e.g., insurance claims), and (iii) assessed the prevalence of symptom exaggeration using a known group design or multi-modal determination method [19,20]. As there is no singular reliable and valid criteria (reference standard) in the literature that is used to assess for symptom exaggeration, we included known group study designs that defined their reference standard based on criteria incorporating both clinical findings and performance on psychometric testing to classify individuals as exaggerating (within diagnostic test terminology, the target positive group), or not exaggerating (the target negative group) their symptoms [21,22].

Examples of two commonly used known group designs are the Slick, Sherman, and Iverson criteria for malingered neurocognitive dysfunction [23] and the Bianchini, Greve, & Glynn criteria for malingered pain-related disability [24]. We excluded studies that used only beyond-chance scores on symptom validity tests as an indicator of symptom exaggeration, since beyond-chance scores are infrequent and likely to result in underestimates [2527]. We restricted our focus to North America as there may be important differences between IMEs conducted within North America where social insurance for disability is limited and Europe where social insurance is prominent. In cases where multiple studies had population overlap, we included only the study with the larger sample size.

Data extraction and risk of bias assessment

Teams of paired reviewers abstracted data independently and in duplicate from all eligible studies using standardized, pre-tested forms. We prefaced data abstraction with calibration exercises to optimize consistency and accuracy of extractions. For all identified studies, the reviewers abstracted the following data: name of first author, year of publication, participant demographics, referral source(s), criteria for establishing symptom exaggeration and reference standard, and the prevalence of symptom exaggeration. After completing training and calibration, pairs of reviewers independently evaluated risk of bias for each included study. They used key criteria tailored to known-group designs, which were developed and pre-tested in collaboration with research methodologists. These criteria included: (i) representativeness of the study population, (ii) validity of outcome assessment (including whether the index test was administered without knowledge of the reference standard, and confidence in the reference standard), (iii) whether those with and without symptom exaggeration were similar across age groups and education level, and (iv) loss to follow-up (≥20% was considered high risk of bias). The response options for all the above risk of bias items included “definitely yes”, “probably yes”, “probably no” and “definitely no”. Also, we evaluated whether the criteria for establishing symptom exaggeration had been shown reliable and valid. We resolved disagreements by consensus or with the help of a third senior reviewer.

We categorized the reference standard and rated our confidence in it as either: (i) ‘weak’ when the study declared a known-group design, however its only criterion for identifying symptom exaggeration was below-chance performance on forced-choice symptom validity testing without any corroborating clinical observations or inconsistencies in medical records. For example, a patient with a mild ankle sprain labeled as exaggerating exclusively because they failed a below‐chance forced‐choice test of pain threshold, with no clinical exam or review of documented pain or functional abilities; (ii) ‘moderate’ where most patients exaggerating symptoms were identified by forced symptom validity testing results, but some cases could be confirmed using other credible indicators. For example, a claimant insists they cannot remember simple details of their daily routine (e.g., the route to their kitchen), yet is casually observed navigating complex tasks with no apparent cognitive difficulty; or (iii) ‘strong’ where exaggeration was determined by either forced symptom validity testing results or other credible clinical evidence. For example, a clinical finding that would classify a patient presenting with persistent post-concussive complaints after a very mild head injury as exaggerating symptoms would include claims of remote memory loss (e.g., loss of spelling ability).

Data synthesis and analysis and certainty in the evidence assessment

We used a random-effects model to pool data for the prevalence of symptom exaggeration among IME examinees and a Freeman-Tukey double arcsine transformation to stabilize the variance [28,29]. This transformation avoids producing confidence intervals (CIs) that include values lower than 0% or greater than 100% [28,29]. We used the DerSimonian and Laird method [30] to pool estimates of symptom exaggeration based on the transformed values and their variances, and then the harmonic mean of sample sizes for back‐transformation to the original units of proportions [31].

We assessed the certainty of evidence based on the Grading of Recommendations Assessment, Development, and Evaluation (GRADE) approach [32]. This approach considers risk of bias, indirectness, inconsistency, imprecision, and small study effects, to appraise the overall certainty of evidence as high, moderate, low, or very low [32]. We estimated that if 20% of IME attendees presented with symptom exaggeration, that would be sufficiently frequent to justify formal evaluation for exaggeration by IME evaluators. Therefore, we rated down for imprecision if the 95%CI associated with the prevalence of symptom exaggeration included 20%. When there were at least 10 studies contributing to meta-analysis, we evaluated small study effects by visual inspection of the funnel plot for asymmetry and calculation of Egger’s test [33].

Subgroup analyses, meta-regression, and sensitivity analyses

We assessed heterogeneity across studies contributing to our pooled estimate of symptom exaggeration using both a statistical test and visual inspection of forest plots. We did not calculate I2 as it can be misleading in cases where the estimates of precision are very narrow due to large sample sizes. Instead, we estimated the between-study variance with tau-squared (τ2), which provides an absolute measure of heterogeneity. We considered τ2 < 0.05 as low, between 0.05–0.1 as moderate, and >0.1 as substantial heterogeneity [34].

We assessed the variability between studies based on five hypotheses. We assumed a higher prevalence of symptom exaggeration with: (1) greater strength of the reference standard, (2) higher proportion of female participants, (3) older age, (4) lower level of formal education, and (5) higher risk of bias on a component-by-component basis. We also explored for subgroup effects based on type of clinical condition but did not pre-specify an anticipated direction of association. We conducted subgroup analyses if there were two or more studies in each subgroup, and evaluated credibility of significant subgroup effects using ICEMAN criteria [35].

We performed meta-regression to explore the relationship between the proportion of women, severity of the presenting complaint, mean age, and years of formal education, with the prevalence of symptom exaggeration. If meta-regression suggested an association, we used visual inspection of the associated scatterplot to estimate a threshold and conducted subgroup analysis. We performed all analyses using Stata software version 16.0 [36]. All comparisons were two-tailed, with a threshold P-value of 0.05.

Ethics approval and consent to participate

We did not require ethics approval for this systematic review and meta-analysis due to our sole use of already published data.

Systematic review update

Considering the speed at which studies exploring the prevalence of symptom exaggeration among IME attendees are published, we plan to update this review within the next five years [37].

Results

Of 20,405 unique citations identified in our search, 44 English-language studies that reported on 46 cohorts and 9,794 patients were eligible for review. (Fig 1). None of the studies had overlapping cohorts. In S5 Table we detail the included and excluded studies with reasons at full text screening. Of the 46 cohorts, 67% (n = 31) reported on patients with traumatic brain injuries (TBI) with or without mixed neurological diseases, 24% (n = 11) on chronic pain patients, and 9% (n = 4) on other populations including toxic exposure (n = 1) [38], personal injury claimants that were not described (n = 1) [39], patients with memory impairment (n = 1) [13] and claimants reporting cognitive dysfunction following exposure to occupational and environmental substances (n = 1) [40]. In terms of criteria used to identify individuals who were exaggerating symptoms, 61% (n = 28) of cohorts relied on the Slick, Sherman, and Iverson criteria for probable malingered neurocognitive dysfunction [23], 24% (n = 11) on the Bianchini criteria [24], and 15% (n = 7) used other criteria such as those proposed by Greiffenstein, Gola, and Baker [41], Nies and Sweet [22] or Lees-Haley methods [42] (Table 1).

Fig 1. PRISMA flow chart.

Fig 1

Table 1. Study Characteristics.

First author, year Study Design (N) Sampling Method Study population Age (mean) % Female Education (mean/
years)
Method used to assess symptom exaggeration
Lees-Haley, 1991 [ 39 ] Prospective cohort
(N = 45)
NR Personal Injury claimants 37.8 58% NR Credibility scale: 100-item true-false scale that provides a sample of the claimant’s behavior during the evaluation session and compares that behavior with both the professional experience of the clinician and with the scores of normative samples.
Greiffenstein, 1995 [ 41 ] Retrospective cohort
(N = 177)
Consecutive TBI patients 35.4 NR 12.1 4 Criteria: (1) improbable symptom histories, (2) improbably poor neuropsychological test scores not accounted for by physical or sensory limitations, (3) claims of subjective remote memory loss and (4) total disability in at least one major social role
* Two or more of the features had to be present.
Suhr, 1997 [ 43 ] Retrospective cohort
(N = 96)
NR TBI patients 35.3 46% 13.2 Criteria by Greiffenstein et al. (1994) a [41] with a modification that poor performance on neuropsychological tests could not be used solely to make a definitive criterion for symptom exaggeration.
Costa, 1999 [ 13 ] Retrospective cohort
(N = 42)
Consecutive Patients with memory impairment 40.7 55% 11.5 Criteria: (1) below chance scores on Victoria Test, (2) less than two rows on Rey-15, (3) digits forward <=3 and digits backward >=4, (4) endorsement of one or more of the following improbable procedural memory deficits: forgetting the order of walking, chewing, and swallowing movements, forgetting how to speak, constantly forgetting the way home or in own home, (5) endorsement of one or more implausible/incorrect items judged to be easy even for the moderately impaired patients with genuine amnesia, and (6) evidence of current, gainful employment
* One or more of the criteria had to be present.
Van Gorp, 1999 [ 44 ] Retrospective cohort
(N = 81)
NR TBI patients 36.6 NR 13.3 Criteria: (1) improbable symptom history; (2) total disability in work or a major social role after 1 year from a mild closed-head injury in which loss of consciousness was less than 1 hr; (3) claims of remote or autobiographical memory loss; and (4) at least one failure on one or more neuropsychological malingering tests
* One or more of the criteria had to be present.
Sweet, 2000 [ 45 ] Retrospective cohort
(N = 63)
NR TBI patients 37.4 NR 13.8 Criteria: (1) poor effortful performance on MDMT and/or Rey 15 Item, (2) evidence of insufficient effort on one or more traditional neuropsychological measures for which valid criteria have been established, (3) plus absence of credible history of neurotrauma, blatant discrepancy between potential injury and patient complaints, (4) blatant discrepancy between type of potential disorder and neuropsychological presentation, and (5) exaggerated patient presentation within a context of litigation or disability application
* More than one of the above criteria and lacking a plausible alternative explanation for patient behavior had to present
Greve, 2003 [ 46 ] Retrospective cohort (N = 151) NR TBI patients 36.6 34% 12.8 Criteria by Slick et al. (1999) b [23]
Lu, 2003 [ 47 ] Retrospective cohort (N = 128) Consecutive TBI and mixed neurological conditions 42.5 44% 12.9 Criteria by Greiffenstein et al. 1994 [41] and van Gorp et al., 1999 [44]: (1) involvement in litigation or seeking to obtain or maintain disability benefits for reported symptoms and impairments at the time of evaluation, (2) evidence of noncredible cognitive symptoms drawn from at least two of six tests designed to discreetly assess motivation and cooperation, and (3) at least one of six “external” criteria or behavioral presentations that are often observed by clinicians as signs of noncredible symptomatology
* All three criteria had to be present.
Barrash, 2004 [ 48 ] Retrospective cohort (N = 108) NR TBI and mixed neurological conditions 45.2 54% 12.9 Criteria by Greiffenstein et al. 1995 [41]: (1) Minimal brain injury (loss of consciousness and post traumatic amnesia <5min; No CT/MRI indications of brain injury and Glasgow Coma Scale>=13), (2) Clear issues of secondary gain (financial compensation, formal accommodations in work or school setting and adjudication issues), and (3) Evidence of dissimulation independent of neuropsychological performances, as indicated by at least two of the following: marked disability in a major psychosocial role, contradiction between patient and collateral sources of information, complaints of remote memory loss or other symptoms that are rarely seen as a consequence of mild head injury.
Heinly, 2005 [ 49 ] Retrospective cohort (N = 344) NR TBI patients 39.6 30% 12.1 Criteria by Slick et al. (1999) b [23]
Curtis, 2006 [ 50 ] Retrospective cohort
(N = 275)
NR TBI patients 38.7 28% 12.3 Criteria by Slick et al. (1999) b [23]
Etherton, 2006a [ 51 ] Retrospective cohort
(N = 81)
NR Chronic pain patients 43.3 36% 11.9 Criteria by Bianchini et al. (2005) c [24]
Greve, 2006a [ 52 ] Not reported
(N = 259)
NR TBI patients 38.7 29% 12.5 Criteria by Slick et al. (1999) b [23]
Greve, 2006b [ 53 ] Not reported
(N = 161)
NR TBI patients 39.3 27% 12.3 Criteria by Slick et al. (1999) b [23]
Greve, 2006c [ 54 ] Not reported
(N = 262)
NR TBI patients 38.3 27% 12.2 Criteria by Slick et al. (1999) b [23]
Greve, 2006d [ 40 ] Retrospective cohort
(N = 128)
NR Cognitive dysfunction upon exposure to occupational and environmental substances 40.8 28% 12 Criteria by Slick et al. (1999) b [23]
Ardolf, 2007 [ 55 ] Retrospective cohort
(N = 105)
NR TBI and mixed neurological conditions 40.1 0% 10.5 Criteria by Slick et al. (1999) b [23]
Greve, 2007 [ 56 ] Prospective cohort
(N = 206)
NR TBI patients 39.0 30% 12.6 Criteria by Slick et al. (1999) b [23]
Henry, 2007 [ 57 ] Retrospective cohort
(N = 54)
NR TBI and mixed neurological conditions 39.8 46% 14.3 Criteria by Slick et al. (1999) b [23]
O’Bryant, 2007 [ 58 ] Retrospective cohort
(N = 329)
Consecutive TBI and mixed neurological conditions 41 33% 12.7 Criteria by Slick et al. (1999) b [23]
Greve, 2007a [ 38 ] Retrospective cohort
(N = 123)
NR Toxic exposure patients 41.3 29% 12.0 Criteria by Slick et al. (1999) b [23]
Aguerrevere, 2008 [ 59 ] Retrospective cohort
(N = 185)
NR TBI patients 37.8 28% 12.4 Criteria by Slick et al. (1999) b [23]
Curtis, 2008 [ 60 ] Prospective cohort
(N = 204)
NR TBI patients 39.6 29% 12.3 Criteria by Slick et al. (1999) b [23]
Greve, 2008 [ 61 ] Prospective cohort
(N = 211)
NR TBI patients 38.3 28% 12.1 Criteria by Slick et al. (1999) b [23]
Ord, 2008 [ 62 ] Not reported
(N = 93)
NR TBI patients 36.2 36% 12.7 Criteria by Slick et al. (1999) b [23]
Greve, 2008b [ 63 ] Not reported
(N TB = 109; N Chronic pain = 228)
NR TBI and Chronic pain patients TBI:
40.35
Chronic pain: 42.5
TBI: 24%
Chronic pain: 35%
TBI: 12.27
Chronic pain: 11.8
Criteria by Slick et al. (1999) b [23]
Henry, 2009 [ 64 ] Retrospective cohort
(N = 161)
Consecutive TBI and mixed neurological conditions 42.0 41% 13.83 Criteria by Slick et al. (1999) b [23]
Greve, 2009 [ 12 ] Retrospective cohort
(N = 282)
NR TBI patients 37.7 27% 11.7 Criteria by Slick et al. (1999) b [23]
Greve, 2009a [ 65 ] Retrospective cohort
(N = 318)
Random Chronic pain patients 41.2 35% 11.8 Criteria by Bianchini et al. (2005) c [24]
Greve, 2009b [ 66 ] Prospective cohort
(N TB = 442; N Chronic pain = 378)
NR TBI and Chronic pain patients TBI:
38.7
Chronic pain: 42.4
TBI:
29%
Chronic pain: 37%
TBI:
12.3
Chronic pain: 11.6
Criteria by Bianchini et al. (2005) c [24]
Greve, 2009c [ 67 ] Retrospective cohort
(N = 604)
Consecutive Chronic pain patients 42.3 36% 11.7 Criteria by Bianchini et al. (2005) c [24]
Greve, 2009d [ 68 ] Retrospective cohort
(N = 508)
Consecutive Chronic pain patients 42.1 35% 11.6 Criteria by Bianchini et al. (2005) c [24]
Bortnik, 2010 [ 69 ] Retrospective cohort
(N = 188)
NR TBI and mixed neurological conditions 42.7 49% 11.9 Criteria by Slick et al. (1999) b [23]
Curtis, 2010 [ 70 ] Retrospective cohort
(N = 74)
Consecutive TBI and mixed neurological conditions 36.3 35% 13 Criteria by Slick et al. (1999) b [23]
Greve, 2010 [ 71 ] Retrospective cohort
(N = 612)
Consecutive Chronic pain patients 41.1 35% 11.7 Criteria by Bianchini et al. (2005) c [24]
Ord, 2010 [ 72 ] Retrospective cohort
(N = 84)
NR TBI patients 39.4 37% 13.0 Criteria by Slick et al. (1999) b [23]
Aguerrevere, 2011 [ 25 ] Prospective cohort
(N = 108)
Consecutive TBI patients 39.9 26% 12.4 Criteria by Slick et al. (1999) b [23]
Roberson, 2013 [ 73 ] Retrospective cohort
(N = 315)
NR TBI and mixed neurological conditions 43.1 44% 13.1 Criteria by Slick et al. (1999) b [23]
Bianchini, 2014 [ 74 ] Retrospective cohort
(N = 328)
NR Chronic pain patients 43.3 35% 12.1 Criteria by Bianchini et al. (2005) c [24]
Guise, 2014 [ 75 ] Retrospective cohort
(N = 119)
Consecutive TBI patients 38.3 31% 12.6 Criteria by Slick et al. (1999) b [23]
Patrick, 2014 [ 76 ] Retrospective cohort
(N = 52)
Consecutive TBI and mixed neurological conditions 43.5 17% 13.1 Criteria by Slick et al. (1999) b [23]
Aguerrevere, 2017 [ 77 ] Retrospective cohort
(N = 348)
NR Chronic pain patients 43.1 36% 11.7 Criteria by Bianchini et al. (2005) c [24]
Bianchini, 2018 [ 78 ] Retrospective cohort
(N = 501)
Consecutive Chronic pain patients 42.3 NR 11.3 Criteria by Bianchini et al. (2005) c [24]
Curtis, 2019 [ 79 ] Retrospective cohort
(N = 219)
NR Chronic pain patients 43.5 32% 12.2 Criteria by Bianchini et al. (2005) c [24]

NR = not reported; MDMT = Medical Symptom Validity Test; PDRT = Portland Digit Recognition Test; TOMM = Test of Memory Malingering; RDS = Reliable Digit Span; FBS = Fake Bad Scale; MI = Malingering Index; WMI = Working Memory Index; PSI = Processing Speed Index

a Criteria by Greffeinstein et al. 1994 [41] include (1) improbable poor performance on more than two neuropsychological measures, (2) total disability in a major social role, (3) contradiction between collateral sources and symptom history, and (4) remote memory loss

b Criteria by Slick et al. (1999) [23] include (A) presence of substantial external incentive, (B) evidence from neuropsychological testing, (C) evidence from self-report, and (D) behaviors meeting the necessary B and C criteria are not fully accounted for by psychiatric, neurological, or developmental factors. External incentive (Criterion A) plus Criterion B and/or C evidence had to present for a diagnosis of malingering. Criterion B behaviors are sufficient for a diagnosis of malingering on their own.

c Criteria by Bianchini et al. (2005) [24] reflects a modification of the criteria of Slick et al. (1999) [23] and includes external incentive and meeting one of the following four conditions: (1) positive findings on either [PDRT or TOMM or RDS] and positive findings on either [FBS or MI]; (2) positive findings on [WMI and PSI] and positive findings on either [FBS or MI]; (3) positive findings on either [PDRT or TOMM] and [WMI and PSI]; or (4) significantly below chance on either [PDRT or TOMM].

Risk of bias

Of the 32% of studies that described their sampling method (14 of 44), 13 used consecutive sampling and one used random sampling methods to identify IME referrals. All studies reported minimal missing data (<5%). Most studies (n = 29, 64%) showed similar age and education characteristics between exaggerating and non-exaggerating groups. No study explicitly stated that IME assessors administered the index test without knowledge of the reference standard. We had moderate confidence in the reference standard used by most studies (n = 35, 80%). None of the known group designs used to evaluate symptom exaggeration provided evidence of reliability and validity testing; however, there has been formal evaluation of psychometric properties of forced-choice tests that were administered in eligible studies (See S4 Table in supplementary material for details). (S2 Table).

Prevalence of symptom exaggeration and additional analyses

The prevalence of symptom exaggeration ranged from 17% to 67%, median 33% (inter-quartile range: 25–44), and the pooled prevalence was 35% (95% confidence interval [CI]: 31–39) (low certainty evidence) (Fig 2). However, we found a significant subgroup effect, of low to moderate credibility, that studies with a higher proportion of women (≥40% vs. < 40%) may be associated with higher rates of exaggeration: 47% (95%CI 36–58) vs. 31% (95%CI 28–35) (test of interaction p = 0.02; Fig 2, Tables 2 and S3). We did not detect any evidence of small study effects for the overall prevalence of symptom exaggeration (Egger’s test P = 0.13; S2 Fig) nor for the subgroup of studies with <40% women (Egger’s test P = 0.16; S2 Fig).

Fig 2. Forest plot for prevalence by proportion of females (P = 0.02).

Fig 2

Table 2. GRADE evidence profile: prevalence of symptom exaggeration among IME attendees in North America.

# of studies # of patients Risk of bias Inconsistency a Indirectness b Imprecision c Publication bias Prevalence (95% CI) Overall certainty in the evidence
Prevalence of symptom exaggeration in studies with >40% female participants
9 1,137 Serious d Serious e Not serious Not serious Could not be assessed as number of included studies <10 46.9% (95% CI: 35.6–58.3) Low
Prevalence of symptom exaggeration in studies with <40% female participants
33 7,891 Serious d Serious e Not serious Not serious Undetected; symmetric funnel plot; Egger’s test P = 0.16 31.4% (95% CI: 28.1–34.6) Low

a Inconsistency refers to variability in effect estimates across studies (i.e., heterogeneity) that could not be adequately explained.

b Indirectness results if the intervention, patients or outcomes are different from the research question under investigation

c In this review, serious imprecision is based on the position of the confidence interval relative to a 20% threshold for symptom exaggeration and if the effect on the patient, or clinical action, would differ depending on whether the upper or the lower boundary of the confidence interval represented the truth.

d We downgraded 1 level for risk of bias as of none the criteria used to evaluate symptom exaggeration have been formally validated.

e We downgraded one level for inconsistency due to a wide range in prevalence among eligible studies (17.4% to 67.3%), which was partially explained by participant sex. Specifically, studies enrolling a higher proportion of women (≥40% vs. < 40%) were associated with higher rates of symptom exaggeration: 47% (95%CI 36–58) vs. 31% (95%CI 28–35; test of interaction p = 0.02).

We found no significant subgroup effects for type of clinical condition (mild TBI versus chronic pain versus other conditions), confidence in the reference standard, age, or education (S3S5 Figs). Meta-regression showed no association between prevalence of symptom exaggeration and age, level of education, or severity of presenting complaint, but did suggest an association with the proportion of female participants (S1, S6 and S7 Figs). We present all extracted data per study in S6 Table.

Discussion

Our systematic review and meta-analysis of observational studies found low certainty evidence, rated down due to risk of bias and inconsistency, that symptom exaggeration may be common among individuals attending for IMEs in North America, affecting approximately 1 in 3 assessments. The prevalence of symptom exaggeration was higher in studies that enrolled a greater proportion of female attendees (47%) vs. a lower proportion of female attendees (31%).

Relation to other studies

This is the first systematic review to summarize the extent of symptom exaggeration among IME attendees in North America. A previous survey of 131 US board-certified neuropsychologists conducting forensic work found that, on average, they estimated 30% of examinees claiming personal injury, disability, or workers’ compensation presented with symptom exaggeration. However, estimated prevalence ranged considerably by diagnosis – from an average of 41% for mild head injuries to 2% for vascular dementia [80]. Our review found no evidence for differences in the prevalence of symptom exaggeration based on clinical condition, but most patients among studies eligible for our review presented with either mild TBI or chronic pain.

Although our review focused on IMEs in North America, data from other regions also suggest high rates of symptom exaggeration. An observational study in Spain reported that of 1,003 participants (61.5% female), drawn from unselected undergraduates, advanced psychology students, the general population, forensic psychologists, and forensic/legal medicine physicians, one-third reported having feigned symptoms or illness [81]. Data from Germany and the Netherlands suggest that one‐fifth to one‐third of clients in forensic or insurance contexts exhibit symptom overreporting [82]. Further, a Swiss study found that 28% to 34% of individuals undergoing medico‐legal evaluations demonstrated probable or definite symptom exaggeration [83].

Our finding suggesting that women are more likely to exaggerate symptoms vs. men is supported by a systematic review of 175 studies that found women report more bodily distress and more numerous, more intense, and more frequent somatic symptoms than men [84]. Reasons for this discrepancy are uncertain, but may include biological differences, greater bodily vigilance and awareness, and higher rates of negative affectivity vs. men [84]. When symptoms are disproportionate to objective pathology, clinicians should inquire about other factors. For example, women are more likely to experience intimate partner violence than men [85,86], and pain patients who report lifetime traumatic events experience greater pain severity [87].

Studies eligible for our review used different strategies and approaches for assessing the prevalence of symptom exaggeration. The National Academy of Neuropsychologists (NAN) and American Academy of Clinical Neuropsychology (AACN) have emphasized the use of a multimethod approach to assess symptom and performance validity. These include clinical interviews, medical records, medical investigations in certain cases, behavioural observations, and symptom and performance validity tests [88]. Specific guidance is not provided on which symptom and performance validity tests should be used, when they should be conducted, and how they should be interpreted [89].

Strengths and limitations

Our study has several methodological strengths including (1) restricting our eligibility criteria to studies employing a known group design or multi-modal approach to assess symptom exaggeration, (2) subgroup analysis and assessment consistent with current best practices [35,90], and (3) use of the GRADE approach to evaluate the certainty of evidence.

In terms of limitations, we restricted our review to IMEs conducted in North America and eligible studies focused mainly on chronic pain and TBI. The generalizability of our findings to other jurisdictions, contexts, and clinical conditions, is uncertain. We were unable to explore the effect of cultural variability on the prevalence of symptom exaggeration as we found no studies within our inclusion criteria that addressed this issue. We did not find evidence for a subgroup effect based on confidence in the refence standard; however, there may have been insufficient variability to identify an association as almost all studies used a reference standard in which we rated moderate confidence. Another limitation of our review is the absence of a compelling reference standard for symptom exaggeration. Furthermore, even within the same reference standard, operationalization can be variable, which may affect prevalence. Another limitation of the primary studies is the lack of stratification of prevalence of symptom exaggeration according to possible effect modifiers, such as sex. Doing so would facilitate within-study subgroup analysis, which are less subject to confounding than between-study subgroup analysis. Another major limitation of the current evidence is that none of the known group approaches for evaluating symptom exaggeration have undergone reliability and validity testing.

Implications for future research and practice

Failure to identify the contribution of symptom exaggeration towards examinee’s complaints not only compromises the reliability and validity of independent assessments but may also adversely impact patient care by medicalizing psychosocial issues [9193]. Our findings suggest that symptom exaggeration is common among patients attending for IMEs; however, we rated down the certainty of evidence due to uncertain psychometric properties of the criteria used to evaluate exaggeration. An urgent research priority is the evaluation of inter-rater reliability of known group and multi-modal systems to appraise symptom exaggeration. Validation of such assessment systems is also critical and extremely challenging, but indirect evidence of validity could be acquired by evaluating accuracy in distinguishing between volunteers who were or were not exaggerating symptoms.

Future research should investigate how cultural factors affect IME outcomes, with attention to language barriers, health beliefs, and potential biases among both examinees and assessors. Another research priority is the development and validation of a structured and comprehensive approach to identify symptom exaggeration in IME assessments. Such an approach should consider observed versus reported abilities, findings of other providers, self-reported history that is discrepant with documented history, and administration of validated tests. A further consideration for research and practice is the use of symptom validity tests that focus on malingering (e.g., Test of Memory Malingering [TOMM], Lees-Haley Fake Bad Scale [FBS]), which imply intent. Clinicians are, understandably and appropriately, hesitant to assign a label of malingering; reasons include the challenges associated with determining intent and the risk of litigation [94]. To circumvent these issues, we would suggest the use of the less value-laden term ‘symptom exaggeration’.

Conclusion

Symptom exaggeration may occur in almost 50% of women and in approximately a third of men undergoing IMEs. Assessors should evaluate symptom exaggeration when conducting IMEs using a multi-modal approach that includes both clinical findings and validated tests of performance effort, and avoid conflation with malingering which presumes intent. Priority areas for future research include establishing the reliability and validity of current evaluation criteria for symptom exaggeration, and development of a structured IME assessment approach that includes consideration of symptom exaggeration.

Supporting information

S1 Table. Search strategies.

(DOCX)

pone.0324684.s001.docx (30.4KB, docx)
S2 Table. Risk of bias assessment.

(DOCX)

pone.0324684.s002.docx (34.4KB, docx)
S3 Table. ICEMAN criteria to assess credibility of subgroup effect of female % and prevalence.

(DOCX)

pone.0324684.s003.docx (28.7KB, docx)
S4 Table. Psychometric properties of tests included in symptom exaggeration criteria with list of references.

(DOCX)

pone.0324684.s004.docx (66KB, docx)
S5 Table. Included and excluded studies at full text screening with reasons.

(DOCX)

pone.0324684.s005.docx (87.8KB, docx)
S6 Table. Data extracted from included studies.

(DOCX)

pone.0324684.s006.docx (36.9KB, docx)
S1 Fig. Meta-regression for proportion of females among 42 studies (p = 0.16).

(DOCX)

pone.0324684.s007.docx (172.6KB, docx)
S2 Fig. a- Funnel plots of overall prevalence (Egger’s test p = 0.13) and b- prevalence in subgroup of studies with female proportion <40% (Egger’s test p = 0.16).

(DOCX)

pone.0324684.s008.docx (170.7KB, docx)
S3 Fig. Subgroup analysis for type of conditions (test of interaction p = 0.95).

(DOCX)

pone.0324684.s009.docx (209.5KB, docx)
S4 Fig. Subgroup analysis for confidence in reference standard (test of interaction p = 0.84).

(DOCX)

pone.0324684.s010.docx (176KB, docx)
S5 Fig. Subgroup analysis for similar age and/or education between groups (test of interaction p = 0.47).

(DOCX)

pone.0324684.s011.docx (205.4KB, docx)
S6 Fig. Meta-regression for average age among 46 cohorts (p = 0.18).

(DOCX)

pone.0324684.s012.docx (142.1KB, docx)
S7 Fig. Meta-regression for average education level among 45 cohorts (p = 0.65).

(DOCX)

pone.0324684.s013.docx (125.7KB, docx)
S1 Checklist. Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) Checklist.

(DOCX)

pone.0324684.s014.docx (35.6KB, docx)
S2 Checklist. Meta-analysis of Observational Studies in Epidemiology (MOOSE) checklist.

(DOCX)

pone.0324684.s015.docx (1.3MB, docx)

Acknowledgments

We would like to thank Michael Bagby from the Departments of Psychology and Psychiatry at University of Toronto for his contributions to the initial discussions around conceptualization and design of this study. No financial compensation was provided to any of these individuals.

Data Availability

All relevant data are within the paper and its Supporting Information files.

Funding Statement

The author(s) received no specific funding for this work.

References

  • 1.Canada S. Canadian Survey on Disability, 2017 to 2022; 2023. Available from: https://www150.statcan.gc.ca/n1/daily-quotidien/231201/dq231201b-eng.htm [Google Scholar]
  • 2.Disability and Health Data System (DHDS) [Internet]; 2020. [cited 2023 Jan 16]. Available from: https://dhds.cdc.gov/SP?LocationId=59&CategoryId=DISEST&ShowFootnotes=true&showMode=&IndicatorIds=STATTYPE,AGEIND,SEXIND,RACEIND,VETIND&pnl0=Chart,false,YR5,CAT1,BO1,AGEADJPREV&pnl1=Chart,false,YR5,DISSTAT,PREV&pnl2=Chart,false,YR5,DISSTAT,AGEADJPREV&pnl3=Chart,false,YR5,DISSTAT,AGEADJPREV&pnl4=Chart,false,YR5,DISSTAT,AGEADJPREV [Google Scholar]
  • 3.Martin DW. Independent medical evaluation: a practical guide. Springer; 2018. [DOI] [PubMed] [Google Scholar]
  • 4.Ebrahim S, Sava H, Kunz R, Busse JW. Ethics and legalities associated with independent medical evaluations. CMAJ. 2014;186(4):248–9. doi: 10.1503/cmaj.131509 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Gill D, Green P, Flaro L, Pucci T. The role of effort testing in independent medical examinations. Med Leg J. 2007;75(Pt 2):64–71. doi: 10.1258/rsmmlj.75.2.64 [DOI] [PubMed] [Google Scholar]
  • 6.Mæland S, Holmås TH, Øyeflaten I, Husabø E, Werner EL, Monstad K. What is the effect of independent medical evaluation on days on sickness benefits for long-term sick listed employees in Norway? A pragmatic randomised controlled trial, the NIME-trial. BMC Public Health. 2022;22(1):400. doi: 10.1186/s12889-022-12800-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Barth J, de Boer WE, Busse JW, Hoving JL, Kedzia S, Couban R. Inter-rater agreement in evaluation of disability: systematic review of reproducibility studies. BMJ. 2017;356. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Kunz R, von Allmen DY, Marelli R, Hoffmann-Richter U, Jeger J, Mager R, et al. The reproducibility of psychiatric evaluations of work disability: two reliability and agreement studies. BMC Psychiatry. 2019;19(1):205. doi: 10.1186/s12888-019-2171-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Bachmann M, de Boer W, Schandelmaier S, Leibold A, Marelli R, Jeger J, et al. Use of a structured functional evaluation process for independent medical evaluations of claimants presenting with disabling mental illness: rationale and design for a multi-center reliability study. BMC Psychiatry. 2016;16:271. doi: 10.1186/s12888-016-0967-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Boskovic I, Gallardo CT, Vrij A, Hope L, Merckelbach H. Verifiability on the run: an experimental study on the verifiability approach to malingered symptoms. Psychiatr Psychol Law. 2018;26(1):65–76. doi: 10.1080/13218719.2018.1483272 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Rumschik SM, Appel JM. Malingering in the psychiatric emergency department: prevalence, predictors, and outcomes. Psychiatr Serv. 2019;70(2):115–22. doi: 10.1176/appi.ps.201800140 [DOI] [PubMed] [Google Scholar]
  • 12.Greve KW, Heinly MT, Bianchini KJ, Love JM. Malingering detection with the Wisconsin Card Sorting Test in mild traumatic brain injury. Clin Neuropsychol. 2009;23(2):343–62. doi: 10.1080/13854040802054169 [DOI] [PubMed] [Google Scholar]
  • 13.Costa D. Psychiatric detection of exaggeration in reports of memory impairment. J Nerv Ment Dis. 1999;187(7):446–8. doi: 10.1097/00005053-199907000-00010 [DOI] [PubMed] [Google Scholar]
  • 14.Walczyk JJ, Sewell N, DiBenedetto MB. A review of approaches to detecting malingering in forensic contexts and promising cognitive load-inducing lie detection techniques. Front Psychiatry. 2018;9:700. doi: 10.3389/fpsyt.2018.00700 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Moher D, Shamseer L, Clarke M, Ghersi D, Liberati A, Petticrew M, et al. Preferred reporting items for systematic review and meta-analysis protocols (PRISMA-P) 2015 statement. Syst Rev. 2015;4(1):1. doi: 10.1186/2046-4053-4-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Stroup DF, Berlin JA, Morton SC, Olkin I, Williamson GD, Rennie D, et al. Meta-analysis of observational studies in epidemiology: a proposal for reporting. Meta-analysis Of Observational Studies in Epidemiology (MOOSE) group. JAMA. 2000;283(15):2008–12. doi: 10.1001/jama.283.15.2008 [DOI] [PubMed] [Google Scholar]
  • 17.BSCI I. Open Science Framework.
  • 18.Distiller S. Data management software. Ottawa (ON): Evidence Partners; 2011. [Google Scholar]
  • 19.Rogers RE. Clinical assessment of malingering and deception. Guilford Press; 2008. [Google Scholar]
  • 20.Rogers R, Kropp PR, Bagby RM, Dickens SE. Faking specific disorders: a study of the Structured Interview of Reported Symptoms (SIRS). J Clin Psychol. 1992;48(5):643–8. doi: 10.1002/1097-4679(199209)48:5<643::aid-jclp2270480511>3.0.co;2-2 [DOI] [PubMed] [Google Scholar]
  • 21.Heilbronner RL, Sweet JJ, Morgan JE, Larrabee GJ, Millis SR, 1 CP. American Academy of Clinical Neuropsychology Consensus Conference Statement on the neuropsychological assessment of effort, response bias, and malingering. Clin Neuropsychol. 2009;23(7):1093–129. [DOI] [PubMed] [Google Scholar]
  • 22.Nies KJ, Sweet JJ. Neuropsychological assessment and malingering: a critical review of past and present strategies. Arch Clin Neuropsychol. 1994;9(6):501–52. doi: 10.1093/arclin/9.6.501 [DOI] [PubMed] [Google Scholar]
  • 23.Slick DJ, Sherman EM, Iverson GL. Diagnostic criteria for malingered neurocognitive dysfunction: proposed standards for clinical practice and research. Clin Neuropsychol. 1999;13(4):545–61. doi: 10.1076/1385-4046(199911)13:04;1-Y;FT545 [DOI] [PubMed] [Google Scholar]
  • 24.Bianchini KJ, Greve KW, Glynn G. On the diagnosis of malingered pain-related disability: lessons from cognitive malingering research. Spine J. 2005;5(4):404–17. doi: 10.1016/j.spinee.2004.11.016 [DOI] [PubMed] [Google Scholar]
  • 25.Aguerrevere LE, Greve KW, Bianchini KJ, Ord JS. Classification accuracy of the Millon Clinical Multiaxial Inventory-III modifier indices in the detection of malingering in traumatic brain injury. J Clin Exp Neuropsychol. 2011;33(5):497–504. doi: 10.1080/13803395.2010.535503 [DOI] [PubMed] [Google Scholar]
  • 26.Cook RJ, Farewell VT. Conditional inference for subject‐specific and marginal agreement: two families of agreement measures. Can J Stat. 1995;23(4):333–44. doi: 10.2307/3315378 [DOI] [Google Scholar]
  • 27.Rogers R. Clinical assessment of malingering and deception. (No Title). 2009.
  • 28.Freeman MF, Tukey JW. Transformations related to the angular and the square root. Ann Math Statist. 1950;21(4):607–11. doi: 10.1214/aoms/1177729756 [DOI] [Google Scholar]
  • 29.Nyaga VN, Arbyn M, Aerts M. Metaprop: a Stata command to perform meta-analysis of binomial data. Arch Public Health. 2014;72:1–10. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.DerSimonian R, Laird N. Meta-analysis in clinical trials. Control Clin Trials. 1986;7(3):177–88. doi: 10.1016/0197-2456(86)90046-2 [DOI] [PubMed] [Google Scholar]
  • 31.Miller JJ. The inverse of the freeman-tukey double arcsine transformation. Am Stat. 1978;32(4):138. doi: 10.2307/2682942 [DOI] [Google Scholar]
  • 32.Guyatt G, Oxman AD, Akl EA, Kunz R, Vist G, Brozek J, et al. GRADE guidelines: 1. Introduction-GRADE evidence profiles and summary of findings tables. J Clin Epidemiol. 2011;64(4):383–94. doi: 10.1016/j.jclinepi.2010.04.026 [DOI] [PubMed] [Google Scholar]
  • 33.Egger M, Davey Smith G, Schneider M, Minder C. Bias in meta-analysis detected by a simple, graphical test. BMJ. 1997;315(7109):629–34. doi: 10.1136/bmj.315.7109.629 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Rücker G, Schwarzer G, Carpenter JR, Schumacher M. Undue reliance on I(2) in assessing heterogeneity may mislead. BMC Med Res Methodol. 2008;8:79. doi: 10.1186/1471-2288-8-79 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Schandelmaier S, Briel M, Varadhan R, Schmid CH, Devasenapathy N, Hayward RA, et al. Development of the Instrument to assess the Credibility of Effect Modification Analyses (ICEMAN) in randomized controlled trials and meta-analyses. CMAJ. 2020;192(32):E901–6. doi: 10.1503/cmaj.200077 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.StataCorp L. Stata statistical software: release 16. College Station (TX): StataCorp; 2019. [Google Scholar]
  • 37.Garner P, Hopewell S, Chandler J, MacLehose H, Akl EA, Beyene J, et al. When and how to update systematic reviews: consensus and checklist. BMJ. 2016;354. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Greve KW, Springer S, Bianchini KJ, Black FW, Heinly MT, Love JM. Malingering in toxic exposure: classification accuracy of Reliable Digit Span and WAIS-III Digit Span scaled scores. Assessment. 2007;14(1):12–21. [DOI] [PubMed] [Google Scholar]
  • 39.Lees-Haley PR, English LT, Glenn WJ. A Fake Bad Scale on the MMPI-2 for personal injury claimants. Psychol Rep. 1991;68(1):203–10. doi: 10.2466/pr0.1991.68.1.203 [DOI] [PubMed] [Google Scholar]
  • 40.Greve KW, Bianchini KJ, Black FW, Heinly MT, Love JM, Swift DA, et al. The prevalence of cognitive malingering in persons reporting exposure to occupational and environmental substances. Neurotoxicology. 2006;27(6):940–50. doi: 10.1016/j.neuro.2006.06.009 [DOI] [PubMed] [Google Scholar]
  • 41.Greiffenstein MF, Gola T, Baker WJ. MMPI-2 validity scales versus domain specific measures in detection of factitious traumatic brain injury. Clin Neuropsychol. 1995;9(3):230–40. doi: 10.1080/13854049508400485 [DOI] [Google Scholar]
  • 42.Lees-Haley PR. Provisional normative data for a credibility scale for assessing personal injury claimants. PR. 1990;66(3):1355. doi: 10.2466/pr0.66.3.1355-1360 [DOI] [PubMed] [Google Scholar]
  • 43.Suhr J, Tranel D, Wefel J, Barrash J. Memory performance after head injury: contributions of malingering, litigation status, psychological factors, and medication use. J Clin Exp Neuropsychol. 1997;19(4):500–14. doi: 10.1080/01688639708403740 [DOI] [PubMed] [Google Scholar]
  • 44.van Gorp WG, Humphrey LA, Kalechstein AL, Brumm VL, McMullen WJ, Stoddard MA, et al. How well do standard clinical neuropsychological tests identify malingering? A preliminary analysis. J Clin Exp Neuropsychol. 1999. Apr;21(2):245–50. doi: 10.1076/jcen.21.2.245.933 [DOI] [PubMed] [Google Scholar]
  • 45.Sweet JJ, Wolfe P, Sattlberger E, Numan B, Rosenfeld JP, Clingerman S, et al. Further investigation of traumatic brain injury versus insufficient effort with the California Verbal Learning Test. Arch Clin Neuropsychol. 2000;15(2):105–13. doi: 10.1093/arclin/15.2.105 [DOI] [PubMed] [Google Scholar]
  • 46.Greve KW, Bianchini KJ, Mathias CW, Houston RJ, Crouch JA. Detecting malingered performance on the Wechsler Adult Intelligence Scale. Validation of Mittenberg’s approach in traumatic brain injury. Arch Clin Neuropsychol. 2003;18(3):245–60. doi: 10.1093/arclin/18.3.245 [DOI] [PubMed] [Google Scholar]
  • 47.Lu PH, Boone KB, Cozolino L, Mitchell C. Effectiveness of the Rey-Osterrieth Complex Figure Test and the Meyers and Meyers recognition trial in the detection of suspect effort. Clin Neuropsychol. 2003;17(3):426–40. doi: 10.1076/clin.17.3.426.18083 [DOI] [PubMed] [Google Scholar]
  • 48.Barrash J, Suhr J, Manzel K. Detecting poor effort and malingering with an expanded version of the Auditory Verbal Learning Test (AVLTX): validation with clinical samples. J Clin Exp Neuropsychol. 2004;26(1):125–40. doi: 10.1076/jcen.26.1.125.23928 [DOI] [PubMed] [Google Scholar]
  • 49.Heinly MT, Greve KW, Bianchini KJ, Love JM, Brennan A. WAIS digit span-based indicators of malingered neurocognitive dysfunction: classification accuracy in traumatic brain injury. Assessment. 2005;12(4):429–44. doi: 10.1177/1073191105281099 [DOI] [PubMed] [Google Scholar]
  • 50.Curtis KL, Greve KW, Bianchini KJ, Brennan A. California verbal learning test indicators of Malingered Neurocognitive Dysfunction: sensitivity and specificity in traumatic brain injury. Assessment. 2006;13(1):46–61. doi: 10.1177/1073191105285210 [DOI] [PubMed] [Google Scholar]
  • 51.Etherton JL, Bianchini KJ, Ciota MA, Heinly MT, Greve KW. Pain, malingering and the WAIS-III Working Memory Index. Spine J. 2006;6(1):61–71. doi: 10.1016/j.spinee.2005.05.382 [DOI] [PubMed] [Google Scholar]
  • 52.Greve KW, Bianchini KJ, Love JM, Brennan A, Heinly MT. Sensitivity and specificity of MMPI-2 validity scales and indicators to malingered neurocognitive dysfunction in traumatic brain injury. Clin Neuropsychol. 2006;20(3):491–512. doi: 10.1080/13854040590967144 [DOI] [PubMed] [Google Scholar]
  • 53.Greve KW, Bianchini KJ, Doane BM. Classification accuracy of the test of memory malingering in traumatic brain injury: results of a known-groups analysis. J Clin Exp Neuropsychol. 2006;28(7):1176–90. doi: 10.1080/13803390500263550 [DOI] [PubMed] [Google Scholar]
  • 54.Greve KW, Bianchini KJ. Classification accuracy of the Portland Digit Recognition Test in traumatic brain injury: results of a known-groups analysis. Clin Neuropsychol. 2006;20(4):816–30. doi: 10.1080/13854040500346610 [DOI] [PubMed] [Google Scholar]
  • 55.Ardolf BR, Denney RL, Houston CM. Base rates of negative response bias and malingered neurocognitive dysfunction among criminal defendants referred for neuropsychological evaluation. Clin Neuropsychol. 2007;21(6):899–916. doi: 10.1080/13825580600966391 [DOI] [PubMed] [Google Scholar]
  • 56.Greve KW, Bianchini KJ, Roberson T. The Booklet Category Test and malingering in traumatic brain injury: classification accuracy in known groups. Clin Neuropsychol. 2007;21(2):318–37. doi: 10.1080/13854040500488552 [DOI] [PubMed] [Google Scholar]
  • 57.Henry GK, Enders C. Probable malingering and performance on the Continuous Visual Memory Test. Appl Neuropsychol. 2007;14(4):267–74. doi: 10.1080/09084280701719245 [DOI] [PubMed] [Google Scholar]
  • 58.O’Bryant SE, Engel LR, Kleiner JS, Vasterling JJ, Black FW. Test of memory malingering (TOMM) trial 1 as a screening measure for insufficient effort. Clin Neuropsychol. 2007;21(3):511–21. doi: 10.1080/13854040600611368 [DOI] [PubMed] [Google Scholar]
  • 59.Aguerrevere LE, Greve KW, Bianchini KJ, Meyers JE. Detecting malingering in traumatic brain injury and chronic pain with an abbreviated version of the Meyers Index for the MMPI-2. Arch Clin Neuropsychol. 2008;23(7–8):831–8. doi: 10.1016/j.acn.2008.06.008 [DOI] [PubMed] [Google Scholar]
  • 60.Curtis KL, Thompson LK, Greve KW, Bianchini KJ. Verbal fluency indicators of malingering in traumatic brain injury: classification accuracy in known groups. Clin Neuropsychol. 2008;22(5):930–45. doi: 10.1080/13854040701563591 [DOI] [PubMed] [Google Scholar]
  • 61.Greve KW, Lotz KL, Bianchini KJ. Observed versus estimated IQ as an index of malingering in traumatic brain injury: classification accuracy in known groups. Appl Neuropsychol. 2008;15(3):161–9. doi: 10.1080/09084280802324085 [DOI] [PubMed] [Google Scholar]
  • 62.Ord JS, Greve KW, Bianchini KJ. Using the Wechsler Memory Scale-III to detect malingering in mild traumatic brain injury. Clin Neuropsychol. 2008;22(4):689–704. doi: 10.1080/13854040701425437 [DOI] [PubMed] [Google Scholar]
  • 63.Greve KW, Ord J, Curtis KL, Bianchini KJ, Brennan A. Detecting malingering in traumatic brain injury and chronic pain: a comparison of three forced-choice symptom validity tests. Clin Neuropsychol. 2008;22(5):896–918. doi: 10.1080/13854040701565208 [DOI] [PubMed] [Google Scholar]
  • 64.Henry GK, Heilbronner RL, Mittenberg W, Enders C, Domboski K. Comparison of the MMPI-2 restructured Demoralization Scale, Depression Scale, and Malingered Mood Disorder Scale in identifying non-credible symptom reporting in personal injury litigants and disability claimants. Clin Neuropsychol. 2009;23(1):153–66. doi: 10.1080/13854040801969524 [DOI] [PubMed] [Google Scholar]
  • 65.Greve KW, Bianchini KJ, Etherton JL, Ord JS, Curtis KL. Detecting malingered pain-related disability: classification accuracy of the Portland Digit Recognition Test. Clin Neuropsychol. 2009;23(5):850–69. doi: 10.1080/13854040802585055 [DOI] [PubMed] [Google Scholar]
  • 66.Greve KW, Curtis KL, Bianchini KJ, Ord JS. Are the original and second edition of the California Verbal Learning Test equally accurate in detecting malingering? Assessment. 2009;16(3):237–48. doi: 10.1177/1073191108326227 [DOI] [PubMed] [Google Scholar]
  • 67.Greve KW, Etherton JL, Ord J, Bianchini KJ, Curtis KL. Detecting malingered pain-related disability: classification accuracy of the test of memory malingering. Clin Neuropsychol. 2009;23(7):1250–71. doi: 10.1080/13854040902828272 [DOI] [PubMed] [Google Scholar]
  • 68.Greve KW, Ord JS, Bianchini KJ, Curtis KL. Prevalence of malingering in patients with chronic pain referred for psychologic evaluation in a medico-legal context. Arch Phys Med Rehabil. 2009;90(7):1117–26. doi: 10.1016/j.apmr.2009.01.018 [DOI] [PubMed] [Google Scholar]
  • 69.Bortnik KE, Boone KB, Marion SD, Amano S, Ziegler E, Cottingham ME, et al. Examination of various WMS-III logical memory scores in the assessment of response bias. Clin Neuropsychol. 2010;24(2):344–57. doi: 10.1080/13854040903307268 [DOI] [PubMed] [Google Scholar]
  • 70.Curtis KL, Greve KW, Brasseux R, Bianchini KJ. Criterion groups validation of the Seashore Rhythm Test and Speech Sounds Perception Test for the detection of malingering in traumatic brain injury. Clin Neuropsychol. 2010;24(5):882–97. doi: 10.1080/13854041003762113 [DOI] [PubMed] [Google Scholar]
  • 71.Greve KW, Bianchini KJ, Etherton JL, Meyers JE, Curtis KL, Ord JS. The Reliable Digit Span test in chronic pain: classification accuracy in detecting malingered pain-related disability. Clin Neuropsychol. 2010;24(1):137–52. doi: 10.1080/13854040902927546 [DOI] [PubMed] [Google Scholar]
  • 72.Ord JS, Boettcher AC, Greve KW, Bianchini KJ. Detection of malingering in mild traumatic brain injury with the Conners’ Continuous Performance Test-II. J Clin Exp Neuropsychol. 2010;32(4):380–7. doi: 10.1080/13803390903066881 [DOI] [PubMed] [Google Scholar]
  • 73.Roberson CJ, Boone KB, Goldberg H, Miora D, Cottingham M, Victor T, et al. Cross validation of the b Test in a large known groups sample. Clin Neuropsychol. 2013;27(3):495–508. doi: 10.1080/13854046.2012.737027 [DOI] [PubMed] [Google Scholar]
  • 74.Bianchini KJ, Aguerrevere LE, Guise BJ, Ord JS, Etherton JL, Meyers JE, et al. Accuracy of the Modified Somatic Perception Questionnaire and Pain Disability Index in the detection of malingered pain-related disability in chronic pain. Clin Neuropsychol. 2014;28(8):1376–94. doi: 10.1080/13854046.2014.986199 [DOI] [PubMed] [Google Scholar]
  • 75.Guise BJ, Thompson MD, Greve KW, Bianchini KJ, West L. Assessment of performance validity in the Stroop Color and Word Test in mild traumatic brain injury patients: a criterion-groups validation design. J Neuropsychol. 2014;8(1):20–33. doi: 10.1111/jnp.12002 [DOI] [PubMed] [Google Scholar]
  • 76.Patrick RE, Horner MD. Psychological characteristics of individuals who put forth inadequate cognitive effort in a secondary gain context. Arch Clin Neuropsychol. 2014;29(8):754–66. doi: 10.1093/arclin/acu054 [DOI] [PubMed] [Google Scholar]
  • 77.Aguerrevere LE, Calamia MR, Greve KW, Bianchini KJ, Curtis KL, Ramirez V. Clusters of financially incentivized chronic pain patients using the Minnesota Multiphasic Personality Inventory-2 Restructured Form (MMPI-2-RF). Psychol Assess. 2018;30(5):634–44. doi: 10.1037/pas0000509 [DOI] [PubMed] [Google Scholar]
  • 78.Bianchini KJ, Aguerrevere LE, Curtis KL, Roebuck-Spencer TM, Frey FC, Greve KW, et al. Classification accuracy of the Minnesota Multiphasic Personality Inventory-2 (MMPI-2)-Restructured form validity scales in detecting malingered pain-related disability. Psychol Assess. 2018;30(7):857–69. doi: 10.1037/pas0000532 [DOI] [PubMed] [Google Scholar]
  • 79.Curtis KL, Aguerrevere LE, Bianchini KJ, Greve KW, Nicks RC. Detecting malingered pain-related disability with the pain catastrophizing scale: a criterion groups validation study. Clin Neuropsychol. 2019;33(8):1485–500. doi: 10.1080/13854046.2019.1575470 [DOI] [PubMed] [Google Scholar]
  • 80.Mittenberg W, Patton C, Canyock EM, Condit DC. Base rates of malingering and symptom exaggeration. J Clin Exp Neuropsychol. 2002;24(8):1094–102. doi: 10.1076/jcen.24.8.1094.8379 [DOI] [PubMed] [Google Scholar]
  • 81.Puente-López E, Pina D, López-López R, Ordi HG, Bošković I, Merten T. Prevalence estimates of symptom feigning and malingering in Spain. Psychol Inj Law. 2023;16(1):1–17. doi: 10.1007/s12207-022-09458-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 82.Merten T, Dandachi-FitzGerald B, Hall V, Bodner T, Giromini L, Lehrner J, et al. Symptom and performance validity assessment in European countries: an update. Psychol Inj Law. 2022;15(2):116–27. doi: 10.1007/s12207-021-09436-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 83.Plohmann AM, Hurter M. Prevalence of poor effort and malingered neurocognitive dysfunction in litigating patients in Switzerland. Z Neuropsychol. 2017. [Google Scholar]
  • 84.Barsky AJ, Peekna HM, Borus JF. Somatic symptom reporting in women and men. J Gen Intern Med. 2001;16(4):266–75. doi: 10.1046/j.1525-1497.2001.00229.x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 85.Lövestad S, Krantz G. Men’s and women’s exposure and perpetration of partner violence: an epidemiological study from Sweden. BMC Public Health. 2012;12:945. doi: 10.1186/1471-2458-12-945 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 86.Umubyeyi A, Mogren I, Ntaganira J, Krantz G. Women are considerably more exposed to intimate partner violence than men in Rwanda: results from a population-based, cross-sectional study. BMC Womens Health. 2014;14:99. doi: 10.1186/1472-6874-14-99 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 87.Nicol AL, Sieberg CB, Clauw DJ, Hassett AL, Moser SE, Brummett CM. The association between a history of lifetime traumatic events and pain severity, physical function, and affective distress in patients with chronic Pain. J Pain. 2016;17(12):1334–48. doi: 10.1016/j.jpain.2016.09.003 [DOI] [PubMed] [Google Scholar]
  • 88.Bush SS, Heilbronner RL, Ruff RM. Psychological assessment of symptom and performance validity, response bias, and malingering: official position of the Association for Scientific Advancement in Psychological Injury and Law. Psychol Inj Law. 2014;7(3):197–205. doi: 10.1007/s12207-014-9198-7 [DOI] [Google Scholar]
  • 89.Sweet JJ, Heilbronner RL, Morgan JE, Larrabee GJ, Rohling ML, Boone KB, et al. American Academy of Clinical Neuropsychology (AACN) 2021 consensus statement on validity assessment: Update of the 2009 AACN consensus conference statement on neuropsychological assessment of effort, response bias, and malingering. Clin Neuropsychol. 2021;35(6):1053–106. doi: 10.1080/13854046.2021.1896036 [DOI] [PubMed] [Google Scholar]
  • 90.Sun X, Briel M, Walter SD, Guyatt GH. Is a subgroup effect believable? Updating criteria to evaluate the credibility of subgroup analyses. BMJ. 2010;340:c117. doi: 10.1136/bmj.c117 [DOI] [PubMed] [Google Scholar]
  • 91.Häuser W, Fitzcharles MA. Facts and myths pertaining to fibromyalgia. Dialogues Clin Neurosci. 2022. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 92.Koesling D, Bozzaro C. Chronic pain as a blind spot in the diagnosis of a depressed society: on the implications of the connection between depression and chronic pain for interpretations of contemporary society. Med Health Care Philos. 2022:1–10. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 93.Burke MJ, Silverberg ND. New framework for the continuum of concussion and functional neurological disorder. BMJ Publishing Group Ltd and British Association of Sport and Exercise Medicine; 2024. [DOI] [PubMed] [Google Scholar]
  • 94.Weiss KJ, Van Dell L. Liability for diagnosing malingering. J Am Acad Psychiatry Law. 2022;45:339–47. [PubMed] [Google Scholar]

Decision Letter 0

Thiago Fernandes

>PONE-D-24-31136>>Prevalence of Symptom Exaggeration Among North American Independent Medical Evaluation Examinees: A systematic review of observational studies>>PLOS ONE

Dear Dr. Busse,

Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process.

Thank you for your submission. I read it with enthusiasm, as it shows notable merits. Nevertheless, I noticed some concerns that I'd like the authors to address. As follows:

1) Please double-check grammar (e.g. punctuation and verb tense)

2) Please double-check refs (e.g. keep the requested style, Vancouver, and also correct formatting, punctuation and remove repetitives)

3) The Abstract is well-structured, but could benefit from conciseness. The inclusion of multiple phrases, such as "symptom exaggeration" and "observational studies," makes it somewhat repetitive. I'd suggest reducing redundant terms to improve readability;

- The logical transitions between the background, methods, and results are somewhat abrupt. For instance, it's unclear how the analysis of "symptom exaggeration" directly links to "inter-rater reliability" or whether the gender-specific findings are derived from a subset analysis or the primary outcomes. Please provide a clearer link between the methodological approach and the primary research question;

- The stats parameters (e.g. p-values, effect sizes, CIs) are provided, but more detail on the significance of these values will enhance the depth of the findings. For example, how do these rates of symptom exaggeration in IMEs impact policy, clinical practice, or future research directions? Offering these insights would broaden the reliability;

- Sample size details are provided, but the inclusion of additional information, such as variations across cohorts or specific diagnostic categories, would be useful. For instance, were patients with traumatic brain injuries or chronic pain assessed differently? This would offer a clearer understanding of subgroup analysis and generalisation;

- A brief explanation of whether the studies had high or low certainty is important;

- The abstract briefly touches on gender differences, but would benefit from elaborating on why women show higher rates of symptom exaggeration in these studies. Consider a brief sentence;

- The findings could be expanded to include more specific quantitative details, such as the range of exaggeration percentages across the cohorts or any significant moderators or predictors found in the meta-analysis;

- The conclusion is concise but could be expanded to include a brief part on how symptom exaggeration influences clinical decision-making during IMEs. Additionally, mentioning how these findings could influence future research or policy changes is essential;

- The abstract would be improved by providing the main strengths, such as prevalence rates for symptom exaggeration or differences in exaggeration across different conditions;

- The term "symptom exaggeration" is mentioned without any clear definition. Please elaborate;

4) The transition between the background on disabilities and the introduction of IMEs feels somewhat confusing. Providing a smoother transition into how symptom exaggeration affects IMEs will improve your argument;

- There are relevant stats about disabilities in North America, which effectively establishes the context. However, there is a need for a more explicit link between this background information and the objective. For example, while the IME process is mentioned, the transition to the issue of symptom exaggeration is quite abrupt

5) The Introduction touches on symptom exaggeration as a factor affecting IME outcomes. Nevertheless, there is the need to state how the authors would measure it or what variables will be examined (e.g. mental health, pain severity).

6) Explicitly state how your study addresses gaps or contradictions found in previous studies. For example, why past standardisation efforts failed. More emphasis on the relevance and novelty of the study is essential;

7) While the IME concept and symptom exaggeration are introduced, there is little detail about the methodology or variables that will be assessed;

- Please articulate your hypotheses with the rationale;

- To improve flow, consider adding a sentence explaining why IMEs are prone to variability, and how symptom exaggeration may influence;

- Summarise the objectives and hypotheses clearly in one sentence;

8) Overall, there are some worrying aspects in Introduction:

- The Introduction needs a clearer flow. Consider providing the background and context of IMEs, then move into the issue of symptom exaggeration. After that, it's important to debate what previous studies have found and highlight their gaps. You can then explain how this study will ease those. Concluding with the rationale, objectives, and hypotheses will give the Introduction a strong structure. Consider to mention why the study is important since it'll make more engaging and relevant for a broader audience;

- The text shifts abruptly between topics, like disability stats and symptom exaggeration, without clearly connecting them or stating the rationale. Brief mentioning the methods used to measure symptom exaggeration is really important;

- The authors could provide more context to strengthen their argument. IMEs often focus too narrowly on the biomedical model, overlooking psychosocial factors and work-related conditions, which influences the fairness and consistency of evaluations. The lack of standardisation and unified guidelines contributes to these issues. Additionally, biases in IME practices, such as favouring employer-provided information over patient accounts, lead to underdiagnosis and lower disability ratings compared to treating physicians. This bias fosters distrust and adds to the inconsistencies in the IME process. For more details, see 10.1097/PEC.0000000000000487 and 10.1007/978-3-319-71906-1;

- Another aspect contributing to IME variability is the difference in approaches, prognoses, and standards among evaluators. Studies show that diagnostic discordance rates can exceed 80%, with higher discordance linked to increased mortality (OR ≈ 1.2), underscoring the need for standardised protocols. Inconsistent assessments can lead to delayed treatment, denied disability benefits, and declining patient quality of life. Standardisation would improve accuracy and consistency, reducing patient burden and enhancing trust in IMEs. Furthermore, the focus on the biomedical model overlooks the role of psychological and social factors, suggesting that adopting a biopsychosocial approach could yield more accurate evaluations;

- The Introduction mentions IMEs and symptom exaggeration, but it fails to clearly state the objectives or what the study aims to explore. Providing specific objectives and hypotheses would make it easier for readers to understand the focus of the research;

- Explain how the variability in IME outcomes could be influenced by symptom exaggeration, tying it to the broader context of disability assessments;

- Strengthen the rationale by incorporating previous findings on symptom exaggeration in IMEs;

- The study focuses on symptom exaggeration as a source of variability in IMEs, an underexplored area. While past findings checked the reliability of IMEs, this work could provide important insights on how symptom exaggeration influences assessments of disability. Please consider providing a clearer explanation of why this is relevant;

- Providing more context on how symptom exaggeration might vary by condition is important; for example, variations in symptom exaggeration between conditions like traumatic brain injury and chronic pain will make the argument more cohese;

- Please clarify the relevance of the European results - what are the similarities or differences between IME in continents? And expand this part. Consider that healthcare systems are different, especially due to different geographical frameworks;

- I'd highly suggest to to introduce the concept of symptom exaggeration within the context of the study;

The Introduction is brief and lacks depth to establish a clear framework. Without the mentioned components, the Introduction feels underdeveloped and won't effectively set ground for the study. This can make the interpretation challenging without a solid rationale of the study;

9) Overall comments on Methods:

- The details about PRISMA and MOOSE are not provided. For example, the authors didn't report the refs for each checklist, and also didn't specify what workflow (e.g. screening, extraction, etc.) was followed. Including detailed information from both checklists, such as principal items, would enhance the presentation. The overall impression is, "How did the authors follow these guidelines?" Finally, why wasn't the systematic review registered on PROSPERO, the widely recommended registry for such studies?;

- The search strategy, while reported in the Sup. file, could be included in the main text to enhance clarity and robustness;

- Why include studies in other languages? Are the terms and interpretations the same across different contexts?;

- Please refine the verb tense for consistency;

- Several aspects of the design require clarification, particularly concerning the definitions, variable selection, and operationalisation. There is also a lack of explanation for the focus on specific countries. The justification for study exclusions, and the criteria for classifying symptom exaggeration could be detailed. Additionally, the handling of population overlap needs to be explained. Furthermore, if six authors conducted the screening, how was inter-rater reliability assessed? Failing to address these concerns risks making the study appear biased and unreliable;

- The eligibility criteria lack a clear rationale for including certain types of studies while excluding others. For instance, the inclusion of studies that "assessed the prevalence of symptom exaggeration using a known group design or multi-modal determination method" seems arbitrary without explaining these specific designs (c.f. 10.1212/CPJ.0000000000000092; and why they are superior or necessary for the systematic review. Please elaborate on why these aspects were chosen;

- If there is no reliable and valid standard for assessing symptoms, it's somewhat concerning to understand the reliability of the categorisation of the quality of the included studies. Actually, some studies report strategies for a reliable assessment of symptom exaggeration. Consider mentioning symptom validity tests. I think providing a detailed assessment criterion or more than one tool would be important (c.f. 10.1002/nur.10092). Consider using the Cochrane risk of bias or other tools (c.f. 10.1136/bmj.d5928; 10.1016/j.acn.2005.02.002). Also, explain how the authors ensured the reliability of the findings and what kind of psychometric assessments were performed (c.f. 10.1007/s12207-021-09436-8; 10.1016/j.acn.2005.02.002);

- The preference for including studies with larger sample size when population overlap exists introduces a potential for selection bias. Larger sample sizes do not always equate to better quality for meta-analysis. It'd be interesting to refer to relevant literature on this matter (e.g. 10.1002/sim.1186; 10.1136/bmj.d7762). I'd highly suggeest conducting robust sensitivity analyses to assess the influence of including or not certain studies (c.f. 10.1016/j.jclinepi.2004.01.<wbr style="color: rgb(34, 34, 34); font-family: Arial, Helvetica, sans-serif; font-size: small;" />018; 10.1111/rssc.12440);

10) Overall comments on Results:

- The authors could consider some approaches: a. run both fixed- and random-effects and scrutinise if there are substantial differences in the effects (i.e. if found, heterogeneity would be a really worrying aspect), b. identify studies with a high risk of bias, or small sample, remove them, and conduct sensitivity analysis (e.g. using l-o-o or cumulative sensitivity analysis). Then check the results, specifically evaluating the existence of differences in effects, CIs overlap, and changes in I² values. This would also be important when analysing subgroups;

- While beyond-chance scores are rare, they can still provide valuable insights into clinical findings (c.f. 10.1016/j.spinee.2004.11.016; 10.1186/s12874-016-0108-4). For instance, some studies have employed complementary approaches to assess symptom exaggeration (c.f. 10.1097/00001199-200004000-<wbr style="color: rgb(34, 34, 34); font-family: Arial, Helvetica, sans-serif; font-size: small;" />00006; 10.1007/BF01874896); Please elaborate;

- In studies where symptom exaggeration is being evaluated, the lack of blinding could significantly influence the results. Therefore, I would highly encourage the authors to expand this part;

- The databases used in the search are comprehensive. However, expanding the search to include other databases, such as Scopus, could ensure that no additional relevant refs are missed. Could the authors clarify why Scopus or similar databases were not included in the search strategy?;

- Additionally, more details regarding the filters applied during the search would be essential. For instance, were specific study designs or types prioritised during the screening process, and if so, which ones?;

- As recommended by Cochrane, systematic reviews should be updated to maintain their relevance and accuracy. Could the authors provide an approximate timeline for when they plan to refresh the search;

- How were duplicates handled during the search process?

- The search strategy is well-constructed and covers several important terms across databases. However, there are some concerning aspects. The use of broad terms such as 'validity' and 'disability' may result in many extraneous results, diminishing the relevance of the retrieved studies and also increasing the overall screening effort. Additionally, applying filters (e.g. specific diagnostic criteria or populations) could help enhance the reliability and focus on more relevant studies. Please clarify the rationale for not using such refinements;

- Please clarify the relevance of the transformation using F-T arcsine in your study. This is really important based on your design. The concern is to provide an explanation of parameters and the rational, enhancing clarity for readers outside the field;

- The trim-and-fill method could provide a more accurate adjustment for potential bias. Overall, after assessing any asymmetry in the funnel plot, this trims and fills missing data points to 'reach' symmetry, recalculating the overall effect size. Could the authors consider applying this approach and debating the influence on their analysis?;

- Please consider expanding the analyses by assessing Tau² (i.e. between-study variance), which could provide better insights into heterogeneity (c.f. 10.1136/bmj.327.7414.557);

- When conducting stats analyses (e.g. meta-regression, subgroup analyses), it’s essential to be cautious of small-study effects, which can influence the robustness of your findings;

- The Results are intriguing. However, the observed high variability and low certainty, as well as the association between women and higher symptom exaggeration, raise concerns about potential confounding factors. Differences in populations, diagnostic tools, and study designs likely contribute to this variability. The authors are encouraged to explore these issues more thoroughly, perhaps expanding the meta-regression to include more DVs for multivariate analysis. Consider additional approaches like structural equation modeling to reduce the number of variables, along with PCA and clustering. This would enhance the depth of your analyses, results, and overall presentation;

- Please specify test parameters (e.g. effect sizes, confidence intervals) alongside p-values for a more comprehensive understanding of the results;

- Ensure consistency in decimal places throughout the results section;

- The CIs for symptom exaggeration in women should be carefully examined, as overlapping CIs could indicate heterogeneity or residual variance;

- Given these shortcomings, I suggest the authors: a) Check for studies with small samples or wide CIs and identify possible outliers, b) Carefully consider the influence of outliers and residuals, c) Reporting I² and Tau² stats to assess heterogeneity. If heterogeneity is high, outliers could be driving the variation between studies;

11) The Discussion is well-structured but lacks conciseness, particularly in the “Relation to Other Studies” and “Implications for Future Research” sections. Streamlining these parts would enhance clarity and flow;

- While the study finds low-certainty evidence of symptom exaggeration in IMEs, the authors didn't provide sufficient arguments about what drives this low certainty. The sources of uncertainty (e.g., inconsistency, bias, imprecision) could be more clearly outlined. Also, consider debating the influence of heterogeneity;

- Some examples of how different studies operationalised symptom exaggeration differently and how this could be 'solved' seem a very interesting approach;

- Suggest practical steps for developing and validating new assessment tools for IMEs;

- Additionally, scrutinise CIs more carefully for potential overlap, as this could indicate heterogeneity;

- There is a concerning aspect regarding the connection to other findings.The arguments lack depth and focus mainly on 'correlation' with somatic symptoms, which may not fully align with IMEs or the study's objectives. Clarify the main differences between studies, focussing on the authors. The link between gender and symptom exaggeration could be simplified, and the data more sharply refined;

- Please provide a detailed explanation of how the multimethod approaches varied across studies and how the differences may have influenced the outcomes of the meta-analysis;

- The limitations are mentioned, but other biases, such as cultural or healthcare factors, could be potential confounders. Also, the lack of standardization and the absence of reliable assessments should be addressed;

- The link between Results and the Discussion could be clearer. Highlighting specific findings from the results, such as effect sizes or CIs, would enhance the reliability and make the conclusions more solid;

- The terminology distinguishing symptom exaggeration from malingering is unclear. Please clarify this distinction;

12) Tables and illustrations:

- The search strategy appears to include terms that may retrieve numerous irrelevant studies. The authors could consider refining the search by removing broad terms and avoiding overuse of wildcards. This can reduce the retrieval of unnecessary results and save future researchers from conducting an excessively lengthy and unfocused search. For example, terms like "independent" and linked terms such as "MMPI" and "work capacity" could be refined or excluded to enhance specificity. In using the current strategy, I encountered studies from adjacent fields, such as business or diagnostics for conditions unrelated to the focus of the review;

- The risk of bias Table lacks clarity and cohesion, as the studies are not organized alphabetically or by any discernible pattern, making it hard to follow. The use of consecutive sampling raises concerns about potential selection bias, which could influence the outcomes and, consequently, generalisation. Additionally, the 5% cutoff for missing data, though comprehensive, should be supported by relevant evidence to justify. While age and education are important demographics, it's unclear why other variables, such as gender, were not included. Expanding the demographic could provide a more comprehensive understanding of their influence on the findings;

- The Table on psychometric data raises concerns as it presents tests used in completely different settings, likely increasing heterogeneity. The variability in sample sizes, demographics, and the use of unreliable tools diminishes the potential for generalisation of the findings. Additionally, there are several outdated refs that should be updated to reflect current practices. Lastly, please double-check the grammar, as there are issues with punctuation and formatting that need correction;

- It's essential to note that the funnel plot is asymmetric, indicating heterogeneity. This reinforces the importance of using the trim-and-fill method to ensure robustness. The authors could also check carefully the sensitivity analysis and assess how the funnel looks after refinement. At a glance, there seem to be small-study effects, given the number of studies on left side;

- The graph on meta-regression exhibits dispersion and variability that isn't related to demographics (i.e. as indicated by the regression line). Please implement additional metrics like CIs or residuals for clarity and include important stats like the p-values, I², and other important metrics;

- The Table on study characteristics is missing important variables, like the ratio or percentage of women, including incorrect reporting (e.g. sampling being 'unclear'), and utilises confusing criteria for assessing symptom exaggeration. Please consider restructuring the Table to enhance clarity;

Overall, this is a well-conducted study, and I commend the endeavour. The findings presented are of notable importance. While this might sound lengthy, it is provided with the best intentions, aiming to align with the high standards expected in scientific communication.

While I found this study to be promising and well-conceived, I believe its current state may require extensive rounds, which could complicate the review process. To help streamline, I’ve provided some suggestions for the authors to consider. I encourage the authors to consider or check these and choose how they wish to proceed. Should they choose to resubmit after corrections, I'd be happy to read the ms again. I apologise if this isn’t the most positive news, but I see great potential in this work.

==============================

Please submit your revised manuscript by Nov 07 2024 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org . When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

Please include the following items when submitting your revised manuscript:>

  • A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'.

  • A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.

  • An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols . Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols .

We look forward to receiving your revised manuscript.

Kind regards,

Thiago P. Fernandes, PhD

Academic Editor

PLOS ONE

Journal Requirements:

When submitting your revision, we need you to address these additional requirements.

1. Please ensure that your manuscript meets PLOS ONE's style requirements, including those for file naming. The PLOS ONE style templates can be found at 

https://journals.plos.org/plosone/s/file?id=wjVg/PLOSOne_formatting_sample_main_body.pdf and 

https://journals.plos.org/plosone/s/file?id=ba62/PLOSOne_formatting_sample_title_authors_affiliations.pdf

2. Please include captions for your Supporting Information files at the end of your manuscript, and update any in-text citations to match accordingly. Please see our Supporting Information guidelines for more information: http://journals.plos.org/plosone/s/supporting-information.

3. As required by our policy on Data Availability, please ensure your manuscript or supplementary information includes the following: 

A numbered table of all studies identified in the literature search, including those that were excluded from the analyses.  

For every excluded study, the table should list the reason(s) for exclusion.  

If any of the included studies are unpublished, include a link (URL) to the primary source or detailed information about how the content can be accessed. 

A table of all data extracted from the primary research sources for the systematic review and/or meta-analysis. The table must include the following information for each study: 

Name of data extractors and date of data extraction 

Confirmation that the study was eligible to be included in the review.  

All data extracted from each study for the reported systematic review and/or meta-analysis that would be needed to replicate your analyses. 

If data or supporting information were obtained from another source (e.g. correspondence with the author of the original research article), please provide the source of data and dates on which the data/information were obtained by your research group. 

If applicable for your analysis, a table showing the completed risk of bias and quality/certainty assessments for each study or outcome.  Please ensure this is provided for each domain or parameter assessed. For example, if you used the Cochrane risk-of-bias tool for randomized trials, provide answers to each of the signalling questions for each study. If you used GRADE to assess certainty of evidence, provide judgements about each of the quality of evidence factor. This should be provided for each outcome.  

An explanation of how missing data were handled. 

This information can be included in the main text, supplementary information, or relevant data repository. Please note that providing these underlying data is a requirement for publication in this journal, and if these data are not provided your manuscript might be rejected. 

[Note: HTML markup is below. Please do not edit.]

Reviewers' comments:

Reviewer's Responses to Questions

>Comments to the Author

1. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented. >

Reviewer #1: Yes

**********

>2. Has the statistical analysis been performed appropriately and rigorously? >

Reviewer #1: Yes

**********

>3. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.>

Reviewer #1: Yes

**********

>4. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.>

Reviewer #1: Yes

**********

>5. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)>

Reviewer #1: Dear Author(s),

I read your work with great interest, and I am pleased to congratulate you on your contribution to scientific research.

I believe that the article is novel and interesting, that it has a sufficient impact, and that it adds to the knowledge base. Plagiarism was not detected. The study appears to follow relevant guidelines and provides an original contribution to the existing scientific literature. There are no flaws in the data presented, and there are no misleading or false conclusions.

The current study is scientifically valid. The applied methodology is adequate. The reasons for performing the study are clear. I recommend the article for publication in PLOS ONE in its current form.

Sincerely,

Reviewer

**********

>6. PLOS authors have the option to publish the peer review history of their article (what does this mean? ). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy .>

Reviewer #1: No

**********

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/ . PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at figures@plos.org . Please note that Supporting Information files do not need this step.

PLoS One. 2025 Jun 25;20(6):e0324684. doi: 10.1371/journal.pone.0324684.r003

Author response to Decision Letter 1


8 Jan 2025

Academic Editor

Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process.

Reply: Thank you for your feedback, and that of the reviewers. We have considered all recommendations and comments and have addressed them in our line-by-line responses below.

General Comments:

1. Thank you for your submission. I read it with enthusiasm, as it shows notable merits. Nevertheless, I noticed some concerns that I'd like the authors to address.

Reply 1: Thank you for your review and comments. We have addressed your concerns line-by-line below.

2. Please double-check grammar (e.g. punctuation and verb tense)

Reply 2: Thank you for your comment. We have reviewed the manuscript for grammatical errors.

3. Please double-check refs (e.g. keep the requested style, Vancouver, and correct formatting, punctuation and remove repetitives)

Reply 3: Thank you for this observation. We made the appropriate revisions to our references using tracked changes in the attached manuscript.

Abstract related comments

4. The Abstract is well-structured but could benefit from conciseness. The inclusion of multiple phrases, such as "symptom exaggeration" and "observational studies," makes it somewhat repetitive. I'd suggest reducing redundant terms to improve readability.

Reply 4: Thank you for your comment. We have taken your suggestion and made our abstract more concise including the removal of redundant terms and phrases.

5. The logical transitions between the background, methods, and results are somewhat abrupt. For instance, it's unclear how the analysis of "symptom exaggeration" directly links to "inter-rater reliability" or whether the gender-specific findings are derived from a subset analysis or the primary outcomes. Please provide a clearer link between the methodological approach and the primary research question

Reply 5: We appreciate your feedback. We undertook an assessment of the prevalence of symptom exaggeration, as if it was common it may help explain the poor inter-rater reliability seen in IME assessments. By quantifying the prevalence of symptom exaggeration, clinicians and policymakers can better understand the scope of this issue and work toward improving the objectivity and reliability of IMEs. We added the following to the abstract under the introduction section:

“This study explored the prevalence of symptom exaggeration among IME examinees in North America, which if common would represent an important consideration for improving the reliability of IMEs.”

As for clarifying the methods related to the gender-specific findings we added the following in the abstract under the methods section:

“We …. explored potential subgroup effects for sex, age, education, clinical condition, and confidence in the reference standard.”

6. The stats parameters (e.g. p-values, effect sizes, CIs) are provided, but more detail on the significance of these values will enhance the depth of the findings. For example, how do these rates of symptom exaggeration in IMEs impact policy, clinical practice, or future research directions? Offering these insights would broaden the reliability.

Reply 6: Thank you. The low certainty in the findings highlighted the need for further research which we added to the conclusion of the abstract which now reads as follows:

“The high prevalence of symptom exaggeration among IME attendees provides a compelling rationale for IME evaluators to formally explore this issue. Future research should establish the reliability and validity of current evaluation criteria for symptom exaggeration and develop a structured IME assessment approach.”

7. Sample size details are provided, but the inclusion of additional information, such as variations across cohorts or specific diagnostic categories, would be useful. For instance, were patients with traumatic brain injuries or chronic pain assessed differently? This would offer a clearer understanding of subgroup analysis and generalisation.

Reply 7: Thank you. We have added clarifications, while attempting to remain concise, under the methods and results sections of the abstract. In terms of methods the added text is as noted in our response to comment 5. As for additions to the results, it now reads as follows:

“We found no significant subgroup effects for type of clinical condition, confidence in the reference standard, age, or education.”

8. A brief explanation of whether the studies had high or low certainty is important.

Reply 8: Thank you for your comment. We made clarifications to the results section by (1) describing the certainty of the body of evidence for our outcome of interest and (2) using informative statements to communicate our findings as per GRADE guidelines 26 (https://www.sciencedirect.com/science/article/pii/S0895435619304160). The language was revised in the abstract under the results section as follows:

“We found low certainty evidence suggesting that studies with a greater proportion of women (≥40% vs. <40%) may be associated with higher rates of exaggeration: 47% (95%CI 36 to 58) vs. 31% (95%CI 28 to 35; test of interaction p=0.02).”

9. The abstract briefly touches on gender differences but would benefit from elaborating on why women show higher rates of symptom exaggeration in these studies. Consider a brief sentence.

Reply 9: Thank you for this observation. We added a sentence to the Results section of the abstract as follows:

“This difference may be due to biological differences, greater bodily awareness, or higher rates of negative affectivity.”

10. The findings could be expanded to include more specific quantitative details, such as the range of exaggeration percentages across the cohorts or any significant moderators or predictors found in the meta-analysis.

Reply 10: We agree. We have revised and added more findings to our results section of the abstract as follows:

“Prevalence of symptom exaggeration across studies ranged from 17% to 67%. We found low certainty evidence suggesting that studies with a greater proportion of women (≥40% vs. <40%) may be associated with higher rates of exaggeration: 47% (95%CI 36 to 58) vs. 31% (95%CI 28 to 35; test of interaction p=0.02). We found no significant subgroup effects for type of clinical condition, confidence in the reference standard, age, or education.”

11. The conclusion is concise but could be expanded to include a brief part on how symptom exaggeration influences clinical decision-making during IMEs. Additionally, mentioning how these findings could influence future research or policy changes is essential;

Reply 11: As noted in our response to comment 6 above, we added information to the conclusion section to expand on the implications of the findings on research and practice.

12. The abstract would be improved by providing the main strengths, such as prevalence rates for symptom exaggeration or differences in exaggeration across different conditions;

Reply 12: We appreciate this feedback and have added information on the range of prevalence rates across studies and noted that we did not find a credible subgroup effect between the overall prevalence and different clinical conditions, confidence in the reference standard, age, or education. This was added to the results section of the abstract as described above in response to comment number 10.

13. The term "symptom exaggeration" is mentioned without any clear definition. Please elaborate

Reply 13: We added a short statement to the background of the abstract as follows:

“…This may be affected by symptom exaggeration where patients may feel pressure to fully convey their level of impairment to qualify for incentives.”

Introduction and objective related comments

14. The transition between the background on disabilities and the introduction of IMEs feels somewhat confusing. Providing a smoother transition into how symptom exaggeration affects IMEs will improve your argument;

- There are relevant stats about disabilities in North America, which effectively establishes the context. However, there is a need for a more explicit link between this background information and the objective. For example, while the IME process is mentioned, the transition to the issue of symptom exaggeration is quite abrupt

Reply 14: We have considered your feedback and added text (that is underlined below) to further clarify our objective. The last paragraph of our introduction now reads as follows:

“Patients referred for IMEs often present with subjective complaints (e.g., mental illness, chronic pain) and may feel pressure to emphasize their level of impairment to qualify for wage replacement benefits, receiving time off work, or other incentives (3, 9, 10). Whether or not IME assessors consider symptom exaggeration therefore has the potential to lead to very different conclusions; however, the prevalence of exaggeration among IME attendees is uncertain and individual studies report rates as low as 18% or as high as 68%. To address this gap in the literature, we undertook the first systematic review of observational studies to explore the prevalence of symptom exaggeration among IME examinees in North America.”

15. The Introduction touches on symptom exaggeration as a factor affecting IME outcomes. Nevertheless, there is the need to state how the authors would measure it or what variables will be examined (e.g. mental health, pain severity).

Reply 15: We agree and described in our methods section the types of studies and criteria used that we considered eligible in this study as noted below:

“As there is no singular reliable and valid criteria (reference standard) in the literature that is used to assess for symptom exaggeration, we included known group study designs that defined their reference standard based on criteria incorporating both clinical findings and performance on psychometric testing to classify individuals as exaggerating (within diagnostic test terminology, the target positive group), or not exaggerating (the target negative group) their symptoms (17, 18). Examples of two commonly used known group designs are the Slick, Sherman, and Iverson criteria for malingered neurocognitive dysfunction (19) and the Bianchini, Greve, & Glynn criteria for malingered pain-related disability (20).”

16. Explicitly state how your study addresses gaps or contradictions found in previous studies. For example, why past standardisation efforts failed. More emphasis on the relevance and novelty of the study is essential.

Reply 16: We undertook the first systematic review of symptom exaggeration among IME attendees. This has been explicitly stated and clarified in our introduction.

17. While the IME concept and symptom exaggeration are introduced, there is little detail about the methodology or variables that will be assessed;

- Please articulate your hypotheses with the rationale;

- To improve flow, consider adding a sentence explaining why IMEs are prone to variability, and how symptom exaggeration may influence;

- Summarise the objectives and hypotheses clearly in one sentence.

Reply 17: Thank you for your feedback. As noted in the reply to comments 14 we have now added the below to further clarify our objective and reasons for it:

“Whether or not IME assessors consider symptom exaggeration therefore has the potential to lead to very different conclusions; however, the prevalence of exaggeration among IME attendees is uncertain and individual studies report rates as low as 18% or as high as 68%.”

18. Overall, there are some worrying aspects in Introduction:

The Introduction needs a clearer flow. Consider providing the background and context of IMEs, then move into the issue of symptom exaggeration. After that, it's important to debate what previous studies have found and highlight their gaps. You can then explain how this study will ease those. Concluding with the rationale, objectives, and hypotheses will give the Introduction a strong structure. Consider to mention why the study is important since it'll make more engaging and relevant for a broader audience

Reply 18: Thank you for your feedback. We have made changes accordingly to ensure clarity and better flow as noted in our responses to comments 14 and 17.

19. The text shifts abruptly between topics, like disability stats and symptom exaggeration, without clearly connecting them or stating the rationale. Brief mentioning the methods used to measure symptom exaggeration is really important

Reply 19: The methods for measuring symptom exaggeration that we considered eligible are defined in the methods section under eligibility criteria as noted in our response to comment 15.

20. The authors could provide more context to strengthen their argument. IMEs often focus too narrowly on the biomedical model, overlooking psychosocial factors and work-related conditions, which influences the fairness and consistency of evaluations. The lack of standardisation and unified guidelines contributes to these issues. Additionally, biases in IME practices, such as favouring employer-provided information over patient accounts, lead to underdiagnosis and lower disability ratings compared to treating physicians. This bias fosters distrust and adds to the inconsistencies in the IME process. For more details, see 10.1097/PEC.0000000000000487 and 10.1007/978-3-319-71906-1

Reply 20: Thanks for this citation, which we have now included the justify the following added statement:

“Independent evaluators may focus too narrowly on a biomedical model to explain symptoms, without sufficient attention to psychosocial and work-related factors.”

21. Another aspect contributing to IME variability is the difference in approaches, prognoses, and standards among evaluators. Studies show that diagnostic discordance rates can exceed 80%, with higher discordance linked to increased mortality (OR ≈ 1.2), underscoring the need for standardised protocols. Inconsistent assessments can lead to delayed treatment, denied disability benefits, and declining patient quality of life. Standardisation would improve accuracy and consistency, reducing patient burden and enhancing trust in IMEs. Furthermore, the focus on the biomedical model overlooks the role of psychological and social factors, suggesting that adopting a biopsychosocial approach could yield more accurate evaluations

Reply 21: In terms of the need for standardised protocols, as noted in our introduction, a recent systematic review suggested that standardization of the assessment process may improve the reliability of IMEs; however, two subsequent studies have failed to support this hypothesis (Reference: Kunz R, von Allmen DY, Marelli R, Hoffmann-Richter U, Jeger J, Mager R, et al. The reproducibility of psychiatric evaluations of work disability: two reliability and agreement studies. BMC psychiatry. 2019;19(1):1-15.). This is what led us to explore another potential source of variability in IME assessments which was the prevalence of symptom exaggeration.

22. The Introduction mentions IMEs and symptom exaggeration, but it fails to clearly state the objectives or what the study aims to explore. Providing specific objectives and hypotheses would make it easier for readers to understand the focus of the research

Reply 22: We have clarified this and added to the text as noted in our responses to comments 14, 17 and 18.

23. Explain how the variability in IME outcomes could be influenced by symptom exaggeration, tying it to the broader context of disability assessments

Reply 23: Thank you for your feedback. As noted in the reply to comments 14 and 17 we have now added the text underlined below to further clarify our objective and reasons for it:

“Another potential source of variability in IME assessments is symptom exaggeration (3). Patients referred for IMEs often present with subjective complaints (e.g., mental illness, chronic pain) and may feel pressure to emphasize their level of impairment to qualify for wage replacement benefits, receiving time off wo

Attachment

Submitted filename: IME_ Response to Reviewers_(Nov 21_2024).docx

pone.0324684.s016.docx (78.7KB, docx)

Decision Letter 1

Thiago Fernandes

>PONE-D-24-31136R1>>Prevalence of Symptom Exaggeration Among North American Independent Medical Evaluation Examinees: A systematic review of observational studies>>PLOS ONE

Dear Dr. Busse,

Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process.

==============================

Thank you for your valuable submission.

That said, there are two main concerns in this round.

Before all, I’d like to mention that I read with enthusiasm and think the study has merits to reach a broad audience. This is the main - and most important - aspect to consider in this lengthy round. As noticed while re-reading (and as pointed out by the reviewer), many of the previously raised concerns remain unaddressed.

It appears that certain important adjustments were not fully considered. The reviewer's comments highlight the need for a more detailed rebuttal, particularly clarifying and addressing all the raised concerns. Although lengthy rounds are not common at this point, I trust that with continued effort, the ms has potential and can be refined to meet the required standards.

I understand this is not the best news, but in its current form, the ms cannot proceed. The concerns raised require considerable adjustments and further work. If the authors are willing to address these issues, I would be happy to reassess the ms after the necessary changes are made.

==============================

Please submit your revised manuscript by Apr 05 2025 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org . When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

Please include the following items when submitting your revised manuscript:>

  • A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'.

  • A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.

  • An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols . Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols .

We look forward to receiving your revised manuscript.

Kind regards,

Thiago P. Fernandes, PhD

Academic Editor

PLOS ONE

[Note: HTML markup is below. Please do not edit.]

Reviewers' comments:

Reviewer's Responses to Questions

>Comments to the Author

1. If the authors have adequately addressed your comments raised in a previous round of review and you feel that this manuscript is now acceptable for publication, you may indicate that here to bypass the “Comments to the Author” section, enter your conflict of interest statement in the “Confidential to Editor” section, and submit your "Accept" recommendation.>

Reviewer #2: (No Response)

Reviewer #3: (No Response)

**********

>2. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented. >

Reviewer #2: Partly

Reviewer #3: Yes

**********

>3. Has the statistical analysis been performed appropriately and rigorously? >

Reviewer #2: Yes

Reviewer #3: Yes

**********

>4. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.>

Reviewer #2: No

Reviewer #3: Yes

**********

>5. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.>

Reviewer #2: Yes

Reviewer #3: Yes

**********

>6. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)>

Reviewer #2: The current manuscript is a revision about the prevalence of symptom exaggeration among IME examinees. I did not review the initial version of this manuscript, but reviewed the revision. I have checked my comments with the initial review comments, please see below.

Major:

Symptom overreporting may have multiple causes (see Merckelbach et al., 2019). Symptom exaggeration due to external incentives is related to malingering. Your reply to point 64 states: "We noticed that most studies use the terms interchangeably and often label their criteria as criteria to assess malingering. However, when using the term malingering one presumes intent. Thus we felt it was an important suggestion to note in our discussion that considering how ‘clinicians are, understandably and appropriately, hesitant to assign a label of malingering…. (45)’ the term symptom exaggeration may be a more appropriate replacement." Indeed, malingering equals symptom overreporting, but symptom overreporting does not necessarily mean malingering. In the context of IME, malingering might be more obvious to use for the manuscript, though in clinical practice one should not always assume but rather take malingering into account.

Both the introduction and discussion are very short, relating to the point 29 raised in the previous round. In my opinion, these comments are not sufficiently dealt with. Also the transitions between different topics within the Introduction and Discussion could be improved.

It is stated that French and Spanish studies were included to reduce language bias, but what about culture bias? You state this later on, but this is something that could be taken into account at an earlier level.

Risk of bias assessment: also relating to point 36> why were no tools used to assess risk of bias? E.g., Cochrane's? Also, what was the measure of agreement (e.g., Kappa)? The selection seems very subjective.

Table 2: what do the levels of inconsistency mean, especially 'serious inconsistency'? How does this relate to your conclusions?

It is concluded that the review found no evidence for differences in the prevalence of symptom exaggeration based on clinical condition, but most patients among studies eligible for our review presented with either mild TBI or chronic pain'. This was not a (preregistered) research question, and how to conclude this based on only two clinical conditions? Other clinical diagnoses might have symptom exaggeration, but are not included in the current review.

Minor:

Some reasoning errors in conclusions; if women report higher rates of symptom exaggeration, then it is logically that studies that included more women show higher raters.

There are still some punctuation errors (extra comma use, delete space) throughout the manuscript. At times the phrasing (e.g., p.23 line 72 ('Such concerns...) could be improved.

For research transparency, what were the exact key words? I know they are stated in supplemental materials and that the apparently take 3 pages to explain, but I do believe these could be stated clearly in the manuscript.

Reviewer #3: Dear authors of the article, I want to thank you for a well-prepared article.

The undoubted advantages of the article are:

The text of the article is logically structured and understandable to the reader. The language meets the level and requirements of the scientific language. The article uses modern data analysis, adequate to the purpose of the study. The research is of high importance for science and practice. The authors openly and adequately note in the article limitations, implications for future research and practice.

I certainly recommend this article for publication in the journal, but I have some questions and recommendations for improvement.:

1. why were only these databases CINAHL, EMBASE, MEDLINE and PsycINFO used for analysis?

2. In the discussion section, it is necessary to go beyond America and compare scientific data with other countries and continents

3. At the end of the paragraph "Eligible studies", clarify what does "large sample size" mean (what is the sample size in numbers)?

4. In the paragraph "We rated confidence in the reference standard as either..." describe in more detail (if possible, give examples) what ‘weak’, ‘moderate’ or ‘strong’ mean?

**********

>7. PLOS authors have the option to publish the peer review history of their article (what does this mean? ). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy .>

Reviewer #2: Yes:  Sanne Houben

Reviewer #3: Yes:  Olga Terekhina

**********

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/ . PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at figures@plos.org . Please note that Supporting Information files do not need this step.

PLoS One. 2025 Jun 25;20(6):e0324684. doi: 10.1371/journal.pone.0324684.r005

Author response to Decision Letter 2


2 Apr 2025

Reviewer 2 Comments

1. Symptom overreporting may have multiple causes (see Merckelbach et al., 2019). Symptom exaggeration due to external incentives is related to malingering. Your reply to point 64 states: "We noticed that most studies use the terms interchangeably and often label their criteria as criteria to assess malingering. However, when using the term malingering one presumes intent. Thus we felt it was an important suggestion to note in our discussion that considering how ‘clinicians are, understandably and appropriately, hesitant to assign a label of malingering…. (45)’ the term symptom exaggeration may be a more appropriate replacement." Indeed, malingering equals symptom overreporting, but symptom overreporting does not necessarily mean malingering. In the context of IME, malingering might be more obvious to use for the manuscript, though in clinical practice one should not always assume but rather take malingering into account.

Reply 2: Thank you for raising this important point regarding the distinction between malingering and symptom exaggeration or overreporting. In our review, we avoided the term malingering because it presupposes evidence of intent, which the included studies did not investigate. None of the primary sources explicitly assessed whether participants deliberately feigned symptoms, so we adopted the more neutral phrase symptom exaggeration to describe the findings without attributing conscious deception. We trust that this explanation provides a clearer rationale for our choice of terminology. We added the following material to the background:

“Also, terminology such as exaggeration, malingering, or over-reporting are defined inconsistently across studies, making it difficult to distinguish intentional deception from psychological amplification of distress.”

3. Both the introduction and discussion are very short, relating to the point 29 raised in the previous round. In my opinion, these comments are not sufficiently dealt with. Also the transitions between different topics within the Introduction and Discussion could be improved.

Reply 3: We expanded and revised our introduction and discussion sections and reviewed both to improve transitions.

4. It is stated that French and Spanish studies were included to reduce language bias, but what about culture bias? You state this later on, but this is something that could be taken into account at an earlier level.

Reply 4: Thank you for raising the issue of culture bias. Although we included studies in English, French, and Spanish to reduce language bias, we acknowledge that restricting our review to a North American setting does not fully address cultural variability. Indeed, cultural factors (e.g., varying attitudes toward disability) may influence how patients report symptoms or interact with examiners. This may have an effect on both the prevalence of symptom exaggeration and assessors’ interpretations. Our review did not find studies addressing this issue specifically within our inclusion criteria; therefore, we could not explore these possible differences. We agree that future research should investigate how cultural diversity affects IME outcomes, with attention to factors such as language barriers, health beliefs, and potential biases among both examinees and examiners. We have now clarified in the discussion section that this is a limitation of the included studies (and therefore our review), and highlighted exploration of the impact of cultural bias on IME findings for future research.

5. Risk of bias assessment: also relating to point 36> why were no tools used to assess risk of bias? E.g., Cochrane's? Also, what was the measure of agreement (e.g., Kappa)? The selection seems very subjective.

Reply 5: Studies eligible for our review were observational in design, and thus the Cochrane tool (designed for Randomized Controlled Trials) was not appropriate. Instead, we collaborated with content experts and methodologists to develop key criteria specific to known group designs (e.g., sample representativeness, reference standard confidence, and missing data). After piloting these criteria, reviewers underwent several rounds of training and independently and in duplicate assessed risk of bias, resolving disagreements by consensus or a third reviewer. Although we did not calculate a kappa statistic, we have clarified these steps in our Methods to illustrate how we increased reliability for our risk-of-bias judgments.

6. Table 2: what do the levels of inconsistency mean, especially 'serious inconsistency'? How does this relate to your conclusions?

Reply 6: In GRADE terminology, “inconsistency” refers to the variation in results across studies (ie, heterogeneity). When GRADE tables state “serious inconsistency,” it means there was substantial variability in effect estimates that we could not adequately explain (for example, no clear reasons for differing results emerged in our subgroup analyses). Unexplained between-study inconsistency lowers our confidence in the pooled estimate, which is reflected in our GRADE assessment of the certainty of evidence. We have expanded on this in a footnote under Table 2.

7. It is concluded that the review found no evidence for differences in the prevalence of symptom exaggeration based on clinical condition, but most patients among studies eligible for our review presented with either mild TBI or chronic pain'. This was not a (preregistered) research question, and how to conclude this based on only two clinical conditions? Other clinical diagnoses might have symptom exaggeration, but are not included in the current review.

Reply 7: We agree. Although our review did not detect differences in the prevalence of symptom exaggeration across clinical conditions, most included studies recruited individuals with mild TBI or chronic pain. Because few other clinical populations were studied, we cannot draw definitive conclusions for other diagnoses. We have included language in the results (prevalence of symptom exaggeration and additional analyses section) and strengths and limitations section of our review clarifying the populations that were represented and highlighting that there are other patient groups that were underrepresented or absent in our review and where more research is needed.

“We found no significant subgroup effects for type of clinical condition (mild TBI versus chronic pain versus other conditions)…”

“In terms of limitations, we restricted our review to IMEs conducted in North America and eligible studies focused mainly on chronic pain and TBI. The generalizability of our findings to other jurisdictions, contexts, and clinical conditions, is uncertain.”

8. Some reasoning errors in conclusions; if women report higher rates of symptom exaggeration, then it is logically that studies that included more women show higher raters.

Reply 8: We anticipated that studies eligible for our review would report a range regarding the prevalence of symptom exaggeration among individuals presenting for independent medical evaluations. Accordingly, we conducted several subgroup analyses to see if we could identify one or more factors that helped explain between-study variability. Only one of our subgroup analyses, the proportion of female participants, showed a credible subgroup effect. This result supported the reporting of separate rates of symptom exaggeration between men and women.

9. There are still some punctuation errors (extra comma use, delete space) throughout the manuscript. At times the phrasing (e.g., p.23 line 72 ('Such concerns...) could be improved.

Reply 9: Thank you. We have reviewed the manuscript more closely and corrected any errors.

10. For research transparency, what were the exact key words? I know they are stated in supplemental materials and that the apparently take 3 pages to explain, but I do believe these could be stated clearly in the manuscript.

Reply 10: We have added keywords used in our literature search to the manuscript, and it now reads as:

The search strategies were developed using a validation set of known relevant articles and included a combination of MeSH headings and free text key words, such as malinger* or litigation or litigant or "insufficient effort" and “independent medical examination” or “independent medical evaluation” or “disability” or “classification accuracy”.

Review 3 Comments

11. Dear authors of the article, I want to thank you for a well-prepared article.

The undoubted advantages of the article are: The text of the article is logically structured and understandable to the reader. The language meets the level and requirements of the scientific language. The article uses modern data analysis, adequate to the purpose of the study. The research is of high importance for science and practice. The authors openly and adequately note in the article limitations, implications for future research and practice.

I certainly recommend this article for publication in the journal, but I have some questions and recommendations for improvement.

Reply 11: Thank you. We appreciate you taking the time to review our manuscript and address your questions below.

12. Why were only these databases CINAHL, EMBASE, MEDLINE and PsycINFO used for analysis?

Reply 12: We selected these four databases (CINAHL, EMBASE, MEDLINE, and PsycINFO) because they collectively capture the key biomedical, nursing/allied health, and psychological/psychiatric literature most relevant to our review topic. Also, we evaluated, with the help of a senior medical librarian, a validation set of known eligible studies to confirm our search strategy and database coverage. When our search retrieved all studies in our validation set, we were reassured that our choice of databases and search approach was sufficiently comprehensive to capture the relevant literature in this area. We further supplemented our search by scanning the reference lists of included articles, and having experts review the list of included studies, providing additional checks for completeness.

13. In the discussion section, it is necessary to go beyond America and compare scientific data with other countries and continents

Reply 13: Thank you for your suggestion. We have now added to the discussion section data comparing prevalence of symptom exaggeration beyond America. The text reads as such:

“Although our review focused on IMEs in North America, data from other regions also suggest high rates of symptom exaggeration. An observational study in Spain reported that of 1,003 participants (61.5% female), drawn from unselected undergraduates, advanced psychology students, the general population, forensic psychologists, and forensic/legal medicine physicians, one-third reported having feigned symptoms or illness (44). Data from Germany and the Netherlands suggest that one‐fifth to one‐third of clients in forensic or insurance contexts exhibit symptom overreporting (45). Further, a Swiss study found that 28% to 34% of individuals undergoing medico‐legal evaluations demonstrated probable or definite symptom exaggeration (46).”

14. At the end of the paragraph "Eligible studies", clarify what does "large sample size" mean (what is the sample size in numbers)?

Reply 14: Regarding this sentence: “In cases where multiple studies had population overlap, we included only the study with the larger sample size.” We were concerned that some datasets may be used by multiple investigators, and the risk of double-counting the same patients in such instances. As such, we set a rule to avoid using data from the same patients more than once by excluding studies that reported on the same dataset. In order to pick a study to exclude (if this situation arose) we decided to exclude the study that used less data (the study with a smaller sample size). As this was a relative decision (in some cases, a smaller study could involve 200 individuals, and in others 1000 individuals) we did not pre-specify a sample size threshold. As it turned out, none of our studies had overlapping cohorts.

15. In the paragraph "We rated confidence in the reference standard as either..." describe in more detail (if possible, give examples) what ‘weak’, ‘moderate’ or ‘strong’ mean?

Reply 15: We have revised this section to provide more detail and examples. It now reads as:

“ We categorized the reference standard and rated our confidence in it as either : (i) ‘weak’ when the study declared a known-group design, however its only criterion for identifying symptom exaggeration was below-chance performance on forced-choice symptom validity testing without any corroborating clinical observations or inconsistencies in medical records. For example, a patient with a mild ankle sprain is labeled as exaggerating exclusively because they fail a below‐chance forced‐choice test of pain threshold, with no clinical exam or review of documented pain or functional abilities; (ii) ‘moderate’ where most patients exaggerating symptoms were identified by forced symptom validity testing results, but some cases could be confirmed using other credible indicators. For example, a claimant insists they cannot remember simple details of their daily routine (e.g., the route to their kitchen), yet is casually observed navigating complex tasks with no apparent cognitive difficulty; or (iii) ‘strong’ where exaggeration was determined by either forced symptom validity testing results or other credible clinical evidence. For example, a clinical finding that would classify a patient presenting with persistent post-concussive complaints after a very mild head injury as exaggerating symptoms would include claims of remote memory loss (e.g., loss of spelling ability).”

Attachment

Submitted filename: IME_Response to reviewers_April 2_2025.docx

pone.0324684.s017.docx (33.8KB, docx)

Decision Letter 2

Thiago Fernandes

Prevalence of Symptom Exaggeration Among North American Independent Medical Evaluation Examinees: A systematic review of observational studies

PONE-D-24-31136R2

Dear Dr. Busse,

We’re pleased to inform you that your manuscript has been judged scientifically suitable for publication and will be formally accepted for publication once it meets all outstanding technical requirements.

Within one week, you’ll receive an e-mail detailing the required amendments. When these have been addressed, you’ll receive a formal acceptance letter and your manuscript will be scheduled for publication.

An invoice will be generated when your article is formally accepted. Please note, if your institution has a publishing partnership with PLOS and your article meets the relevant criteria, all or part of your publication costs will be covered. Please make sure your user information is up-to-date by logging into Editorial Manager at Editorial Manager®  and clicking the ‘Update My Information' link at the top of the page. If you have any questions relating to publication charges, please contact our Author Billing department directly at authorbilling@plos.org.

If your institution or institutions have a press office, please notify them about your upcoming paper to help maximize its impact. If they’ll be preparing press materials, please inform our press team as soon as possible -- no later than 48 hours after receiving the formal acceptance. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

Kind regards,

Thiago P. Fernandes, PhD

Academic Editor

PLOS ONE

Additional Editor Comments (optional):

Reviewers' comments:

Reviewer's Responses to Questions

>Comments to the Author

1. If the authors have adequately addressed your comments raised in a previous round of review and you feel that this manuscript is now acceptable for publication, you may indicate that here to bypass the “Comments to the Author” section, enter your conflict of interest statement in the “Confidential to Editor” section, and submit your "Accept" recommendation.>

Reviewer #3: All comments have been addressed

**********

>2. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented. >

Reviewer #3: Yes

**********

>3. Has the statistical analysis been performed appropriately and rigorously? >

Reviewer #3: Yes

**********

>4. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.>

Reviewer #3: Yes

**********

>5. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.>

Reviewer #3: Yes

**********

>6. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)>

Reviewer #3: The authors correctly and competently provided answers to all questions and comments, specifying in detail the corrections that were made to the text. The authors have done a lot of work in the process of preparing the article. The article turned out to be interesting for the professional community and has theoretical and practical significance.The authors openly and adequately note in the article limitations, implications for future research and practice. I certainly recommend this article for publication in the journal

**********

>7. PLOS authors have the option to publish the peer review history of their article (what does this mean? ). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy .>

Reviewer #3: Yes:  Olga Terekhina

**********

Acceptance letter

Thiago Fernandes

PONE-D-24-31136R2

PLOS ONE

Dear Dr. Busse,

I'm pleased to inform you that your manuscript has been deemed suitable for publication in PLOS ONE. Congratulations! Your manuscript is now being handed over to our production team.

At this stage, our production department will prepare your paper for publication. This includes ensuring the following:

* All references, tables, and figures are properly cited

* All relevant supporting information is included in the manuscript submission,

* There are no issues that prevent the paper from being properly typeset

You will receive further instructions from the production team, including instructions on how to review your proof when it is ready. Please keep in mind that we are working through a large volume of accepted articles, so please give us a few days to review your paper and let you know the next and final steps.

Lastly, if your institution or institutions have a press office, please let them know about your upcoming paper now to help maximize its impact. If they'll be preparing press materials, please inform our press team within the next 48 hours. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

If we can help with anything else, please email us at customercare@plos.org.

Thank you for submitting your work to PLOS ONE and supporting open access.

Kind regards,

PLOS ONE Editorial Office Staff

on behalf of

Dr. Thiago P. Fernandes

Academic Editor

PLOS ONE

Associated Data

    This section collects any data citations, data availability statements, or supplementary materials included in this article.

    Supplementary Materials

    S1 Table. Search strategies.

    (DOCX)

    pone.0324684.s001.docx (30.4KB, docx)
    S2 Table. Risk of bias assessment.

    (DOCX)

    pone.0324684.s002.docx (34.4KB, docx)
    S3 Table. ICEMAN criteria to assess credibility of subgroup effect of female % and prevalence.

    (DOCX)

    pone.0324684.s003.docx (28.7KB, docx)
    S4 Table. Psychometric properties of tests included in symptom exaggeration criteria with list of references.

    (DOCX)

    pone.0324684.s004.docx (66KB, docx)
    S5 Table. Included and excluded studies at full text screening with reasons.

    (DOCX)

    pone.0324684.s005.docx (87.8KB, docx)
    S6 Table. Data extracted from included studies.

    (DOCX)

    pone.0324684.s006.docx (36.9KB, docx)
    S1 Fig. Meta-regression for proportion of females among 42 studies (p = 0.16).

    (DOCX)

    pone.0324684.s007.docx (172.6KB, docx)
    S2 Fig. a- Funnel plots of overall prevalence (Egger’s test p = 0.13) and b- prevalence in subgroup of studies with female proportion <40% (Egger’s test p = 0.16).

    (DOCX)

    pone.0324684.s008.docx (170.7KB, docx)
    S3 Fig. Subgroup analysis for type of conditions (test of interaction p = 0.95).

    (DOCX)

    pone.0324684.s009.docx (209.5KB, docx)
    S4 Fig. Subgroup analysis for confidence in reference standard (test of interaction p = 0.84).

    (DOCX)

    pone.0324684.s010.docx (176KB, docx)
    S5 Fig. Subgroup analysis for similar age and/or education between groups (test of interaction p = 0.47).

    (DOCX)

    pone.0324684.s011.docx (205.4KB, docx)
    S6 Fig. Meta-regression for average age among 46 cohorts (p = 0.18).

    (DOCX)

    pone.0324684.s012.docx (142.1KB, docx)
    S7 Fig. Meta-regression for average education level among 45 cohorts (p = 0.65).

    (DOCX)

    pone.0324684.s013.docx (125.7KB, docx)
    S1 Checklist. Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) Checklist.

    (DOCX)

    pone.0324684.s014.docx (35.6KB, docx)
    S2 Checklist. Meta-analysis of Observational Studies in Epidemiology (MOOSE) checklist.

    (DOCX)

    pone.0324684.s015.docx (1.3MB, docx)
    Attachment

    Submitted filename: IME_ Response to Reviewers_(Nov 21_2024).docx

    pone.0324684.s016.docx (78.7KB, docx)
    Attachment

    Submitted filename: IME_Response to reviewers_April 2_2025.docx

    pone.0324684.s017.docx (33.8KB, docx)

    Data Availability Statement

    All relevant data are within the paper and its Supporting Information files.


    Articles from PLOS One are provided here courtesy of PLOS

    RESOURCES