Skip to main content
Sage Choice logoLink to Sage Choice
. 2026 Aug 17;46(7):819–824. doi: 10.1177/0272989X261471178

The Undertesting Bias in Clinical Decision Making: The Role of Test Dichotomization

Murat C Mungan 1,✉
PMCID: PMC13601696  PMID: 42605268

Abstract

Threshold models, formalized by Pauker and Kassirer, guide clinical decisions about testing and treatment by partitioning pretest probabilities into three action zones using two threshold pretest probabilities: the test threshold and the test–treatment threshold. This derivation implicitly assumes that diagnostic tests possess fixed sensitivity and specificity, effectively treating them as binary predictors. However, many diagnostic tests yield continuous or ordinal outputs, in which case the optimal sensitivity–specificity pair varies with the pretest probability. I demonstrate that abstracting from this dependency leads to an undertesting bias: the testing threshold is inflated and/or the test–treatment threshold is deflated, resulting in a narrower than optimal testing window. This bias systematically undervalues continuous diagnostic tests and leads to their underuse. Clinical decision-making models and guidelines should therefore recognize that optimal test score cutoffs depend on pretest probabilities to avoid this systematic underuse of diagnostic testing.

Keywords: threshold models, diagnostic tests, pretest probabilities, Pauker-Kassirer model, clinical decision making, dichotomization

Introduction

Threshold models are widely used to analyze clinical decisions about whether to treat, test, or withhold intervention under diagnostic uncertainty.1–4 Pauker and Kassirer 1 formalized this approach by explicitly deriving a “testing threshold” (denoted Tt ) and a “test–treatment threshold” (denoted Ttrx>Tt ), which partition pretest probabilities (ie, the clinician’s assessment of the patient having the disease prior to conducting the diagnostic test) into 3 regions. When the pretest probability is below the testing threshold Tt , it is optimal to withhold both testing and treatment, and when it is above the test–treatment threshold Ttrx , it is optimal to treat without testing. In the intermediate range where the pretest probability is between these two values, it is optimal to conduct a diagnostic test to decide whether to treat the patient. Figure 1 depicts these thresholds.

Figure 1.

Figure 1.

Pretest probability action thresholds. The pretest disease probability, p , is on the horizontal axis.

This model has been very influential, both in the theoretical and applied literatures on medical decision making, and it continues to be treated as one of the major advances in the field in recent discussions of decision threshold models.4,5 Variants and extensions of it continue to be developed and applied, for example, in work that introduces treatment thresholds into diagnostic test evaluation or uses threshold analysis to structure decision tables,6,7 as well as in recent extensions that incorporate therapeutic risk. 8

Despite this interest, one simplifying assumption in this approach has not yet received much attention. Specifically, Pauker and Kassirer assume that the diagnostic test possesses a fixed sensitivity and fixed specificity, which is equivalent to it being a binary predictor,9 –11 that is, one yielding a positive or negative result, as in the case of simple SARS-CoV-2 rapid antigen tests. 12 However, numerous laboratory assays, imaging modalities, multivariable risk scores, and signal-to-cutoff–based assays yield continuous or ordinal outputs, such as biomarker concentrations; imaging severity scores; predicted risk percentages; or the cycle threshold value ( Ct ) in COVID polymerase chain reaction tests, rather than dichotomous positive/negative results.13 –15 The binary predictor assumption is thus a form of dichotomization—a process explicitly cautioned against in this journal 16 —which can inadvertently lead to misleading conclusions.

Here, I demonstrate that when the diagnostic test is continuous, employing Pauker and Kassirer’s simplifying assumption systematically leads to an undertesting bias, that is, an inefficient tendency to underutilize diagnostic tests. More precisely, a clinician who adopts Pauker and Kassirer’s thresholds will use a narrower-than-optimal testing range. The rationale is that, compared with a clinician forced to use a test fixed at a particular sensitivity–specificity setting, a clinician who can set sensitivity and specificity anywhere on the test’s ROC curve can extract more value from the test: for patients with a high pretest probability of disease, they can increase sensitivity while decreasing specificity; for patients with a low pretest probability, they can do the reverse—making testing more worthwhile over a broader range of pretest probability values (and therefore for more patients). This result is formalized in an appendix in the supplementary material (henceforth ”Appendix”), because the rationale behind it is intuitive, as explained in further detail next.

When the test is continuous, its sensitivity and specificity can be adjusted. In fact, the efficient decision frontiers associated with continuous tests are described by what are often called receiver operating characteristic (ROC) curves,17 –19 which have been studied and used frequently in the medical literature.13 –15,20,21 This causes a third threshold, which is absent in Pauker and Kassirer’s model, to emerge: a test score, as opposed to a threshold probability, determining when the clinician will treat the patient upon administering the diagnostic test. i This threshold, which I call the “treat–withhold threshold” and denote stw , is distinct from the test–treat threshold. The test–treat threshold is a pretest probability that determines whether a diagnostic test should be conducted or whether the patient should be treated without further testing. In contrast, the treat–withhold threshold is a test score that determines whether the patient should be treated in light of all information available to the clinician and is therefore often also called the “test score cutoff.” When a clinician uses the diagnostic test optimally, this threshold is chosen in a utility-maximizing manner, and its value depends on test characteristics as well as the pretest probability; the derivation of this threshold within a model where the clinician chooses between the three options listed in Figure 1 is described in the Appendix. I also developed a simple program that calculates the optimal values of all three thresholds as well as optimal sensitivity–specificity pairs as a function of pretest probability for equal-variance binormal tests. 24 That program can be used to generate examples, in addition to those presented below, to illustrate the points made here.

Table 1 provides a summary of the 3 thresholds for clarity.

Table 1.

Description of the 3 Thresholds.

Threshold Notation Threshold Type Decision Guided
Test Tt Pretest probability Test versus withhold treatment without testing
Test–treat Ttrx Pretest probability Test versus treat without testing
Treat–withhold stw Test score Treat versus withhold based on the test

The treat–withhold threshold has been studied in the medical literature, but in simpler models.5,6,20,25 Although these models derive the optimal treat–withhold threshold as I do here, they focus on the clinician’s binary decision of whether to treat a patient in light of all available information (e.g., a test result), as opposed to the Pauker and Kassirer framework, in which the clinician’s antecedent decision of whether to conduct a test is also studied. ii However, because Pauker and Kassirer consider a binary predictor, the treat–withhold threshold is absent in their more complicated decision problem: there, the clinician receives a test result that is only either positive or negative and decides to treat the patient only when it is positive. 1 Thus, models that study the treat–withhold threshold do not consider the clinician’s initial decision problem, and models that incorporate this initial decision problem do not study the treat–withhold threshold. However, the economic framework I recently proposed 26 can be used to combine the initial testing decision and the optimal threshold adjustment as explained here.

Quite importantly, and as noted in the medical decision-making literature, the treat–withhold threshold is a function of the pretest probability.2,3,20,27 –29 This is because a high pretest probability combined with a moderate test score can yield the same posttest probability of disease as a low pretest probability combined with a high test score. Therefore, the optimal treat–withhold threshold must adjust to reflect the information already available to the clinician. Consistent with this understanding, both empirical observations and normative analyses indicate that clinicians routinely apply different decision thresholds to the same continuous test across clinical contexts.2,4,5,29 Reviews and methodological papers explicitly argue that the optimal treat–withhold threshold depends on disease prevalence and other factors, leading to different recommended cutoffs in screening, diagnostic, and confirmatory settings.6,7,15,20

Theoretical and empirical analyses thus suggest that the treat–withhold threshold is a function of the pretest probability. When one assumes that the diagnostic test has a fixed specificity and sensitivity, one removes this dependency. This, in turn, artificially deflates the value of a test, by assuming that it is incapable of being optimally adjusted in light of other information that the clinician has.

To illustrate, consider an equal variance binormal diagnostic test. The test scores, s , are normally distributed both when the patient has the disease and when he does not, but with different means. If the mean of the disease-distribution is larger than the mean of the no-disease-distribution, it follows that larger test scores are more indicative of the disease. Thus, after a test is conducted, the clinician may choose a cutoff test score, s^ , and recommend treatment only when s>s^ . This test score can be adjusted optimally based on the clinician’s pretest probability estimate or it may be fixed. In the latter case, the diagnostic test has a fixed sensitivity and specificity, and in the former its specificity and sensitivity optimally adjust to the pretest probability.

Figure 2a and b compare the optimal sensitivity and specificity of a test iii against the specificity–sensitivity pair (0.95,0.90) it achieves through a fixed cutoff score ( s^≈1.645 ). This fixed sensitivity–specificity pair is the one used in an example considered by Pauker and Kassirer. The dashed line indicates the pretest probability ( ps^≈0.072 ) for which s^ is also the optimal cutoff score, that is, s^=stw(ps^) . In other words, the fixed test characteristics assumed in calculations coincide with the optimal test characteristics only when the pretest probability happens to equal ps^ ; for other pretest values, these test characteristics diverge. iv Figure 2a and b illustrate that this divergence grows as the pretest probability moves away from ps^ .

Figure 2.

Figure 2.

(a, b) The optimal sensitivity ( sens*(p) ) and specificity ( spec*(p) ) as a function of the pretest probability p are plotted against the fixed sensitivity and specificity assumed (0.9 and 0.95). The optimal values are obtained when the diagnostic test is the unique equal variance binormal test capable of generating the assumed fixed test characteristics in an example in Pauker and Kassirer. 1 The parameters are as assumed in the same example: the benefit of correct treatment is 19 times the cost of administering the test, and the cost of treating a patient without a disease is 2.5 times the cost of conducting a test.

More generally, clinicians who reflect carefully on their pretest assessments can use diagnostic tests in different ways depending on patient characteristics. This is exemplified through the sensitivity–specificity variations that can be optimally adopted by clinicians in Figures 2a and b. Relying on rigid rules of thumb when interpreting diagnostic tests across diverse cases can therefore undermine the clinical process. In fact, assuming one must use a fixed test cutoff across all cases reduces the perceived value of the test. Therefore, the difference between the actual value of the test and the perceived value of the test under the assumption that it must be used with a fixed cutoff is increasing in the gap between the pretest probability and ps^ . This is illustrated in Figure 3.

Figure 3.

Figure 3.

The difference (denoted Δ(p) ) in the utility from using the test when its cutoff is chosen optimally versus when it is set to achieve sensitivity 0.90 and specificity 0.95. The optimal values are obtained using the same assumptions as in Figures 2a and b and are a function of the pretest probability p on the horizontal axis.

This exercise illustrates how using a fixed sensitivity–specificity pair for calculation purposes leads to a low estimated value of using a diagnostic test and a narrower than optimal range between the test and test–treat thresholds. This naturally leads clinicians to use a test less often than is optimal, regardless of which fixed sensitivity–specificity pair is chosen for the calculation. This is illustrated in Figure 4a and b, which depict the optimal test and test–treat thresholds (the dashed curves) versus the Pauker–Kassirer thresholds. The latter are obtained by choosing a particular specificity–sensitivity pair associated with the diagnostic test, and the specificity is varied along the horizontal axis.

Figure 4.

Figure 4.

(a, b) Optimal thresholds (dashed) are plotted against the Pauker–Kassirer thresholds (solid) obtained by assuming the diagnostic test specificity equals the level on the horizontal axis.

These observations demonstrate that using a fixed specificity–sensitivity pair as in Pauker and Kassirer 1 will yield a larger test threshold (i.e., Tt ) and/or a smaller test–treat threshold (i.e., Ttrx ) than is optimal, which implies a narrower than optimal testing window.

This undertesting bias is general and not limited to the particular examples provided here. It follows whenever the diagnostic test is continuous. Neither the specifics of the diagnostic test nor the precise choice of the fixed specificity–sensitivity pair affect this conclusion. The rationale, as explained here, is that assuming the continuous diagnostic test must be used with a fixed sensitivity–specificity pair amounts to undervaluing it and thus leads to its underuse. Because the formalization of these claims provides little further insight, it is relegated to the Appendix.

The bias identified here can be corrected by anticipating the dependency of the diagnostic test’s sensitivity and specificity on the clinician’s pretest assessments. In practice, this insight can be leveraged to develop decision support tools for personalized medicine that integrate individualized pretest probability estimates with optimally adjusted test thresholds.

Supplemental Material

sj-pdf-1-mdm-10.1177_0272989X261471178 – Supplemental material for The Undertesting Bias in Clinical Decision Making: The Role of Test Dichotomization

Supplemental material, sj-pdf-1-mdm-10.1177_0272989X261471178 for The Undertesting Bias in Clinical Decision Making: The Role of Test Dichotomization by Murat C. Mungan in Medical Decision Making

i.

This assumes that the test score is monotonically indicative of the disease. If it is not, the test score can be transformed into a score that has this property (e.g., a likelihood ratio) as described in prior work.22,23

ii.

As explained in the Appendix, the treat–withhold threshold is derived assuming a test is conducted and can be calculated even for pretest probabilities for which it is optimal to not conduct a test. Thus, the standard approach to deriving the treat–withhold threshold is unaffected by the presence of Pauker and Kassirer’s antecedent decision problem of whether to conduct a test.

iii.

Specifically, the scores are distributed normally with equal variance (normalized to 1) and mean 0 when there is no disease and mean ≈2.9265 when the disease is present. This describes the unique set of equal-variance binormal tests that generate the (0.95, 0.90) specificity–sensitivity pair assumed in Pauker and Kassirer. 1

iv.

These observations can also be illustrated on the ROC curve associated with the test: while indifference curves obtained when the pretest probability is ps^ achieve tangency with the ROC curve at the specificity–sensitivity pair (0.95, 0.90) that sits on the ROC curve, the tangency condition is not met for indifference curves obtained with any other pretest probability. Thus, the fixed specificity–sensitivity pair is optimal only when the pretest probability equals ps^ .

Footnotes

ORCID iD: Murat C. Mungan Inline graphic https://orcid.org/0000-0003-1948-6488

Data, Consent, and Ethical Considerations: This technical note presents a theoretical analysis and does not involve human subjects, data collection, or empirical datasets. As such, ethical considerations, consent to participate, consent for publication, and data availability statements are not applicable to this work.

Funding: The author received no financial support for the research, authorship, and/or publication of this article.

The author declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.

Supplemental Material: Supplemental material for this article is available online.

References

  • 1. Pauker SG, Kassirer JP. The threshold approach to clinical decision making. N Engl J Med. 1980;302(20):1109–1117. [DOI] [PubMed] [Google Scholar]
  • 2. Hunink MG, et al. Decision making in health and medicine: integrating evidence and values. 2nd ed. Cambridge University Press; 2014. [Google Scholar]
  • 3. Sox HC, Higgins MC, Owens DK. Medical decision making. 2nd ed. Wiley-Blackwell; 2013. [Google Scholar]
  • 4. Scarffe A, Coates A, Brand K, Michalowski W. Decision threshold models in medical decision making: a scoping literature review. BMC Med Inform Decis Mak. 2024;24:273. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5. Djulbegovic B, Hozo I, Mayrhofer T, Van den Ende J, Guyatt G. Expected utility versus expected regret theory versions of decision curve analysis do not lead to the same results. Med Decis Making. 2014;34(4):567–579. [Google Scholar]
  • 6. Moons KGM, et al. Treatment thresholds in diagnostic test evaluation: an alternative approach. Med Decis Making. 1997;17(4):447–454. [DOI] [PubMed] [Google Scholar]
  • 7. Young MJ, et al. Threshold analysis of decision tables. Med Decis Making. 1990;10(4): 289–297. [DOI] [PubMed] [Google Scholar]
  • 8. Felder S, Mayrhofer T. Threshold analysis in the presence of both the diagnostic and the therapeutic risk. Eur J Health Econ. 2018;19(7):1035–1043. 10.1007/s10198-017-0951-1 [DOI] [PubMed] [Google Scholar]
  • 9. Vickers AJ, Elkin EB. Decision curve analysis: a novel method for evaluating prediction models. Med Decis Making. 2006;26(6):565–574. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10. Dobler CC, Murad MH. Interpreting diagnostic tests with continuous results and no gold standard. BMJ Evid Based Med. 2017;22(6):213–215. [DOI] [PubMed] [Google Scholar]
  • 11. Cantor SB, Kattan MW. Determining the area under the ROC curve for a binary diagnostic test. Med Decis Making. 2000;20(4):468–470. [DOI] [PubMed] [Google Scholar]
  • 12. Manten K, et al. Clinical accuracy of instrument-based SARS-COV-2 antigen diagnostic tests: a systematic review and meta-analysis. Virol J. 2024;21(1):99. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. Margaret Sullivan Pepe. The statistical evaluation of medical tests for classification and prediction. Oxford University Press; 2003. [Google Scholar]
  • 14. Zhou X-H, Obuchowski NA, McClish DK. Statistical methods in diagnostic medicine. 2nd ed. Wiley; 2011. [Google Scholar]
  • 15. Bossuyt PMM, Cohen JF, Gatsonis CA, Korevaar DA. Receiver operating characteristic curve analysis in diagnostic research. Turk J Emerg Med. 2022;22(4):247–256. [Google Scholar]
  • 16. Dawson NV, Weiss R. Dichotomizing continuous variables in statistical analysis: a practice to avoid. Med Decis Making. 2012;32(2):225–226. 10.1177/0272989X12437605 [DOI] [PubMed] [Google Scholar]
  • 17. Egan JP. Signal detection theory and ROC analysis. Academic Press; 1975. [Google Scholar]
  • 18. Weber T. Simple methods for evaluating and comparing binary experiments. Theory Decis. 2010;69:257–288. [Google Scholar]
  • 19. Fluet C, Mungan MC. Oriented data-generating processes: a categorization of ROC curves. Theory Decis. 2026;100:1–37. [Google Scholar]
  • 20. Irwin RJ, Irwin TC. A principled approach to setting optimal diagnostic thresholds: where roc and indifference curves meet. Eur J Intern Med. 2011;22(3):230–234. [DOI] [PubMed] [Google Scholar]
  • 21. Krzanowski WJ, Hand DJ. ROC curves for continuous data. Chapman & Hall/CRC; 2009. [Google Scholar]
  • 22. Mungan MC. Symmetry, presumptions, and the judges design. PLoS One. 2026;21(1):e0340446. 10.1371/journal.pone.0340446 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23. Mungan MC. The Blackstone ratio, modified. J Theor Polit. 2025;38(2):133–155. 10.1177/09516298251403406 [DOI] [Google Scholar]
  • 24. Mungan MC. Equal-variance binormal pretest probability threshold calculator (version 1.0.0) [computer software], 2026. 10.5281/zenodo.21014984 [DOI]
  • 25. Hilden J, Gerds TA. A note on the evaluation of novel biomarkers: do not rely on integrated discrimination improvement and net reclassification index. Stat Med. 2014;33(19):3405–3414. [DOI] [PubMed] [Google Scholar]
  • 26. Mungan MC. Abbreviated judgments and asymmetries, 2025. SSRN working paper. https://ssrn.com/abstract=5384730.
  • 27. Fagan TJ. Letter: nomogram for bayes theorem. N Engl J Med. 1975;293(5):257. [DOI] [PubMed] [Google Scholar]
  • 28. American Society for Microbiology. Why pretest and posttest probability matter in the time of COVID-19. ASM; 2020. [Google Scholar]
  • 29. Ebell MH, Locatelli I, Senn N, Courvoisier DS. Accuracy of practitioner estimates of probability of diagnosis before and after testing in primary care: a randomized clinical trial. JAMA Intern Med. 2021;181(7):987–995. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

sj-pdf-1-mdm-10.1177_0272989X261471178 – Supplemental material for The Undertesting Bias in Clinical Decision Making: The Role of Test Dichotomization

Supplemental material, sj-pdf-1-mdm-10.1177_0272989X261471178 for The Undertesting Bias in Clinical Decision Making: The Role of Test Dichotomization by Murat C. Mungan in Medical Decision Making


Articles from Medical Decision Making are provided here courtesy of SAGE Publications

RESOURCES