Skip to main content
Clinical Orthopaedics and Related Research logoLink to Clinical Orthopaedics and Related Research
letter
. 2021 Jan 4;479(3):636–638. doi: 10.1097/CORR.0000000000001632

Reply to the Letter to the Editor: Can Predictive Modeling Tools Identify Patients at High Risk of Prolonged Opioid Use after ACL Reconstruction?

Ashley B Anderson 1,2,3, George C Balazs 1,2,3, Clare F Grazal 1,2,3, Benjamin K Potter 1,2,3, Jonathon F Dickens 1,2,3, Jonathan A Forsberg 1,2,3,
PMCID: PMC7899541  PMID: 33399402

To the Editor,

Thank you for giving us the opportunity to respond to the letter to the editor [3] concerning our paper that evaluated machine learning models designed to estimate the likelihood of prolonged opioid use after ACL reconstruction [1]. We thank Drs. Wang and Zhu for their interest in our research and for their questions. We are humbled by them, and we welcome the opportunity to respond.

First, the comment by Drs. Wang and Zhu regarding the need to depict all decision curves is well taken. Our intent was to evaluate each of the four modeling techniques using a series of objective criteria, including decision curve analysis. They correctly noted that two of our four models performed similarly, and we did not do a good job supporting our conclusion that the gradient boosting machine model was most suitable for a clinical application. As they suggested, it may be helpful for readers to visualize all of the decision curves in order to directly compare them. We considered this during the manuscript revision process, but determined that including all of the decision curve analysis graphs may be unnecessary. Based on your comments, we now see the error.

In response, please see the four decision curve analysis graphs (Fig. 1). Any improvement in net benefit is clinically relevant, so even subtle differences can be important. In addition, decision curve analyses allow for the evaluation of model performance across a broad range of threshold probabilities, which is one of its many strengths. Readers may recall that by using decision curve analysis, it is possible to identify threshold values above which it might be appropriate to initiate more-intensive opioid counseling and monitoring. For this study we evaluated each model across threshold probabilities from 0.10 to 0.40, which we believed, albeit subjectively, to be clinically relevant for estimating the likelihood of postoperative prolonged opioid use. In doing so, the gradient boosting machine model demonstrated a generalized net benefit of 8% over the logistic regression model. We acknowledge differences such as this can be difficult to discern by simple inspection of the decision curve analysis curves, and we also consider it important to provide objective measures of performance. In future publications comparing model performance, we plan to calculate the area under the decision curve analysis curve across clinically relevant ranges of threshold probabilities. Further, we may employ techniques to determine whether the probability of net benefit of a particular model across is superior to its comparators; however, this particular technique remains in development.

Fig. 1.

Fig. 1

A-D A comparison of decision curve analysis graphs is shown, including (A) random forest, (B) Bayesian belief network, (C) logistic regression, and (D) gradient boosting machine.

Second, while we appreciate Drs. Wang and Zhu’s opinion that “the calibration curve of the logistic regression seems better,” we do not agree with this assessment because there is, in fact, no consensus on what is and isn’t calibrated. As such, we prefer to rely more heavily on decision curve analysis than calibration curves in order to determine which model, if any, is best suitable for clinical use. Still, their comment raises an important point about the subjective evaluation of calibration curves, and we appreciate that they brought it up.

We also appreciate their comment regarding “black box” algorithms, and believe they should be strictly avoided in healthcare applications. It is not enough to determine whether a particular model may be useful in the clinical setting. One must also answer the question “how does it work?” We’ve done our best to describe how the gradient boosting machine uses the features by describing in detail the results of the gradient boosting machine feature-selection and modeling processes using the local interpretable model-agnostic explanations library in R© (Figs. 5 and 6 of the referenced manuscript). In doing so, we present both the relative influence of the features as well as their directionality of influence in estimating the likelihood of opioid dependence in this specific patient population. We believe that visualizing data in this manner is not only helpful to readers but is also a critical step toward understanding the relative contributions of each feature and ensuring each feature is relevant to the clinical focus area.

Further, we appreciate their comment that logistic regression models are “much easier” to implement, and we agree that paper-based tools and nomograms can be handy. We considered paper-based nomograms for use in the clinic but have now transitioned to cloud-based solutions. As Box and Draper tell us, “All models are wrong, but some are useful” [2]. The U.S. Food and Drug Administration (FDA) will soon evaluate tools such as the one we proposed using a new regulatory framework. When considering software as a medical device, the FDA emphasized the importance of version control and iterative improvements in models derived from “real-world learning or adaptation.” By producing software as a medical device solution, we are committed not only to developing models that are clinically relevant by rigorous statistical standards, but we also commit to improving the models over time as more data become available, as new prognostic features are discovered, and as treatment philosophies change. For this to be successful we believe a cloud-based environment rather than a paper-based system is necessary.

Finally, we acknowledge the comment by Drs. Wang and Zhu regarding the importance of external validation. As military physicians, we are responsible for the care and well-being of 9.6 million beneficiaries worldwide. Although our use of the military data repository may limit the use of such a tool in civilian patient populations, it is likely that one or more of the models described by the current manuscript will undergo prospective validation in US military beneficiaries. In the event prospective validation fails, we will have a larger sample size on which to base future modeling efforts. Until we received your letter, we had not considered the possibility of external validation in foreign military or civilian patient populations. Doing so may help us identify suitable surrogate features for those listed, including regions and military ranks, as you proposed. If agreeable, we would welcome the opportunity for international collaboration.

Our team is delighted to receive international interest and feedback. We greatly appreciate your comments, and hope that our response is sufficient.

Footnotes

(RE: Wang Q, Zhu H. Letter to the Editor: Can Predictive Modeling Tools Identify Patients at High Risk of Prolonged Opioid Use After ACL Reconstruction? Clin Orthop Relat Res. 2021;479:634-635.The institution of one or more of the authors (GCB) has received, during the study period, funding from the Society of Military Orthopaedic Surgeons (SOMOS) for grant sponsorship of this work.

Each author certifies that neither he nor she, nor any member of his or her immediate family, has funding or commercial associations (consultancies, stock ownership, equity interest, patent/licensing arrangements, etc.) that might pose a conflict of interest in connection with the submitted article.

All ICMJE Conflict of Interest Forms for authors and Clinical Orthopaedics and Related Research® editors and board members are on file with the publication and can be viewed on request.

Each author certifies that his or her institution approved the human protocol for this investigation and that all investigations were conducted in conformity with ethical principles of research.

This work was performed at Walter Reed National Military Medical Center, Bethesda, MD, USA.

References

  • 1.Anderson AB, Grazal CF, Balazs GC, Potter BK, Dickens JF, Forsberg JA. Can predictive modeling tools identify patients at high risk of prolonged opioid use after ACL reconstruction? Clin Orthop Relat Res. 2020;478:1610-1618. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Box GEP, Draper NR. Empirical Model-Building and Response Surfaces. John Wiley & Sons; 1987. [Google Scholar]
  • 3.Wang Q, Zhu H. Can predictive modeling tools identify patients at high risk of prolonged opioid use after ACL reconstruction? Clin Orthop Relat Res. 2021;479:634-635. [DOI] [PMC free article] [PubMed] [Google Scholar]

Articles from Clinical Orthopaedics and Related Research are provided here courtesy of The Association of Bone and Joint Surgeons

RESOURCES