Skip to main content
Journal of Spine Surgery logoLink to Journal of Spine Surgery
. 2026 Jul 21;12(8):138. doi: 10.21037/jss-2026-0130

AI parameters for enhancing spine surgery outcomes: a narrative review

Jaden Wise 1,#, Isabella Merem 1,#, Justin Gumbs 1, Zachary Comella 1, Angela Rodio 1, Victoria Lin 2, Victoria Verrengia 3, Rudy Paul 2, Min Shi 4, Maohua Lin 2,✉, Frank D Vrionis 3,✉
PMCID: PMC13554307  PMID: 42719262

Abstract

Background and Objective

Artificial intelligence (AI) is increasingly applied in spine surgery for diagnosis, operative planning, intraoperative guidance, and outcome prediction. However, reported technical performance does not always translate into clinical reliability. This narrative review evaluates how parameter-level decisions influence the reliability and clinical translation of AI models in spine surgery.

Methods

A structured literature search of PubMed, Scopus, and Google Scholar was performed for English-language studies published from January 2015 through March 2025. Search terms combined spine surgery concepts with AI, machine learning, deep learning, hyperparameters, learning rate, feature selection, regularization, model validation, and training strategy. Studies were synthesized narratively according to recurring parameter domains and clinical implementation context.

Key Content and Findings

Across preoperative, intraoperative, and postoperative applications, learning rate, feature selection, regularization, model architecture, optimization strategy, and evaluation metrics shaped convergence behavior, overfitting risk, interpretability, calibration, and external validity. Reported performance varied substantially by task, dataset, outcome definition, and validation strategy. Common limitations included retrospective single-center datasets, inconsistent parameter reporting, limited external validation, and challenges related to bias, interpretability, workflow integration, and model drift.

Conclusions

AI reliability in spine surgery depends on parameter-level development choices rather than algorithm selection alone. More standardized reporting, calibration assessment, prospective validation, and lifecycle monitoring are needed before AI tools can be integrated safely into spine surgery workflows.

Keywords: Artificial intelligence (AI), spine surgery, machine learning, hyperparameters, model validation

Introduction

Spine surgery requires integration of anatomical, technical, and patient-specific factors across diagnosis, surgical planning, and postoperative management. The anatomical complexity of the spine presents persistent challenges during both diagnosis and surgery. Variations such as transitional vertebrae, congenital anomalies, and altered vertebral morphology can complicate image interpretation and increase intraoperative risk (1,2). In some cases, these anatomic variants may also obscure or mimic other spinal pathologies, further complicating diagnosis and reinforcing the need for more individualized planning approaches (1-5).

Patient-specific factors add another layer of complexity to perioperative decision making. Complications such as adjacent segment disease, hardware failure, and postoperative neurological deficits contribute to outcome heterogeneity and make patient selection more challenging, particularly when reporting standards and treatment thresholds vary across institutions and surgeons (2). Advanced age has been associated with increased morbidity and mortality following spine surgery (6-9). Medical comorbidities may further obscure presentation, increase perioperative risk, and complicate recovery (7,8,10). These overlapping risk factors highlight the importance of careful preoperative optimization and individualized risk assessment (11-13).

Body habitus also influences outcomes. Obesity, particularly morbid obesity, has been associated with wound complications, longer operative times, prolonged hospitalization, and postoperative pulmonary complications after spine surgery (14-19).

Operative and postoperative challenges further contribute to the complexity of spine care. Case-specific technical difficulty, ergonomic demands, cerebrospinal fluid leakage, bleeding, hematoma formation, thromboembolic events, neurological injury, and revision or deformity surgery can all increase perioperative risk (7,20-25).

Hardware-related complications remain another major source of failure after spine surgery. Screw loosening, rod fracture, and implant migration may compromise construct stability and require further intervention (26). Adjacent segment disease is a frequent cause of recurrent symptoms and reoperation after fusion procedures (22,27-30). Other complications, including pseudarthrosis and surgical site infection, may prolong treatment and reduce quality of life, while proximal junctional kyphosis and junctional failure remain particularly important concerns in adult spinal deformity surgery (22,27-31).

In response to these challenges, artificial intelligence (AI) is increasingly being investigated as a decision-support tool in spine surgery. Machine learning algorithms have been applied to preoperative risk prediction, surgical planning, and intraoperative navigation, with several studies reporting improved accuracy in complication prediction and procedural optimization (32-34). AI-assisted robotic systems have also demonstrated improved pedicle screw placement accuracy while reducing radiation exposure compared with conventional approaches (32,35).

Despite these advances, important barriers to clinical adoption remain. Many AI models depend on large, high-quality datasets that are difficult to obtain in surgical practice and may introduce bias when patient populations are not adequately represented (2). Limited interpretability also continues to restrict clinical trust, as the rationale underlying some AI-generated predictions is not always transparent (2,36). In addition, high implementation costs and increased operative times associated with navigation and robotic systems continue to limit widespread use (25,32).

AI applications now span diagnostic refinement, preoperative planning, intraoperative navigation, and postoperative risk prediction. Across these settings, models have been used to identify spinal deformity and degenerative pathology, support surgical planning, improve pedicle screw placement accuracy, reduce radiation exposure, and predict complications such as infection, wound events, and thromboembolism (32,37-39). Recent spine-specific work on risk prediction and digital twin frameworks has further emphasized the potential for multimodal AI to support surgical decision-making, while also reinforcing the importance of calibration, validation, interpretability, and workflow integration before clinical deployment (40,41). However, the clinical message for surgeons remains less clear than the reported technical performance of individual models. A model may achieve strong internal accuracy, but its usefulness in practice depends on whether it can support decisions that occur before, during, or after surgery.

For spine surgeons, AI should therefore be evaluated according to the clinical problem it addresses. Preoperative applications may assist with diagnosis, risk stratification, patient selection, and surgical planning. Intraoperative applications may support navigation, robotic assistance, localization, neuromonitoring, and real-time decision support. Postoperative applications may help predict complications, monitor recovery, identify patients at risk for readmission, and detect implant-related failure. Across each phase, parameter-level decisions such as learning rate, feature selection, regularization, model architecture, optimization strategy, and validation approach remain important because they influence convergence stability, overfitting risk, interpretability, calibration, bias, and clinician trust (42-45).

This review therefore examines AI in spine surgery through the perioperative workflow, emphasizing how model-development choices influence reliability across preoperative, intraoperative, and postoperative care. By organizing the literature around surgical decision-making, this review aims to clarify what AI models can teach spine surgeons, where current tools remain limited, and what evidence is needed before these systems can be safely integrated into spine surgery workflows. Figure 1 presents an end-to-end AI workflow for spine surgery, showing how clinical and imaging data may be transformed into decision support across the perioperative continuum. We present this article in accordance with the Narrative Review reporting checklist (available at https://jss.amegroups.com/article/view/10.21037/jss-2026-0130/rc).

Figure 1.

Figure 1

Clinical artificial intelligence workflow for spine surgery. The figure illustrates the translational pipeline from clinical and imaging data acquisition through preprocessing, segmentation or annotation, model training, parameter optimization, internal validation, external validation, clinical deployment, and post-deployment monitoring. Parameter-level decisions influence reliability at each stage, including convergence stability during training, overfitting control during validation, calibration before deployment, and model drift surveillance after clinical implementation. AI, artificial intelligence; CT, computed tomography; MRI, magnetic resonance imaging.

Methods

This manuscript was prepared as a narrative review examining how core AI and machine learning parameters are used in spine surgery research. A narrative approach was chosen because the existing literature is highly variable in study design, data sources, outcome definitions, and model construction. Differences in datasets, validation strategies, and reporting practices make formal quantitative synthesis difficult and, in many cases, misleading. The search strategy summary is provided in Table 1.

Table 1. Search strategy summary.

Items Specification
Date of search Searches were conducted from April 12, 2025 through April 21, 2025
Databases and other sources searched PubMed, Scopus, and Google Scholar were searched. Reference lists of high-impact and sentinel articles were manually reviewed to identify additional relevant studies
Search terms used (“spine surgery” OR “spinal fusion” OR “spinal deformity” OR “pedicle screw”) AND (“artificial intelligence” OR “machine learning” OR “deep learning”) AND (“learning rate” OR “hyperparameter*” OR “feature selection” OR “regularization” OR “model validation” OR “training strategy”)
Timeframe English-language studies published from January 2015 through March 2025 were considered
Inclusion and exclusion criteria Eligible studies applied AI or machine learning to spine surgery or spine-related clinical decision making and reported hyperparameter configuration, model training strategy, feature selection methodology, validation framework, or performance calibration approach. English-language full-text studies were considered. Studies focused exclusively on non-surgical spine conditions, technical algorithm development without clinical linkage, editorials, non-English articles, and conference abstracts were excluded
Selection process Titles and abstracts were screened for relevance to AI applications in surgical spine care. Full-text review was performed for preliminarily eligible studies. Screening and eligibility assessment were performed by a single reviewer (J.W.) because of the narrative design
Additional considerations Studies were synthesized according to recurring parameter domains and clinical implementation contexts. Formal quantitative synthesis was not performed because of heterogeneity in datasets, outcomes, validation strategies, and model reporting

AI, artificial intelligence.

Search strategy and study identification

Relevant studies were identified through systematic searches of PubMed, Scopus, and Google Scholar. Searches were conducted from April 12, 2025 through April 21, 2025, and focused on English-language literature published from January 2015 through March 2025. Database-specific search strings were constructed using combinations of controlled vocabulary and free-text terms related to AI and spine surgery. Representative search syntax included: (“spine surgery” OR “spinal fusion” OR “spinal deformity” OR “pedicle screw”) AND (“artificial intelligence” OR “machine learning” OR “deep learning”) AND (“learning rate” OR “hyperparameter*” OR “feature selection” OR “regularization” OR “model validation” OR “training strategy”). Reference lists of high-impact and sentinel articles were manually reviewed to identify additional studies not captured in the primary search. Following peer review, a targeted literature update was performed to identify recent spine-specific AI studies, digital twin literature, reporting standards, appraisal tools, and clinical evaluation frameworks that directly supported the manuscript’s existing focus on reliability, calibration, validation, interpretability, workflow integration, and lifecycle monitoring.

Titles and abstracts were screened for relevance to AI applications in surgical spine care. Full-text review was performed for studies meeting preliminary inclusion criteria. When methodological details regarding parameter selection or model training were unclear, manuscripts were reviewed in full to determine eligibility. Screening and eligibility assessment were performed by a single reviewer given the narrative design of the study.

Eligibility criteria

Studies were considered eligible if they applied AI or machine learning techniques to spine surgery or spine-related clinical decision making and reported at least one of the following methodological elements: hyperparameter configuration, model training strategy, feature selection methodology, validation framework, or performance calibration approach. Studies focused exclusively on non-surgical spine conditions, technical algorithm development without clinical linkage, editorials, and conference abstracts were excluded. When multiple publications analyzed overlapping patient cohorts or datasets, the most methodologically comprehensive study was prioritized for parameter analysis.

Thematic synthesis framework

Rather than cataloguing all AI applications in spine surgery, studies were synthesized according to recurring parameter domains and clinical implementation contexts. Grouping frameworks were developed iteratively based on model architecture, parameter tuning methodology, validation structure, and surgical task application. Emphasis was placed on identifying recurring model-development patterns, reporting deficiencies, and sources of instability that may limit translational reliability.

Methodological appraisal

Given heterogeneity in datasets, outcomes, and reporting practices, formal risk-of-bias meta-analysis was not feasible. However, a structured methodological appraisal was performed examining dataset origin, sample size, validation strategy, parameter transparency, and performance reporting completeness. These domains were used to contextualize reliability and generalizability across included studies. When applicable, reporting and appraisal considerations were also interpreted in relation to current AI-focused guidance for prediction models and imaging-based AI, including TRIPOD+AI, PROBAST+AI, and CLAIM (46-48).

To strengthen the parameter-focused synthesis, included studies were additionally examined for reported relationships between parameter choices and model behavior, including convergence stability, overfitting risk, interpretability, generalizability, and predictive performance. Parameter domains included learning rate, feature selection, regularization, model architecture, and optimization strategy. Because formal meta-analysis was limited by heterogeneity in datasets, outcome definitions, model architectures, and reporting practices, findings were summarized narratively and organized across the parameter domains discussed below, with model-specific parameter sensitivities summarized in Table 2.

Table 2. Model-specific parameter sensitivities relevant to spine surgery artificial intelligence.

Model type Common spine surgery applications Most influential parameters Reliability concern if poorly tuned
CNNs/deep learning MRI/CT segmentation, deformity assessment, fusion prediction, imaging-based surgical planning Learning rate, batch size, dropout, data augmentation, number of layers, optimizer selection Overfitting, unstable segmentation boundaries, reduced performance across scanners or imaging protocols
Random forest Complication prediction, risk stratification, tabular clinical outcome modeling Number of trees, tree depth, minimum samples per split, feature subset size, class weighting Overfitting in small datasets, reduced generalizability, unstable feature importance
XGBoost/
LightGBM
Postoperative outcome prediction, complication modeling, personalized surgical planning Learning rate, maximum tree depth, number of estimators, subsampling, regularization terms High internal discrimination with poor calibration or limited external validity
Support vector machine Spinal abnormality classification, structured clinical prediction tasks Kernel selection, C parameter, gamma, feature scaling, feature selection Poor performance when features are not scaled, kernel is mismatched, or dimensionality is excessive
Reinforcement learning Pedicle screw planning, robotic/navigation support, adaptive intraoperative decision-making Reward function, exploration strategy, learning rate, safety constraints, update frequency Unsafe or unstable recommendations if the reward structure does not reflect clinically meaningful safety constraints
AutoML frameworks Automated model selection, hyperparameter optimization, clinical prediction workflows Search space, optimization metric, validation scheme, stopping criteria Selection of models optimized for internal accuracy rather than calibration, interpretability, or external reliability
AI-FEA/surrogate biomechanical models Patient-specific stress prediction, implant loading, subsidence modeling, construct stability, digital twin development Mesh quality, material property assignment, boundary conditions, loading assumptions, neural network architecture, loss-function weighting Biomechanically implausible outputs, poor transferability across anatomy or implants, unstable prediction of stress or deformation fields

AI, artificial intelligence; AutoML, automated machine learning; CNN, convolutional neural network; CT, computed tomography; FEA, finite element analysis; MRI, magnetic resonance imaging.

To further summarize recurring limitations, identified challenges were grouped into major domains and corresponding subdomains. Major domains included AI parameter constraints, AI methodology constraints, model-specific limitations, broader data and interpretability challenges, regulatory and ethical considerations, and operational barriers. Because formal pooling was not appropriate given heterogeneity in study design and reporting, Figure 2 presents a semi-quantitative visualization based on the number of discrete limitation subdomains identified within each major category rather than pooled study-level prevalence.

Figure 2.

Figure 2

Semi-quantitative distribution of limitation domains in spine surgery artificial intelligence. Bars represent the number of discrete limitation subdomains identified within each major category in the structured limitations synthesis. Categories are not mutually exclusive and do not represent pooled study-level prevalence.

AI parameters in context

While core machine learning parameters are mathematically consistent across disciplines, their behavior in spine surgery datasets reflects adaptations to cohort size constraints, imaging heterogeneity, deformity anatomy, and surgical outcome complexity. Accordingly, parameter optimization strategies reported in spine literature reflect task-specific modifications rather than generalized AI conventions.

Key AI parameters play a central role in determining how machine learning models behave when applied to spine surgery. Learning rate, feature selection, regularization, architecture, optimization strategy, and evaluation metrics collectively determine how models learn from limited and heterogeneous spine surgery data. In spine surgery applications, where datasets are often limited in size and heterogeneous in structure, these choices can substantially affect performance and generalizability.

Learning rate and optimization dynamics

Learning rate governs the magnitude of parameter updates during model training and is a primary determinant of convergence behavior. While higher learning rates may accelerate training, they can introduce instability, particularly in spine surgery datasets that are often limited in size and exhibit substantial heterogeneity (44,49). In such settings, large gradient updates may cause the model to overshoot optimal solutions, resulting in oscillatory or non-convergent training dynamics.

Conversely, lower learning rates promote stable convergence but may significantly increase training time and increase susceptibility to local minima, especially in high-dimensional parameter spaces. These challenges are amplified in clinical datasets with noise, class imbalance, and limited sample sizes, where small perturbations can disproportionately influence model updates.

Adaptive optimization strategies, including Adam and RMSProp, dynamically adjust learning rates during training and have demonstrated improved convergence stability and performance in imaging-based spine applications (50,51). These methods are particularly advantageous when dataset variability limits the effectiveness of fixed learning rate schedules.

Importantly, learning rate selection does not operate in isolation. Its effects are closely coupled with model architecture and regularization strategies. For example, high learning rates combined with insufficient regularization may exacerbate overfitting, while overly aggressive regularization may suppress meaningful learning altogether. Effective optimization therefore requires coordinated tuning of learning rate alongside complementary parameters to balance convergence speed, stability, and generalizability (33,52).

Feature selection and dimensionality considerations

Feature selection plays a critical role in determining model performance, interpretability, and generalizability in spine surgery applications. Clinical datasets often contain a large number of variables relative to sample size, increasing the risk of overfitting and reducing model stability (44,53).

Manual feature selection guided by clinical expertise can enhance interpretability and ensure that model inputs align with known physiological and surgical factors. However, this approach may fail to capture complex nonlinear relationships present in high-dimensional data. In contrast, automated feature selection methods, such as principal component analysis (PCA) and recursive feature elimination (RFE), can improve predictive performance by reducing dimensionality and filtering noise. These approaches have been associated with modest improvements in model performance across classification tasks but often reduce transparency, limiting clinical interpretability (53,54).

The effectiveness of feature selection strategies is also dependent on model architecture. Tree-based methods, such as random forests, inherently perform feature selection during training, whereas neural networks rely more heavily on preprocessing and input representation (55,56). As a result, inappropriate feature selection may disproportionately affect certain model types, leading to either information loss or unnecessary model complexity.

In practice, optimal feature selection requires balancing predictive performance with interpretability, particularly in clinical settings where understanding model behavior is essential for adoption and trust.

Regularization and overfitting control

Regularization is essential for controlling model complexity and preventing overfitting, particularly in spine surgery datasets that are often small, heterogeneous, and high-dimensional (44,49). Without appropriate regularization, models may memorize training data rather than generalize to new patients, resulting in degraded real-world performance.

L1 regularization promotes sparsity by penalizing the absolute magnitude of model coefficients, effectively performing implicit feature selection. This can improve interpretability but may reduce predictive performance if important variables are suppressed. L2 regularization, by contrast, penalizes the squared magnitude of coefficients and promotes more stable weight distributions, improving generalization without eliminating features entirely (52,57).

In deep learning models, techniques such as dropout further mitigate overfitting by randomly deactivating subsets of neurons during training, encouraging the model to learn more robust representations (50,57). However, excessive regularization, whether through high penalty terms or aggressive dropout, may lead to underfitting, particularly in models with limited capacity or insufficient training data.

The impact of regularization is strongly influenced by model architecture and dataset size. Deeper or more complex models typically require stronger regularization to maintain stability, whereas simpler models may require minimal constraint. Consequently, regularization strategies must be carefully tuned in conjunction with model design to achieve an appropriate balance between bias and variance.

Model architecture and data modality alignment

Model architecture plays a central role in determining how effectively patterns can be extracted from different types of data in spine surgery applications. The choice of architecture must align with the structure and dimensionality of the input data to achieve optimal performance.

Traditional machine learning models, including random forests and support vector machines, have demonstrated stable and reliable performance in tabular clinical datasets, where structured variables such as demographics, comorbidities, and operative factors are predominant (33,55). These models are less prone to overfitting in small datasets and often require fewer computational resources.

In contrast, deep learning architectures, particularly convolutional neural networks (CNNs), are better suited for imaging-based tasks such as segmentation, deformity assessment, and anatomical classification (50,51). These models can capture complex spatial relationships but require substantially larger datasets and are more sensitive to parameter tuning and regularization strategies.

Importantly, architectural complexity interacts with other parameters, including learning rate and regularization. More complex models may achieve higher performance but are also more susceptible to instability and overfitting, particularly in limited datasets (33,49). Therefore, model selection should be guided not only by task requirements but also by the availability and quality of data.

Because parameter sensitivity varies substantially by model architecture, optimization strategies cannot be generalized across all AI systems. CNN-based models are often sensitive to learning rate, dropout, data augmentation, and architecture depth, whereas tree-based approaches depend heavily on tree depth, feature selection, and regularization (33,50,51,57,58). Reinforcement learning systems introduce additional dependence on reward function design, exploration strategy, and safety constraints (59,60). Table 2 summarizes model-specific parameter sensitivities relevant to spine surgery AI applications.

Hyperparameter optimization strategies

Effective hyperparameter optimization is essential for achieving reliable model performance. Traditional approaches, such as grid search, systematically explore predefined parameter combinations but are computationally expensive and may be inefficient in high-dimensional parameter spaces (33,52).

More advanced techniques, including Bayesian optimization, offer a more efficient alternative by modeling the relationship between parameter configurations and model performance, allowing for iterative refinement of the search process. These methods are particularly advantageous in spine surgery applications, where limited datasets necessitate efficient use of available training data (44,52).

Despite their importance, hyperparameter optimization strategies are inconsistently reported across studies, limiting reproducibility and comparison (49,61). Future work should prioritize systematic evaluation of optimization methods to determine their impact on model stability and clinical applicability.

Parameter interactions and dataset constraints

A critical but often underemphasized consideration is the interaction between parameters, particularly in the context of small and heterogeneous clinical datasets. In spine surgery applications, limited sample sizes amplify the effects of parameter misconfiguration, making coordinated tuning essential (44,49).

For example, high model complexity combined with insufficient regularization may lead to rapid overfitting, whereas overly aggressive regularization may suppress meaningful signal (52,57). Similarly, learning rate selection interacts with both model architecture and dataset size, influencing convergence stability and generalizability.

These interactions highlight that parameter tuning should not be approached in isolation. Instead, effective model development requires a holistic strategy that accounts for dataset characteristics, clinical context, and the interplay between multiple hyperparameters (33,61).

Representative failure modes illustrate why parameter selection has direct clinical relevance. In small CNN-based imaging datasets, insufficient regularization, limited data augmentation, or excessive architectural complexity may produce strong internal performance while memorizing scanner-specific, institution-specific, or preprocessing-related features rather than clinically meaningful anatomy. This risk is particularly relevant in spine surgery, where many datasets remain retrospective, single-center, and limited in size (33,44,49-51). Similarly, an overly high learning rate may produce unstable convergence, whereas excessive regularization may suppress clinically meaningful signal and contribute to underfitting (33,52,57). These problems become most apparent during external validation, where models trained on one institution’s imaging protocols, operative techniques, or patient population may demonstrate reduced calibration, discrimination, or clinical utility when applied elsewhere (39,43,45,62). Therefore, parameter selection should be interpreted as a determinant of clinical reliability rather than a purely technical step in model optimization.

Evaluation frameworks and learning paradigms

Evaluation metrics and clinical relevance

Evaluation metrics are essential for assessing model performance but must be interpreted within the context of clinical applicability. Commonly reported metrics in spine surgery studies include accuracy, precision, recall, F1 score, and area under the curve (AUC). While these metrics provide complementary perspectives, reliance on accuracy alone may be misleading, particularly in imbalanced datasets where class prevalence skews performance estimates.

Importantly, strong performance on traditional metrics does not necessarily translate to clinical usefulness. A model may achieve a high AUC yet provide poorly calibrated risk estimates, limiting its ability to inform decision-making at clinically meaningful thresholds. As such, evaluation should extend beyond discrimination to include calibration and clinical utility, particularly in applications such as complication prediction and surgical risk stratification.

Calibration is particularly important in spine surgery risk prediction because clinicians often rely on estimated probabilities when counseling patients, selecting operative candidates, or determining the intensity of postoperative monitoring (45,62). A model with strong discrimination may correctly rank high-risk and low-risk patients but still overestimate or underestimate absolute risk, which may lead to inappropriate reassurance, unnecessary intervention, or inefficient resource allocation. Similarly, decision thresholds should be evaluated in relation to clinically meaningful actions, such as whether a predicted complication risk would alter surgical planning, trigger preoperative optimization, justify closer postoperative surveillance, or change patient counseling. Future studies should therefore report not only discrimination metrics such as AUC, but also calibration measures, threshold-based performance, and evidence that model use improves decision-making compared with standard clinical assessment (39,45). Decision curve analysis may further help determine whether a prediction model provides net clinical benefit across clinically meaningful risk thresholds, rather than only demonstrating statistical discrimination (63).

Reported performance across spine surgery applications demonstrates variability depending on task and dataset. For example, random forest models have achieved AUC values of approximately 0.85 in predicting postoperative neurological outcomes, while similar approaches have demonstrated classification accuracy exceeding 85% in scoliosis cohorts (64,65). However, these outcomes are influenced not only by model choice but also by parameter selection, dataset characteristics, and evaluation methodology.

Because accuracy, AUC, F1 score, and Dice coefficient measure different aspects of model performance, values across diagnostic classification, segmentation, and outcome prediction tasks should not be interpreted as directly comparable. Reported performance across spine surgery AI studies is therefore heterogeneous and task-dependent rather than uniform. Broad reviews of machine learning applications in spine surgery have described average predictive accuracies in the mid-70% range across mixed clinical prediction tasks, whereas selected task-specific models, such as scoliosis classification and metastatic spinal tumor outcome prediction, have reported stronger discrimination, including accuracy above 85% or AUC values near 0.85–0.90 (56,64,65).

Learning paradigms and model behavior

Learning paradigms further shape model behavior and performance. Supervised learning remains the most widely used approach in spine surgery, relying on labeled datasets to support classification and outcome prediction. These models are highly sensitive to parameter choices, including learning rate, regularization, and feature selection, particularly in small datasets where overfitting risk is elevated (66,67).

Unsupervised learning methods, including clustering and dimensionality reduction, aim to identify latent patterns without predefined labels. These approaches are useful for exploring heterogeneity in spinal pathologies, such as subgroup identification in scoliosis populations, but are heavily dependent on feature representation and model design. Poorly selected features or inappropriate dimensionality reduction can obscure clinically meaningful structure, limiting interpretability and applicability (66,67).

Reinforcement learning represents an emerging paradigm in spine surgery, particularly in intraoperative and robotic-assisted applications. By optimizing decision-making through iterative feedback, reinforcement learning systems can adapt surgical planning in real time. The SafeRPlan framework, for example, applies deep reinforcement learning to pedicle screw placement, demonstrating improvements in procedural accuracy and safety (59,60). However, these systems require robust evaluation frameworks and conservative parameter tuning to ensure reliability in clinical environments.

Automated machine learning (AutoML) and integrated workflows

AutoML frameworks aim to streamline model development by integrating parameter selection, model choice, and optimization into unified workflows. These approaches reduce user-dependent variability and improve reproducibility, which is particularly valuable in clinical settings where expertise in machine learning may be limited. Expert-augmented AutoML systems have demonstrated improved consistency in tasks such as spinal cord injury outcome prediction (53,68).

Despite these advantages, AutoML does not eliminate the need for careful evaluation and parameter oversight. Model performance remains dependent on underlying data quality, parameter constraints, and evaluation strategy. As such, these frameworks should be viewed as tools to assist, rather than replace, informed model development.

Applications in spine surgery

The practical value of AI in spine surgery is best understood by the point in care at which the model is intended to assist the surgeon. Preoperative models may help identify risk, define anatomy, select candidates, or plan correction. Intraoperative models may support navigation, localization, robotic planning, neuromonitoring, or real-time safety checks. Postoperative models may help anticipate complications, monitor recovery, identify patients at risk for readmission, and detect implant-related failure (32,37-39,59,60,69-72). This workflow-based framework connects each model to a specific surgical decision rather than presenting AI methods as isolated technical categories.

Preoperative applications

Preoperative AI applications are most relevant when they help surgeons define anatomy, estimate patient-specific risk, or plan treatment before entering the operating room. In the reviewed literature, these applications include classification of spinal abnormalities, surgical candidacy assessment, complication prediction, and support for individualized surgical planning (42,53,71-73). During preoperative evaluation, learning rate influences convergence efficiency in models used for diagnostic classification, surgical candidacy assessment, and complication risk prediction. In diagnostic settings, appropriate tuning of learning rate has been shown to improve classification accuracy in models such as support vector machines and logistic regression when identifying spinal abnormalities (53). Similar considerations apply to surgical planning, where stable convergence supports more reliable prediction of treatment suitability and perioperative risk. Poorly tuned learning rates in these settings may lead to unstable predictions that limit clinical usability (32,38).

Feature selection is similarly important during the preoperative phase, particularly when models incorporate large numbers of clinical, imaging, and demographic variables. By reducing dimensionality and emphasizing relevant predictors, feature selection can improve both diagnostic accuracy and surgical planning. Prior studies have shown that feature selection may strengthen personalized decision making and improve classification of spinal abnormalities (42,66). In diagnostic applications, techniques such as univariate feature selection and PCA have improved classification performance (53). In planning applications, feature selection has also been used to identify predictors of progression and operative necessity in lumbar spine pathology (73). Learning rate, feature selection, and regularization strategies are summarized in Table 3.

Table 3. Overview of key AI parameters in spine surgery.
Parameter Definition Techniques/models used Applications in spine surgery Representative performance findings and interpretation
Learning rate Controls the magnitude of model weight updates during training and influences convergence speed and stability Neural networks, CNNs, XGBoost, gradient boosting models Diagnostic imaging, surgical planning, complication prediction, outcome prediction Learning rate tuning was associated with improved convergence stability and model discrimination in selected spine AI applications. Reported performance varied by task and metric; broad reviews describe average predictive accuracy in the mid-70% range across heterogeneous spine surgery ML applications, while selected imaging-based studies report Dice coefficients approaching 0.90 (51,73)
Feature selection Identifies clinically and statistically relevant variables to reduce dimensionality and improve model stability Boruta, recursive feature elimination, PCA, univariate selection, clinician-guided selection Spinal abnormality diagnosis, surgical planning, risk stratification Feature selection improved model stability and classification performance in spinal abnormality diagnosis by reducing noise and overfitting. However, reported gains varied by selection method, model type, and validation strategy (48,49,51)
Regularization Penalizes model complexity to reduce overfitting and improve generalizability L1/LASSO, L2/Ridge, dropout, manifold regularization Outcome prediction, image analysis, intraoperative navigation, high-dimensional clinical datasets Regularization supported generalizability in high-dimensional datasets, particularly when predictors were correlated or sample sizes were limited. Excessive regularization may contribute to underfitting, and reported performance varied by dataset size, model complexity, and validation design (47,51,52)

AI, artificial intelligence; CNN, convolutional neural network; LASSO, least absolute shrinkage and selection operator; ML, machine learning; XGBoost, extreme gradient boosting.

Intraoperative applications

In the intraoperative setting, model usefulness depends on the ability to process limited but reliable data in real time. Feature selection contributes to intraoperative decision support by restricting inputs to variables that can be consistently acquired and interpreted during surgery. This has been associated with improved accuracy in localization tasks and guidance systems, supporting safer and more precise interventions (54). Regularization further supports stability by constraining overly complex models and reducing sensitivity to noise. In automated measurement and classification systems, regularization has improved reliability by limiting overfitting (57). Similar principles apply to intraoperative and perioperative prediction models, where regularization may preserve meaningful predictors while suppressing less informative variation (55). In monitoring applications such as motor evoked potential surveillance, this stability is especially important because model sensitivity and specificity must remain consistent enough to detect meaningful neurologic change without producing excessive false alarms (69).

Supervised learning remains the most widely used framework across spine surgery applications, including imaging analysis, segmentation, and outcome prediction. In these settings, careful tuning of learning rate and regularization improves generalizability and reduces overfitting, particularly when datasets are limited in size or highly heterogeneous (66). Representative predictive studies have reported mean overall accuracies in the mid-70% range across spine surgery tasks, although performance varies substantially by dataset and endpoint (56).

Reinforcement learning is also relevant to intraoperative and robotic-assisted spine surgery. SafeRPlan, for example, applies deep reinforcement learning to intraoperative pedicle screw placement planning and illustrates how AI may optimize screw trajectory while incorporating safety constraints (59,60). In this setting, conservative parameter tuning and robust evaluation are essential because model outputs must be reliable, interpretable, and actionable during surgery.

Postoperative applications

Postoperative applications increasingly focus on complication prediction, recovery trajectories, and subgroup identification. Unsupervised learning methods are especially useful for exploratory analysis, where clustering and dimensionality reduction may reveal subgroups within heterogeneous spine populations and support more personalized postoperative management (49). These approaches, however, remain highly dependent on feature representation and parameter selection. Prior studies have used unsupervised methods to examine heterogeneity in spinal pathology and postoperative recovery patterns, including prolonged recovery and complication-associated trajectories (44,66). Similar approaches have also been applied to the evaluation of learning curves and postoperative outcomes in minimally invasive lumbar procedures (70). For spine surgeons, these postoperative models may be most useful when they identify patients who require closer surveillance, earlier intervention, or more individualized follow-up after surgery.

Model architecture also shapes performance across phases of care. CNNs are widely used for imaging-based tasks such as segmentation, deformity assessment, and preoperative planning (21,42,58). CNN-based models have also demonstrated strong performance in predicting fusion success after anterior cervical discectomy and fusion (50). More complex deep learning frameworks that integrate multiple data types have likewise been applied to deformity correction planning in adolescent idiopathic scoliosis (51). Traditional machine learning models such as support vector machines and random forests remain particularly useful for tabular clinical data, where they have been used to classify spinal pathologies, predict outcomes, and identify important risk factors (42,44,58). Gradient boosting approaches, including XGBoost and LightGBM, have also demonstrated strong predictive performance in postoperative outcome modeling and personalized surgical planning (71,72).

Across preoperative, intraoperative, and postoperative applications, these studies collectively show that performance in spine surgery is not determined by algorithm choice alone. Rather, clinical utility depends on how model development choices shape convergence, robustness, interpretability, and external validity.

Current limitations

Several limitations continue to restrict the clinical translation of AI in spine surgery. At the parameter level, poorly selected learning rates may cause unstable convergence, while excessive feature reduction or imbalanced regularization may produce underfitting, overfitting, or loss of clinically relevant information. These concerns are amplified in spine surgery datasets, which are often small, retrospective, heterogeneous, and high dimensional.

Methodological limitations are also common. Many published models rely on single-center datasets, inconsistent endpoint definitions, limited external validation, and incomplete reporting of training strategy, calibration, and hyperparameter selection. As a result, strong internal performance does not necessarily imply clinical generalizability across different hospitals, surgeons, imaging protocols, patient populations, or operative workflows.

Model-specific constraints further affect reliability. CNNs may perform well in imaging-based applications but require large datasets and substantial computational resources, whereas tree-based models and support vector machines may be more practical for tabular data but remain vulnerable to class imbalance and overfitting. Regulatory, ethical, and implementation barriers, including data privacy, algorithmic bias, clinician trust, workflow disruption, cost, and model drift, must also be addressed before routine clinical deployment. These limitations are summarized in Table 4 and illustrated in Figure 2.

Table 4. Key limitations and mitigation strategies for artificial intelligence in spine surgery. Limitations are grouped into major domains to summarize recurring barriers to clinical reliability, generalizability, and implementation.

Limitation domain Key challenges Clinical relevance Potential mitigation
Parameter-related limitations Poor learning rate selection, excessive feature reduction, or imbalanced regularization may cause instability, underfitting, or overfitting Misconfigured models may produce unreliable predictions, especially in small or high-dimensional spine datasets Use cross-validation, adaptive learning rates, early stopping, systematic hyperparameter tuning, and transparent reporting of model configuration
Data quality and generalizability Small sample sizes, retrospective datasets, single-center cohorts, missing data, and limited demographic diversity Models may perform well internally but fail across different institutions, imaging protocols, surgeons, and patient populations Develop multi-center datasets, standardize data collection, perform external validation, and use prospective evaluation when possible
Model-specific constraints CNNs require large imaging datasets and computational resources; RF/SVM models may overfit tabular or imbalanced data Model choice may limit reliability if not matched to data type, dataset size, and clinical task Match architecture to data modality, use regularization, address class imbalance, and evaluate performance across independent cohorts
Interpretability and clinician trust Deep learning models may function as “black boxes”, making predictions difficult to explain Poor transparency limits surgeon confidence and may reduce adoption in high-stakes decision-making Apply explainable AI methods such as SHAP, saliency maps, feature importance analysis, and model-agnostic interpretability tools
Regulatory, ethical, and bias concerns Privacy risks, unclear regulatory pathways, algorithmic bias, and underrepresentation of minority populations AI tools may worsen disparities or raise liability concerns if deployed without oversight Use secure data-sharing protocols, fairness-aware modeling, subgroup performance auditing, and regulatory-aligned validation frameworks
Clinical integration barriers Cost, specialized equipment, workflow disruption, clinician training needs, and potential increases in operative time Even accurate models may not be adopted if they are expensive, inefficient, or difficult to integrate into surgical workflows Prioritize user-centered design, interoperability with existing systems, cost-effectiveness studies, clinician education, and phased implementation
Lifecycle management Model drift may occur as patient populations, surgical techniques, implants, and imaging protocols evolve Previously validated models may lose accuracy after deployment Monitor post-deployment performance, recalibrate models periodically, audit for drift, and update models using new data

AI, artificial intelligence; CNN, convolutional neural network; RF, random forest; SHAP, SHapley Additive exPlanations; SVM, support vector machine.

Clinical implementation strategies

Reliable implementation will require standardized data capture, multi-institutional collaboration, transparent model reporting, and validation strategies that reflect real-world spine surgery practice. De-identified datasets should preserve clinically meaningful variables, including demographics, comorbidities, imaging findings, operative details, and postoperative outcomes, while using interoperable formats that facilitate cleaning, harmonization, and analysis. External validation and prospective evaluation should become routine before AI systems are incorporated into operative planning, intraoperative navigation, or postoperative risk prediction.

Clinical integration should also account for interpretability, regulation, cost, and lifecycle management. Explainable AI methods may help surgeons assess whether model outputs are biologically plausible and clinically actionable. Surgeon-facing AI literacy is also important because clinical adoption depends on whether users can understand a model’s intended use, limitations, calibration, and appropriate role in decision-making (74). Systems intended to influence diagnosis, surgical planning, or intraoperative guidance should clearly define intended use, degree of autonomy, validation procedures, and post-deployment monitoring. Because surgical techniques, implant systems, imaging protocols, and patient populations evolve, deployed models should be audited for calibration, subgroup performance, dataset shift, and model drift.

Future directions

Future progress in spine surgery AI should prioritize not only more complex architectures, but also better data quality, transparent parameter reporting, external validation, and equitable model development. Multi-center datasets and federated learning approaches may improve generalizability while preserving patient privacy. More rigorous comparisons of hyperparameter optimization strategies, including grid search and Bayesian optimization, are needed to determine which methods produce the most stable clinical predictions across heterogeneous spine datasets.

Integration of AI with finite element analysis (FEA) and biomechanical modeling

A particularly important future direction is the integration of AI with FEA and biomechanical simulation. Traditional FEA provides mechanistic insight into stress distribution, implant loading, endplate interaction, construct stability, and adjacent segment biomechanics, but patient-specific modeling may be limited by segmentation demands, mesh generation, assumptions regarding material properties, boundary condition selection, and computational cost (75,76). AI-assisted workflows may help address these limitations by accelerating image segmentation, automating model generation, estimating biomechanical parameters, and approximating simulation outputs that would otherwise require time-intensive computational analysis (76).

Physics-informed neural networks, surrogate modeling, and digital twin frameworks may be especially relevant for spine surgery because they offer a potential bridge between data-driven prediction and biomechanical plausibility. Physics-informed neural networks incorporate governing physical relationships into model training, allowing neural networks to learn from data while remaining constrained by known physical laws (77). In spine surgery, this concept could support patient-specific prediction of implant performance, screw loosening, cage subsidence, adjacent segment loading, or deformity correction mechanics. However, AI-FEA models introduce additional parameter interactions beyond conventional machine learning. Model behavior may depend not only on learning rate, regularization, architecture, and feature selection, but also on mesh quality, material property assignment, loading conditions, boundary constraints, and assumptions used to construct the biomechanical model (75-77). Future AI-FEA research should therefore evaluate how both computational and biomechanical parameters influence prediction stability, interpretability, and clinical reliability. Recent work on dynamic spine stabilization and embedded biomechanical feedback systems further illustrates how sensor-enabled constructs and longitudinal biomechanical data may eventually support patient-specific AI and digital twin models (78).

Future studies should also integrate clinical, imaging, laboratory, and perioperative data while using feature selection and regularization strategies that preserve interpretability. Prospective data collection, calibration assessment, threshold-based evaluation, subgroup auditing, and post-deployment monitoring will be essential for determining whether AI tools improve decision-making beyond standard clinical assessment. These priorities may help move AI in spine surgery from retrospective proof-of-concept studies toward reproducible, equitable, and clinically dependable systems.

AI systems should also be evaluated using a lifecycle framework that includes transparent parameter reporting, calibration assessment, external validation, prospective evaluation, post-deployment monitoring, and periodic recalibration rather than relying solely on retrospective discrimination metrics. Early clinical evaluation and trial testing should align with emerging AI reporting frameworks, including DECIDE-AI for early-stage clinical evaluation and SPIRIT-AI and CONSORT-AI when AI systems are evaluated in clinical trial protocols or trial reports (79-81). These practices are especially important because performance may degrade when clinical workflows, imaging protocols, implant systems, or patient populations change after deployment.

Strengths and limitations of this review

This review is strengthened by its parameter-focused framework, which shifts attention from broad algorithm categories toward the specific development choices that influence model stability, interpretability, and clinical translation in spine surgery AI. The synthesis also integrates technical model-development considerations with clinically relevant phases of care, including preoperative planning, intraoperative guidance, and postoperative outcome prediction. However, the review is limited by the narrative design, heterogeneity of included studies, inconsistent reporting of hyperparameter configurations across the literature, and lack of formal quantitative pooling. Because many included studies were retrospective, single-center, and task-specific, the findings should be interpreted as a framework for evaluating clinical reliability rather than as pooled evidence of comparative model effectiveness.

Conclusions

AI has become an increasingly visible component of spine surgery research, with applications spanning diagnosis, surgical planning, intraoperative guidance, and postoperative outcome prediction. While many studies report improvements in predictive performance or technical accuracy, their clinical value depends on whether model development choices produce reliable, interpretable, and externally valid predictions.

Learning rate, feature selection, regularization strategies, model architecture, and optimization methods all influence model behavior, particularly in spine surgery datasets that are frequently limited in size and heterogeneous in composition. In tabular clinical prediction tasks, tree-based methods such as XGBoost and random forests have shown strong performance in several spine applications, whereas CNNs remain especially useful in imaging-based tasks such as segmentation and deformity assessment. Across both settings, poorly configured models may exhibit unstable convergence, overfitting, and limited external validity, whereas well-constrained models are more likely to produce reliable and interpretable predictions.

Different learning paradigms further influence how these parameters affect clinical outcomes. Supervised learning remains the most widely applied approach and continues to demonstrate utility across imaging analysis and outcome prediction tasks. Unsupervised learning offers exploratory insight into disease heterogeneity, while reinforcement learning introduces new possibilities for adaptive intraoperative decision-making. Across these approaches, careful model development remains essential for reliability and clinical relevance.

As AI continues to mature within spine surgery, future progress will likely depend less on increasingly complex algorithms alone and more on rigorous validation, transparent reporting, and careful attention to how models are trained and constrained (75,76). Aligning technical model development with clinical validation may help narrow the gap between retrospective performance and real-world implementation.

By focusing on the role of core AI parameters across the surgical continuum, this review highlights considerations that are essential for translating AI from experimental tools into reliable clinical support systems. The most clinically useful AI models in spine surgery will not necessarily be those with the highest retrospective accuracy, but those that remain calibrated, interpretable, externally valid, and stable across institutions, surgeons, imaging protocols, implant systems, and patient populations (39,43-46,62,82). Addressing these factors may support safer implementation, improve clinician trust, and ultimately contribute to more precise and individualized spine surgery care.

Supplementary

The article’s supplementary files as

jss-12-08-138-rc.pdf (155.3KB, pdf)
DOI: 10.21037/jss-2026-0130
jss-12-08-138-coif.pdf (858.8KB, pdf)
DOI: 10.21037/jss-2026-0130

Acknowledgments

None.

Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved.

Footnotes

Reporting Checklist: The authors have completed the Narrative Review reporting checklist. Available at https://jss.amegroups.com/article/view/10.21037/jss-2026-0130/rc

Funding: This work was supported by the Marcus Neuroscience Institute (No. C-24-335) and the Helene and Stephen Weicholz Foundation (No. GT-005511).

Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://jss.amegroups.com/article/view/10.21037/jss-2026-0130/coif). The authors have no conflicts of interest to declare.

References

  • 1.Kato S. Complications of thoracic spine surgery - Their avoidance and management. J Clin Neurosci 2020;81:12-7. 10.1016/j.jocn.2020.09.012 [DOI] [PubMed] [Google Scholar]
  • 2.Mensah EO, Chalif JI, Baker JG, et al. Challenges in Contemporary Spine Surgery: A Comprehensive Review of Surgical, Technological, and Patient-Specific Issues. J Clin Med 2024;13:5460. 10.3390/jcm13185460 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Diebo BG, Henry J, Lafage V, et al. Sagittal deformities of the spine: factors influencing the outcomes and complications. Eur Spine J 2015;24 Suppl 1:S3-15. 10.1007/s00586-014-3653-8 [DOI] [PubMed] [Google Scholar]
  • 4.Takenaka S, Kashii M, Iwasaki M, et al. Risk factor analysis of surgery-related complications in primary cervical spine surgery for degenerative diseases using a surgeon-maintained database. Bone Joint J 2021;103-B:157-63. 10.1302/0301-620X.103B1.BJJ-2020-1226.R1 [DOI] [PubMed] [Google Scholar]
  • 5.Walcott BP, Coumans JV, Kahle KT. Diagnostic pitfalls in spine surgery: masqueraders of surgical spine disease. Neurosurg Focus 2011;31:E1. 10.3171/2011.7.FOCUS11114 [DOI] [PubMed] [Google Scholar]
  • 6.Nakashima H, Tetreault LA, Nagoshi N, et al. Does age affect surgical outcomes in patients with degenerative cervical myelopathy? Results from the prospective multicenter AOSpine International study on 479 patients. J Neurol Neurosurg Psychiatry 2016;87:734-40. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Schoenfeld AJ, Ochoa LM, Bader JO, et al. Risk factors for immediate postoperative complications and mortality following spine surgery: a study of 3475 patients from the National Surgical Quality Improvement Program. J Bone Joint Surg Am 2011;93:1577-82. 10.2106/JBJS.J.01048 [DOI] [PubMed] [Google Scholar]
  • 8.Soroceanu A, Burton DC, Oren JH, et al. Medical Complications After Adult Spinal Deformity Surgery: Incidence, Risk Factors, and Clinical Impact. Spine (Phila Pa 1976) 2016;41:1718-23. 10.1097/BRS.0000000000001636 [DOI] [PubMed] [Google Scholar]
  • 9.Zileli M, Dursun E. How to Improve Outcomes of Spine Surgery in Geriatric Patients. World Neurosurg 2020;140:519-26. 10.1016/j.wneu.2020.04.060 [DOI] [PubMed] [Google Scholar]
  • 10.Diltz ZR, West EJ, Colatruglio MR, et al. Perioperative Management of Comorbidities in Spine Surgery. Orthop Clin North Am 2023;54:349-58. 10.1016/j.ocl.2023.02.007 [DOI] [PubMed] [Google Scholar]
  • 11.Khechen B, Haws BE, Bawa MS, et al. The Impact of Comorbidity Burden on Complications, Length of Stay, and Direct Hospital Costs After Minimally Invasive Transforaminal Lumbar Interbody Fusion. Spine (Phila Pa 1976) 2019;44:363-8. 10.1097/BRS.0000000000002834 [DOI] [PubMed] [Google Scholar]
  • 12.Kuo CC, Soliman MAR, Aguirre AO, et al. Risk factors of early complications after thoracic and lumbar spinal deformity surgery: a systematic review and meta-analysis. Eur Spine J 2023;32:899-913. 10.1007/s00586-022-07486-3 [DOI] [PubMed] [Google Scholar]
  • 13.Yagi M, Fujita N, Okada E, et al. Impact of Frailty and Comorbidities on Surgical Outcomes and Complications in Adult Spinal Disorders. Spine (Phila Pa 1976) 2018;43:1259-67. 10.1097/BRS.0000000000002596 [DOI] [PubMed] [Google Scholar]
  • 14.Adindu E, Singh D, Geck M, et al. The impact of obesity on postoperative and perioperative outcomes in lumbar spine surgery: a systematic review and meta-analysis. Spine J 2025;25:1081-95. 10.1016/j.spinee.2024.12.006 [DOI] [PubMed] [Google Scholar]
  • 15.Goyal A, Elminawy M, Kerezoudis P, et al. Impact of obesity on outcomes following lumbar spine surgery: A systematic review and meta-analysis. Clin Neurol Neurosurg 2019;177:27-36. 10.1016/j.clineuro.2018.12.012 [DOI] [PubMed] [Google Scholar]
  • 16.Barrie U, Reddy RV, Elguindy M, et al. Impact of obesity on complications and surgical outcomes after adult degenerative scoliosis spine surgery. Clin Neurol Neurosurg 2023;226:107619. 10.1016/j.clineuro.2023.107619 [DOI] [PubMed] [Google Scholar]
  • 17.Harrop JS, Mohamed B, Bisson EF, et al. Congress of Neurological Surgeons Systematic Review and Evidence-Based Guidelines for Perioperative Spine: Preoperative Surgical Risk Assessment. Neurosurgery 2021;89:S9-S18. 10.1093/neuros/nyab316 [DOI] [PubMed] [Google Scholar]
  • 18.Mohamed B, Wang MC, Bisson EF, et al. Congress of Neurological Surgeons Systematic Review and Evidence-Based Guidelines for Perioperative Spine: Preoperative Pulmonary Evaluation and Optimization. Neurosurgery 2021;89:S33-41. 10.1093/neuros/nyab319 [DOI] [PubMed] [Google Scholar]
  • 19.Puvanesarajah V, Werner BC, Cancienne JM, et al. Morbid Obesity and Lumbar Fusion in Patients Older Than 65 Years: Complications, Readmissions, Costs, and Length of Stay. Spine (Phila Pa 1976) 2017;42:122-7. 10.1097/BRS.0000000000001692 [DOI] [PubMed] [Google Scholar]
  • 20.Alostaz M, Bansal A, Gyawali P, et al. Ergonomics in Spine Surgery: A Systematic Review. Spine (Phila Pa 1976) 2024;49:E250-61. 10.1097/BRS.0000000000005055 [DOI] [PubMed] [Google Scholar]
  • 21.Katsevman GA, Daffner SD, Brandmeir NJ, et al. Complexities of spine surgery in obese patient populations: a narrative review. Spine J 2020;20:501-11. 10.1016/j.spinee.2019.12.011 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Linzey JR, Lillard J, LaBagnara M, et al. Complications and Avoidance in Adult Spinal Deformity Surgery. Neurosurg Clin N Am 2023;34:665-75. 10.1016/j.nec.2023.06.012 [DOI] [PubMed] [Google Scholar]
  • 23.Bohl DD, Webb ML, Lukasiewicz AM, et al. Timing of Complications After Spinal Fusion Surgery. Spine (Phila Pa 1976) 2015;40:1527-35. 10.1097/BRS.0000000000001073 [DOI] [PubMed] [Google Scholar]
  • 24.Hohenberger C, Albert R, Schmidt NO, et al. Incidence of medical and surgical complications after elective lumbar spine surgery. Clin Neurol Neurosurg 2022;220:107348. 10.1016/j.clineuro.2022.107348 [DOI] [PubMed] [Google Scholar]
  • 25.Romero-Muñoz LM, Segura-Fragoso A, Talavera-Díaz F, et al. Neurological injury as a complication of spinal surgery: incidence, risk factors, and prognosis. Spinal Cord 2020;58:318-23. 10.1038/s41393-019-0367-0 [DOI] [PubMed] [Google Scholar]
  • 26.Planas Gil A, Chárlez Marco A, Loste Ramos A, et al. Acute complications in open/miss primary and revision thoracolumbar spine surgery: a descriptive study of the most common complications and treatment of choice. Int Orthop 2024;48:555-61. 10.1007/s00264-023-06047-7 [DOI] [PubMed] [Google Scholar]
  • 27.Cerpa M, Zuckerman SL, Lenke LG, et al. Long-term follow-up of non‑neurologic and neurologic complications after complex adult spinal deformity surgery: results from the Scoli-RISK-1 study. Eur Spine J 2025;34:1790-800. 10.1007/s00586-025-08683-6 [DOI] [PubMed] [Google Scholar]
  • 28.Cho KJ, Suk SI, Park SR, et al. Complications in posterior fusion and instrumentation for degenerative lumbar scoliosis. Spine (Phila Pa 1976) 2007;32:2232-7. 10.1097/BRS.0b013e31814b2d3c [DOI] [PubMed] [Google Scholar]
  • 29.Kim HJ, Iyer S, Zebala LP, et al. Perioperative Neurologic Complications in Adult Spinal Deformity Surgery: Incidence and Risk Factors in 564 Patients. Spine (Phila Pa 1976) 2017;42:420-7. 10.1097/BRS.0000000000001774 [DOI] [PubMed] [Google Scholar]
  • 30.Lange N, Stadtmüller T, Scheibel S, et al. Analysis of risk factors for perioperative complications in spine surgery. Sci Rep 2022;12:14350. 10.1038/s41598-022-18417-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Smith JS, Buell TJ, Shaffrey CI, et al. Prospective multicenter assessment of complication rates associated with adult cervical deformity surgery in 133 patients with minimum 1-year follow-up. J Neurosurg Spine 2020;33:588-600. 10.3171/2020.4.SPINE20213 [DOI] [PubMed] [Google Scholar]
  • 32.Bcharah G, Gupta N, Panico N, et al. Innovations in Spine Surgery: A Narrative Review of Current Integrative Technologies. World Neurosurg 2024;184:127-36. 10.1016/j.wneu.2023.12.124 [DOI] [PubMed] [Google Scholar]
  • 33.Rodrigues AJ, Schonfeld E, Varshneya K, et al. Comparison of Deep Learning and Classical Machine Learning Algorithms to Predict Postoperative Outcomes for Anterior Cervical Discectomy and Fusion Procedures With State-of-the-art Performance. Spine (Phila Pa 1976) 2022;47:1637-44. 10.1097/BRS.0000000000004481 [DOI] [PubMed] [Google Scholar]
  • 34.Yagi M, Yamanouchi K, Fujita N, et al. Revolutionizing Spinal Care: Current Applications and Future Directions of Artificial Intelligence and Machine Learning. J Clin Med 2023;12:4188. 10.3390/jcm12134188 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Qi Z, Da H, Yanming F, et al. Current status and prospects of robot-assisted spine surgery. Expert Rev Med Devices 2025;22:187-92. 10.1080/17434440.2025.2467779 [DOI] [PubMed] [Google Scholar]
  • 36.Mehmet S, Elmarawany MN, Harding I, et al. AI versus the spinal surgeons in the management of controversial spinal surgery scenarios. Eur Spine J 2025;34:3736-46. 10.1007/s00586-025-08825-w [DOI] [PubMed] [Google Scholar]
  • 37.Hopkins BS, Mazmudar A, Driscoll C, et al. Using artificial intelligence (AI) to predict postoperative surgical site infection: A retrospective cohort of 4046 posterior spinal fusions. Clin Neurol Neurosurg 2020;192:105718. 10.1016/j.clineuro.2020.105718 [DOI] [PubMed] [Google Scholar]
  • 38.Kim JS, Merrill RK, Arvind V, et al. Examining the Ability of Artificial Neural Networks Machine Learning Models to Accurately Predict Complications Following Posterior Lumbar Spine Fusion. Spine (Phila Pa 1976) 2018;43:853-60. 10.1097/BRS.0000000000002442 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Hurkmans C, Bibault JE, Brock KK, et al. A joint ESTRO and AAPM guideline for development, clinical validation and reporting of artificial intelligence models in radiation therapy. Radiother Oncol 2024;197:110345. 10.1016/j.radonc.2024.110345 [DOI] [PubMed] [Google Scholar]
  • 40.Salman S, Phadke R, Kumar R, et al. Risk prediction in spine surgery: a scoping review of traditional models, artificial intelligence, and the challenge of clinical translation. Spine Deform 2026. [Epub ahead of print]. doi: . 10.1007/s43390-026-01365-3 [DOI] [PubMed] [Google Scholar]
  • 41.Salman SG, Phadke R, Kumar R, et al. Digital twins and multimodal artificial intelligence in spine care: a scoping review of concepts, evidence, and translational barriers. Spine Deform 2026. [Epub ahead of print]. doi: . 10.1007/s43390-026-01397-9 [DOI] [PubMed] [Google Scholar]
  • 42.Charles YP, Lamas V, Ntilikina Y. Artificial intelligence and treatment algorithms in spine surgery. Orthop Traumatol Surg Res 2023;109:103456. 10.1016/j.otsr.2022.103456 [DOI] [PubMed] [Google Scholar]
  • 43.Hornung AL, Hornung CM, Mallow GM, et al. Artificial intelligence and spine imaging: limitations, regulatory issues and future direction. Eur Spine J 2022;31:2007-21. 10.1007/s00586-021-07108-4 [DOI] [PubMed] [Google Scholar]
  • 44.Khan JA, Houdane A, Hajja A, et al. Assessing the Utility and Challenges of Machine Learning in Spinal Deformity Management: A Systematic Review. Spine (Phila Pa 1976) 2025. [Epub ahead of print]. doi: . 10.1097/BRS.0000000000005236 [DOI] [PubMed] [Google Scholar]
  • 45.Azad TD, Ehresman J, Ahmed AK, et al. Fostering reproducibility and generalizability in machine learning for clinical prediction modeling in spine surgery. Spine J 2021;21:1610-6. 10.1016/j.spinee.2020.10.006 [DOI] [PubMed] [Google Scholar]
  • 46.Collins GS, Moons KGM, Dhiman P, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ 2024;385:e078378. 10.1136/bmj-2023-078378 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Moons KGM, Damen JAA, Kaul T, et al. PROBAST+AI: an updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods. BMJ 2025;388:e082505. 10.1136/bmj-2024-082505 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Tejani AS, Klontzas ME, Gatti AA, et al. Checklist for Artificial Intelligence in Medical Imaging (CLAIM): 2024 Update. Radiol Artif Intell 2024;6:e240300. 10.1148/ryai.240300 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Lopez CD, Boddapati V, Lombardi JM, et al. Artificial Learning and Machine Learning Applications in Spine Surgery: A Systematic Review. Global Spine J 2022;12:1561-72. 10.1177/21925682211049164 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Park S, Kim JK, Chang MC, et al. Assessment of Fusion After Anterior Cervical Discectomy and Fusion Using Convolutional Neural Network Algorithm. Spine (Phila Pa 1976) 2022;47:1645-50. 10.1097/BRS.0000000000004439 [DOI] [PubMed] [Google Scholar]
  • 51.Chen K, Zhai X, Chen Z, et al. Deep learning based decision-making and outcome prediction for adolescent idiopathic scoliosis patients with posterior surgery. Sci Rep 2025;15:3389. 10.1038/s41598-025-87370-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Guan Y, Li Y, Ke Z, et al. Learning-Assisted Fast Determination of Regularization Parameter in Constrained Image Reconstruction. IEEE Trans Biomed Eng 2024;71:2253-64. 10.1109/TBME.2024.3367762 [DOI] [PubMed] [Google Scholar]
  • 53.Raihan-Al-Masud M, Mondal MRH. Data-driven diagnosis of spinal abnormalities using feature selection and machine learning algorithms. PLoS One 2020;15:e0228422. 10.1371/journal.pone.0228422 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Rajpurohit V, Danish SF, Hargreaves EL, et al. Optimizing computational feature sets for subthalamic nucleus localization in DBS surgery with feature selection. Clin Neurophysiol 2015;126:975-82. 10.1016/j.clinph.2014.05.039 [DOI] [PubMed] [Google Scholar]
  • 55.Difazio RL, Strout TD, Vessey JA, et al. Comparison of two modeling approaches for the identification of predictors of complications in children with cerebral palsy following spine surgery. BMC Med Res Methodol 2024;24:236. 10.1186/s12874-024-02360-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Tragaris T, Benetos IS, Vlamis J, et al. Machine Learning Applications in Spine Surgery. Cureus 2023;15:e48078. 10.7759/cureus.48078 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Pang S, Su Z, Leung S, et al. Direct automated quantitative measurement of spine by cascade amplifier regression network with manifold regularization. Med Image Anal 2019;55:103-15. 10.1016/j.media.2019.04.012 [DOI] [PubMed] [Google Scholar]
  • 58.Johnson GW, Chanbour H, Ali MA, et al. Artificial Intelligence to Preoperatively Predict Proximal Junction Kyphosis Following Adult Spinal Deformity Surgery: Soft Tissue Imaging May Be Necessary for Accurate Models. Spine (Phila Pa 1976) 2023;48:1688-95. 10.1097/BRS.0000000000004816 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Ao Y, Esfandiari H, Carrillo F, et al. SafeRPlan: Safe deep reinforcement learning for intraoperative planning of pedicle screw placement. Med Image Anal 2025;99:103345. 10.1016/j.media.2024.103345 [DOI] [PubMed] [Google Scholar]
  • 60.Datta S, Li Y, Ruppert MM, et al. Reinforcement learning in surgery. Surgery 2021;170:329-32. 10.1016/j.surg.2020.11.040 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.McDonnell JM, Evans SR, McCarthy L, et al. The diagnostic and prognostic value of artificial intelligence and artificial neural networks in spinal surgery : a narrative review. Bone Joint J 2021;103-B:1442-8. 10.1302/0301-620X.103B9.BJJ-2021-0192.R1 [DOI] [PubMed] [Google Scholar]
  • 62.Wondra JP, 2nd, Kelly MP, Greenberg J, et al. Validation of Adult Spinal Deformity Surgical Outcome Prediction Tools in Adult Symptomatic Lumbar Scoliosis. Spine (Phila Pa 1976) 2023;48:21-8. 10.1097/BRS.0000000000004416 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Vickers AJ, Holland F. Decision curve analysis to evaluate the clinical benefit of prediction models. Spine J 2021;21:1643-8. 10.1016/j.spinee.2021.02.024 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Maki S, Shiratani Y, Orita S, et al. Predicting Postoperative Neurological Outcomes in Metastatic Spinal Tumor Surgery Using Machine Learning. Spine (Phila Pa 1976) 2026;51:100-6. 10.1097/BRS.0000000000005322 [DOI] [PubMed] [Google Scholar]
  • 65.Yu J, Lahoti YS, McCandless KC, et al. Automated Scoliosis Cobb Angle Classification in Biplanar Radiograph Imaging With Explainable Machine Learning Models. Spine (Phila Pa 1976) 2025;50:E259-67. 10.1097/BRS.0000000000005312 [DOI] [PubMed] [Google Scholar]
  • 66.Hornung AL, Hornung CM, Mallow GM, et al. Artificial intelligence in spine care: current applications and future utility. Eur Spine J 2022;31:2057-81. 10.1007/s00586-022-07176-0 [DOI] [PubMed] [Google Scholar]
  • 67.Pantanowitz L, Pearce T, Abukhiran I, et al. Nongenerative Artificial Intelligence in Medicine: Advancements and Applications in Supervised and Unsupervised Machine Learning. Mod Pathol 2025;38:100680. 10.1016/j.modpat.2024.100680 [DOI] [PubMed] [Google Scholar]
  • 68.Chou A, Torres-Espin A, Kyritsis N, et al. Expert-augmented automated machine learning optimizes hemodynamic predictors of spinal cord injury outcome. PLoS One 2022;17:e0265254. 10.1371/journal.pone.0265254 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Zhuang Q, Wang S, Zhang J, et al. How to make the best use of intraoperative motor evoked potential monitoring? Experience in 1162 consecutive spinal deformity surgical procedures. Spine (Phila Pa 1976) 2014;39:E1425-32. [DOI] [PubMed] [Google Scholar]
  • 70.Saravi B, Zink A, Ülkümen S, et al. Artificial intelligence-based analysis of associations between learning curve and clinical outcomes in endoscopic and microsurgical lumbar decompression surgery. Eur Spine J 2024;33:4171-81. 10.1007/s00586-023-08084-7 [DOI] [PubMed] [Google Scholar]
  • 71.Guo Z, Wang P, Ye S, et al. Interpretable Machine Learning Models Based on Shapley Additive Explanations for Predicting the Risk of Cerebrospinal Fluid Leakage in Lumbar Fusion Surgery. Spine (Phila Pa 1976) 2024;49:1281-93. 10.1097/BRS.0000000000005087 [DOI] [PubMed] [Google Scholar]
  • 72.Wang D, Wang Q, Cui P, et al. Machine-learning models for the prediction of ideal surgical outcomes in patients with adult spinal deformity. Bone Joint J 2025;107-B:337-45. 10.1302/0301-620X.107B3.BJJ-2024-1220.R1 [DOI] [PubMed] [Google Scholar]
  • 73.Xie N, Wilson PJ, Reddy R. Use of machine learning to model surgical decision-making in lumbar spine surgery. Eur Spine J 2022;31:2000-6. 10.1007/s00586-021-07104-8 [DOI] [PubMed] [Google Scholar]
  • 74.Kumar R, Phadke R, Salman S. Advancing AI literacy in Canadian orthopedic education: a framework for equitable and inclusive training. Can Med Educ J 2025;16:36-8. 10.36834/cmej.82734 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 75.Beaulieu E, Wise J, Merem I, et al. A Review of Finite Element Analysis in Spine Surgery Decision-Making. J Clin Med 2026;15:2584. 10.3390/jcm15072584 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76.Ahmadi M, Chen H, Lin M, et al. Streamlined and efficient patient-specific modeling for lumbar spine segmentation and finite element analysis. Sci Rep 2025;15:35619. 10.1038/s41598-025-19664-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77.Raissi M, Perdikaris P, Karniadakis GE. Physics-informed neural networks: a deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations. J Comput Phys 2019;378:686-707. [Google Scholar]
  • 78.Kumar R, Bouras A, Phadke R, et al. Dynamic spine stabilization through mechanically tuned constructs and embedded biomechanical feedback systems: a narrative review. Discov Sens 2026;2:26. [Google Scholar]
  • 79.Vasey B, Nagendran M, Campbell B, et al. Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Nat Med 2022;28:924-33. 10.1038/s41591-022-01772-9 [DOI] [PubMed] [Google Scholar]
  • 80.Cruz Rivera S, Liu X, Chan AW, et al. Guidelines for clinical trial protocols for interventions involving artificial intelligence: the SPIRIT-AI extension. Nat Med 2020;26:1351-63. 10.1038/s41591-020-1037-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 81.Liu X, Cruz Rivera S, Moher D, et al. Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: the CONSORT-AI extension. Nat Med 2020;26:1364-74. 10.1038/s41591-020-1034-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 82.Kore A, Abbasi Bavil E, Subasri V, et al. Empirical data drift detection experiments on real-world medical imaging data. Nat Commun 2024;15:1887. 10.1038/s41467-024-46142-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 83.Shastry P, Sonawane B, Mohan K, et al. AI and Deep Learning for Automated Segmentation and Quantitative Measurement of Spinal Structures in MRI. arXiv 2025. arXiv:2503.11281.

Associated Data

    This section collects any data citations, data availability statements, or supplementary materials included in this article.

    Supplementary Materials

    The article’s supplementary files as

    jss-12-08-138-rc.pdf (155.3KB, pdf)
    DOI: 10.21037/jss-2026-0130
    jss-12-08-138-coif.pdf (858.8KB, pdf)
    DOI: 10.21037/jss-2026-0130

    Articles from Journal of Spine Surgery are provided here courtesy of OSS Press

    RESOURCES