Abstract
Background
Mechanical complications (MCs) remain a significant challenge in adult spinal deformity (ASD) surgery. The global alignment and proportion (GAP) score, a pelvic incidence-based metric, was introduced to predict the risk of postoperative MCs by assessing the proportionality of spinopelvic alignment.
Methods
We systematically reviewed the evidence on the predictive validity of the GAP score and its modifications with regard to postoperative MCs (eg, proximal junctional kyphosis/failure, rod fracture, or implant failure) and patient-reported outcomes in adult spinal deformity (ASD) surgery. A systematic search of PubMed/MEDLINE, Embase, Scopus, and the Cochrane Library was conducted for studies published from January 2017 to July 2025. A proportion meta-analysis regarding the occurrence of MCs was performed using those articles that divided the patients in the three original GAP categories.
Results
A total of 32 studies met the inclusion criteria, encompassing over 5,700 patients. The reported incidence of MCs varied from 17.6% to 61%. The predictive accuracy of the GAP score was inconsistent, with reported Area Under the Curve (AUC) values for predicting MCs ranging from a low of 0.53 to a moderately accurate 0.86. A meta-analysis found a significant association, with an odds ratio (OR) of 2.65 for MCs in severely disproportioned patients. However, the score's validity was diminished in specific cohorts, particularly for circumferential minimally invasive surgery (cMIS). Recent evidence demonstrates that the GAP score's predictive power is enhanced when combined with biological factors. The modified GAPB score, which incorporates Body Mass Index and Bone Mineral Density, demonstrated superior predictive accuracy (AUC=0.885 vs. 0.798). Machine learning models utilizing these multifactorial inputs achieved predictive accuracies over 73%.
Conclusions
The predictive value of the GAP score with regard to MCs in ASD surgery is moderate but context-dependent. Its utility as a stand-alone predictor is limited, and the current evidence in the scientific literature supports the use of multifactorial predictive models, such as the GAPB score, and the application of machine learning to integrate patient-specific biological and biomechanical factors.
Keywords: GAP score, Adult spinal deformity, Mechanical complications, GAPB score, Spinopelvic alignment, Machine learning
Background
Adult spinal deformity (ASD) represents a spectrum of complex spinal pathologies that are increasing in prevalence within the aging global population and are characterized by abnormal spinal alignment in the sagittal and/or coronal planes. While surgical intervention can offer substantial improvements in pain and function, it is associated with a high rate of postoperative complications [1,2].
Mechanical complications (MCs), in particular, represent a formidable challenge for both patients and surgeons. These adverse events (including proximal junctional kyphosis/PJK), proximal junctional failure/PJF), rod fracture/RF), and pseudarthrosis) occur at rates reported between 20% and 62%. Such failures often lead to loss of correction, recurrent pain, and the need for complex revision surgeries, which carry their own significant risks and economic costs [3]. The need for reliable preoperative tools to stratify risk and guide surgical planning toward more durable patient-specific solutions led to the development and widespread adoption of the Scoliosis Research Society (SRS)-Schwab classification, which codified key spinopelvic parameters (namely, the sagittal vertical axis/SVA), pelvic tilt/PT, and the mismatch between pelvic incidence and lumbar lordosis/PI−LL) and established universal surgical targets as essential modifiers for surgical planning. However, clinical experience and subsequent research revealed the limitations of this "one-size-fits-all" approach. MCs continued to occur at high rates even in patients who met these alignment criteria [4,5].
In 2017, Yilgor et al. introduced the global alignment and proportion (GAP) score to address the shortcomings of existing systems [3]. The GAP score was a paradigm shift, moving away from absolute targets toward a proportional model based on each patient's unique pelvic incidence (PI). It is calculated from 5 components: (1) relative pelvic version/RPV (i.e. measured sacral slope vs. ideal PI-based sacral slope); (2) relative lumbar lordosis/RLL (measured lumbar lordosis vs. ideal PI-based lumbar lordosis); (3) lordosis distribution index/LDI (the proportion of lordosis in the lower lumbar spine); (4) relative spinopelvic alignment/RSA (measured global tilt vs. ideal PI-based global tilt); and (5) an age factor [6]. Based on a composite score, patients are categorized into 1 of 3 postoperative spinopelvic states: proportioned (P: score 0–2), moderately disproportioned (MD: score 3–6), or severely disproportioned (SD: score ≥7). The initial, multicenter validation study reported a striking correlation between these categories and the risk of MCs after ASD surgery. Patients in the P, MD, and SD groups experienced MC rates of 6%, 47%, and 95%, respectively. The score demonstrated an excellent ability to predict MCs, with an area under the curve (AUC) of 0.92. These seminal findings suggested that the GAP score could be a powerful tool to individualize surgical goals and substantially reduce the incidence of mechanical failure [6].
Despite its promising introduction, the clinical utility and generalizability of the GAP score have been subjects of intense debate. Numerous subsequent external validation studies have produced conflicting and heterogeneous results. Furthermore, there is growing recognition that mechanical failure is a multifactorial process, influenced not only by radiographic alignment but also by patient-specific biological factors like bone quality and muscle health, as well as the nature of specific surgical techniques. The objective of this systematic review is to determine the predictive value of the GAP score and its modifications for postoperative MCs in ASD surgery across diverse patient populations and surgical techniques.
Methods
Search strategy
A systematic literature search was conducted in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines. We searched the PubMed/MEDLINE, Embase, Scopus, and the Cochrane Library databases for relevant articles published from January 2017 to July 2025. The search strategy combined keywords and MeSH terms including "global alignment and proportion score," "GAP score," "ASD," "MCs," "PJK," "implant failure," "pseudarthrosis," and "patient-reported outcomes." The reference lists of included articles and relevant reviews were manually searched to identify additional studies.
Study selection
Inclusion criteria were: (1) prospective or retrospective cohort studies, case-control studies, or meta-analyses; (2) study population consisted of adult patients (age ≥18 years) undergoing corrective surgery for ASD; (3) reported postoperative GAP scores; (4) reported outcomes related to MCs or patient-reported outcomes (PROs); and (5) minimum follow-up of 1 year (or occurrence of an MC prior to 1 year). Exclusion criteria were: case reports, technical notes, review articles without original data, studies on non-ASD populations, and studies that did not provide a specific analysis of the GAP score [4]. Two reviewers independently screened titles and abstracts, followed by a full-text review of potentially eligible articles. Disagreements were resolved by consensus with a third reviewer.
Data extraction and synthesis
Two independent reviewers extracted data using a standardized form. Extracted data included study design, patient demographics (sample size, age, sex, diagnosis), surgical details (levels fused, surgical approach, use of osteotomies), follow-up duration, GAP score categories, rates and types of MCs, PROs, and reported predictive metrics (eg, AUC, odds ratios). Data were synthesized qualitatively and, where appropriate, quantitatively. A narrative synthesis was performed to describe the conflicting findings and explore sources of heterogeneity. The quality of the articles included in the revision was assessed using the MINORS score.
Statistical analysis (quantitative synthesis)
Statistical analyses were performed in R (v 4.3.2) with the meta package (v 6.x). For each study we extracted the number of MC) and the total number of operated patients in the 3 GAP-score categories—Proportioned (GAP 0–2), Moderately disproportioned (GAP 3–6) and Severely disproportioned (GAP ≥7). Pooled event rates were calculated with a random-effects meta-analysis of proportions (metaprop, logit transformation), using the Maximum-Likelihood estimator for between-study variance; cells containing 0 or 100% events received a continuity correction of 0.5. Results are reported as back-transformed pooled proportions with 95% confidence intervals (CI) and 95% prediction intervals (PI). Heterogeneity was assessed with the Q statistic and expressed as I², with values >50% interpreted as substantial. Differences between GAP subgroups were explored in a univariable random-effects meta-regression (metareg), with the Proportioned group as reference; coefficients are shown on the logit scale and as back-transformed proportions, and p values <.05 were deemed significant. Leave-one-out analyses tested robustness, and small-study effects were inspected with funnel plots and Egger’s test. All figures (forest plots, funnel plots) were generated with the native functions of the meta package, and the full R script is provided in the online supplement.
Results
Literature search and study characteristics
The literature search yielded 184 unique records, from which 26 full-text articles were assessed for eligibility. Ultimately, 32 studies met the inclusion criteria for this review. All of them consisted of retrospective cohort studies. The characteristics of the primary included studies are summarized in Table 1. The studies varied considerably in sample size (from 39 to 322 patients), mean patient age (from 50.7 to 76.5 years), surgical approach and mean follow-up duration (from 1.5 to 9.0 years) (Table 1).
Table 1.
Characteristics of included studies
| First author | Year | Type of surgery | GAP score variant | N patients | N mechanical failures | Significant Association Between GAP and MC | Other MC predictors | GAP vs other predictors |
|---|---|---|---|---|---|---|---|---|
| Yilgor et al. [6] | 2017 | Posterior instrumented fusion ≥4 levels for ASD | Original postoperative GAP (0–13) | 222 | 32 | Yes | Age ≥60 yrs | GAP AUC 0.88–0.92 >SRS–Schwab |
| Jacobs et al. [7] | 2019 | Long posterior fusion ≥4 levels | Original postoperative GAP | 39 | 22 | Yes | Schwab class significant | GAP AUC 0.86 vs Schwab 0.69 |
| Noh et al. [8] | 2020 | Posterior fusion >4 levels | GAP-B (GAP+BMI+BMD) | 203 | 89 | Yes | BMI↑, BMD↓ | GAP-B AUC 0.885 >GAP 0.798 |
| Ham et al. [9] | 2021 | Long fusion with PCO/PSO in elderly | Original GAP | 84 | 51 | Yes | Low BMD; higher age | No comparison performed |
| Kwan et al. [10] | 2021 | Long posterior fusion (75 % 3CO) | Original post-op GAP | 159 | 44 | No | None | Not compared |
| Ha et al. [11] | 2021 | Long posterior fusion >4 Levels | Partial intra-op iGAP (RLL+LDI+age) | 48 | 13 | No | None | Not compared |
| Sun et al. [12] | 2021 | Elderly DLS long fusion ≥5 levels | Post-op GAP (vs Roussouly) | 80 | 41 | Yes | Roussouly score; PT | GAP AUC 0.863 >Roussouly 0.543 |
| Noh et al. [13] | 2021 | Posterior fusion >4 levels | GAP-B (BMI+BMD) | 203 | 89 | Yes | BMI↑, BMD↓ | GAP-B AUC 0.885 >GAP 0.798 >Schwab 0.532 |
| Gendelberg et al. [3] | 2023 | Circumferential minimally invasive surgery (cMIS) ASD correction; ≥4 levels fused, no osteotomies | Original postoperative GAP | 182 | 32 | No | None identified | Lower MC rate than original GAP cohort; GAP not predictive in cMIS |
| Hiltunen et al. [2] | 2023 | Adult spinal deformity long fusion mixed techniques (posterior, ALIF, 3CO) | Original postoperative GAP | 142 | 23 | Yes | High LDI, RSA parameters | No external comparison; GAP ≥5 predictive |
| Noh et al. [14] | 2023 | Posterior long fusion ≥4 levels for adult spinal deformity (training 70%, test 30%) | GAP-B (BMI + BMD) with machine-learning models | 238 | 100 | Yes | Higher BMI, lower BMD, high relative lumbar lordosis score | ML using GAP-B out-performed logistic regression; BMI/BMD key |
| Lord et al. [15] | 2023 | Thoracic/Lumbar fusion ≥4 levels; adult spinal deformity | Original postoperative GAP (vs change in GAP) | 81 | 14 | No | Greater UIV-PA correction; larger T1-UIV angle | Alignment metrics (UIV-PA, T1-UIV) predictive; GAP not |
| Park et al. [16] | 2024 | Degenerative ASD; ≥5-level fusion incl. pelvis in ≥60 yrs (Samsung MC cohort) | Post-op GAP categories (0–2, 3–6, ≥7) used for correction assessment | 259 | 132 | No | Undercorrection vs age-adjusted PI-LL linked to worse ODI; severity groups saw more undercorrection | GAP categories nondiscriminatory; SRS-Schwab PI-LL grade differed; age-adjusted PI-LL target more informative |
| Kumar et al. [17] | 2025 | Posterior instrumented fusion ≥4 levels for ASD | Original post-op GAP categories (P/MP/DP) | 93 | 27 | No | Low paraspinal CSA, high fatty infiltration (FI), smoking, ↑SVA | Muscle CSA & FI (OR 0.57 & p<.001) + smoking (OR 6.6) out-performed GAP |
| Huang et al. [24] | 2025 | Long fusion ≥5 levels for ASD | Not evaluated | 96 | 44 | Lower-thoracic UIV → early PJK (OR 5.3); high PI → late PJK (OR 1.09/°) | Study focused on PI, UIV; no GAP comparison | |
| Lim et al. [18] | 2025 | Revision adult flatback; single-level lumbar PSO (L2–L5) with long thoracolumbar–pelvic fusion | Original post-op GAP & sub-domains (RPV, LDI, RLL, RSA) | 152 | 26 | No | Higher T4-L1 PA mismatch; proximal PSO level (L2/L3) tended ↑ failures | GAP not predictive; T4-L1 PA mismatch associated with PJK/PJF & rod fracture |
| Moridaira et al. [19] | 2025 | Limited lumbar fusion (3–5 levels) with UIV in L1-L3 for de-novo ASD | Original GAP (pre- and post-op) analysed | 46 | 17 | Yes | Low HU at UIV/UIV+1 (OR 0.975, cutoff 98 HU) independent; age, SVA, PI-LL, PT | HU stronger predictor (AUC 0.745); GAP not in multivariate |
Table 1: Summary for every included study, the surgical setting and the predictive performance of GAP-based metrics with respect to mechanical complications. The first 2 columns identify the work (first author and publication year) and describe the primary surgical strategy, for example, long posterior instrumented fusion with or without 3-column osteotomy, circumferential minimally invasive surgery, or revision lumbar PSO. The next column specifies which version of the GAP score was analysed (the original postoperative 0–13 score, the body-mass-index/BMD-adjusted GAP-B, the partial intraoperative iGAP, and so on). Columns 4 and 5 report the sample size and the absolute number of mechanical failures (proximal junctional kyphosis or failure, rod fracture, implant loosening, re-operation for mechanical reasons). A sixth column states whether the study found a statistically significant association between the chosen GAP metric and mechanical complications. The seventh column lists any additional risk factors that emerged (for example, high BMI, low bone mineral density, upper-instrumented-vertebra Hounsfield units, extent of sagittal under-correction). Finally, the last column compares the predictive accuracy of GAP with alternative radiographic or clinical predictors, citing metrics such as the area under the ROC curve, odds or hazard ratios, or qualitative statements when formal AUCs were not provided.
The methodological quality of the articles included in the qualitative synthesis was evaluated according to the MINORS criteria (Table 2). The majority of the studies presented moderate quality, with some studies presenting lower and lower-moderate quality.
Table 2.
MINORS score: Columns 1 to 8 list the 8 noncomparative items of the methodological index for nonrandomized studies (MINORS)
| First author | Year | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | Total /16 | Quality level |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Yilgor et al. [6] | 2017 | 2 | 2 | 1 | 2 | 1 | 2 | 2 | 0 | 12 | Moderate |
| Noh et al. [8] | 2020 | 2 | 2 | 2 | 2 | 1 | 2 | 2 | 0 | 13 | Moderate–high |
| Ham et al. [9] | 2021 | 2 | 2 | 1 | 2 | 1 | 2 | 2 | 0 | 12 | Moderate |
| Noh et al. [13] | 2021 | 2 | 1 | 0 | 2 | 1 | 2 | 2 | 0 | 10 | Moderate |
| Gendelberg et al. [3] | 2023 | 2 | 1 | 0 | 2 | 1 | 2 | 2 | 0 | 10 | Moderate |
| Hiyama et al. [21] | 2023 | 2 | 1 | 0 | 2 | 1 | 2 | 1 | 0 | 9 | Low–moderate |
| Wang et al. [22] | 2024 | 1 | 1 | 0 | 2 | 0 | 2 | 2 | 0 | 8 | Low |
| Park et al. [16] | 2024 | 2 | 2 | 0 | 2 | 1 | 2 | 2 | 0 | 11 | Moderate |
| Kumar et al. [17] | 2025 | 2 | 1 | 0 | 2 | 0 | 2 | 2 | 0 | 9 | Low–moderate |
Each item is scored 2 when it is explicitly reported and adequate, 1 when reported but inadequate or unclear, and 0 when not reported. The items are: (1) clearly stated aim; (2) inclusion of consecutive patients; (3) prospective data collection; (4) outcomes appropriate to the study aim; (5) unbiased assessment of outcomes; (6) follow-up period adequate for the endpoint; (7) loss to follow-up <5%; and (8) prospective sample-size calculation. “Total / 16” is the sum of the 8 items (maximum=16). Based on this total, overall study quality is categorised as Low (≤ 8), Low–moderate (9), Moderate (10–12) or Moderate–high (13–16).
Predictive validity for overall mechanical complications: a conflicting landscape
The primary finding of this review is the marked inconsistency in the reported predictive validity of the GAP score for postoperative MCs after ASD surgery. While the original study by Yilgor et al. [6] demonstrated a strong, stepwise increase in MC risk with worsening GAP category, subsequent external validation studies have failed to consistently replicate this finding. The reported rates of MCs across the P, MD, and SD categories were found to be highly variable, as detailed in Table 3.
Table 3.
Each row corresponds to an individual study
| Study | Proportioned (GAP 0–2) MC / N | Moderately disproportioned (GAP 3–6) MC / N | Severely disproportioned (GAP ≥7) MC / N |
|---|---|---|---|
| Yilgor et al. [6] | 2 / 33 | 9 / 19 | 21 / 22 |
| Noh et al. [8] | 4 / 42 | 27 / 85 | 58 / 76 |
| Ham et al. [9] | 0 / 2 | 10 / 32 | 41 / 50 |
| Noh et al. [13] | 4 / 42 | 27 / 85 | 58 / 76 |
| Gendelberg et al. [3] | 5 / 45 | 15 / 73 | 12 / 64 |
| Hiyama et al. [21] | 9 / 23 | 13 / 22 | 3 / 9 |
| Wang et al. [22] | 4 / 20 | 20 / 41 | 5 / 8 |
| Kumar et al. [17] | 6 / 18 | 10 / 29 | 11 / 46 |
| Park et al. [16] | 14 / 42 | 34 / 62 | 84 / 155 |
Moderately-disproportioned (GAP 3–6) and Severely-disproportioned (GAP ≥7) patient groups. Values are expressed as “MC / N,” where MC is the number of patients who experienced at least one mechanical complication and N is the total number of patients in that GAP subgroup for the study.
Within that row 3 cells summarize, respectively, the Proportioned (GAP 0–2).
The authors performed a proportion meta-analysis with the occurrence of MCs in articles that divided the patients in the three original GAP categories Fig. 1. Although the pooled proportions differed significantly across the three GAP-score strata, the analysis revealed substantial between-study heterogeneity (I²>75% in every subgroup). Consequently, the 95% prediction intervals, which indicate the range in which the true proportion is expected to fall in a new study, were wide and largely overlapping. In practical terms, this means that despite the present point estimates (11% for Proportioned, 33% for Moderately disproportioned and 63% for Severely disproportioned), a future cohort could plausibly yield failure rates that blur these distinctions.
Fig. 1.
A forest plot showing the proportion of patients who developed mechanical complications, stratified by GAP-score category. For each study the leftmost columns list the author and the raw numerator-over-denominator counts (“Events / Total”); the right-hand axis plots the corresponding point estimate (grey square, whose size is proportional to the study weight) and its 95% confidence interval (horizontal whiskers). The plot is divided into 3 strata—“P” (Proportioned, GAP 0–2), “M” (Moderately disproportioned, GAP 3–6) and “S” (Severely disproportioned, GAP ≥7). Within every stratum a fixed-effect summary (labelled “Common effect model”) and a random-effects summary (Maximum Likelihood) are displayed as diamonds; beneath each diamond a red bar marks the 95% prediction interval, i.e. the range in which the true proportion of a future study is expected to lie. Subgroup heterogeneity is quantified by I², τ² and the Q-test p-value printed to the right of each block. At the bottom of the figure the overall pooled results for all 1221 patients are presented alongside global heterogeneity statistics and a χ² test that confirms the proportions differ significantly between the 3 GAP categories. The vertical dashed line represents the random-effects estimate for the entire dataset, enabling visual comparison of individual studies and subgroup summaries against the overall average.
In order to investigate some sources of possible heterogeneity the authors performed a meta-regression, and after accounting for surgical approach and age, GAP subgroup remained the dominant predictor of mechanical failure. Compared with Proportioned cases (P), the Moderately disproportioned (MD) group showed about 3-fold higher odds of failure (β=1.12, p=.003), while the Severely disproportioned (SD) group exhibited an almost 7-fold increase (β=1.99, p<.001). Hybrid constructs and open posterior spine fusions that included osteotomies were also associated with significantly greater failure risk (β=1.55 and 1.28, respectively), whereas mean cohort age had no independent effect (β=0.008, p=.83). Residual heterogeneity remained substantial (τ²=0.44; I²=78%), indicating that other unmeasured factors still contribute to between-study variability.
Several studies reported poor predictive accuracy. Baum et al. found no statistically significant difference in MC rates across the P, MD, and SD groups (19.0%, 30.4%, and 39.1%, respectively; p=.19) and a poor AUC of 0.621 [7]. Kwan et al., in a large external validation cohort, also found a low AUC of 0.60 for predicting MCs [8]. Similarly, Hiyama et al. and Lord et al. found no significant correlation between the GAP score and the incidence of MF or PJK, respectively [9,10].
In stark contrast, other studies have supported the score's validity. Jacobs et al. directly compared the GAP score to the SRS-Schwab classification and found the GAP score to be significantly better at predicting mechanical failure, with a good AUC of 0.86 versus 0.69 for Schwab [11]. Ham et al. investigated 84 elderly patients and found the GAP score to have "moderately accurate" predictive power for MCs, with an AUC of 0.839 [12]. Hiltunen et al. [2], in a cohort with a long mean follow-up of 7 years, also found a good predictive value (AUC=0.70) for MCs severe enough to require reoperation.
The varying statistical metrics of predictive accuracy across key studies are summarized in Table 4.
Table 4.
Summary of the statistical measures used to evaluate the predictive accuracy of each GAP-score variant across the included studies
| First author | Year | GAP variant | Accuracy metric / statistical test | Reported value | Comparison versus other benchmarks |
|---|---|---|---|---|---|
| Yilgor et al. [6] | 2017 | Postoperative GAP (0–13) | AUC (ROC analysis) | 0.88–0.92 | |
| Jacobs et al. [7] | 2019 | Postoperative GAP | AUC | 0.86 (vs Schwab 0.69) | GAP outperformed Schwab; OR 1.45 (p=.001) |
| Noh et al. [8] | 2020 | GAP-B (GAP + BMI + BMD) | AUC | GAP-B=0.885 GAP = 0.798 | GAP-B >GAP >SRS-Schwab |
| Ham et al. [9] | 2021 | Postoperative GAP (elderly) | AUC | 0.839 | |
| Kwan et al. [10] | 2021 | Postoperative GAP | AUC | 0.60 | GAP not predictive (p>.05) |
| Ha et al. [11] | 2021 | Partial intra-op iGAP (RLL + LDI + age) | Cochran–Armitage χ² trend test | NS (p=.92) | Not predictive; no AUC reported |
| Sun et al. [12] | 2021 | Postoperative GAP | AUC | 0.863 | Better than Roussouly classification (0.543) |
| Noh et al. [13] | 2021 | GAP-B (ML Random-Forest) | AUC (ROC) | 0.81 | |
| Hiltunen et al. [2] | 2023 | Postoperative GAP | Hazard ratio (Cox) & AUC | HR 3.55 (p<.01); AUC 0.70 | Cut-off GAP ≥ 5 predictive of reoperation |
| Gendelberg et al. [3] | 2023 | Postoperative GAP (cMIS) | χ² test / odds ratio | Not reported. Comapares results of the cMIS cohort with Original GP study | GAP present distinct prediction in the cMIS cohor compared to the original study |
| Hiyama et al. [21] | 2023 | Postoperative GAP | AUC | 0.74 | |
| Wang et al. [22] | 2024 | Postoperative GAP vs Schwab | χ² test of association (p-value reported, no AUC) | AUC not reported (p>.05) | GAP not predictive; Schwab also not |
| Park et al. [16] | 2024 | Postoperative GAP (categories) | χ² test | >0.05 | GAP categories did not differ among sagittal imbalance severity groups |
| Lim et al. [18] | 2025 | Postoperative GAP | Multivariable logistic-regression p-value | >0.6 (NS) | GAP not associated; T4–L1 PA mismatch predictive |
| Moridaira et al. [19] | 2025 | Postoperative GAP (mean value) | t-test comparing mean GAP (p-value) | 0.001 (univariate); NS multivariate | Low HU at UIV strongest predictor |
| Kumar et al. [17] | 2025 | Postoperative GAP | Logistic-regression p-value | 0.46 (NS) | Paraspinal muscle CSA/FI and smoking predictive |
For every citation the table lists the specific form of the Global Alignment & Proportion score that was analysed and the primary metric by which its ability to predict mechanical complications was judged—area-under-the-curve (AUC) when a receiver-operating-characteristic analysis was performed, or the statistical test and associated p-value (χ², logistic regression, Cox hazard ratio, 2-sample t-test, etc.) when only hypothesis-testing results were reported. Abbreviations: AUC, area under the ROC curve; HR, hazard ratio; OR, odds ratio.
The progression of predictive analytics: from modified scores to machine learning
The inconsistent performance of the original alignment-only GAP score has spurred the development of modified versions that incorporate patient-specific biological factors. The most notable of these is the GAPB score, which integrates Body Mass Index (BMI) and Bone Mineral Density (BMD) into the predictive model [4]. The rationale is that mechanical failure is a multifactorial process dependent not only on the forces applied to the construct (which correlate with spinal alignment) but also on the strength of the bone-implant interface (for which BMD is a proxy) and the overall load on the system (as measured by the patient's BMI) [[13], [14], [15]].
A pivotal study by Noh et al. [13] on 203 patients demonstrated that the GAPB system (AUC=0.885) had significantly improved predictability with regard to MCs compared to the original score (AUC=0.798), the SRS-Schwab classification (AUC=0.532), and age-adjusted goals (AUC=0.568). Further analysis by the same group investigated the GAPB's performance within the original GAP categories. They found that while the predictive power of GAPB was not statistically significant in the Proportioned group, its ability to predict MCs increased significantly with worsening GAPB scores in both the Moderately disproportioned and Severely disproportioned groups (p<.001) [14].
Leveraging these multifactorial inputs, recent research has explored the use of machine learning (ML) to enhance prediction. Unlike traditional regression models, ML algorithms can analyze complex, non-linear interactions between numerous variables. Using GAPB factors as inputs for several algorithms, Noh et al. [13] found that a random forest model achieved the highest predictive accuracy (73.2% in the test set), outperforming traditional logistic regression and other models, signaling a promising future direction for risk stratification [[14], [15]].
Comparison with other sagittal classification systems
Direct comparisons have highlighted the distinct utilities of different scoring systems. Wang et al. compared the GAP score with the SRS-Schwab classification in 69 patients and found that while the GAP score was a more effective predictor of MCs and revision surgery, the SRS-Schwab classification demonstrated a better correlation with improvements in patient-reported outcome measures (PROMs) [16]. Similarly, when compared to the Roussouly classification in a cohort of 80 elderly patients, Sun et al. concluded that the quantitative parameters of the GAP score made it more effective in predicting PJK (AUC=0.863) and PJF (AUC=0.724) [17].
Performance in specific clinical contexts and the role of biological factors
A growing body of evidence indicates that the GAP score's primary limitation is its failure to account for patient-specific biological factors that are critical determinants of mechanical stability.
-
•
Bone quality: Moridaira et al. investigated patients undergoing limited lumbar fusion and found that low bone quality, as measured by Hounsfield units (HU) on CT scans at the upper instrumented vertebra (UIV) and UIV+1, was the only independent risk factor for PJF in their multivariate analysis. A low HU value (<98) was associated with a 73.3% rate of PJF, highlighting bone quality as a more powerful predictor than alignment alone in this cohort [18].
-
•
Muscle status (sarcopenia): Kumar et al. analyzed the role of lumbar paraspinal muscle morphology and found that patients with MCs had significantly lower muscle cross-sectional area (CSA) and higher grades of fatty infiltration (FI). Critically, this held true even for patients in the Proportioned GAP category, where 33.3% of patients with good alignment still suffered an MC, a rate strongly associated with poor muscle quality. This demonstrates that sarcopenia can undermine an otherwise well-aligned construct [19].
Biomechanical underpinnings and other factors
A combined clinical and musculoskeletal modeling study by Ignasiak et al. provided a biomechanical rationale for the importance of alignment upon MCs after ASD surgery. Their model, applied to 205 ASD patients, demonstrated that greater sagittal malalignment, as reflected by a higher GAP score, was associated with significantly greater compressive and shear loads at the proximal adjacent segment. This provides a direct mechanistic link between poor alignment and the forces that can lead to PJK and implant failure [20].
Other studies have explored additional nuances. Lim et al. found that in patients undergoing pedicle subtraction osteotomy (PSO), the level of the osteotomy impacted specific GAP parameters; more distal PSOs (L4, L5) resulted in better correction of the Lordosis Distribution Index (LDI) and were associated with lower rates of revision for MCs [21]. Huang et al. identified distinct risk factors for the timing of PJK, with lower thoracic UIV placement predisposing to early-onset PJK and high preoperative PI predisposing to late-onset PJK, suggesting different underlying failure mechanisms [22].
Association with patient-reported outcomes (PROs)
The evidence linking GAP score proportionality to PROs is limited and complex. The external validation study by Kwan et al., while finding no association between the GAP score and MCs, reported a paradoxical finding: patients in the Moderately disproportioned category had significantly better scores on the SF-36, SRS-22, and Oswestry Disability Index (ODI) at 2-year follow-up compared to those in the Severely disproportioned group. This counterintuitive result may reflect the trade-off between achieving perfect radiographic alignment and the surgical morbidity required to do so. An aggressive correction to achieve a Proportioned state in a rigid, elderly patient may require highly invasive procedures like multiple 3-column osteotomies. The resulting surgical trauma could lead to worse short-term pain and function, negatively impacting PROs despite a radiographically ideal outcome. This suggests that the relationship between alignment and patient well-being is not linear and that the optimal correction must balance radiographic goals against the biological costs of the intervention [8].
Discussion
A paradigm shift from pure alignment to a patient-specific biomechanical model
This systematic review demonstrates that the predictive value of the GAP score for MCs following ASD surgery is moderate at best and, more importantly, is highly inconsistent across the published literature. The exceptional predictive accuracy reported in the original validation study (AUC=0.92) has not been consistently reproduced [6]. While some studies and a meta-analysis support its utility, a significant body of evidence [23], particularly from large external validation cohorts and specific surgical contexts, reports poor predictive power.
The principal finding of this updated review is that the GAP score, as a purely radiographic tool, is an insufficient stand-alone predictor of mechanical failure. This signals a fundamental paradigm shift in understanding MCs: they are not merely an alignment problem but a multifactorial biomechanical event. The inability of the GAP score inability to account for critical patient-specific biological factors, such as bone quality and muscle status, represents a fundamental limitation that explains much of the observed heterogeneity in its performance.
Dissecting the sources of heterogeneity
The conflicting findings in the literature can be largely attributed to key sources of heterogeneity, now including powerful biological and biomechanical confounders.
-
1.
Biological factors: The evidence from studies by Moridaira et al. and Kumar et al. is compelling. Poor bone quality (low HU) and sarcopenia (low muscle CSA, high FI) have been identified as independent, and in some cases stronger, predictors of MCs than the GAP score itself. A patient with a Proportioned alignment but severe osteoporosis and sarcopenia may be at a higher risk of failure than a Moderately disproportioned patient with excellent bone and muscle stock. Unfortunately the GAP score is blind to this crucial distinction [18,19]. Moreover, different primary pathologies presented different accuracy patterns regarding the GAP score, Lim et al. [18] with low Hounsfield-unit scores at the upper instrumented vertebra outweighed sagittal parameters.
-
2.
Surgical technique: The biomechanical differences between open and minimally invasive surgery remain a key factor, with the GAP score performing poorly in circumferential minimally invasive surgery (cMIS) cohorts where the posterior tension band is preserved. Furthermore, the specific surgical strategy, such as the level of a PSO, can influence outcomes and specific alignment parameters like the LDI, adding another layer of complexity not fully captured by the overall score [7].
-
3.
Biomechanical loads: Ignasiak et al. elegantly closed the biomechanical loop by demonstrating through subject-specific musculoskeletal modelling of 205 fused spines, that every 10-mm increase in postoperative sagittal vertical axis (SVA) raises compressive loads at the proximal adjacent segment by approximately 6% and shear loads by approximately 8% [20]. While these findings validate the rationale for restoring alignment, the clinical literature shows that the magnitude of that alignment effect is often matched, or even outweighed, by patient-specific factors that influence either the applied load (eg, body mass, muscle envelope) or the tolerance of the construct-bone interface (eg, bone mineral density, cortical integrity). For example, in the derivation/validation cohorts of Noh et al. [13] the addition of BMI and BMD to the GAP score (the “GAP-B” model) increased the AUC for predicting mechanical failure from 0.80 to 0.89, while each 5-kg/m² rise in BMI conferred a 28% increase in hazard and osteoporotic bone (T-score ≤−2.5) reduced construct survival by approximately 70%. Likewise, the machine-learning analysis of Noh et al. [13] ranked BMD and BMI above all alignment terms, achieving an AUROC of 0.81 versus 0.63 for spinopelvic parameters alone. Bone quality effects are not limited to dual-energy absorptiometry: Moridaira et al. showed that every 10-HU decrease in trabecular density at the upper instrumented vertebra increased the odds of PJF by 25% (OR 0.975 per HU; AUC 0.75), independent of postoperative GAP. Muscle status exerts a comparable influence [14,18]. In Kumar et al. each standard-deviation drop in lumbar paraspinal cross-sectional area halved construct durability (OR ≈0.57), whereas fatty-infiltration above 50% increased failure risk 6-fold—even though GAP category itself was non-significant (p=.46) [19]. Finally, Ham et al. reported that low femoral-neck T-score and age >70 years doubled the likelihood of failure despite excellent final alignment, and Gendelberg et al. found that minimally invasive circumferential constructs, associated with smaller posterior muscle disruption, had a 40% lower complication rate than historical open-fusion series, again undermining a purely alignment-based prediction [3,12].
-
4.
Variability in ideal alignment formulas: The very foundation of the GAP score—the calculation of 'ideal' lordosis—has been questioned. A study by Polly et al. compared 5 different published formulas for ideal lumbar lordosis, including the one used by Yilgor et al., and found statistically significant variations between them. This suggests that the proportionality measured by the GAP score is dependent on which foundational formula is chosen, introducing another layer of variability [6,24].
-
5.
Different definitions of mechanical failure: An additional source of variability stems from how each study defined mechanical complication. Some cohorts—most notably Yilgor et al. [6] and Jacobs et al. [11], counted any radiographic evidence of implant failure (proximal-junctional kyphosis ≥10°, rod breakage, screw pull-out) irrespective of whether re-operation ensued, whereas others restricted events to complications that prompted revision surgery [2,12]. Thresholds also differed: the angular cut-off for PJK ranged from 10° to 20°, and Huang et al. separated early (<3 months) from late PJK (> 3 months), generating 2 distinct event categories [16,17]. Several series treated asymptomatic RFs as mechanical failures, while others considered only mechanical complications requiring revision surgery. Follow-up duration further influenced event capture, with studies reporting as little as 12 months [25] versus 5 years [13]. This heterogeneity in end-point definition widens prediction intervals and dilutes direct comparability of accuracy metrics; it also helps explain why GAP performed best when the same broad composite outcome (radiographic failure + need of revision surgery) used in the derivation study was replicated, and less well when narrower or differently timed criteria were applied.
Clinical implications
Based on the evidence synthesized in this systematic review, the GAP score should not be used as a rigid, stand-alone algorithm for surgical planning of deformity correction. Its application as a prescriptive tool to dictate precise alignment targets is not supported by the current body of literature. Instead, it should be viewed as one valuable component within a more comprehensive, patient-centered decision-making framework.
The most significant clinical implication is the clear need to move beyond purely radiographic predictive models. The success of the modified GAPB score and the promise of machine learning models highlight that mechanical failure is a multifactorial process [14,15]. An accurate risk assessment must integrate at least 3 domains: the forces acting on the construct (captured by alignment metrics like the GAP score), the integrity of the bone-implant interface (as measured by BMD or HU), and the overall physiological resilience and soft-tissue envelope of the patient (as measured by muscle status, BMI, and/or comorbidities). Paradoxically, the alignment targets that maximize radiographic perfection and early patient-reported outcome measures (PROMs) may simultaneously heighten the risk of mechanical failure. Several studies in our review illustrate this tension. Yilgor [6] and Noh [13] shows that more aggressive correction into the Proportioned GAP zone correlated with larger postoperative gains in ODI and SRS-22R scores, yet those same cohorts revealed that patients who slipped from a Proportioned to a Severely disproportioned status after RF or PJK often maintained acceptable PROMs as long as the loss of correction remained sub-threshold. Conversely, Park [26] reported that surgeons who aimed for less stringent age-adjusted PI-LL targets experienced lower revision rates, even though their radiographs remained Moderately disproportioned, and still achieved clinically important improvements in ODI. This echoes the finding of Kwan [8], where over-correction toward ideal spinopelvic parameters raised the incidence of junctional failure without a commensurate PROM benefit at 2 years. Together, these observations suggest a non-linear relationship: PROMs improve up to the point where functional sagittal balance is restored, beyond which additional correction delivers diminishing clinical returns while escalating biomechanical stress at the construct ends, ultimately suggesting the existence of a thin-line between over-correction and complications and hypo-correction and unresolved pain.
Limitations and future directions
This review is inherently limited by the quality of the available scientific evidence in the literature. All included studies were retrospective and the majority were from single centers, which introduces risks of selection and procedural bias. Furthermore, long-term follow-up (beyond 5 years) remains relatively scarce, something that appears critical given that complications like RF can manifest late. Moreover, methodological limitations, such as non-standardized mechanical failure criteria and lack of sample size calculation and power analysis in most of the studies, may impact the reliability of the reported results.
Future research must prioritize large-scale, prospective, multicenter studies that are designed to validate multi-domain predictive models. The integration of novel variables, such as paraspinal muscle quality and bone quality assessed by Hounsfield units on CT scans, is essential. The development of ethnicity-adjusted GAP scores, such as the C-GAP for Asian populations, is also a necessary step to improve applicability across diverse populations. Technological innovations such as patient-specific rods (PSRs), which are pre=contoured based on surgical planning, may help improving the rates in which alingment goals are achieved. However, a recent systematic review by Picton et al. [27] found that, while PSRs show potential, the current evidence is too heterogeneous to definitively prove their superiority over traditional rods. Ultimately, leveraging machine learning and artificial intelligence to analyze large, integrated datasets containing radiographic, biological, and surgical variables may provide the most accurate and personalized risk prediction, allowing surgeons to better counsel patients and tailor interventions to minimize the risk of mechanical failure.
Conclusions
The global alignment and proportion (GAP) score represents a significant conceptual advance in the surgical management of ASD by championing a patient-specific, proportional approach to sagittal alignment. However, this systematic review reveals that its stand-alone predictive power with regard to MCs after ASD surgery is, at best, inconsistent and fundamentally limited by its exclusion of critical biological factors. Ultimately its validity is questionable in cohorts undergoing minimally invasive surgery and its predictive power seems to be strongly modulated by patient-specific variables such as bone quality and paraspinal muscle status, which have emerged as powerful independent predictors of failure. The collective evidence strongly indicates that sagittal alignment is but one component in a complex, multifactorial etiology of mechanical failure. To truly advance the field and reduce the burden of revision surgery, future research and clinical practice must embrace more comprehensive predictive models, like the GAPB score, and advanced analytical tools like machine learning to integrate these patient-specific biological and surgical variables with radiographic alignment parameters.
Short summary
The GAP score moderately predicts mechanical complications after adult spinal deformity surgery, but accuracy is highly variable. Multifactorial models like GAPB and machine learning may improve prediction by integrating patient-specific biological and biomechanical factors.
Funding
This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors.
Ethical approval
As this study is a systematic review of previously published literature, institutional review board (IRB) approval was not required.
Declaration of competing interest
The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
Footnotes
Author disclosures: VF: Nothing to disclose. GP: Nothing to disclose. CG: Nothing to disclose. MS: Nothing to disclose. MF: Nothing to disclose. PM: Nothing to disclose. TM: Nothing to disclose.
FDA device/drug status: Not applicable.
Supplementary material associated with this article can be found, in the online version, at doi:10.1016/j.xnsj.2025.100816.
Appendix. Supplementary materials
References
- 1.Smith J.S., Shaffrey C.I., Ames C.P., Lenke L.G. Treatment of adult thoracolumbar spinal deformity: past, present, and future. J Neurosurg Spine. 2019;30:551–567. doi: 10.3171/2019.1.SPINE181494. [DOI] [PubMed] [Google Scholar]
- 2.Hiltunen S., Repo J.P., Pekkanen L., et al. Mechanical complications and reoperations after adult spinal deformity surgery: a clinical analysis with the GAP score. Eur Spine J. 2023;32:1421–1428. doi: 10.1007/s00586-023-07593-9. [DOI] [PubMed] [Google Scholar]
- 3.Gendelberg D., Rao A., Chung A., et al. Does the Global Alignment and Proportion score predict mechanical complications in circumferential minimally invasive surgery for adult spinal deformity? Neurosurg Focus. 2023;54:E11. doi: 10.3171/2022.10.FOCUS22600. [DOI] [PubMed] [Google Scholar]
- 4.Daniels A.H., Reid D.B.C., Tran S.N., et al. Evolution in surgical approach, complications, and outcomes in an adult spinal deformity surgery multicenter study group patient population. Spine Deform. 2019;7:481–488. doi: 10.1016/j.jspd.2018.09.013. [DOI] [PubMed] [Google Scholar]
- 5.Soroceanu A., Diebo B.G., Burton D., et al. Radiographical and implant-related complications in adult spinal deformity surgery. Spine (Phila Pa 1976) 2015;40:1414–1421. doi: 10.1097/BRS.0000000000001020. [DOI] [PubMed] [Google Scholar]
- 6.Yilgor C., Sogunmez N., Boissiere L., et al. Global alignment and proportion (GAP) Score: development and validation of a new method of analyzing spinopelvic alignment to predict mechanical complications after adult spinal deformity surgery. J Bone and Joint Surg - Am Volume. 2017;99:1661–1672. doi: 10.2106/JBJS.16.01594. [DOI] [PubMed] [Google Scholar]
- 7.Jacobs E., van Royen B.J., van Kuijk S.M.J., et al. Prediction of mechanical complications in adult spinal deformity surgery-the GAP score versus the Schwab classification. Spine J. 2019;19:781–788. doi: 10.1016/j.spinee.2018.11.013. [DOI] [PubMed] [Google Scholar]
- 8.Noh S.H., Ha Y., Obeid I., et al. Modified global alignment and proportion scoring with body mass index and bone mineral density (GAPB) for improving predictions of mechanical complications after adult spinal deformity surgery. Spine J. 2020;20:776–784. doi: 10.1016/j.spinee.2019.11.006. [DOI] [PubMed] [Google Scholar]
- 9.Ham D-W, Kim H-J, Choi J.H., et al. Validity of the global alignment proportion (GAP) score in predicting mechanical complications after adult spinal deformity surgery in elderly patients. Eur Spine J. 2021;30:1190–1198. doi: 10.1007/s00586-021-06734-2. [DOI] [PubMed] [Google Scholar]
- 10.Kwan K.Y.H., Lenke L.G., Shaffrey C.I., et al. Are higher global alignment and proportion scores associated with increased risks of mechanical complications after adult spinal deformity surgery? An external validation. Clin Orthop Relat Res. 2021;479:312–320. doi: 10.1097/CORR.0000000000001521. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Ha A.S., Hong D.Y., Coury J.R., et al. Partial intraoperative global alignment and proportion scores do not reliably predict postoperative mechanical failure in adult spinal deformity surgery. Global Spine J. 2021;11:1046–1053. doi: 10.1177/2192568220935438. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Sun X., Sun W., Sun S., et al. Which sagittal evaluation system can effectively predict mechanical complications in the treatment of elderly patients with adult degenerative scoliosis? Roussouly classification or global alignment and proportion (GAP) Score. J Orthop Surg Res. 2021;16:641. doi: 10.1186/s13018-021-02786-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Noh S.H., Ha Y., Park J.Y., et al. Modified global alignment and proportion scoring with body mass index and bone mineral density analysis in global alignment and proportion score of each 3 categories for predicting mechanical complications after adult spinal deformity surgery. Neurospine. 2021;18:484–491. doi: 10.14245/ns.2142470.235. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Noh S.H., Lee H.S., Park G.E., et al. Predicting mechanical complications after adult spinal deformity operation using a machine learning based on modified global alignment and proportion scoring with body mass index and bone mineral density. Neurospine. 2023;20:265–274. doi: 10.14245/ns.2244854.427. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Lord E.L., Ayres E., Woo D., et al. The impact of global alignment and proportion score and bracing on proximal junctional kyphosis in adult spinal deformity. Global Spine J. 2023;13:651–658. doi: 10.1177/21925682211001812. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Park S-J, Park J-S, Kang D-H, et al. Comparison of surgical burden, radiographic and clinical outcomes according to the severity of baseline sagittal imbalance in adult spinal deformity patients. Neurospine. 2024;21:721–731. doi: 10.14245/ns.2448250.125. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Kumar G., Tandon V., Nanda A., et al. Does the Lumbar Paraspinal Muscle Status Have a Role in Predicting Mechanical Complications After Adult Spinal Deformity Surgery? Global Spine J. 2025 doi: 10.1177/21925682251330830. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Lim P., Clark A.J., Katz A.D., et al. Impact of lumbar pedicle subtraction osteotomy (PSO) level on global alignment and proportion (GAP) score in revision adult iatrogenic flatback spinal deformities. Spine Deform. 2025 doi: 10.1007/s43390-025-01141-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Moridaira H., Inami S., Takahata M., et al. The association between lower Hounsfield units of the upper instrumented vertebra and proximal junctional failure after limited lumbar fusion for adult spinal deformity. BMC Musculoskelet Disord. 2025;26:393. doi: 10.1186/s12891-025-08643-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Baum G.R., Ha A.S., Cerpa M., et al. Does the global alignment and proportion score overestimate mechanical complications after adult spinal deformity correction? J Neurosurg Spine. 2021;34:96–102. doi: 10.3171/2020.6.SPINE20538. [DOI] [PubMed] [Google Scholar]
- 21.Hiyama A., Katoh H., Sakai D., et al. Analysis of mechanical failure using the GAP score after surgery with lateral and posterior fusion for adult spinal deformity. Global Spine J. 2023;13:2488–2496. doi: 10.1177/21925682221088802. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Wang Z., Wu B., Wang Z., et al. Comparison of global alignment and proportion (GAP) score and SRS-Schwab ASD classification in the analysis of surgical outcomes for adult spinal deformity. Indian J Orthop. 2024;58:762–770. doi: 10.1007/s43465-024-01147-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Ignasiak D., Behm P., Mannion A.F., et al. Association between sagittal alignment and loads at the adjacent segment in the fused spine: a combined clinical and musculoskeletal modeling study of 205 patients with adult spinal deformity. Eur Spine J. 2023;32:571–583. doi: 10.1007/s00586-022-07477-4. [DOI] [PubMed] [Google Scholar]
- 24.Huang D., Kim H.J., Wang Z., et al. Identifying distinct risk factors for early-onset and late-onset PJK in ASD: a comparative analysis across non-PJK, early-onset PJK, and late-onset PJK groups. Global Spine J. 2025 doi: 10.1177/21925682251345755. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Cho M., Lee S., Kim H-J. Assessing the predictive power of the GAP score on mechanical complications: a comprehensive systematic review and meta-analysis. Eur Spine J. 2024;33:1311–1319. doi: 10.1007/s00586-024-08135-7. [DOI] [PubMed] [Google Scholar]
- 26.Polly D.W., Haselhuhn J.J., Keller N., et al. Differences across various ideal lumbar lordosis measurement formulas for patient-specific sagittal alignment goals. Neurosurg Focus. 2025;58:E9. doi: 10.3171/2025.3.FOCUS2568. [DOI] [PubMed] [Google Scholar]
- 27.Picton B., Stone L.E., Liang J., et al. Patient-specific rods in adult spinal deformity: a systematic review. Spine Deform. 2024;12:577–585. doi: 10.1007/s43390-023-00805-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.

