Abstract
Purpose
To develop mapping algorithms from the EORTC-QLQ-C30 to EQ-5D-5L utilities for gynecological cancer survivors in follow-up care, enabling cost-effectiveness analyses when EQ-5D-5L utilities are unavailable.
Methods
We used data from a Norwegian multicentre longitudinal study of gynecological cancer survivors in post-treatment follow-up (LETSGO), including 663 patients with 2,624 observations. Ten model types were estimated for three EQ-5D-5L value sets (Norwegian, UK, US): ordinary least squares, Tobit, beta and fractional logistic regression, linear mixed models, an adjusted limited dependent variable mixture model, and two-part models combining logistic regression for the probability of perfect health with OLS, mixed, beta or fractional logistic regression for utilities below 1. Each was fitted with three prespecified covariate sets comprising all EORTC-QLQ-C30 scales, with and without age, comorbidities and treatment type. Stratified five-fold cross-validation at the patient level was used for internal validation. Performance was assessed using mean absolute error (MAE), root mean squared error (RMSE), proportion of predictions with absolute error (AE) ≤ 0.05 and Lin’s concordance correlation coefficient (CCC) and summarized using an average ranking value.
Results
A two-part fractional logistic model with all covariates ranked highest, with MAE 0.0539–0.0660, RMSE 0.0806–0.0960, AE ≤ 0.05 for 54.8–67.1% of observations and CCC 0.7559–0.8034. Differences between the best-performing specifications were small. Accuracy declined at low utility levels, though overall mean bias was minimal.
Conclusions
The algorithm enables estimation of EQ-5D-5L utilities from EORTC-QLQ-C30 scores in populations with severity profiles similar to the LETSGO cohort. Accuracy is reduced in poor health states.
Supplementary Information
The online version contains supplementary material available at https://doi.org/10.1007/s11136-026-04399-2.
Keywords: Mapping, Crosswalk, EQ-5D-5L, EORTC-QLQ-C30, Gynecological cancer
Plain language summary
Gynecological cancers were diagnosed in nearly 1.5 million women worldwide in 2022. As survival improves, more women live for many years after treatment, often with lasting effects on their health and daily lives. Deciding how to spend limited healthcare resources requires comparing the benefits of different treatments and services. These comparisons use a specific measure of health-related quality of life, but many clinical studies collect other questionnaires instead. When the required measure is missing, researchers can use a statistical method called mapping to estimate it from the questionnaires that were collected. Many mapping methods exist, but none have been developed for women who have completed treatment for gynecological cancer. We therefore developed and tested such a method, using data from 663 women followed for up to three years after treatment in Norway. We compared ten different statistical approaches and found one that predicted quality-of-life values accurately for most women. Predictions were less accurate for women in the poorest health, so results should be interpreted with care when this group is important. The method allows researchers to include gynecological cancer survivors in health economic evaluations even when the required questionnaire was not used.
Supplementary Information
The online version contains supplementary material available at https://doi.org/10.1007/s11136-026-04399-2.
Introduction
Gynecological cancers (ovarian, uterine, cervical and vulvar/vaginal malignancies) account for 15% of new cancer diagnoses affecting women globally [1], with 1,473,427 new cases and 680,372 deaths in 2022 [2]. An expected increase in the number of new cases over the next decades [2] together with improved survival rates over time [3, 4], means more women will be living with disease- and treatment related late effects, highlighting the need for evidence-based, cost-effective survivorship care and health economic evaluations.
Health economic evaluations inform decisions on allocating scarce resources by comparing costs and health outcomes across interventions. Many health technology assessment agencies recommend cost-effectiveness analysis based on quality-adjusted life years (QALYs), which combine survival with health-state utility values (HSUVs) derived from generic preference-based measures such as the EuroQol five-dimension questionnaire (EQ-5D) [5].
Clinical studies frequently collect disease-specific instruments, such as the European Organisation for Research and Treatment of Cancer Quality of Life Questionnaire Core 30 [6] (EORTC-QLQ-C30; hereafter QLQ-C30) for cancer patients, to assess patient-reported outcomes. Because these non-preference-based measures are not suited for economic evaluation, numerous studies have estimated algorithms to translate health-related quality of life (HRQoL) scores into HSUVs such as EQ-5D. Although mapping introduces additional uncertainty compared to direct measurement of EQ-5D and direct measurement of HSUVs is always recommended [7], the use of mapping algorithms is frequently the only feasible approach for deriving utility values for cost-effectiveness analysis when the preferred instrument was not collected in the original study. The ISPOR guidelines for good practice in mapping studies [7] and The MAPS Reporting Statement for Studies Mapping onto Generic Preference-Based Outcome Measures [8] (Online Resource 1) provide methodological guidance for the development, validation, and reporting of such algorithms.
Despite the growing number of mapping studies, gynecological cancers remain underrepresented. The Health Economics Research Centre (HERC) database of mapping studies version 10 [9] lists 275 studies mapping to EQ-5D, including 28 from QLQ-C30, but none based exclusively on gynecological cancer patients. As highlighted by Longworth and Rowen in their note on studies mapping from QLQ-C30 to EQ-5D, there is a need for mapping algorithms specific to distinct cancer patient populations [10]. Pooling diagnoses reflects the intended use of the algorithms in mixed survivorship clinics.
The primary objective of the present study was to develop and internally validate mapping algorithms from QLQ-C30 to utility values derived from the five-level version of the EQ-5D (EQ-5D-5L) [11], using Norwegian, UK, and US value sets, for gynecological cancer survivors in follow-up care, enabling cost-effectiveness analyses when EQ-5D-5L has not been collected.
Methods
Estimation sample
The estimation sample was obtained from the Lifestyle and Empowerment Techniques in Survivorship of Gynaecologic Oncology (LETSGO) study [12] (ClinicalTrials.gov: NCT04122235), a multicentre quasi-experimental study evaluating an alternative follow-up strategy for gynecological cancer survivors compared with standard follow-up. 755 participants were recruited at five university hospitals and seven local hospitals across Norway, of which 14 were screening failures, leaving 741 eligible patients included in the study. The questionnaire-package included QLQ-C30, EQ-5D-5L and the Self-Administered Comorbidity Questionnaire [13]. Questionnaires were collected at baseline (when patients entered follow-up after end of treatment) and then at 3, 6, 12, 24 and 36 months. Information on cancer diagnosis and primary treatment was collected from medical records. Further details on the LETSGO study are available elsewhere [12, 14].
Source and target measures
EQ-5D-5L is a generic HRQoL questionnaire with five dimensions; mobility, self-care, usual activities, pain/discomfort and anxiety/depression, each with five response levels. Dimension responses define health profiles, which are converted to utility values using country-specific value sets, ranging from 1 (perfect health) to values below 0 (worse than dead). We used EQ-5D-5L utilities calculated based on the Norwegian value set [15], the recently published UK value set [16] and the US value set [17] as target measures.
The QLQ-C30 is a cancer-specific HRQoL instrument, widely used in clinical trials. It consists of 30 items, divided into 15 functional items, 12 symptom items, 2 global health status/quality of life items and 1 item on financial difficulties. Each functional-, symptom- and the financial difficulties item has four levels: not at all, a little, quite a bit and very much. The global health status and quality of life items are rated on a scale from 1 to 7, with 1 indicating “very poor” and 7 indicating “excellent”. The items are grouped into 5 functional scales, 9 symptom scales/items (including financial difficulties) and one global health status and quality of life scale (QL). We transformed the Likert-type responses into linear scale scores ranging from 0 to 100, in accordance with the EORTC scoring manual [18]. A high score for functional scales represents high levels of functioning, a high score for QL represents high HRQoL, while high scores for symptom scales/items represent a high level of problems.
We examined conceptual overlap between EQ-5D-5L utilities and each of the 15 QLQ-C30 scales using Spearman correlations.
Missing data
Complete-case analysis was pre-specified as the primary approach. To characterize non-response, expected assessments were derived from each patient's registration date and the protocol schedule, excluding time points falling after study exit (due to death, disease progression, or withdrawal) or after the data extraction date. Observations with incomplete responses to EQ-5D-5L dimensions and/or QLQ-C30 items, resulting in the inability to compute EQ-5D utilities or one or more QLQ-C30 scale scores in accordance with the EORTC scoring manual were excluded in the primary analysis. To assess the robustness of the complete-case approach, we used multiple imputation to impute missing EQ-5D dimensions and QLQ-C30 scale scores among returned, incomplete questionnaires and re-estimated the preferred model (Online Resource 2).
Validation methods
We used stratified five-fold cross-validation at the patient level for all models, with the same fold assignment applied to all analyses. Fold construction was based solely on EQ-5D-5L domain responses, independent of value set. For each patient we derived an indicator of ever reporting perfect health and the median Level Sum Score over time; patients were stratified by these measures, randomly ordered within strata (fixed seed) and allocated to folds to achieve approximately equal fold sizes and a balanced severity distribution. All repeated measures from a given patient were assigned to the same fold, with each fold used once for validation and four times for training.
Modelling approaches
We estimated several direct mapping models for each value set to characterize the relationship between QLQ-C30 scores and EQ-5D-5L utilities: ordinary least squares (OLS), linear mixed models (LMM), Tobit, fractional logistic and beta regressions, two-part models combining a logistic regression for P(Perfect health) with OLS, LMM, fractional logistic or beta regression for utilities conditional on imperfect health and adjusted limited dependent variable mixture models (ALDVMM). OLS models were included as a benchmark because they are widely used in the mapping literature [19], despite potentially predicting values outside the EQ-5D-5L range. LMMs with a patient level random intercept were included to account for the repeated measures structure of the data. Tobit models, although designed for censored outcomes rather than EQ-5D-5L utilities, are also frequently reported and were therefore considered for comparison [19]. Fractional logistic regression constrains predictions to the closed interval [0, 1], while beta regression constrains predictions to the open interval (0, 1), both of which correspond closely to the typical range of EQ-5D-5L utilities. Because a considerable proportion of patients answering EQ-5D-5L report no problems on all items, indicating “perfect health”, we also fitted two-part models to explicitly model the probability of perfect health and the expected utility conditional on imperfect health. ALDVMMs were included as they are frequently recommended for mapping to EQ-5D, particularly where the utility distribution exhibits a spike at full health and multimodality below it [20]. ALDVMMs were estimated with two components; models with additional components did not converge. A 1-inflated beta model was additionally explored but produced results almost identical to the two-part beta model and is not reported further.
For models using fractional logistic and beta regression, observed utilities were transformed linearly to the appropriate interval before estimation, following Smithson & Verkulien [21]. Utilities were first mapped from ( [L,1]) to ( [0,1]) using
![]() |
where (L) is the lower bound of the value set. For beta regression, these values were further transformed to the open interval ((0,1)) as
![]() |
where N is the number of observations in the estimation folds. All models were estimated in Stata with cluster-robust standard errors to account for repeated measures per patient.
We also explored response-mapping models, in which each EQ-5D-5L dimension is modelled separately, but concluded this was inappropriate because of sparse data in the most severe levels of several dimensions, which can compromise model stability [7].
As a sensitivity analysis, the selected model was re-estimated with quadratic terms for the two QLQ-C30 scales most strongly correlated with the EQ-5D index, to assess whether the assumption of linearity materially affected predictive performance.
Covariates considered for mapping
For each model, we estimated three covariate sets: (1) all QLQ-C30 scale scores; (2) QLQ-C30 plus age; and (3) QLQ-C30 plus age, number of comorbidities and a binary indicator included to distinguish patients who received surgery only (1) from those who received any additional or alternative treatment, including radiotherapy, chemotherapy, or hormonal therapy (0)). Age was included because of its well-established association with HRQoL and recommendations to include it when available [7]. Comorbidities and treatment type were added a priori, based on the hypothesis that additional diseases and systemic treatments are associated with lower HRQoL during follow-up [22]. To assess potential multicollinearity from including all QLQ-C30 scales, we calculated variance inflation factors.
Estimation of predicted utilities
Predicted utilities were obtained by applying the estimated model coefficients to each observation’s covariate values, using Stata’s predict postestimation command. Predictions from models using fractional logistic- and beta regressions were transformed back to the original scale (for details; see Online Resource 3). For the LMMs, predictions were obtained by setting the random intercept to zero, yielding the population-averaged expectation. For the OLS, LMM and the two-part models using OLS and LMM conditional on imperfect health, predicted values exceeding 1 were set to 1; no predictions fell below the lower bound for any value set. Predicted utility index values from the two-part models were estimated as
![]() |
where P(Perfect health) denotes the predicted probability of having perfect health and U the predicted utility conditional on imperfect health.
Measures of model performance
Out-of-fold predictive performance was evaluated using mean absolute error (MAE), root mean squared error (RMSE), proportion of predictions with absolute error (AE) ≤ 0.05 and Lin’s concordance correlation coefficient (CCC). For each value set, all model specifications were ranked on each metric, and an average ranking value (ARV) was calculated across metrics to summarize overall performance and select preferred models. We also report mean bias (predicted – observed utility), the range of predicted utilities and the calibration intercept, slope and
from calibration regressions of observed on predicted utilities, estimated on out-of-fold predictions with cluster-robust standard errors at the patient level. Predictive performance was additionally assessed stratified on cancer diagnosis, treatment groups and levels of observed utility.
Confidence intervals for MAE, RMSE and the proportion of predictions with AE ≤ 0.05 were obtained by cluster bootstrap with patient-level resampling (1,000 replications, percentile intervals), conditioning on the estimated models; intervals for CCC and for the calibration statistics were obtained analytically. Differences in predictive performance between the selected model and each alternative specification were estimated using a paired cluster bootstrap on the same resamples, accounting for the correlation between prediction errors across models.
Results
Across all 741 eligible patients, 3,836 assessments were expected, of which 1,014 (32%) were not returned; of these, 77% were from patients who completed follow-up as planned or were still in active follow-up at data extraction, 10% from patients who later discontinued follow-up due to progression or death, and 13% from patients who later discontinued follow-up for other or unrecorded reasons. 67 patients returned neither the EQ-5D-5L nor the QLQ-C30 at any time point and were therefore not available for the mapping analyses, leaving 674 patients.
Final sample size
In total, 198 observations were excluded from the primary analysis due to incomplete response, and 11 patients were excluded entirely because all their observations had incomplete responses. Complete case exclusion resulted in a final sample of 663 patients and 2,624 observations at different time points (number of observations per patient: mean 3.96; median 4; minimum 1; maximum 6). The robustness of this exclusion is examined in a multiple imputation sensitivity analysis (Online Resource 2).
Descriptive information of final sample
At baseline (Table 1), a majority of patients were aged between 45 and 74 years. Educational level was generally medium or high, and most patients were married or cohabiting and not in paid work. Uterine and ovarian cancers were the most common diagnoses, reflecting the general distribution of gynaecological cancer types in Norway. Most patients underwent surgery as part of their treatment (92.2%), and 55% of patients were treated with surgery only. 71.9% reported at least one comorbidity.
Table 1.
Baseline characteristics of 663 patients
| Variables | N | % |
|---|---|---|
| Age at baseline | ||
| < 45 | 74 | 11.2 |
| 45–59 | 185 | 27.9 |
| 60–74 | 304 | 45.9 |
| ≥ 75 | 100 | 15.1 |
| Education | ||
| Missing | 6 | 0.9 |
| Low | 89 | 13.4 |
| Medium | 294 | 44.3 |
| High | 274 | 41.3 |
| Family status | ||
| Missing | 4 | 0.6 |
| Unmarried | 71 | 10.7 |
| Married/cohabitant | 454 | 68.5 |
| Divorced | 68 | 10.3 |
| Widowed | 66 | 9.9 |
| Currently in paid work | ||
| No | 420 | 63.3 |
| Yes | 243 | 36.7 |
| Cancer diagnosis | ||
| Vulva | 10 | 1.5 |
| Vagina | 2 | 0.3 |
| Cervix | 105 | 15.8 |
| Uterus | 376 | 56.7 |
| Ovarii/perineum | 170 | 25.6 |
| Treatment | ||
| Hormonal treatment | 14 | 2.1 |
| Chemotherapy | 284 | 42.8 |
| Radiation therapy | 56 | 8.4 |
| Surgery | 611 | 92.2 |
| Surgery only | 365 | 55.0 |
| Number of comorbidities | ||
| None | 186 | 28.1 |
| 1 | 193 | 29.3 |
| 2 | 155 | 23.4 |
| 3 or more | 129 | 19.3 |
Mean observed EQ-5D-5L utilities across all observations were relatively high for both the Norwegian, UK, and US value sets (0.870, 0.881 and 0.842, respectively) (Table 2). Functioning scale scores were also high, with low mean symptom scores. Fatigue, insomnia and pain were the most prominent symptoms.
Table 2.
Source and target measures, across all observations
| Variable | Obs | Mean | SD | Min | Max |
|---|---|---|---|---|---|
| EQ-5D-5L Norwegian value set | 2,624 | 0.870 | 0.137 | − 0.053 | 1.000 |
| EQ-5D-5L UK value set | 2,624 | 0.881 | 0.141 | − 0.217 | 1.000 |
| EQ-5D-5L US value set | 2,624 | 0.842 | 0.167 | − 0.309 | 1.000 |
| Global health status (QL score) | 2,624 | 71.637 | 19.808 | 0.000 | 100.000 |
| Physical functioning (PF score) | 2,624 | 81.311 | 18.910 | 6.667 | 100.000 |
| Role functioning (RF score) | 2,624 | 76.359 | 27.025 | 0.000 | 100.000 |
| Emotional functioning (EF score) | 2,624 | 82.120 | 20.105 | 0.000 | 100.000 |
| Cognitive functioning (CF score) | 2,624 | 82.698 | 20.515 | 0.000 | 100.000 |
| Social functioning (SF score) | 2,624 | 78.163 | 24.369 | 0.000 | 100.000 |
| Fatigue (FA score) | 2,624 | 33.587 | 25.503 | 0.000 | 100.000 |
| Nausea and vomiting (NV score) | 2,624 | 4.300 | 10.706 | 0.000 | 100.000 |
| Pain (PA score) | 2,624 | 25.197 | 25.602 | 0.000 | 100.000 |
| Dyspnoea (DY score) | 2,624 | 17.111 | 23.927 | 0.000 | 100.000 |
| Insomnia (SL score) | 2,624 | 32.876 | 30.855 | 0.000 | 100.000 |
| Appetite loss (AP score) | 2,624 | 8.562 | 19.985 | 0.000 | 100.000 |
| Constipation (CO score) | 2,624 | 19.576 | 27.429 | 0.000 | 100.000 |
| Diarrhoea (DI score) | 2,624 | 17.658 | 24.712 | 0.000 | 100.000 |
| Financial difficulties (FI score) | 2,624 | 6.314 | 17.976 | 0.000 | 100.000 |
Across value sets, utilities were strongly correlated (|p|≥ 0.6) with physical functioning, pain and QL, and additionally with role functioning for the UK and US value sets, and moderately correlated (0.4 ≤|p|< 0.6) with emotional and social functioning and fatigue. The observed Spearman correlations (Online Resource 4) between EQ-5D-5L utilities and QLQ-C30 scales had the expected directions and were moderate to strong for several domains, supporting their inclusion as predictors. The maximum variance inflation factor (role functioning) was 3.15 and the mean was 1.86, indicating no problematic multicollinearity.
Model performance and model selection
Table 3 summarizes out-of-fold performance for all model specifications with the full covariate set, with the ARV of the highest ranking model specification for each value set highlighted in bold. Including age alone as an additional covariate marginally reduced performance for most model types, whereas including age, number of comorbidities and the surgery-only indicator modestly improved all models except the beta model with the UK value set. We therefore focus on model specifications with the full covariate set in Table 3, although ranking (ARV) was performed across all specifications for each value set. Performance measures, including confidence intervals, for all model specifications are shown in Online Resource 5.
Table 3.
Performance metrics. All models with the full set of covariates
| Model | MAE | RMSE | AE ≤ 0.05 (%) | CCC | ARV | Mean bias | Range |
|
|---|---|---|---|---|---|---|---|---|
| Norwegian value set | ||||||||
| OLS 3 | 0.0557 | 0.0832 | 59.8 | 0.7721 | 15.50 | − 0.0002 | 0.433–1.000 | 0.629 |
| Linear mixed model 3 | 0.0559 | 0.0836 | 60.3 | 0.7589 | 18.00 | − 0.0008 | 0.467–1.000 | 0.628 |
| Tobit 3 | 0.0564 | 0.0820 | 62.3 | 0.7809 | 11.00 | − 0.0014 | 0.393–0.987 | 0.639 |
| Beta 3 | 0.0580 | 0.0858 | 62.3 | 0.7903 | 16.75 | 0.0008 | 0.060–0.975 | 0.627 |
| Fractional logit 3 | 0.0568 | 0.0826 | 60.2 | 0.7872 | 14.75 | 0.0002 | 0.101–0.966 | 0.637 |
| TPM OLS 3 | 0.0566 | 0.0842 | 59.4 | 0.7687 | 23.00 | 0.0002 | 0.441–1.000 | 0.620 |
| TPM Mixed model 3 | 0.0556 | 0.0841 | 59.8 | 0.7648 | 18.75 | 0.0064 | 0.467–1.000 | 0.623 |
| TPM Beta 3 | 0.0550 | 0.0810 | 62.2 | 0.7795 | 7.75 | − 0.0016 | 0.263–0.995 | 0.649 |
| TPM Fractional logit 3 | 0.0545 | 0.0806 | 62.7 | 0.7897 | 2.50 | 0.0003 | 0.210–0.995 | 0.651 |
| ALDVMM 3 | 0.0556 | 0.0834 | 63.0 | 0.7507 | 13.25 | 0.0052 | 0.474–0.987 | 0.637 |
| UK value set | ||||||||
| OLS 3 | 0.0571 | 0.0929 | 61.8 | 0.7221 | 19.75 | − 0.0007 | 0.454–1.000 | 0.565 |
| Linear mixed model 3 | 0.0568 | 0.0931 | 62.8 | 0.7098 | 20.75 | − 0.0011 | 0.483–1.000 | 0.565 |
| Tobit 3 | 0.0577 | 0.0910 | 64.9 | 0.7345 | 16.50 | − 0.0047 | 0.408–0.987 | 0.584 |
| Beta 3 | 0.0561 | 0.0901 | 68.2 | 0.7659 | 7.50 | − 0.0007 | 0.066–0.977 | 0.599 |
| Fractional logit 3 | 0.0555 | 0.0888 | 67.8 | 0.7606 | 6.00 | 0.0003 | 0.084–0.972 | 0.604 |
| TPM OLS 3 | 0.0578 | 0.0936 | 61.7 | 0.7209 | 24.00 | 0.0002 | 0.456–1.000 | 0.559 |
| TPM Mixed model 3 | 0.0565 | 0.0933 | 62.6 | 0.7153 | 20.25 | 0.0046 | 0.483–1.000 | 0.563 |
| TPM Beta 3 | 0.0547 | 0.0890 | 66.2 | 0.7370 | 8.50 | − 0.0024 | 0.271–0.995 | 0.605 |
| TPM Fractional logit 3 | 0.0539 | 0.0881 | 67.1 | 0.7559 | 5.25 | 0.0004 | 0.191–0.996 | 0.609 |
| ALDVMM 3 | 0.0550 | 0.0944 | 68.2 | 0.6715 | 15.75 | 0.0084 | 0.531–0.989 | 0.580 |
| US value set | ||||||||
| OLS 3 | 0.0666 | 0.0989 | 52.9 | 0.7875 | 14.00 | − 0.0003 | 0.300–1.000 | 0.649 |
| Linear mixed model 3 | 0.0669 | 0.0993 | 53.0 | 0.7757 | 18.00 | − 0.0012 | 0.340–1.000 | 0.649 |
| Tobit 3 | 0.0677 | 0.0977 | 54.3 | 0.7941 | 13.00 | − 0.0020 | 0.252–0.984 | 0.658 |
| Beta 3 | 0.0704 | 0.1028 | 53.0 | 0.7985 | 20.50 | 0.0010 | − 0.105–0.970 | 0.640 |
| Fractional logit 3 | 0.0693 | 0.0988 | 49.6 | 0.7983 | 17.75 | 0.0001 | − 0.064–0.961 | 0.652 |
| TPM OLS 3 | 0.0672 | 0.0998 | 53.0 | 0.7853 | 18.50 | 0.0002 | 0.308–1.000 | 0.643 |
| TPM Mixed model 3 | 0.0663 | 0.0997 | 53.3 | 0.7810 | 15.00 | 0.0073 | 0.340–1.000 | 0.645 |
| TPM Beta 3 | 0.0665 | 0.0962 | 54.1 | 0.7964 | 7.00 | − 0.0022 | 0.095–0.994 | 0.669 |
| TPM Fractional logit 3 | 0.0660 | 0.0960 | 54.8 | 0.8034 | 2.25 | 0.0002 | 0.049–0.994 | 0.669 |
| ALDVMM 3 | 0.0668 | 0.0992 | 56.0 | 0.7692 | 13.50 | 0.0058 | 0.346–0.987 | 0.655 |
MAE = mean absolute error, RMSE = root mean squared error, AE = absolute error, CCC = Lin’s concordance correlation, ARV = average ranking value
Across value sets, the two-part fractional logistic model with the full covariate set (TPM Fractional logit 3) had the most favorable ARV and was selected as the preferred model. Pairwise comparisons of the selected model and each alternative model show that differences between the top performing models were small in absolute terms, although confidence intervals for several comparisons excluded zero (Online Resource 6). The selected model performed at least as well as all alternatives on most measures. The same table reports a version of the selected model with quadratic terms for physical functioning and pain; improvements were small and inconsistent across value sets, with the largest difference in MAE being 0.0008.
A visual assessment of fit for the preferred model (TPM Fractional 3) is provided by Fig. 1. For each value set, the figure shows density plots of observed and predicted utilities, calibration plots of observed versus predicted utilities (with a fitted calibration line and a dashed 45° reference line for perfect prediction), and Bland–Altman plots. The predicted utility distribution closely resembled the observed distributions across the range of values, although the model did not reproduce values of exactly 1.0. Calibration slopes ranged from 0.99 to 1.01 across value sets, with confidence intervals including 1 and intercepts including zero, indicating close average agreement between predicted and observed utilities. The Bland–Altman plots indicate systematic bias, with overestimation at lower observed utilities and slight underestimation near 1.0, a pattern commonly reported in mapping studies [23] and largely attributable to regression to the mean in combination with the bounded and skewed distribution of EQ-5D-5L utilities. The calibration slopes indicate that this bias is largely confined to the tails of the distribution and does not reflect systematic shrinkage across the full range.
Fig. 1.

Predictive performance of TPM Fractional logistic regression 3. From top to bottom: Density plots of observed and predicted utilities, calibration plots of observed versus predicted utilities (with a fitted calibration line and a dashed 45° reference line for perfect prediction) and Bland–Altman plot
Predictive performance deteriorated at low levels of observed utility, while performance was broadly consistent across the three cancer diagnoses and between patients receiving surgery alone and combination therapy (Table 4).
Table 4.
Performance metrics of TPM Fractional logistic regression 3, Norwegian value set, by subgroup
| Subgroup | MAE | RMSE | AE ≤ 0.05 (%) | CCC | Mean bias | N |
|---|---|---|---|---|---|---|
| Observed utility | ||||||
| < 0.60 | 0.1908 | 0.2200 | 9.5 | 0.1726 | 148 (84 patients) | |
| 0.60–0.79 | 0.0750 | 0.0967 | 41.2 | 0.0275 | 352 (195 patients) | |
| 0.80–0.99 | 0.0435 | 0.0584 | 67.4 | − 0.0094 | 1,583 (546 patients) | |
| = 1.00 | 0.0362 | 0.0473 | 77.6 | − 0.0362 | 541 (220 patients) | |
| Cancer type | ||||||
| Cervix | 0.0577 | 0.0865 | 59.7 | 0.7675 | 0.0076 | 429 (105 patients) |
| Uterus | 0.0544 | 0.0810 | 63.0 | 0.7903 | − 0.0003 | 1,544 (376 patients) |
| Ovarii/perineum | 0.0535 | 0.0766 | 63.3 | 0.8029 | − 0.0017 | 608 (170 patients) |
| Treatment group | ||||||
| Surgery alone | 0.0535 | 0.0799 | 64.1 | 0.7958 | 0.0002 | 1,494 (365 patients) |
| Combination therapy | 0.0560 | 0.0819 | 60.6 | 0.7802 | 0.0005 | 1,072 (282 patients) |
MAE = mean absolute error, RMSE = root mean squared error, AE = absolute error, CCC = Lin’s concordance correlation. CCC is not reported by utility band, where variation in observed utility is restricted by construction
To assess prediction accuracy across the distribution of observed QL, Fig. 2 plots the mean and the 5th and 95th percentiles of observed and predicted EQ-5D-5L utilities, as well as the number of observations, conditional on QL. For QL scores above 30, mean predicted utilities closely follow mean observed utilities. Across all QL levels, predicted utilities show less dispersion around the mean than observed utilities, and the spread decreases with increasing QL and number of observations for both observed and predicted values.
Fig. 2.

Observed and predicted EQ-5D-5L utilities conditional on QL
Model coefficients
Coefficients and standard errors for the preferred two-part fractional logistic model, refitted on the full sample, are shown in Table 5, and a description of how to incorporate the mapping in practice in Online Resource 3. The variance–covariance matrices for the preferred model are provided in Online Resource 7, to enable probabilistic sensitivity analysis in economic evaluations. In-sample RMSE from the final model estimated on the full dataset was 0.0789 for the Norwegian value set, 0.0857 for the UK value set and 0.0941 for the US value set. Cross-validated RMSE is reported in Table 3.
Table 5.
Model coefficients of preferred model, refitted on the whole sample
| Variable | P(Perfect health) | Utility conditional on imperfect health | ||
|---|---|---|---|---|
| NO value set | UK value set | US value set | ||
| Physical functioning | 0.062030*** | 0.014922*** | 0.017595*** | 0.019390*** |
| (0.011168) | (0.001398) | (0.001634) | (0.001384) | |
| Role functioning | 0.003689 | − 0.000138 | 0.001425 | 0.001902** |
| (0.006398) | (0.000884) | (0.000997) | (0.000848) | |
| Emotional functioning | 0.057609*** | 0.011671*** | 0.008882*** | 0.006655*** |
| (0.007988) | (0.001024) | (0.001169) | (0.000964) | |
| Cognitive functioning | 0.009854 | − 0.000236 | − 0.000235 | 0.000031 |
| (0.006518) | (0.000930) | (0.001141) | (0.000934) | |
| Social functioning | − 0.002711 | 0.001324 | 0.000853 | 0.000888 |
| (0.006520) | (0.000956) | (0.001169) | (0.000961) | |
| Fatigue | 0.014563** | 0.002670*** | 0.002503** | 0.002668*** |
| (0.006120) | (0.000978) | (0.001166) | (0.000986) | |
| Nausea and vomiting | − 0.000458 | 0.000779 | 0.000203 | 0.000941 |
| (0.014647) | (0.001317) | (0.001556) | (0.001262) | |
| Pain | − 0.084610*** | − 0.006886*** | − 0.006513*** | − 0.005841*** |
| (0.008411) | (0.000818) | (0.001021) | (0.000806) | |
| Dyspnoea | − 0.000539 | 0.000457 | 0.000681 | 0.000330 |
| (0.005096) | (0.000718) | (0.000861) | (0.000710) | |
| Insomnia | − 0.004537 | − 0.001193** | − 0.001510** | − 0.000984 |
| (0.003468) | (0.000593) | (0.000726) | (0.000603) | |
| Appetite loss | 0.014152** | 0.000125 | 0.000291 | 0.000025 |
| (0.006582) | (0.000800) | (0.000913) | (0.000761) | |
| Constipation | − 0.002785 | − 0.000178 | − 0.000263 | − 0.000409 |
| (0.004355) | (0.000624) | (0.000727) | (0.000609) | |
| Diarrhoea | − 0.003279 | 0.001128* | 0.000884 | 0.000871 |
| (0.003578) | (0.000667) | (0.000809) | (0.000655) | |
| Financial difficulties | 0.002590 | − 0.002152*** | − 0.002876*** | − 0.002344*** |
| (0.007038) | (0.000814) | (0.001010) | (0.000755) | |
| Global health status/QoL | 0.020858*** | 0.004372*** | 0.004931*** | 0.004681*** |
| (0.007435) | (0.001211) | (0.001365) | (0.001147) | |
| Age | − 0.001277 | 0.001669 | 0.001396 | 0.001365 |
| (0.007437) | (0.001554) | (0.001824) | (0.001627) | |
| Number of comorbidities | − 0.220475** | − 0.045101*** | − 0.045231*** | − 0.049547*** |
| (0.089796) | (0.012892) | (0.014277) | (0.012976) | |
| Surgical treatment only | 0.357811* | − 0.007283 | 0.001608 | 0.046662 |
| (0.186953) | (0.038525) | (0.045306) | (0.038329) | |
| Constant | − 13.91878*** | − 0.118958 | 0.020725 | − 0.373779* |
| (1.623104) | (0.216428) | (0.272235) | (0.217580) | |
| Log-likelihood | − 746.644665 | − 689.237692 | − 613.036300 | − 736.104094 |
| Akaike information criteria | 1531.289331 | 1416.475384 | 1264.072599 | 1510.208188 |
| Bayesian information criteria | 1642.865979 | 1523.665108 | 1371.262324 | 1617.397912 |
Pseudo
|
0.440801 | 0.061544 | 0.068890 | 0.069882 |
| N | 2,624 | 2,083 | 2,083 | 2,083 |
Significance: *p < 0.10, **p < 0.05, ***p < 0.01
As the utility of perfect health is always 1, the logistic regression for P(Perfect health) is identical across value sets, whereas coefficients from the fractional logistic regressions for utilities <1 differ between value sets because of differences in preference weights. Physical functioning, emotional functioning, fatigue, pain, QL and number of comorbidities were statistically significant predictors (p < 0.05) in both the logistic part and the conditional utility regression for all three value sets. Financial difficulties were statistically significant in the regression conditional on imperfect health. Fatigue showed unexpected positive coefficients in both the logistic component predicting perfect health and the fractional logistic regression conditional on imperfect health, while appetite loss showed an unexpected positive coefficient only in the logistic component predicting perfect health. Otherwise, significant coefficients had expected signs: positive for functional scales and QL, and negative for symptom scales and comorbidities.
Coefficient estimates were essentially unchanged under multiple imputation, with all differences well within one standard error (Online Resource 2).
Discussion
We developed mapping algorithms to transform QLQ-C30 scores into EQ-5D-5L utility values for gynecological cancer survivors in follow-up care, enabling health economic evaluations in studies where EQ-5D-5L is unavailable. Across Norwegian, UK and US value sets, the two-part fractional logistic model with the full set of covariates (TPM Fractional 3) was preferred as it achieved best performance on most model performance measures and had the best ARV. Performance deteriorated at low levels of observed utility, but overall bias remained small and calibration plots indicated good average agreement between observed and predicted utilities across most of the range. Stratified five-fold cross-validation ensured that the distribution of EQ-5D-5L utilities and the presence of perfect health were balanced across folds, supporting stable performance estimates with a skewed utility distribution.
Several published mapping studies use data-driven model selection procedures, such as stepwise removal of non-significant variables [24, 25], exclusion of covariates with “unexpected” signs [25], or selective inclusion of polynomial [24] and interaction terms [26]. Such approaches can increase the risk of overfitting and produce unstable models that do not generalize well, and excluding theoretically relevant covariates on the basis of statistical significance alone may lead to misspecification [27, 28]. In line with ISPOR good-practice recommendations for mapping [7], we specified the full set of QLQ-C30 scale scores as predictors a priori and did not use data-driven model selection procedures such as stepwise regression or exclusion based solely on statistical significance or coefficient sign.
Consistent with this approach, fatigue was retained in the final algorithm despite showing unexpected positive coefficients in both model components. The unexpected sign for fatigue likely reflects the difficulty of isolating its independent contribution in the presence of several correlated functional scales. Once physical functioning, role functioning, social functioning and QL are controlled for, the residual variance uniquely attributable to fatigue may be limited and unstable, resulting in an unexpected coefficient sign.
We assessed whether the assumption of linearity restricted performance by adding quadratic terms for physical functioning and pain to the selected model. Improvements were inconsistent across value sets, favoring the addition of quadratic terms in the UK tariff but not the US tariff. That the direction of benefit varied across tariffs applied to identical response data suggests these differences do not reflect a stable departure from linearity, and the pre-specified linear specification was retained.
To our knowledge, this is the first study to map QLQ-C30 to EQ-5D-5L utilities in a sample consisting exclusively of gynecological cancer patients. Direct comparisons with other studies are limited by differences in cancer populations, value sets, and validation strategies; we therefore focused on studies with broadly similar EQ-5D-5L utility levels and direct mapping from QLQ-C30 to EQ-5D-5L.
Gebert et al. [29] developed direct and response-mapping algorithms in metastatic breast cancer using the German value set (mean utility 0.812, SD 0.212) and reported for their preferred adjusted beta regression a MAE of 0.081, RMSE of 0.130 and CCC of 0.781 in the validation dataset. Their preferred model, which included all QLQ-C30 scales, is comparable in structure to our Beta1 specification; in our data, beta regression performed well on CCC but was outperformed by two-part fractional logistic models in overall ARV ranking. Hagiwara et al. [25] mapped QLQ-C30 and FACT-G to EQ-5D-5L in a mixed cancer cohort from Japan (median utility 0.823), reporting a two-part beta regression with exclusion of coefficients with unexpected signs as their preferred direct-mapping model, with MAE 0.077 and RMSE 0.101 in nine-fold cross-validation. In our study, two-part beta models also performed well but were consistently outperformed by the chosen two-part fractional logistic specification across value sets. Huang et al. [30] mapped QLQ-C30 and QLQ H&N35 to EQ-5D-5L and SF-6D in papillary thyroid cancer (mean utility 0.87, SD 0.0939) using the Chinese value set; their preferred beta mixture model for EQ-5D-5L, which included all QLQ-C30 and H&N35 scales plus age and gender, yielded MAE 0.0391, RMSE 0.0553 and AE ≤ 0.05 for 73.14% of predictions.
The set of significant predictors in our preferred mapping algorithms is consistent with those of Hagiwara et al. [25] and Huang et al. [30]. Physical functioning, emotional functioning, pain and QL appear as significant predictors with the expected signs. Financial difficulties was a significant predictor in our study and in Huang et al., as was appetite loss—notably with an unexpected sign in both the P(Perfect health) component in our study and in Huang et al. Predictors that were significant in the comparison studies but not in the present study included nausea and vomiting, insomnia and age (Huang et al.) and role functioning (Hagiwara et al.). Gebert et al. [29] does not include a table of coefficients for their direct mapping algorithm, but they discuss the existence of coefficients with counterintuitive signs, including fatigue.
Overall, our results are consistent with the broader mapping literature in showing that two-part- and flexible non-linear models typically outperform simple OLS [29] and that prediction is least accurate at the extremes of the utility distribution, with systematic overprediction at low utilities and underprediction near full health [23].
This study benefits from a sample that is broadly representative of gynecological cancer survivors in Norway. We tested several regression models, with different features for handling the skewed and bounded nature of EQ-5D-5L utilities, across three value sets and assessed performance based on a range of performance measures. Although performance varied somewhat between value sets, the preferred algorithm earned the highest ranking for all value sets, suggesting that the choice of a preferred model is robust across value sets. Rather than relying on rankings alone, we estimated pairwise differences in performance with confidence intervals. The selected model performed at least as well as all alternatives on most measures of model performance, though differences between the leading specifications were small in absolute terms.
Limitations
In the primary analysis, we handled missing data using a complete-case approach. Multiple imputation of missing domains and scales for incomplete responses did not materially change the estimated coefficients. Of completely missing questionnaires, 77% were from patients who completed the follow-up as planned or were still in active follow-up at the time of data extraction. These findings suggest that any bias from complete-case exclusion and missing assessments is likely to be limited, although we cannot fully exclude bias if missingness is related to unobserved aspects of health status.
Our analyses are based on direct mapping, linking QLQ-C30 responses directly to EQ-5D-5L utility indices. A key limitation of direct mapping is that predictions are tied to specific value sets, whereas many health technology assessment agencies recommend using country-specific value sets for cost-effectiveness analyses [19]. To enhance applicability, we developed algorithms using Norwegian, UK and US value sets as targets.
Mapping algorithms provide a practical solution when preference-based measures have not been collected, but they should generally be considered a second-best option to direct measurement of utilities [7]. Although we consider the LETSGO sample to be representative of gynecological cancer survivors in Norway, until validated on an independent sample, application of these algorithms to other settings should be accompanied by sensitivity analyses exploring alternative assumptions about utilities. Further research should be conducted to validate our proposed algorithm on an external sample.
Conclusions
This study developed algorithms to transform QLQ-C30 scale scores into EQ-5D-5L utility values for patients in follow-up care after gynecological cancer treatment. The preferred two-part fractional logistic model showed good predictive accuracy across three value sets and may be used for mapping when preference-based utility measures have not been collected, particularly in survivorship cohorts and provided that the target population has a distribution of QL similar to that of the LETSGO sample.
Supplementary Information
Below is the link to the electronic supplementary material.
Acknowledgements
The work was supported by Sørlandet Hospital Trust.
Abbreviations
- AE
Absolute error
- ALDVMM
Adjusted limited dependent variable mixture model
- ARV
Average ranking value
- CCC
Lin’s concordance correlation coefficient
- EQ-5D
EuroQol five-dimension
- EQ-5D-5L
EuroQol five-dimension 5-level
- HRQoL
Health-related quality of life
- HSUVs
Health-state utility values
- LETSGO
Lifestyle and Empowerment Techniques in Survivorship of Gynaecologic Oncology
- LMM
Linear mixed model
- MAE
Mean absolute error
- OLS
Ordinary least squares
- QALYs
Quality-adjusted life years
- QL
Global health status and quality of life scale
- QLQ-C30
European Organisation for Research and Treatment of Cancer Quality of Life Questionnaire Core 30
- RMSE
Root mean squared error
- TPM
Two-part model
Author contributions
All authors contributed to the study conception and design. Material preparation and analysis were performed by T.Å. The first draft of the manuscript was written by T.Å. and all authors commented on previous versions of the manuscript. All authors read and approved the final manuscript.
Funding
Open access funding provided by University of Oslo (incl Oslo University Hospital). The work was supported by Sørlandet Hospital Trust.
Data availability
The data underlying this study contain sensitive health information and cannot be made publicly available. Data may be made available upon reasonable request, subject to ethical approval and data protection regulations.
Declarations
Conflict of interest
The authors declare no conflict of interest.
Ethical approval
The study was approved by the Regional Committee of Medical Research Ethics in Norway (2019/11093).
Consent to participate
Written informed consent was obtained from all participants.
Consent to publish
Not applicable.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Bray, F., Laversanne, M., Sung, H., Ferlay, J., Siegel, R. L., Soerjomataram, I., & Jemal, A. (2024). Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA: A Cancer Journal for Clinicians,74(3), 229–263. 10.3322/caac.21834 [DOI] [PubMed] [Google Scholar]
- 2.Zhu, B., Gu, H., Mao, Z., Beeraka, N. M., Zhao, X., Anand, M. P., Zheng, Y., Zhao, R., Li, S., Manogaran, P., Fan, R., Nikolenko, V. N., Wen, H., Basappa, B., & Liu, J. (2024). Global burden of gynaecological cancers in 2022 and projections to 2050. Journal of Global Health,14, 04155. 10.7189/jogh.14.04155 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Tichanek, F., Försti, A., Hemminki, O., Hemminki, A., Hemminki, K., & Hernandez, E. (2023). Survival, incidence, and mortality trends in female cancers in the Nordic countries. Obstetrics and Gynecology International,2023, 1–9. 10.1155/2023/6909414 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Zhou, X., Yang, D., Zou, Y., Tang, D., Chen, J., Li, Z., Shen, Q., Xu, Q., & Xiang, Y. (2024). Long-term survival trend of gynecological cancer: A systematic review of population-based cancer registration data. Biomedical and Environmental Sciences,37(8), 897–921. 10.3967/bes2024.133 [DOI] [PubMed] [Google Scholar]
- 5.Rowen, D., Azzabi Zouraq, I., Chevrou-Severac, H., & van Hout, B. (2017). International regulations and recommendations for utility data for health technology assessment. PharmacoEconomics,35(Suppl 1), 11–19. 10.1007/s40273-017-0544-y [DOI] [PubMed] [Google Scholar]
- 6.Aaronson, N. K., Ahmedzai, S., Bergman, B., Bullinger, M., Cull, A., Duez, N. J., Filiberti, A., Flechtner, H., Fleishman, S. B., d’Haes, J. C. J. M., Kaasa, S., Klee, M., Osoba, D., Razavi, D., Rofe, P. B., Schraub, S., Sneeuw, K., Sullivan, M., & Takeda, F. (1993). The European Organization for Research and Treatment of Cancer QLQ-C30: A quality-of-life instrument for use in international clinical trials in oncology. Journal of the National Cancer Institute,85(5), 365–376. 10.1093/jnci/85.5.365 [DOI] [PubMed] [Google Scholar]
- 7.Wailoo, A. J., Hernandez-Alava, M., Manca, A., Mejia, A., Ray, J., Crawford, B., Botteman, M., & Busschbach, J. (2017). Mapping to estimate health-state utility from non–preference-based outcome measures: An ISPOR good practices for outcomes research task force report. Value in Health,20(1), 18–27. 10.1016/j.jval.2016.11.006 [DOI] [PubMed] [Google Scholar]
- 8.Petrou, S., Rivero-Arias, O., Dakin, H., Longworth, L., Oppe, M., Froud, R., & Gray, A. (2015). The MAPS reporting statement for studies mapping onto generic preference-based outcome measures: Explanation and elaboration. PharmacoEconomics,33(10), 993–1011. 10.1007/s40273-015-0312-9 [DOI] [PubMed] [Google Scholar]
- 9.Dakin, H., Koleva-Kolarova, R., Mallon, K., Burns, R., Yang, Y., & Abel, L. (2025). HERC database of mapping studies, Version 10.0. University of Oxford. [Google Scholar]
- 10.Longworth, L. B. A. M. P., & Rowen, D. B. A. M. P. (2013). Mapping to obtain EQ-5D utility values for use in NICE health technology assessments. Value in Health,16(1), 202–210. 10.1016/j.jval.2012.10.010 [DOI] [PubMed] [Google Scholar]
- 11.Devlin, N., Pickard, S., & Busschbach, J. (2022). The development of the EQ-5D-5L and its value sets. In N. Devlin, B. Roudijk, & K. Ludwig (Eds.), Value sets for EQ-5D-5L: A compendium, comparative review & user guide (pp. 1–12). Springer International Publishing. [PubMed] [Google Scholar]
- 12.Vistad, I., Skorstad, M., Demmelmaier, I., Småstuen, M. C., Lindemann, K., Wisløff, T., van de Poll-Franse, L. V., & Berntsen, S. (2021). Lifestyle and empowerment techniques in survivorship of gynaecologic oncology (LETSGO study): A study protocol for a multicentre longitudinal interventional study using mobile health technology and biobanking. British Medical Journal Open,11(7), e050930–e050930. 10.1136/bmjopen-2021-050930 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Sangha, O., Stucki, G., Liang, M. H., Fossel, A. H., & Katz, J. N. (2003). The self-administered comorbidity questionnaire: A new method to assess comorbidity for clinical and health services research. Arthritis & Rheumatism,49(2), 156–163. 10.1002/art.10993 [DOI] [PubMed] [Google Scholar]
- 14.Vistad, I., Skorstad, M., Berntsen, S., Demmelmaier, I., Bjorge, L., van de Poll-franse, L. V., Lindemann, K., Nilsen, E. B., Haug, A., Bentzen, A. G., Bohlin, T., Stokstad, T., Brummer, T., Liavaag, A., Lomsdal, S., Holstad, A. D., & Hagen, M. C. (2025). Empowerment and quality of life in gynecological cancer survivors: Outcomes from a multicenter quasi-experimental cohort study from Norway (the LETSGO trial). Cancer. 10.1002/cncr.70040 [DOI] [PubMed] [Google Scholar]
- 15.Garratt, A. M., Stavem, K., Shaw, J. W., & Rand, K. (2025). EQ-5D-5L value set for Norway: A hybrid model using cTTO and DCE data. Quality of Life Research,34(2), 417–427. 10.1007/s11136-024-03837-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Rowen, D., Mukuria, C., Bray, N., Carlton, J., Longworth, L., Meads, D., Oluboyede, Y., O’Neill, C., & Yang, Y. (2026). A United Kingdom value set for the EQ-5D-5L. Value Health,29(5), 858–869. 10.1016/j.jval.2026.03.008 [DOI] [PubMed] [Google Scholar]
- 17.Pickard, A. S., Law, E. H., Jiang, R., Pullenayegum, E., Shaw, J. W., Xie, F., Oppe, M., Boye, K. S., Chapman, R. H., Gong, C. L., Balch, A., & Busschbach, J. J. V. (2019). United States Valuation of EQ-5D-5L health states using an international protocol. Value Health,22(8), 931–941. 10.1016/j.jval.2019.02.009 [DOI] [PubMed] [Google Scholar]
- 18.Fayers, P., Aaronson, N. K., Bjordal, K., Grønvold, M., Curran, D., & Bottomley, A. (2001). EORTC QLQ-C30 scoring manual: European Organisation for research and treatment of cancer.
- 19.Mukuria, C., Rowen, D., Harnan, S., Rawdin, A., Wong, R., Ara, R., & Brazier, J. (2019). An updated systematic review of studies mapping (or cross-walking) measures of health-related quality of life to generic preference-based measures to generate utility values. Applied Health Economics and Health Policy,17(3), 295–313. 10.1007/s40258-019-00467-6 [DOI] [PubMed] [Google Scholar]
- 20.Hernández Alava, M., Wailoo, A. J., & Ara, R. (2012). Tails from the Peak District: Adjusted limited dependent variable mixture models of EQ-5D questionnaire health state utility values. Value in Health,15(3), 550–561. 10.1016/j.jval.2011.12.014 [DOI] [PubMed] [Google Scholar]
- 21.Smithson, M., & Verkuilen, J. (2006). A better lemon squeezer? Maximum-likelihood regression with beta-distributed dependent variables. Psychological Methods,11(1), 54–71. 10.1037/1082-989X.11.1.54 [DOI] [PubMed] [Google Scholar]
- 22.Nout, R. A., van de Poll-Franse, L. V., Lybeert, M. L. M., Wárlám-Rodenhuis, C. C., Jobsen, J. J., Mens, J. W. M., Lutgens, L. C. H. W., Pras, B., van Putten, W. L. J., & Creutzberg, C. L. (2011). Long-term outcome and quality of life of patients with endometrial carcinoma treated with or without pelvic radiotherapy in the Post Operative Radiation Therapy in Endometrial Carcinoma 1 (PORTEC-1) trial. Journal of Clinical Oncology,29(13), 1692–1700. 10.1200/JCO.2010.32.4590 [DOI] [PubMed] [Google Scholar]
- 23.Brazier, J. E., Yang, Y., Tsuchiya, A., & Rowen, D. L. (2010). A review of studies mapping (or cross walking) non-preference based measures of health to generic preference-based measures. European Journal of Health Economics,11(2), 215–225. 10.1007/s10198-009-0168-z [DOI] [PubMed] [Google Scholar]
- 24.Zhou, H., Liu, X., Bao, R., Qiu, L., Zhang, Y., Gu, Q., & Yang, Q. (2025). Development of mapping algorithms for gastric cancer: Translating EORTC QLQ-C30 and QLQ-STO22 to EQ-5D-5L health utilities. Quality of Life Research. 10.1007/s11136-025-04063-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Hagiwara, Y., Shiroiwa, T., Taira, N., Kawahara, T., Konomura, K., Noto, S., Fukuda, T., & Shimozuma, K. (2020). Mapping EORTC QLQ-C30 and FACT-G onto EQ-5D-5L index for patients with cancer. Health and Quality of Life Outcomes,18(1), 354. 10.1186/s12955-020-01611-w [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Ameri, H., Yousefi, M., Yaseri, M., Nahvijou, A., Arab, M., & Akbari Sari, A. (2020). Mapping EORTC-QLQ-C30 and QLQ-CR29 onto EQ-5D-5L in colorectal cancer patients. Journal of Gastrointestinal Cancer,51(1), 196–203. 10.1007/s12029-019-00229-6 [DOI] [PubMed] [Google Scholar]
- 27.Steyerberg, E. W., Eijkemans, M. J. C., Harrell, F. E., & Habbema, J. D. F. (2001). Prognostic modeling with logistic regression analysis: In search of a sensible strategy in small data sets. Medical Decision Making,21(1), 45–56. 10.1177/0272989X0102100106 [DOI] [PubMed] [Google Scholar]
- 28.Babyak, M. A. (2004). What you see may not be what you get: A brief, nontechnical introduction to overfitting in regression-type models. Psychosomatic Medicine,66(3), 411–421. 10.1097/01.psy.0000127692.23278.a9 [DOI] [PubMed] [Google Scholar]
- 29.Gebert, P., Hage, A. M., Fischer, F., Klapproth, C. P., Grittner, U., & Karsten, M. M. (2025). Translating the EORTC CAT core and the QLQ-C30 to the EQ-5D-5L in patients with metastatic breast cancer: A comparison of direct and indirect mapping algorithms. European Journal of Health Economics. 10.1007/s10198-025-01824-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Huang, D., Zeng, D., Tang, Y., Jiang, L., & Yang, Q. (2024). Mapping the EORTC QLQ-C30 and QLQ H&N35 to the EQ-5D-5L and SF-6D for papillary thyroid carcinoma. Quality of Life Research,33(2), 491–505. 10.1007/s11136-023-03540-9 [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The data underlying this study contain sensitive health information and cannot be made publicly available. Data may be made available upon reasonable request, subject to ethical approval and data protection regulations.





