Skip to main content
Journal of the American Heart Association: Cardiovascular and Cerebrovascular Disease logoLink to Journal of the American Heart Association: Cardiovascular and Cerebrovascular Disease
. 2025 Jul 30;14(16):e040386. doi: 10.1161/JAHA.124.040386

External Validation of a Risk Prediction Model for Atherosclerotic Cardiovascular Diseases in a Large National Health‐Checkup and Claim Database

Takanori Honda 1,2,, Yoshihiko Furuta 2,3, Akihiro Maezono 2,4, Sanmei Chen 2,5, Yuki Ishida 2,6, Hiroko Furuhashi 2, Masaya Kumamoto 2,3,7, Emi Oishi 2,3, Yasumi Kimura 2,8, Daigo Yoshida 2,9, Toshiharu Ninomiya 1,2
PMCID: PMC12533744  PMID: 40736084

Abstract

Background

The Hisayama risk prediction model for atherosclerotic cardiovascular diseases (ASCVDs) has been featured in the latest Japanese preventive guidelines, yet it lacks external validation.

Methods

A retrospective cohort study of 420 552 ASCVD‐free individuals aged 40–74 years was performed using data from the national health checkup and administrative claim database recorded between April 2015 and November 2020. Incident cases of ASCVD were ascertained from International Classification of Diseases, Tenth Revision (ICD‐10) codes and accompanying medical practice codes.

Results

During a median follow‐up of 4.4 years, 3998 individuals developed ASCVD. The Hisayama ASCVD model had good predictive performance in the setting of the Japanese national health checkup. The original model performed well for estimating relative risks (hazard ratio, 32.73 [95% CI, 24.59–43.55] in the highest versus lowest decile) with a satisfactory discriminative ability (Uno's C‐statistic, 0.759 [95% CI, 0.751–0.766]), while it tended to overestimate the predicted probability of ASCVD for the study population (calibration‐in‐the‐large, 0.42 [95% CI, 0.41–0.44]). Improved performance was observed after recalibration by substituting the baseline survival rate (calibration‐in‐the‐large, 1.00 [95% CI, 0.97–1.04]) and by refitting the model (calibration‐in‐the‐large, 1.00 [95% CI, 0.97–1.03]; Uno's C‐statistic, 0.766 [95% CI, 0.759–0.774]). The predictive performance was good across subgroups of age, sex, and type of insurance.

Conclusions

This external validation study demonstrated that the Hisayama ASCVD model had good predictive performance in the national health checkup setting in Japan, suggesting the use of the Hisayama ASCVD model for efficient risk stratification and risk communication in the nationwide health checkups.

Keywords: clinical prediction model, coronary heart disease, retrospective study, stroke

Subject Categories: Epidemiology, Cardiovascular Disease, Lifestyle


Nonstandard Abbreviations and Acronyms

EPOCH‐JAPAN

Evidence for Cardiovascular Prevention From Observational Cohorts in Japan

JALS

Japan Arteriosclerosis Longitudinal Study

Research Perspective.

What Is New?

  • This study externally validated the Hisayama atherosclerotic cardiovascular disease risk prediction model using a large national health checkup and administrative claim database in Japan, confirming its strong predictive performance while identifying a tendency to overestimate absolute risk, which was improved through recalibration.

  • The results support the Hisayama atherosclerotic cardiovascular disease model's potential for nationwide implementation in health checkups to facilitate risk stratification and communication, enhancing atherosclerotic cardiovascular disease prevention efforts in Japan.

What Question Should Be Addressed Next?

  • To further validate and calibrate the absolute predicted risk of atherosclerotic cardiovascular disease, additional efforts are needed to accurately quantify baseline incidence rates and to include additional validation, including variations in age groups and geographic regions.

Atherosclerotic cardiovascular diseases (ASCVDs) are 1 of the leading causes of death worldwide, imposing a significant economic burden globally. 1 , 2 Several existing models have been incorporated into preventive guidelines and are widely used for facilitating risk stratification and lifestyle modification for ASCVD prevention. 3 , 4 However, most of the existing clinical prediction models have not been externally validated. 5 A systematic review of clinical prediction models for cardiovascular diseases showed that 64% (231 of 363 models) had not been externally validated, with particularly scarce validation of the models developed for Asians. 5 Clinical prediction models are expected to be validated in the applicable racial and ethnic groups, since disease presentation, relevant risk factors and their contributions, and therapeutic requirements and responses may differ. 6 In external validation studies of clinical prediction models, there remain concerns about small sample sizes (specifically, small numbers of events) and inadequate assessment of calibration and clinical utility in real‐world settings. 7

Recently, a new risk prediction tool for the primary prevention of ASCVD was incorporated in Japanese guidelines. 8 , 9 This tool was based on a multivariable clinical prediction model for predicting 10‐year risk of ASCVD for Japanese adults aged ≥40 years (the Hisayama ASCVD model) and was derived from a well‐characterized, population‐based prospective study of a general Japanese population in a community. 10 The Hisayama ASCVD model was developed with myocardial infarction and atherothrombotic stroke as end points, both of which have atherosclerosis as a common mechanistic basis. As such, the model is considered particularly useful for lipid management in clinical practice. 10 However, the Hisayama ASCVD model has not yet been validated in external populations.

In Japan, health checkups are mandatory as part of a nationwide screening program to identify individuals with visceral obesity and a high cardiovascular risk for the prevention of cardiometabolic diseases. 11 , 12 Annual health checkups are conducted by health insurers in a nationally standardized manner. Because all predictors included in the Hisayama ASCVD model are readily covered in the health checkups, real‐world nationwide validation of the Hisayama ASCVD model is possible. Such validation would facilitate nationwide implementation of the model and promote efficient and standardized screening of high‐risk individuals. In this study, therefore, we aimed to validate the Hisayama ASCVD risk prediction model using a large‐scale database consisting of the national health checkup and administrative claim data in Japan.

Methods

Data Availability

The database used in this study (DeSC database) is a commercial database and is not publicly accessible.

Data Sources and Study Population

This was a retrospective cohort study using the DeSC database, which consists of specific health checkup and administrative claims data. 13 In Japan, all adult residents aged 40 to 74 years are subject to a national screening program known as the specific health checkup, which is provided by their health insurers. The specific health checkup includes a series of examinations for the prevention of circulatory and metabolic diseases. The results of the examination are used for risk stratification to determine subsequent interventions, which is called the specific health guidance. Detailed descriptions of the health insurance systems and the nationwide screening and intervention program in Japan (ie, the specific health checkup and the specific health guidance) are provided in Data S1. The DeSC database includes data provided by health insurers of the Employees' Health Insurance system (Kenpo), the National Health Insurance program (Kokuho), and the Late‐Stage Medical Care System for the Elderly (Koki Koreisha Iryo Seido). The database does not contain information that could be used to identify an individual's specific insurers. The DeSC database has been shown to be representative of the general population of Japan on the basis of the distributions of chronic diseases and relevant risk factors. 13 The present analysis included individuals aged 40 to 74 years who were free from ASCVD at baseline and were registered in the DeSC database at least once during the period from April 2015 to November 2020. The study protocol was approved by the Kyushu University Institutional Review Board for Clinical Research (2022–2025). Informed consent from individuals was not required due to the retrospective and anonymized nature of the data. This study was reported in accordance with the Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis statement. 14

Definitions and Follow‐Up of ASCVD

The primary outcome was ASCVD, defined as either ischemic stroke or coronary heart disease (CHD). Incident cases of ASCVD were determined using information on International Classification of Diseases, Tenth Revision (ICD‐10) codes, drug prescriptions, emergency hospitalization, and surgery or rehabilitation. Because identifying coronary diseases and stroke cases solely by ICD‐10 codes could cause a substantially high proportion of false‐positive cases, 15 we used a combination of the ICD‐10 codes and information on medical procedures to extract more definitive cases of ASCVD. The specific definitions and codes used to extract information on diagnoses, prescriptions, surgery, and rehabilitation from the claims data are provided in Table S1. Three authors (T.H., Y.F., and A.M.) including a stroke physician (Y.F.) and a cardiologist (A.M.) discussed and determined the medical procedure specific to each of ischemic stroke and CHD. The end point was determined if an ICD‐10 code was accompanied by the prescription of relevant drug(s). Specifically, diagnosed cases of stroke (ICD‐10: I63) or CHD (ICD‐10: I21‐I25) with the corresponding prescriptions for each disease in the same month as the diagnosis or the following month were considered as an occurrence of the primary outcome. Furthermore, ASCVD hospitalization was defined as the relevant ICD‐10 code accompanied by a code for emergency hospitalization, and ASCVD surgery or rehabilitation as the relevant ICD‐10 code accompanied by a code for surgical treatment or rehabilitation, respectively. We also used total CVD, defined as either developed stroke (ICD‐10: I60‐64) or CHD, as an alternative outcome.

Participants were followed up until the date of ASCVD onset, the end of health insurance eligibility, or the last record in the database (November 2022), whichever came first. The onset date was defined as the day of diagnosis, as documented in the medical record. If an individual experienced >1 ischemic stroke or CHD event, the date of the first event was considered the ASCVD onset date. The last day of insurance coverage eligibility was set as the 15th of each month as the data were recorded monthly.

Measurement of Predictors

The predictors of the Hisayama ASCVD model include age, sex, systolic blood pressure, presence of diabetes, high‐density lipoprotein cholesterol, low‐density lipoprotein cholesterol, presence of proteinuria, current smoking, and absence of a regular exercise habit, as described in our previous report. 10 Physical examination and interview data from the national health checkup database were used to define each predictor. For individuals who had multiple health examination records, the earliest record was used as the baseline.

Measurements in the national health checkups are standardized nationwide; a standard protocol is publicly available. 16 Although the standard protocol was revised once during the study period (in 2018), no substantial changes were made with respect to the measurement of the predictors of interest. Definitions of the predictors were generally in accordance with the report in which the model was developed. 10 Blood pressure is supposed to be measured twice in the standard protocol. However, only 1 measurement was recorded in 77% of those who had blood pressure measurements. Therefore, only the first blood pressure record was used in this study. Examinees of the medical checkup are instructed to fast for at least 10 hours, and fasting or casual plasma glucose and serum high‐density lipoprotein and low‐density lipoprotein cholesterol levels were measured in accordance with standardized methods and procedures. Diabetes was defined as having a fasting plasma blood glucose level of ≥126 mg/dL, a casual plasma blood glucose level of ≥200 mg/dL, a hemoglobin A1c of ≥6.5%, or the use of diabetes medications, based on the American Diabetes Association's standard. 17 Although this definition differed slightly from the definition in the original development cohort, which did not consider hemoglobin A1c, 10 we adopted this definition of diabetes in consideration of the common use of hemoglobin A1c in clinical settings. Proteinuria was defined as 1+ or higher on a urine dipstick test. Smoking and exercise habits were assessed by the standard questionnaire. Respondents were asked if they currently smoked. For exercise habits, respondents were asked whether they had a regular exercise habit of at least 30 minutes twice a week for at least 1 year.

Statistical Analysis

Characteristics of the study participants are presented as mean±SD, frequency, or median with interquartile range. Comparisons were made by sex, age group (<60 or ≥60 years), and health insurer (employee health insurance or national health insurance).

Longitudinal analyses were conducted to evaluate performance measures of the clinical prediction model in a population after excluding individuals with missing values for any of the predictors. The proportion of missing values of predictors at baseline were computed to assess the extent to which individual items could be evaluated. We performed the available‐case analysis to examine the predictive ability under conditions where actual missing data occur at nationwide screening. We calculated the predicted probability of developing ASCVD on the basis of our previously published Hisayama ASCVD prediction model. 10 We used the regression coefficients (β) and the mean of predictor profiles (x¯) for the linear predictors as is in the Hisayama ASCVD prediction model. 10 We recalculated the incidence of ASCVD, defined as ischemic stroke or CHD, over the same period in the same cohort, and the 4‐year incidence rate was further calculated using the following formula: S0(t=4)=[S0(t=10)]4/10. The observed 10‐year probability of ASCVD S0(t=10) in the Hisayama cohort was 0.9527, and the resulting S0(t=4) was calculated to be 0.9808. The predicted probability based on the published prediction model (hereafter the original model) was calculated as P^ = 1–0.9808exp(∑β·x‐6.7963).

The Kaplan–Meier curve was drawn according to the deciles for predicted probability by the original Hisayama ASCVD prediction model. Participants were classified by deciles of predicted incidence probability, and Kaplan–Meier curves were plotted. The crude hazard ratio (HR) and 95% CI were obtained using a Cox proportional hazards model with the lowest decile as the reference group. Since the median follow‐up period for this study was 4.4 years, the predicted values were calculated with a 4‐year time horizon for the evaluation of model performance. We followed the practical guidance of McLernon et al. 18

We evaluated the predictive performance of the original Hisayama ASCVD model (hereinafter the original model). Furthermore, we constructed 2 recalibrated models: a substituted model, in which only the baseline survival was updated (S0(t=4)=0.9808), and a refit model, in which both the baseline survival and the regression coefficients were updated. The S0 substituted model was constructed by updating the baseline survival rate estimated from the DeSC database from the original one. The S0 and regression coefficients for the refit model were estimated by newly fitting a Cox regression model using the same covariates set to the original model.

We examined the calibration, discrimination, and clinical utility of the original, the substituted, and the refit models in accordance with the guidance and the sample code provided by McLernon et al. 18 For calibration measures, calibration‐in‐the‐large and the calibration slope were assessed, and calibration plots were drawn. 19 Calibration‐in‐the‐large is the ratio of the average of the predicted probability to the observed incidence in the population, and is equivalent to the observed‐to‐expected ratio. The calibration slope represents the strength of the effect per unit increase or decrease of the predictor on the observed incidence. The calibration slope was estimated as the coefficient of the linear predictor in a Cox regression model that included the linear predictor calculated from the model as the only covariate. For calibration plots, the observed probabilities computed by product‐limit estimates were plotted against the deciles of the predicted probabilities. A smoothed calibration curve was also drawn using the approach indicated in the aforementioned report, 18 where observed incidence was calculated by fitting a secondary Cox regression model with a cubic spline function of the complementary log–log transformed predicted survival with 3 knots (0.1, 0.5, and 0.9) as a covariate. Uno's C‐statistic was used to evaluate discrimination. 20 The Uno's C‐statistic is a weighted value based on the censoring probability for the widely used Harrell's C‐statistic. We used Uno's C‐statistic because the start of follow‐up differed across individuals and thus caused a large amount of censoring. Decision curve analysis was performed to evaluate the clinical utility of the models. 21 The median and the 2.5th and 97.5th percentile points of each measure were calculated by 200‐bootstrap resampling. Subgroup analyses were performed by sex, age group (<60 or ≥60 years), and the health insurers (employees' health insurance or national health insurance) to examine the potential difference in model performance in specific subgroups.

We performed a sensitivity analysis using an alternative definition of total CVD as the end point, which additionally included codes for cerebral hemorrhage (I60–I62) or unspecified stroke (I64). We also performed a sensitivity analysis using an alternative diagnosis of ASCVD, which used the codes for emergency hospitalization or surgery/rehabilitation.

In addition, we directly compared the discriminative performance with that of relevant clinical prediction models developed in other cohorts. 22 , 23 , 24 , 25 , 26 We presented the discrimination measures of 3 clinical prediction models in Japanese populations, including the Suita Study, 22 EPOCH‐JAPAN (Evidence for Cardiovascular Prevention from Observational Cohorts in Japan), 23 and the JALS (Japan Arteriosclerosis Longitudinal Study), 24 and 2 models developed for Westerners (the Framingham Heart Study 25 and the pooled cohort equations 26 ). We calculated Uno's C‐statistic for each model. The predictors used to calculate incidence probabilities for each model are listed in the Table S2.

Analyses were performed using SAS version 9.4 (SAS Institute, Cary, NC) and Python 3.9.13 (Python Software Foundation, Wilmington, DE). We did not conduct statistical tests based on P values due to the large sample size. Alternatively, uncertainty in the estimates were presented as 95% CIs. Based on the sample code from McLernon et al, 18 we used the following Python packages and modules to assess model performance measures: statsmodels, scipy, lifelines (KaplanMeierFitter, CoxPHFitter), scikit‐survival.metrics (concordance_index_ipcw, brier_score), and dcurves.dca.

Results

A flowchart of participant selection is shown in Figure 1. Of 570 972 eligible individuals, 528 501 individuals were included in the cross‐sectional analysis for the frequency of missing values. Characteristics of study individuals and the proportion of missing values in each predictor are shown in Table 1. The proportion of individuals with missing values in at least 1 of the variables comprising the CVD risk prediction model was 20.04%. The predictor with the highest frequency of missing values was exercise habit (20.04%). The frequencies of missing values for systolic blood pressure, the presence of diabetes, high‐density lipoprotein cholesterol, low‐density lipoprotein cholesterol, and triglycerides were <0.1%, and those for proteinuria and smoking habits were 0.44% and 0.15%, respectively.

Figure 1. Flowchart of participant selection.

Figure 1

Table 1.

Baseline Characteristics of the Eligible Population and the Population With Complete Information on Predictors

Eligible population (n=528 501) Population without missing values on predictors (n=420 552)
Mean±SD, proportion, or median (IQR) Missing value (%) Mean±SD, proportion, or median (IQR)
Age, y 56.8 (10.3) 0.00 56.6 (10.3)
Men, % 49.8 0.00 49.1
Beneficiaries of employees' health insurance, % 49.0 0.00 49.0
Systolic blood pressure, mm Hg 124.7 (18.2) 0.03 124.4 (18.3)
Diastolic blood pressure, mm Hg 75.8 (11.7) 0.03 75.7 (11.8)
Use of antihypertensive drug, % 20.0 0.25 19.0
Hypertension, % 34.8 0.22 33.8
Fasting or casual glucose, mg/dL 97.8 (18.9) 30.64 97.2 (17.6)
Hemoglobin A1c, % 5.7 (0.6) 5.92 5.7 (0.6)
Use of glucose‐lowering drug, % 5.0 0.26 4.7
Diabetes, % 8.6 0.01 8.2
HDL cholesterol, mg/dL 64.0 (16.9) 0.03 64.0 (16.8)
LDL cholesterol, mg/dL 126.0 (31.0) 0.05 125.9 (30.9)
Triglycerides, mg/dL (median, IQR) 92 (66–135) 0.03 91 (65–133)
Proteinuria, % 3.4 0.45 3.1
Current smoking, % 18.2 0.15 18.3
No regular exercise habit, % 68.4 20.04 68.3

HDL indicates high‐density lipoprotein; IQR, interquartile range; and LDL, low‐density lipoprotein.

We evaluated the model performance longitudinally for 420 552 individuals, excluding those with missing measurements in any predictor. The baseline characteristics are shown in Table 1. During a median follow‐up of 4.4 years (interquartile range, 2.3–5.1 years; maximum, 5.6 years), 3998 individuals newly developed ASCVD. The median predicted probability of developing ASCVD in a 4‐year time horizon, estimated by the original model, at baseline was 1.22% (interquartile range, 0.59%–2.44%). Figure 2 shows the Kaplan–Meier survival curves and crude HRs for developing ASCVD according to the deciles of predicted probability of ASCVD. The observed incidence increased log‐linearly for groups with higher predicted probability.

Figure 2. Kaplan–Meier survival curves and crude hazard ratios (95% CIs) according to the deciles of predicted probability by the original model.

Figure 2

Kaplan–Meier curves are represented by a color gradient from blue (first decile) to red (10th decile). D1 to D10 denote the first to the 10th deciles, respectively. The hazard ratios (95% CI) were estimated using a Cox proportional hazard model with the first decile (D1) set as the reference category.

Figure 3 shows the calibration plots with calibration curves for the original model, the substituted model, and the refit model. Table S3 shows the coefficients for each predictor in the refit model. As shown in Figure 3, the original model was approximately on a straight line, but the predicted probabilities by the original model were consistently higher than the actual incidence observed in the study population across the risk levels. For the substituted model, the calibration curve was close to a diagonal line. The refit model showed a nearly perfect calibration curve, with a calibration‐in‐the‐large of 1.00 (95% CI, 0.97–1.03) and a calibration slope of 1.00 (95% CI, 0.97–1.04).

Figure 3. Calibration of the original, substituted, and refit model for predicting atherosclerotic cardiovascular diseases events.

Figure 3

Observed probabilities were estimated using cubic spline regression (Poisson model) with the predicted probability as the explanatory variable and the log cumulative hazard as the offset term.

The original model had satisfactory discrimination (Uno's C‐statistic, 0.759 [95% CI, 0.751–0.766]). The C‐statistic for the substituted model was, of course, identical to that of the original model. The C‐statistic for the refit model was 0.766 (95% CI, 0.759–0.774). Table 2 shows the calibration and discrimination in the substituted model and the refit model by sex, age, and types of insurance. The predictive performance was generally similar across the covariates except in those aged ≥60 years: This subgroup showed a relatively poor discrimination with C‐statistics of 0.687 in the substituted model and 0.696 in the refit model.

Table 2.

Subgroup Analyses for Calibration and Discrimination Measures in the Substituted and the Refit Model

Substituted model Refit model
Calibration‐in‐the‐large (95% CI) Calibration slope (95% CI) C‐statistic (95% CI) Calibration‐in‐the‐large (95% CI) Calibration slope (95% CI) C‐statistic (95% CI)
Sex
Male 0.90 (0.86–0.94) 1.05 (1.00–1.10) 0.746 (0.736–0.755) 0.98 (0.94–1.02) 1.01 (0.96–1.06) 0.754 (0.744–0.763)
Female 1.17 (1.10–1.24) 1.04 (0.96–1.12) 0.746 (0.731–0.761) 0.97 (0.91–1.03) 0.97 (0.90–1.04) 0.747 (0.733–0.762)
Age groups
40–59 y 1.12 (1.06–1.19) 1.12 (1.05–1.20) 0.749 (0.734–0.766) 1.20 (1.11–1.25) 1.16 (1.09–1.23) 0.760 (0.746–0.776)
≥60 y 1.04 (1.04–1.08) 0.89 (0.83–0.94) 0.687 (0.677–0.698) 1.02 (0.98–1.06) 0.95 (0.90–1.00) 0.696 (0.685–0.707)
Type of health insurance
Employee health insurance 1.11 (1.05–1.17) 1.03 (0.97–1.08) 0.767 (0.757–0.777) 1.16 (1.10–1.23) 1.05 (0.99–1.10) 0.776 (0.766–0.787)
National health insurance 1.02 (0.97–1.06) 0.95 (0.90–1.00) 0.728 (0.717–0.737) 0.99 (0.95–1.04) 1.01 (0.96–1.06) 0.735 (0.724–0.744)

Table S4 shows the results of the sensitivity analysis treating total CVD as the outcome instead of ASCVD. The discriminative performances were slightly lower when predicting total CVD than those for ASCVD, while the values were in acceptable ranges. Table S5 shows the results of sensitivity analyses using alternative definitions of ASCVD. Predictive performances were similar when ASCVD was defined by ICD‐10 codes accompanied by emergency hospitalization (ASCVD hospitalization) or by surgery/rehabilitation (ASCVD surgery). Additional analyses of discriminative performance confirmed that the discriminative ability of the clinical prediction model tested in this study was comparable or superior to that of other clinical prediction models (Table S6).

Figure 4 shows the results of decision curve analysis. All models had higher net benefit compared with the treat‐all or treat‐none approaches in the low range of threshold probabilities. However, the original model no longer had a benefit when the threshold probability exceeded 3%. The substituted model and the refit model showed similar net benefit across the threshold probabilities.

Figure 4. Decision curve analysis of the net benefit of using each model.

Figure 4

The gray line represents the treat‐all strategy, and the black (horizontal) line represents the treat‐none strategy. Models where lines are above both the treat‐all and the treat‐none strategies can be interpreted as having net benefit (clinical utility).

Discussion

In this study, we investigated the external validity of a previously developed clinical prediction model for ASCVD using a large, independent database of health checkup and administrative claims data. We found that the originally published Hisayama ASCVD model performed well for the estimation of relative risks with a satisfactory discriminative ability, but it tended to overestimate the absolute risk of ASCVD for the present study population. The predictive performance could be further improved by substituting baseline survival rate. Our findings suggest that the Hisayama ASCVD model could be implemented in the national health checkup setting.

We observed an ordered, log‐linear increase in the risk of developing ASCVD with each decile category increase in predictive value by the original model, and all models examined in this study generally had satisfactory discrimination. In a meta‐analysis of external validation studies, the C‐statistics of the Framingham coronary heart model and the pooled cohort equations were 0.68 to 0.74, and the C‐statistics in the present study were higher than those in the previous studies. 27 In addition, the substituted and refit models showed sufficient calibration, discrimination, and clinical usefulness, indicating the limited benefit of refitting the regression equation. These findings suggested that the set of predictors included in the Hisayama ASCVD model was reasonable and useful for stratifying potential ASCVD risk into relative risk categories.

When estimated by the original model, the predicted absolute risk (ie, predicted incidence probabilities) for ASCVD in the study population was consistently overestimated. The meta‐analysis examining the performance of the Framingham score and the pooled cohort equations also noted that the predicted probabilities tend to be overestimated. 27 The overestimation in this study may be due in part to the ≈30‐year time difference from the baseline examination of the development cohort. Decreases in the incidence of CVD, particularly ischemic stroke, with changes in treatment for those at high risk for CVD (eg, introduction of statins, angiotensin receptor blockers, and aggressive control of risk factors) during this period have been reported. 28 Another potential cause of the overestimation is the differences in methods used to detect ASCVD outcomes. It has been noted that the extraction of diagnosis codes from a single electronic database can result in more missed cases compared with extraction the combination of an electronic database plus a death or registry database. 29 , 30 In this study, only claims data were used, which may have resulted in certain oversights of events and informative censoring that could have potentially associated with the outcomes. The observed improvements in predictive ability and clinical usefulness by recalibration should be carefully interpreted since we cannot confirm the “true” ASCVD incidence in the DeSC database. Such a value could be obtained by extrapolating the absolute risk from epidemiological studies that comprehensively investigated and adjudicated the development of ASCVD. Further efforts are needed to accurately quantify baseline incidence rates to validate and calibrate the absolute predicted risk of ASCVD.

Subgroup analysis showed good predictive performances across the subgroups. However, discrimination was relatively poor in the older‐age subgroup. This finding of poor discrimination in older adults was consistent with a previous report on the QRISK3 CVD prediction algorithm of the UK electronic database 4 and another large validation study of the pooled cohort equations with over 20 million cases. 31 The exact reasons are unknown, but there may be greater unexplained variation in the risk of developing CVD in older people. This hypothesis is supported by the findings of a recent study from the UK Biobank, which showed that the CVD‐risk fraction attributable to modifiable factors decreased with age. 32 It might be necessary to consider additional predictors and reweighting for the older individuals.

In the present study, there were almost no missing values for any of the predictors included in the clinical prediction model except for exercise habits. The specific health checkup uses information on the presence or absence of components of metabolic syndrome and current smoking habits for risk stratification to determine the level of subsequent health guidance, but it does not use information on exercise habits, 16 and thus it is possible that some insurers did not actively ask questions regarding regular exercise. Although it constitutes a trade‐off between accuracy and implementability, the decision to exclude items on exercise habits may be beneficial in terms of implementation. In fact, the Japan Atherosclerosis Society has provided a prediction tool that excluded exercise habits from their prediction tool. 8

The strengths of this study include the representativeness of the sample of the DeSC database and the large sample size, which accounted for ≈2% of the total number of participants in the annual specific health checkup in Japan; the use of a large and representative sample allowed us to provide stable estimates even in the subgroup analyses, and to show the robustness of our findings. Limitations of this study should also be noted. First, the diagnosis of ASCVD was based on the records in the administrative claim information, which limits its accuracy. It should be noted, however, that the sensitivity analysis showed good discriminatory ability even when alternative outcome definitions were used. Second, the study population was limited to enrollees of particular health insurers. However, since the decision to provide information to the DeSC database was made by the insurers, not by individuals, the participation bias at the individual level should be minimal. Finally, because the database was masked to job types and geographic locations, variations in model performance could not be examined for these factors.

Conclusions

This study demonstrated that the Hisayama ASCVD model had good predictive performance in the national health checkup setting in Japan. Our findings support the usefulness of the Hisayama ASCVD model for discriminating between individuals at high and low risk of ASCVD in an entire population across various subpopulations. The Hisayama ASCVD model would contribute to efficient risk stratification and risk communication in the nationwide health checkups.

Sources of Funding

This study was supported by the Health and Labour Sciences Research Grants of the Ministry of Health, Labour and Welfare of Japan (JPMH23FA1006), the Ministry of Education, Culture, Sports, Science and Technology of Japan (JSPS KAKENHI Grant Number JP22K17396, for T.H.), a research grant from the Japan Arteriosclerosis Prevention Fund (for T.H., T.N., D.Y., and S.C.), and a commissioned research fund from the DeSC Healthcare Inc. (for T.N.). The funders had no role in the design and conduct of the study; collection, management, analysis, and interpretation of the data; preparation, review, or approval of the manuscript; and decision to submit the manuscript for publication.

Disclosures

T.N. received a commissioned research fund from DeSC Healthcare Inc. The remaining authors have no disclosures to report.

Supporting information

Data S1. Supplemental Methods

Tables S1–S6

References 33, 34

JAH3-14-e040386-s001.pdf (339.3KB, pdf)

Acknowledgments

The authors thank the study participants. The statistical analyses were carried out using the computer resources offered under the category of General Projects by the Research Institute for Information Technology, Kyushu University. T.H. and T.N. conceptualized and designed this study.

Author contributions: T.H., Y.F., and A.K. contributed to the acquisition, analysis, or interpretation of data. T.H. performed the data analysis and drafted the manuscript. All authors contributed to critical revision of the manuscript and approved the final version. T.H. had full access to all the data in the study and takes responsibility for the integrity of the data and the accuracy of the data analysis. The database used in this study (DeSC database) is a commercial database and is not publicly accessible.

This manuscript was sent to Yen‐Hung Lin, MD, PhD, Associate Editor, for review by expert referees, editorial decision, and final disposition.

For Sources of Funding and Disclosures, see page 9.

References

  • 1. Roth GA, Mensah GA, Johnson CO, Addolorato G, Ammirati E, Baddour LM, Barengo NC, Beaton AZ, Benjamin E, Benziger CP, et al. Global burden of cardiovascular diseases and risk factors, 1990–2019: update from the GBD 2019 study. J Am Coll Cardiol. 2020;76:2982–3021. doi: 10.1016/j.jacc.2020.11.010 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2. Luo F, Chapel G, Ye Z, Jackson SL, Roy K. Labor income losses associated with heart disease and stroke from the 2019 Panel Study of Income Dynamics. JAMA Netw Open. 2023;6:e232658. doi: 10.1001/jamanetworkopen.2023.2658 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3. Muntner P, Colantonio LD, Cushman M, Goff DC Jr, Howard G, Howard VJ, Kissla B, Levitan EB, Lloyed‐Jones DM, Safford MM. Validation of the atherosclerotic cardiovascular disease Pooled Cohort risk equations. JAMA. 2014;311:1406–1415. doi: 10.1001/jama.2014.2630 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4. Hippisley‐Cox J, Coupland C, Brindle P. Development and validation of QRISK3 risk prediction algorithms to estimate future risk of cardiovascular disease: prospective cohort study. BMJ. 2017;357:j2099. doi: 10.1136/bmj.j2099 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5. Damen JAAG, Hooft L, Schuit E, Debray TPA, Collins G, Tzoulaki I, Lassale CM, Siontis GCM, Chiocchia V, Roberts C, et al. Prediction models for cardiovascular disease risk in the general population: systematic review. Br Med J. 2016;353:i2416. doi: 10.1136/bmj.i2416 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6. Dalton ARH, Bottle A, Soljak M, Majeed A, Millett C. Ethnic group differences in cardiovascular risk assessment scores: national cross‐sectional study. Ethn Health. 2014;19:367–384. doi: 10.1080/13557858.2013.797568 [DOI] [PubMed] [Google Scholar]
  • 7. Collins GS, de Groot JA, Dutton S, Omar R, Shanyinde M, Tajar A, Voysey M, Wharton R, Yu LM, Moons KG, et al. External validation of multivariable prediction models: a systematic review of methodological conduct and reporting. BMC Med Res Methodol. 2014;14:40. doi: 10.1186/1471-2288-14-40 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8. Okamura T, Tsukamoto K, Arai H, Fujioka Y, Ishigaki Y, Koba S, Ohmura H, Shoji T, Yokote K, Yoshida H, et al. Japan Atherosclerosis Society (JAS) guidelines for prevention of atherosclerotic cardiovascular diseases 2022. J Atheroscler Thromb. 2024;31:641–853. doi: 10.5551/jat.GL2022 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9. Fujiyoshi A, Kohsaka S, Hata J, Hara M, Kai H, Masuda D, Miyamatsu N, Nishio Y, Ogura M, Sata M, et al. JCS 2023 guideline on the primary prevention of coronary artery disease. Circ J. 2024;88:763–842. doi: 10.1253/circj.CJ-23-0285 [DOI] [PubMed] [Google Scholar]
  • 10. Honda T, Chen S, Hata J, Yoshida D, Hirakawa Y, Furuta Y, Shibata M, Sakata S, Kitazono T, Ninomiya T. Development and validation of a risk prediction model for atherosclerotic cardiovascular disease in Japanese adults: the Hisayama study. J Atheroscler Thromb. 2022;29:345–361. doi: 10.5551/jat.61960 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. Nakao YM, Miyamoto Y, Ueshima K, Nakao K, Nakai M, Nishimura K, Yasuno S, Hosoda K, Ogawa Y, Itoh H, et al. Effectiveness of nationwide screening and lifestyle intervention for abdominal obesity and cardiometabolic risks in Japan: the metabolic syndrome and comprehensive lifestyle intervention study on nationwide database in Japan (MetS ACTION‐J study). PLoS One. 2018;13:e0190862. doi: 10.1371/journal.pone.0190862 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12. Fukuma S, Iizuka T, Ikenoue T, Tsugawa Y. Association of the national health guidance intervention for obesity and cardiovascular risks with health outcomes among Japanese men. JAMA Intern Med. 2020;180:1630–1637. doi: 10.1001/jamainternmed.2020.4334 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. Okada A, Yasunaga H. Prevalence of noncommunicable diseases in Japan using a newly developed administrative claims database covering young, middle‐aged, and elderly people. JMA J. 2022;5:190–198. doi: 10.31662/jmaj.2021-0189 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Collins GS, Reitsma JB, Altman DG, Moons KGM; Members of the TRIPOD group . Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (TRIPOD): the TRIPOD Statement. Eur Urol. 2015;67:1142–1151. doi: 10.1016/j.eururo.2014.11.025 [DOI] [PubMed] [Google Scholar]
  • 15. Fujihara K, Yamada‐Harada M, Matsubayashi Y, Kitazawa M, Yamamoto M, Yaguchi Y, Seida H, Kodama S, Akazawa K, Sone H. Accuracy of Japanese claims data in identifying diabetes‐related complications. Pharmacoepidemiol Drug Saf. 2021;30:594–601. doi: 10.1002/pds.5213 [DOI] [PubMed] [Google Scholar]
  • 16. Ministry of Health, Labour and Welfare . Standard Program of Health Checkup and Health Guidance (2018) . Accessed May 5, 2025. https://www.mhlw.go.jp/stf/seisakunitsuite/bunya/0000194155.html.
  • 17. Elsayed NA, Aleppo G, Aroda VR, Bannuru RR, Brown FM, Bruemmer D, Collins BS, Gaglia JL, Hilliard ME, Isaacs D, et al. 2. Classification and diagnosis of diabetes: standards of care in diabetes—2023. Diabetes Care. 2023;46:S19–S40. doi: 10.2337/dc23-S002 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18. McLernon DJ, Giardiello D, van Calster B, Wynants L, van Geloven N, Smeden M, Therneau T, Steyerberg EW; Topic Groups 6 and 8 of the STRATOS Initiative . Assessing performance and clinical usefulness in prediction models with survival outcomes: practical guidance for Cox proportional hazards models. Ann Intern Med. 2023;176:105–114. doi: 10.7326/M22-0844 [DOI] [PubMed] [Google Scholar]
  • 19. van Calster B, Nieboer D, Vergouwe Y, de Cock B, Pencina MJ, Steyerberg EW. A calibration hierarchy for risk models was defined: from utopia to empirical data. J Clin Epidemiol. 2016;74:167–176. doi: 10.1016/j.jclinepi.2015.12.005 [DOI] [PubMed] [Google Scholar]
  • 20. Uno H, Cai T, Pencina MJ, D'Agostino RB, Wei LJ. On the C‐statistics for evaluating overall adequacy of risk prediction procedures with censored survival data. Stat Med. 2011;30:1105–1117. doi: 10.1002/sim.4154 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21. Vickers AJ, Cronin AM, Elkin EB, Gonen M. Extensions to decision curve analysis, a novel method for evaluating diagnostic tests, prediction models and molecular markers. BMC Med Inform Decis Mak. 2008;8:1–17. doi: 10.1186/1472-6947-8-53 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. Nakai M, Watanabe M, Kokubo Y, Nishimura K, Higashiyama A, Takegami M, Nakao YM, Okamura T, Miyamoto Y. Development of a cardiovascular disease risk prediction model using the Suita Study, a population‐based prospective cohort study in Japan. J Atheroscler Thromb. 2020;28:304. doi: 10.5551/jat.ER48843 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23. Li Y, Yatsuya H, Tanaka S, Iso H, Okayama A, Tsuji I, Sakata K, Miyamoto Y, Ueshima H, Miura K, et al. Estimation of 10‐year risk of death from coronary heart disease, stroke, and cardiovascular disease in a pooled analysis of japanese cohorts: Epoch‐Japan. J Atheroscler Thromb. 2021;28:816–825. doi: 10.5551/jat.58958 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Harada A, Ueshima H, Kinoshita Y, Miura K, Ohkubo T, Asayama K, Ohashi Y; Japan Arteriosclerosis Longitudinal Study Group . Absolute risk score for stroke, myocardial infarction, and all cardiovascular disease: Japan Arteriosclerosis Longitudinal Study. Hypertens Res. 2019;42:567–579. doi: 10.1038/s41440-019-0220-z [DOI] [PubMed] [Google Scholar]
  • 25. D'Agostino RB, Vasan RS, Pencina MJ, Wolf PA, Cobain M, Massaro JM, Kannel WB. General cardiovascular risk profile for use in primary care: the Framingham heart study. Circulation. 2008;117:743–753. doi: 10.1161/CIRCULATIONAHA.107.699579 [DOI] [PubMed] [Google Scholar]
  • 26. Goff DC, Lloyd‐Jones DM, Bennett G, Coady S, D'Agostino RB, Gibbons R, Greenland P, Lackland DT, Levy D, O'Donnell CJ, et al. 2013 ACC/AHA guideline on the assessment of cardiovascular risk. Circulation. 2014;129:49–73. doi: 10.1016/j.jacc.2014.02.606 [DOI] [Google Scholar]
  • 27. Damen JA, Pajouheshnia R, Heus P, Moons KGM, Reitsma JB, Scholten RJPM, Hooft L, Debray TPA. Performance of the Framingham risk models and pooled cohort equations for predicting 10‐year risk of cardiovascular disease: a systematic review and meta‐analysis. BMC Med. 2019;17:109. doi: 10.1186/s12916-019-1340-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28. Hata J, Ninomiya T, Hirakawa Y, Nagata M, Mukai N, Gotoh S, Fukuhara M, Ikeda F, Shikata K, Yoshida D, et al. Secular trends in cardiovascular disease and its risk factors in Japanese: half century data from the Hisayama Study (1961–2009). Circulation. 2013;128:1198–1205. doi: 10.1161/CIRCULATIONAHA.113.002424 [DOI] [PubMed] [Google Scholar]
  • 29. Herrett E, Shah AD, Boggon R, Denaxas S, Smeeth L, van Staa T, Timmis A, Hemingway H. Completeness and diagnostic validity of recording acute myocardial infarction events in primary care, hospital care, disease registry, and national mortality records: cohort study. Br Med J. 2013;346:1–12. doi: 10.1136/bmj.f2350 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30. Davidson J, Banerjee A, Muzambi R, Smeeth L, Warren‐Gash C. Validity of acute cardiovascular outcome diagnoses recorded in european electronic health records: a systematic review. Clin Epidemiol. 2020;12:1095–1111. doi: 10.2147/CLEP.S265619 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31. Kartoun U, Khurshid S, Kwon BC, Patel AP, Batra P, Philippakis A, Khera AV, Ellinor PT, Lubitz SA, Ng K. Prediction performance and fairness heterogeneity in cardiovascular risk models. Sci Rep. 2022;12:12542. doi: 10.1038/s41598-022-16615-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32. Tian F, Chen L, Qian M, Xia H, Zhang Z, Zhang J, Wang C, Vaughn MG, Tabet M, Lin H. Ranking age‐specific modifiable risk factors for cardiovascular disease and mortality: evidence from a population‐based longitudinal study. eClinicalMedicine. 2023;64:102230. doi: 10.1016/j.eclinm.2023.102230 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33. Ministry of Health, Labour and Welfare . Overview of Medical Service Regime in Japan . Accessed May 5, 2025. https://www.mhlw.go.jp/bunya/iryouhoken/iryouhoken01/dl/01_eng.pdf.
  • 34. Ministry of Health, Labour and Welfare . Status of the Specific Health Checkup and Specific Health Guidance in FY2021 . 2023. Accessed May 5, 2025. https://www.mhlw.go.jp/stf/seisakunitsuite/bunya/newpage_00043.html.

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Data S1. Supplemental Methods

Tables S1–S6

References 33, 34

JAH3-14-e040386-s001.pdf (339.3KB, pdf)

Data Availability Statement

The database used in this study (DeSC database) is a commercial database and is not publicly accessible.


Articles from Journal of the American Heart Association: Cardiovascular and Cerebrovascular Disease are provided here courtesy of Wiley

RESOURCES