Abstract
Several decades ago a dramatic leap forward occurred in the development and application of statistical methods for modeling radiation risk at the Radiation Effects Research Foundation (RERF). Poisson regression analysis for grouped person-year cohort data and the linear excess relative risk model were introduced, and subsequently a devoted software system, Epicure® (https://www.hirosoft.com), was developed by researchers at RERF and at the U.S. National Cancer Institute. Numerous advancements in understanding radiation effects on humans were made possible with these methods, which are still the state-of-the-art for risk assessment at RERF and have remained part of the standard toolbox for radiation—and other environmental—epidemiological studies worldwide. Nevertheless, as our understanding of radiation risk has increased, so have the breadth and depth of questions that require answers based on emerging data that are not amenable to these conventional methods. This overview briefly recounts the conventional methods and then describes our recent diversification into the use or development of new statistical approaches to meet the challenges of burgeoning biological data and emerging mechanistic information. We briefly discuss the development and application of new methods, current and planned, that are part of the RERF Statistics Department’s role in supporting institution-wide research, especially in our collaborations involving the Life Span Study, Adult Health Study, and First-generation Offspring Clinical Study. Some approaches to modeling and assessing radiation risk with newer methods mentioned herein have already been published, while some are still in development or are only beginning at the proposal stage.
Keywords: radiation risk, statistical methods, atomic-bomb survivors
Graphical Abstract
Graphical Abstract.
1. Current state of the art
Characterizing the dose–response relationship to facilitate inference about mechanisms of radiation-related disease and to inform decisions about acceptable exposure levels are important goals in evaluating risks from radiation exposure. The primary methods used for radiation risk estimation at the Radiation Effects Research Foundation (RERF) evolved in step with the development of sophisticated statistical methods for risk assessment with epidemiological and biological observational data and increasing computational power several decades ago [1–3]. In this section we describe the current state of principal radiation risk analyses conducted at RERF, focusing on specific statistical and computational issues on which RERF statisticians played a role.
1.1. Poisson regression
Early mortality and cancer-incidence follow-up data on the atomic bomb survivor cohort were analysed with contingency table methods using stratification on key explanatory variables [4]. From the mid-1980s the primary approach has instead used Poisson regression [5], a part of the generalized linear modeling (GLM) framework [6], which was implemented at RERF along with a general excess relative risk (ERR) model that facilitates assessing how risk depends on modifying factors such as sex of an exposed individual, age at the time of radiation exposure, and attained age [7]. To facilitate Poisson regression analyses of cohort follow-up data, RERF researchers also developed preeminent software for constructing finely stratified person-year tables with multiple time scales and time-dependent variables. With GLM, the Poisson mean, which is the product of the incidence or mortality rate and the person-years with stratified person-year table data, is estimated as a log-linear regression function of explanatory variables. Stratification on age and time variables in the person-year table is typically based on 5-year intervals. The background rate is typically modeled with parametric functions of age, sex, calendar time or birth year, and lifestyle factors, although one can avoid assuming parametric models by using stratification based on the person-year strata or appropriate collapsing of them. Poisson regression with person-year data remains the primary approach for standard risk analyses at RERF [8, 9].
1.2. Excess relative risk model and risk modification
The ERR model differs from a standard Poisson regression model in that, whereas the latter models the effects of an explanatory variable as a log-linear function [producing an estimate of the log of the rate ratio (RR) for radiation exposure], the ERR model in its simplest form estimates the effect of exposure as an excess increment in RR per unit dose: ERR = RR − 1, where ERR = βd with dose d and β the ERR per unit dose. This form of dose–response model is based on the linear non-threshold (LNT) model used for radiation protection [10]. With a log-linear model the estimated RR is guaranteed to be positive, but the estimated RR from an ERR model can be negative, leading to negative risk (negative incidence) when βd < −1. The ERR estimate must therefore be constrained to be >−1 times the inverse of the maximum value of dose in the study data, although there is no guarantee that this will result in successful model fit to the data, and how to compute confidence bounds in such situations is an open issue.
Risk modification (effect modification or, more appropriately, statistical interaction [11]), makes it possible to characterize how risk varies with time and other factors. Differences in ERR by sex are accommodated by estimating separate values of β for males and females: i.e. βM and βF. Modification of the ERR by other factors is assessed primarily via log-linear functions multiplying the ERR component: ERR = βd exp{γ'z} for a vector z of modifying variables. Effect modification can also be defined as the sum or product of multiple ERR terms (one for radiation exposure and one or more for interacting variables). Nonlinear functions (such as polynomials and splines) and categorical factors can also be used in the exponential ERR-modifying function; recent examples include the effect of age at exposure on risk of breast or uterine cancer in relation to age at menarche [12, 13] and the effect of smoking on radiation-related risk of lung cancer [14]. Estimating statistical interaction with the ERR is thought to be more reasonable than with the default multiplicative form of the log-linear RR model.
An R package NonlinRiskFunctions that implements the ERR model and its effect modification, for use with Poisson regression via generalized nonlinear modeling with the gnm package in R [15], is available at https://www.rerf.or.jp/en/about/organization-en/chart-e/statisi_e/analysistools_e/. Example code for a typical ERR model with effect modification is shown in Fig. 1.
Figure 1.
Example NonlinRiskFunctions code for Poisson regression with an ERR model including effect modification by the logarithm of attained age centered at age 70 (variable lage70) and age at exposure centered at age 30 (variable agex30). The ellipsis (...) represents additional terms that would be included in the model for the background rate, such as the log of age and its square, sex, birth cohort, etc. PYr, person-years
1.3. Dose estimation error
Retrospective dose reconstruction in environmental epidemiology studies can be challenging. Systematic uncertainties can generally be avoided with careful modeling or knowledge of the factors that determine dose. Measurement error (or random error), on the other hand, is often unavoidable and can result in biased risk regression estimates, loss of power and precision, and masking of the shape of the dose–response [16]. The dose reconstruction (dosimetry) system at RERF is particularly well-documented and of unusually high quality [17–19]. Nevertheless, it is subject to individual (random) errors from uncertainties in input data related to survivors’ memory at the time of data collection and shared (Berkson) errors from the use of average shielding or radiation-transmission parameters when survivors’ individual exposure experiences are not known. The current approach to adjusting for potential measurement-error bias at RERF is the regression calibration method [20], which substitutes an estimate of the expected value of true (but unknown) dose given the dosimetry-system estimate of dose. The estimated expected value of true dose is based on an assumed distribution of true doses and an assumed form of random measurement error in the estimated doses; both assumptions are unverifiable in practice. This topic is described in greater detail later, in Section 5.
1.4. Migration adjustment
Subjects in the RERF Lifespan Study (LSS), comprising ∼120 000 atomic bomb survivors, are followed through regional cancer registries to ascertain incident cancers. Cancers diagnosed outside of the regional catchment areas are not reflected when cancer data are abstracted, which could lead to downward bias in incidence estimates. It is therefore prudent to adjust follow-up time (via person-years within each covariate stratum of the person-year data table) to reflect the time subjects spend living solely within the regional catchment areas. The adjustment factor, estimated from a subset of the LSS for which richer residence history information is available, is the probability of residing in the catchment area at any time for a given city, sex, and attained age. Until recently, analyses of LSS cancer incidence applied discrete adjustments to total time at risk for each cell of the person-year table with predefined strata [9, 21].
Recently several parts of the residence probability estimation procedure were improved. With regard to data processing, expansions of regional registries’ catchment areas over time were taken into account when determining residence status for each data point of residence history. In modeling the residence probabilities, we introduced smoothing splines to the time dimensions, increasing the ability to capture variability in migration patterns by estimating continuous probabilities (Lindner et al., in preparation). As a result, sex- and city-specific residence probability estimates are now smooth functions of calendar year and attained age, rather than piecewise constant discrete values, so they can be incorporated into any configuration of person-year tables (regardless of stratum boundaries) for any endpoint of interest by utilizing person-year weighted values. A forthcoming LSS cancer incidence update will incorporate these refined estimates.
2. Standard statistical models for estimating risk
2.1. Proportional hazards models
These are models in which the effects of covariates act multiplicatively on the baseline hazard (incidence or mortality rate among nonexposed individuals) via a log-linear regression model. With Cox regression, the baseline hazard (a function of the primary time scale, usually age in epidemiological studies [22]) is left unspecified when estimating the effects of covariates, which are parameterized as log hazard ratios. Effects of other time scales can be incorporated with parametric functions in the log-linear model or dealt with via stratification. With Poisson regression, the baseline hazard is typically specified as a parametric function of the primary time scale, but can also be dealt with via stratification. Parametric event-time distributions for the baseline hazard (such as a Weibull model) can also be fit.
2.2. General risk models
A limitation of proportional hazards models is the log-linear form, which implies that the simplest dose–response relationship is exponential. The ERR model, in which the dose response is linear rather than log-linear but still multiplies the background rate, is one alternative among what are known as general risk models [23]. Another is the excess absolute rate (EAR) model, in which the effect of exposure adds to, rather than multiplies, the background rate. The ERR and EAR models are equivalent ways of parameterizing the total expected rate, since {(background rate) times (1 + ERR)} = {(background rate) + EAR} [24]. However, estimates of effect modification generally differ between the two scales, which can be informative (e.g. in applying risk estimates to different populations). Mixtures of relative and additive risk are useful when it is thought that joint effects of exposure and other factors might be neither purely multiplicative nor purely additive [25].
2.3. Noncollapsibility
In the fitting of risk regression models, the risk estimate can change as a function of what other variables are in the model (noncollapsibility [26, 27]). Noncollapsibility seems to be underappreciated because of the mistaken conventional wisdom that the values of estimated regression coefficients do not change according to what other variables are included or omitted (which holds for ordinary linear regression models). Our results with the excess odds ratio in binary regression (the binary-data analog of the rate-data ERR) suggested that the effect of noncollapsibility might be greater than with traditional logistic regression [28]. Work is needed to better understand noncollapsibility properties of the ERR with Poisson regression and other incidence-based models.
3. More complex risk estimation
Insights into complex biological and causal mechanisms of disease that have arisen through years of research pose challenges for standard risk models. One challenge is how to incorporate mechanisms that are complex functions of time that might interact with exposure in complex ways; mechanistic modeling is one way of dealing with such complexities. Another challenge is that radiation effects on disease are not limited to direct mechanisms; there may be multiple pathways, some involving intermediate (mediating) states caused by radiation that are themselves causes of disease. Furthermore, complex diseases can share common causes or pathways in their initiation and development, which means that power can presumably be increased by modeling shared processes; however, standard risk models are designed for single outcomes only. Yet another challenge is that disease processes often involve stages that, although clinically definable, cannot be directly quantified but are instead inferred from physiological or biomarker measurements; this leads to a need to accommodate latent factors. This section describes recent work to address some of these issues using more contemporary statistical methods.
3.1. Mechanistic modeling
With mechanistic modeling, rather than assuming a single model relating the exposure to the ultimate outcome (such as an incident cancer), one assumes that the ultimate outcome is the result of stages in disease development, and separate models specify how the exposure is associated with each stage. Extending the Armitage–Doll multistage carcinogenesis model with the assumption that radiation exposure could affect any mutations during carcinogenesis resulted in attained-age modification of the radiation effect [29]. Mechanistic modeling provided new insights into the age-dependence of carcinogenic risk and supported the hypothesis that radiation-related carcinogenesis might depend on attained age [30, 31]. The multistage model has been further extended by incorporating genomic instability [32, 33]. RERF statisticians have collaborated with modelers in Germany on the extension and application of multi-path multistage models [34, 35].
3.2. Mediation
Mediation occurs when another cause of a radiation-related disease is itself caused by radiation. Examples investigated at RERF include serum estradiol [36], which is related to risk of postmenopausal breast cancer and has been associated with prior radiation exposure, and chronic infection with hepatitis B virus (HBV) [37], which is a cause of hepatocellular carcinoma and has been linked to prior radiation exposure via an inability of exposed survivors to clear HBV infections possibly arising from blood transfusions [38]. The study of mediation by HBV applied contemporary methods of mediation analysis in cohort studies with a proportional hazards regression model for disease occurrence and a logistic regression model for binary mediator [39]. Because of noncollapsibility (described earlier), the difference in risk for radiation with and without adjustment for the mediator in a proportional hazards model does not produce a valid estimate of the proportion of total radiation effect that is mediated. It is therefore necessary to estimate the extent of mediation (the mediation proportion) using indirect and direct path coefficients from separate models: one for radiation effect on the mediator and one for radiation and mediator effects on the disease outcome (Fig. 2). This results in a ratio estimator (risk via the mediated—indirect—path divided by the sum of risks via indirect and direct paths). Making inference for a ratio estimate is challenging but, at least in the case of HBV, there is no question about whether mediation occurs because all causal paths are well-established, so the focus is on estimating the mediation proportion rather than testing whether mediation exists. Work is ongoing to apply mediation analysis with structural equation modeling (SEM) via the Mplus software [40] to assess mediation by cytokines in radiation-associated risk of latent atherosclerosis subtypes [41].
Figure 2.
Simple mediation model with direct and indirect (mediated) effects of radiation. The “direct” effect of radiation includes all potential pathways from exposure to outcome that are not explicitly modeled. The regression model for radiation effect on the mediator (thick dashed line) includes a coefficient for the pathway “A” and adjustment for variables that are confounders of the exposure-mediator association (ZA). The regression model for radiation and mediator effects on the disease outcome (thick solid lines) includes coefficients for the pathways “B” and “C,” as well as adjustment for variables that are confounders of the radiation-disease (ZB) and mediator-disease (ZC) associations
3.3. Multiple outcomes
Two types of multiple-outcome data arise in the atomic bomb survivor studies. One type is where only the first occurrence of several outcomes under study is recorded, after which observation stops (so-called competing risks, which may involve informative censoring if a competing risk is related to radiation exposure). This occurs in analyses of first-primary cancers of individual organ sites (where deaths and cancers at other organ sites are censored) and with analyses of cancer mortality (where deaths from noncancer diseases could be a competing risk) [42]. The other type of multiple-outcome data is where more than one type of outcome is observed. This occurs with analyses of multiple first-primary incident cancers in the LSS [43] and with clinical follow-up for complex adult-onset diseases in clinical studies [44].
In the case of the former (only the first outcome is observed), one way the competing risks situation can be analyzed is with Poisson regression using a joint analysis [45]. This involves stacking the person-year data into K replicates (K blocks), one for each of K distinct outcome types, adding a factor variable representing the outcome type, and specifying the numbers of cases as being those that are of the type in the kth block, k = 1, …, K (Fig. 3). If joint analysis is used only for estimating cause-specific hazards and not for estimating probabilities (e.g. event-free or survival probability), it is valid for estimating risk [46]. Parameters specific to outcome type are estimated by interactions between the corresponding covariates and an indicator variable for outcome type. If an overall intercept is to be estimated instead of type-specific intercepts, the model should adjust for the multiplicity of person years; otherwise, it will estimate a Poisson mean that is K times the true mean. A term [−log(K)] added to the model with coefficient fixed at 1 (an offset) will correct for this. If no type-specific parameters are specified, the result is the same as if a single model were fit to all outcomes combined. If all parameters are type-specific, the results are the same as if a separate analysis were made for each outcome type, with censoring by all other outcomes. One benefit of joint analysis is to be able to estimate common background parameters for cancer types with small numbers if they can be assumed to share similar background effects of certain covariates. Another is to be able to assess and directly test heterogeneity of radiation risk across outcome types. We have reported that sharing background parameter effects across outcome types could lead to more precise estimation of type-specific radiation effects but does not improve precision of other model parameter estimates [47]. This approach has been implemented in recent cancer subtype-specific analyses within the LSS [48, 49].
Figure 3.
Layout of data for joint analysis of multiple outcomes for a person-year data set having N person-year cells with outcome that can be subdivided into K subtypes. All person-year table variables except for number of cases “D” are replicated for each outcome subtype; the numbers of cases of the corresponding outcome type are instead recorded in the number of cases column and an index of outcome subtype is added
Joint analysis is limited in that either a particular parameter is shared by (exactly equal in) two or more types of outcome or is estimated separately in all types. This limitation can be relaxed by using empirical Bayes estimation [50]. This approach allows for certain parameters to arise from a common distribution rather than sharing a common single value. It specifies a “prior” distribution on the set of parameters that are thought to be similar; the estimates are free to vary within a range imposed by the variance of that prior distribution. A prior distribution that is noninformative (flat across the entire range of possible parameter values) will lead to similar inference as if the parameters had been estimated separately by outcome type. A prior distribution that is somewhat informative (such as a normal distribution) will tend to shrink the individual type-specific parameter estimates toward a common mean (in which case the individual estimates borrow strength from one another). The smaller the variance of the prior distribution, the more the individual parameter estimates will shrink toward a common value. Empirical Bayes models are fit by computer-intensive simulation methods (Markov Chain Monte Carlo) and so are not conducive to initial model building and testing. However, once initial exploratory analyses have been performed with ordinary likelihood-based joint analysis, empirical Bayes can be used to more flexibly estimate similar parameters along with corresponding probability intervals (credible intervals). Work is currently underway, in collaboration with the RERF Epidemiology Department, to apply this approach to estimate and test heterogeneity of radiation risk across multiple cancer types while flexibly estimating background parameters in a compromise between the shared-value and separate site-specific approaches.
When more than one type of outcome is observed, a competing risks analysis is not appropriate. As study participants age, they might progress through a series of new disease onsets, tracing what is called a “life history.” Each new disease onset indicates a transition to a new state, so models designed for such data are called “multistate models.” As with competing risks models, estimation of probabilities with standard event-time methods (such as the Kaplan–Meier estimated survival curve) is not appropriate. However, probabilities can be estimated by using the “matrix product integral” or Aalen–Johansen method, which is based on Nelson–Aalen estimators of the cumulative hazards for the state transitions [51]. A simple multistate model, a progressive illness-death model with an additional state to model dropout, is illustrated in Fig. 4. This approach forms the ongoing analysis of risk for parental radiation exposure in the RERF clinical study of multi-factorial adult-onset disease risk in children of atomic bomb survivors. An advantage of multistate modeling is the ability to account for possibly informative dropout and informative competing risk of death. A further advantage is that, rather than estimating risk separately for each disease endpoint of interest (which would be inefficient given correlations due to possible sharing of genetic mechanisms and lifestyle risk factors among disease endpoints), risks for each of many possible disease histories can be estimated and then tested for whether they vary among certain subsets of disease states. A disadvantage is that, as the number of disease outcomes grows, the number of potential multiple-disease states, and hence the number of potential transitions, grows dramatically, resulting in the need to impose constraints to achieve parsimony and ensure computational feasibility.
Figure 4.
Simple multistate model (illness-death model) with possibly informative dropout (loss to follow-up, which might be due to disease-related dropout). States are represented by boxes. Arrows represent transitions from one state (end with no arrowhead) to another (end with arrowhead). The transition hazard functions (hfrom,to) are the instantaneous (age-specific) rates of occurrence of entering a “to” state given being in the “from” state just prior to the transition. “Diseased” means diagnosed with any one or more of the disease outcomes under study. “Healthy” means free of the study diseases (participant may have other, unrelated or perhaps related, maladies)
3.4. Latent variables and joint modeling
Latent variables have come to be recognized as being nearly omnipresent in statistical modeling because of practical issues such as missing data, unobserved heterogeneity as a source of residual within-cluster correlation, difficulty in defining or directly measuring certain variables (hypothetical latent constructs as well as measurement error or misclassification), and counterfactuals or potential outcomes in causal modeling [52].
We have employed SEM to assess radiation risk for unobservable (latent) atherosclerosis subtypes on the basis of measurable atherosclerosis indicators [53]. Adjustment for other important clinical risk factors in SEM leads to the so-called multiple indicators, multiple causes (MIMIC) model [54]. Two advantages of SEM are (i) measurement error in the indicator variables is accounted for and (ii) mediation can be incorporated.
Latent effects can also be used to account for dependence between longitudinal and event-time processes in joint modeling [55]. When investigating the relationship between radiation exposure and longitudinally measured biomarkers, it is crucial to address the impact of informative loss to follow-up due to death. This challenge is typically handled by joint modeling, which simultaneously models the longitudinal and event-time outcomes using latent random effects in the longitudinal outcome to account for unobserved factors possibly influencing the event time. It assumes that, given the observed history, the censoring mechanism and the visiting (participation in longitudinal measurements) process are independent of the true event times and future (unobservable) longitudinal measurements [56]. We have used this approach to evaluate radiation effects on longitudinally measured biomarkers in the RERF biennial-exam-based clinical Adult Health Study (AHS), a subset of the LSS [57, 58].
4. Characterizing dose–response shape with focus on low-dose risk
Although the atomic bomb survivor studies have been considered by some to be high-dose studies given that some survivors in the RERF study cohorts received radiation doses close to the limit of survivability, the dose distribution among survivors comprises mostly low doses because survival after the atomic bombings was roughly correlated with distance from the hypocenter, which is inversely correlated with dose. Knowledge of the effects of low doses is a key objective for setting radiation protection standards, but studying low-dose effects is challenging due to low power to detect small effects. Despite the large size of the atomic bomb survivor cohort and the large number of survivors with low dose estimates, effects of radiation exposure on cancer and other health outcomes are still unclear below ∼100 mGy. Standard risk models are generally based on global models fit to data over the entire dose range, so they are not well suited to focusing on low-dose effects. Various alternative approaches have therefore been employed recently.
4.1. Bayesian and other smoothing approaches
The conventional approach to radiation dose–response estimation, which is based on global parametric models such as the LNT model mentioned earlier, may be misleading for studying low-dose effects. As an alternative, a Bayesian semiparametric model that employs a connected piecewise-linear dose–response function has been utilized [59, 60]. This model provides greater flexibility for capturing potential nonlinear dose–response relationships. In addition, uncertainty in radiation risk estimates at low doses is appropriately accounted for by leveraging the advantages of Bayesian estimation without relying on asymptotic approximations, thereby providing a more robust quantification of uncertainty.
4.2. Machine learning approaches
Artificial intelligence (AI) and machine learning (ML) offer data-driven, nonlinear, and flexible alternatives to traditional parametric models. Deep learning, particularly deep neural networks (DNNs), has driven significant breakthroughs across various fields [61]. To leverage the capabilities of DNNs, we are currently investigating their applicability in low-dose radiation risk assessment using data from the LSS. Our goal is to capture subtle dose–response nuances and enhance the accuracy of radiation risk estimation at low doses. Additionally, we are proposing a hybrid approach that integrates DNNs with parametric risk models, combining the interpretability of parametric methods with the predictive power of AI/ML.
5. Uncertainty in measurements
5.1. Measurement error in dosimetry
The regression calibration method mentioned earlier is regularly used to account for effects of random dosimetry error in RERF risk analyses. However, we are working on improved methods for dealing with dose measurement error. The regression calibration method has been extended to account for another source of measurement error—Berkson or averaging error—in addition to classical random error [62]. Berkson error is often considered to be related to the RERF dosimetry system because of the use of average values as inputs in cases of survivors with incomplete data on exposure circumstances. Whereas random errors are uncorrelated with true values, Berkson errors are uncorrelated with estimated values. Both types of error are present in the RERF dose estimates; however, this extension has yet to be utilized in risk analyses at RERF because the relative contributions of random and averaging errors are unknown. More recent work has been conducted with alternative approaches, such as MIMIC models [63] and the simulation-extrapolation (SIMEX) method [64]. The MIMIC approach is based on instrumental variables (which are difficult to identify), so it is not used routinely in risk analyses at RERF. The SIMEX approach was applied to critically evaluate the key assumption regarding the distribution of true dose in the regression calibration method. SIMEX performed well regardless of the true dose distribution and had less bias than regression calibration. SIMEX can deal with a combination of random and averaging errors and can be used with grouped person-year data, but due to computational burden it has not yet been routinely incorporated in risk analyses at RERF.
5.2. Uncertainty in measured biomarkers
Because biomarker measurements can be affected by factors such as laboratory variability and long-term sample storage, their reliability should be assessed to inform their usefulness in population studies. Intra- and inter-assay variability can be estimated by the fit of a mixed model with random effects for study participants and summarized by the coefficient of variation derived from the mixed model variance components [65]. With these methods, we could identify cytokines that are candidates for assessing mediation of radiation effects on atherosclerosis in the aforementioned study that is based on SEM.
5.3. Missing data
Apart from the AHS subset of the LSS cohort, lifestyle information for modeling background rates and effect modification comes primarily from mail surveys. Missing data are a common problem with survey data; ignoring it—e.g. by using simple methods such as excluding incomplete data records (a complete-records-only, or “complete case,” analysis) or adding missing-value categories—can result in bias [66]. The study on radiation, smoking, and lung cancer [14] was, to our knowledge, the first at RERF to account for a missing covariate (smoking) in a principled way, by using single-value imputation. Analyses using even more principled missing-data methods (such as multiple imputation for missing smoking data [67]) are currently being implemented at RERF. For example, multiple imputation was applied in the mechanistic modeling of lung carcinogenesis [34].
6. Challenges of big data
6.1. Genome-wide association studies and gene-radiation interaction
Genome-wide scans for genomic variants related to disease risk and gene-environment interaction have yet to be introduced in RERF research. Genomic variants have traditionally been assessed with agnostic genome-wide association studies (GWAS) by testing gene-disease associations and gene-environment interactions with individual single nucleotide polymorphisms (SNPs). Because cancer involves complex mechanisms, it is unlikely that all cases of a disease share a common variant as a direct cause or as a cause that interacts with exposure (i.e. complex diseases are not Mendelian). It is more likely that, across the many individuals in a population, many variants might be involved; examining variants one-at-a-time, therefore, can result in loss of power. It can be more efficient to search for causal variants and assess gene-environment interaction via gene-set and pathway analyses. We conducted a SNP-set and pathway analysis of GWAS data in collaboration with the University of Hawaii Cancer Center and University of Southern California [68], based on the Multi-Ethnic Cohort Study, using the R package “SKAT” (sequence kernel association test [69]). An extension of SKAT for testing interactions of gene sets and pathways with exposure (“iSKAT” [70]), or analogous methods, is anticipated to be a key component of analyses of gene-radiation interaction in future GWAS studies at RERF.
6.2. Multi-omics analyses
Identifying a risk allele or risk variant in a population is not necessarily synonymous with risk to an individual: a variant in the exomic part of the genome must be expressed for its effect to be realized. Integrated, multi-omic studies—such as joint analysis of genomic variants, expression, and methylation—can be more informative [71]. An approach based on a combined-data matrix [72] was used in a pilot collaboration between RERF and the University of Hawaii Cancer Center Informatics Core Group [73]. Together with downstream analyses, such as with pathway identification software, the results of such methods can elucidate genomic mechanisms potentially related to disease.
6.3. Machine learning with image data
There has long been interest in investigating low-dose radiation effects on life shortening [74]. Radiation exposure is known to accelerate cellular and molecular aging, and anatomical changes in chest X-ray (CXR) images can serve as predictors of biological age. Building on a large convolutional neural network model trained on CXR images from US cohorts [75], we are in the early stages of investigating its applicability for CXR-age prediction in A-bomb survivors using data from the AHS [76]. Our work includes developing and adapting algorithms for image quality assessment and feature identification. If the existing model proves unsuitable for the Japanese population, we propose retraining it using “transfer learning” (reuse of a pretrained model by training it in a new application). CXR-age estimates and selected features from this study may help stratify survivors into radiosensitive and radioresistant subgroups, ultimately contributing to personalized radiation risk assessment.
7. Conclusion
The use of advanced statistical methodologies at RERF began with Poisson regression of person-year grouped data and risk modeling via an ERR model. This approach remains a key tool in the assessment of radiation risk and has provided an essential foundation for assessing radiation-related health risks. While these techniques have proven reliable over time, the increasing complexity and volume of data, specific mechanistic questions, and evolving research needs often require more advanced methods that transcend traditional approaches. Therefore, the RERF Statistics Department continues to develop and apply novel statistical techniques better suited to addressing emerging needs. By expanding our analytical toolkit and fostering interdisciplinary collaborations, we aim to address new scientific challenges and refine our understanding of radiation effects. These efforts will ensure that RERF remains at the forefront of radiation science, equipped to provide rigorous, evidence-driven insights for public health and policy based on the data collected with the invaluable cooperation of the atomic-bomb survivors. We welcome proposals from interested researchers who wish to collaborate in methodological areas of mutual interest.
Acknowledgments
The authors are grateful to Dr Alina Brenner for her critical reading of an earlier version of the manuscript, and to Ms Sachiyo Funamoto for her skillful handling of data that made possible much of the work described here.
Contributor Information
John Cologne, Department of Statistics, Radiation Effects Research Foundation, Hiroshima 732-0815, Japan.
Munechika Misumi, Department of Statistics, Radiation Effects Research Foundation, Hiroshima 732-0815, Japan.
Hanna Lindner, Department of Statistics, Radiation Effects Research Foundation, Hiroshima 732-0815, Japan.
Zhenqiu Liu, Department of Statistics, Radiation Effects Research Foundation, Hiroshima 732-0815, Japan.
Richard Sposto, Department of Statistics, Radiation Effects Research Foundation, Hiroshima 732-0815, Japan.
Author contributions
All authors contributed to writing and checking the manuscript.
Funding
The Radiation Effects Research Foundation (RERF), Hiroshima and Nagasaki, Japan is a public interest foundation funded by the Japanese Ministry of Health, Labour and Welfare and the US Department of Energy (DOE) with the purpose of conducting research and studies for peaceful purposes on medical effects of radiation and associated diseases in humans, to contribute to the health and welfare of atomic bomb survivors and all humankind. RERF sincerely appreciates the support and cooperation of the atomic bomb survivors and their children and families. The research was also funded in part through DOE award DE-HS0000031 to the US National Academy of Sciences. The views of the authors do not necessarily reflect those of the two governments.
Data availability
Although no new data are presented in this article, the Radiation Effects Research Foundation provides some downloadable summary data on its website: https://www.rerf.or.jp/. Please see the “Library (Papers/Data)” link there.
References
- 1. Moolgavkar SH, Prentice RL (eds). Modern Statistical Methods in Chronic Disease Epidemiology. New York: John Wiley & Sons, 1986. [Google Scholar]
- 2. Breslow NE, Day NE. Statistical Methods in Cancer Research: Volume II (The Design and Analysis of Cohort Studies). Lyon, France: International Agency for Research on Cancer, 1987. [PubMed] [Google Scholar]
- 3. Moolgavkar SH (ed). Scientific Issues in Quantitative Cancer Risk Assessment. Boston: Birkhäuser, 1990. [Google Scholar]
- 4. Kato H, Schull WJ. Studies of the mortality of A-bomb survivors 7. Mortality, 1950–78: part 1. Cancer mortality. Radiat Res 1982;90:395–432. 10.2307/3575716 [DOI] [PubMed] [Google Scholar]
- 5. Frome EL. The analysis of rates using Poisson regression models. Biometrics 1983;39:665–74. 10.2307/2531094 [DOI] [PubMed] [Google Scholar]
- 6. McCullagh P, Nelder JA. Generalized Linear Models. 2nd edn. London: Chapman & Hall, 1989. [Google Scholar]
- 7. Preston DL, Kato H, Kopecky K et al. Studies of the mortality of A-bomb survivors. 8. Cancer mortality, 1950–1982. Radiat Res 1987;111:151–78. 10.2307/3577030 [DOI] [PubMed] [Google Scholar]
- 8. Ozasa K, Shimizu Y, Suyama A et al. Studies of the mortality of atomic bomb survivors, Report 14, 1950–2003: an overview of cancer and noncancer diseases. Radiat Res 2012;177:229–43. 10.1667/RR2629.1 [DOI] [PubMed] [Google Scholar]
- 9. Grant EJ, Brenner A, Sugiyama H et al. Solid cancer incidence among the life span study of atomic bomb survivors: 1958–2009. Radiat Res 2017;187:513–37. 10.1667/RR14492.1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10. Boice JD. The linear nonthreshold (LNT) model as used in radiation protection: an NCRP update. Int J Radiat Biol 2017;93:1079–92. 10.1080/09553002.2017.1328750 [DOI] [PubMed] [Google Scholar]
- 11. Rothman KJ, Greenland S, Walker AM. Concepts of interaction. Am J Epidemiol 1980;112:467–70. 10.1093/oxfordjournals.aje.a113015 [DOI] [PubMed] [Google Scholar]
- 12. Brenner AV, Preston DL, Sakata R et al. Incidence of breast cancer in the Life Span Study of atomic bomb survivors: 1958–2009. Radiat Res 2018;190:433–44. 10.1667/RR15015.1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13. Utada M, Brenner AV, Preston DL et al. Radiation risks of uterine cancer in atomic bomb survivors: 1958–2009. JNCI Cancer Spectr 2019;2:pky081. 10.1093/jncics/pky081 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14. Furukawa K, Preston DL, Lönn S et al. Radiation and smoking effects on lung cancer incidence among atomic bomb survivors. Radiat Res 2010;174:72–82. 10.1667/RR2083.1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15. Turner H, Firth D. Generalized Nonlinear Models in R: An Overview of the GNM Package. PDF vignette. 2023. https://cran.r-project.org/web/packages/gnm/vignettes/gnmOverview.pdf (24 September 2025).
- 16. Carroll RJ, Ruppert D, Stefanski LA et al. Measurement Error in Nonlinear Models: A Modern Perspective. 2nd edn. Boca Raton, FL: Chapman & Hall/CRC, 2006. [Google Scholar]
- 17. Cullings HM, Fujita S, Funamoto S et al. Dose estimation for atomic bomb survivor studies: its evolution and present status. Radiat Res 2006;166:219–54. 10.1667/RR3546.1 [DOI] [PubMed] [Google Scholar]
- 18. Cullings HM, Grant EJ, Egbert SD et al. DS02R1: improvements to atomic bomb survivors’ input data and implementation of dosimetry system 2002 (DS02) and resulting changes in estimated doses. Health Phys 2017;112:56–97. 10.1097/HP.0000000000000598 [DOI] [PubMed] [Google Scholar]
- 19. Griffin K, Paulbeck C, Bolch W et al. Dosimetric impact of a new computational voxel phantom series for the Japanese atomic bomb survivors: children and adults. Radiat Res 2019;191:369–79. 10.1667/RR15267.1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20. Pierce DA, Stram DO, Vaeth M. Allowing for random errors in radiation dose estimates for the atomic bomb survivor data. Radiat Res 1990;123:275–84. 10.2307/3577733 [DOI] [PubMed] [Google Scholar]
- 21. Sposto R, Preston DL. Correcting for catchment area nonresidency in studies based on tumor-registry data. Technical report CR 1-92. Hiroshima, Japan: Radiation Effects Research Foundation, 1992.
- 22. Cologne J, Hsu W-L, Abbott RD et al. Proportional hazards regression in epidemiologic follow-up studies: an intuitive consideration of primary time scale. Epidemiology 2012;21:565–73. 10.1097/EDE.0b013e318253e418 [DOI] [PubMed] [Google Scholar]
- 23. Thomas DC. General relative-risk models for survival time and matched case-control analysis. Biometrics 1981;37:673–86. 10.2307/2530149 [DOI] [Google Scholar]
- 24. Cullings HM, Cologne JB. Risk from ionizing radiation. In: Melnick EL, Everitt BS (eds.), Encyclopedia of Quantitative Risk Analysis and Assessment. West Sussex: John Wiley & Sons, 2008, 1540–6. [Google Scholar]
- 25. Cologne JB, Pawel DJ, Sharp GB et al. Uncertainty in estimating probability of causation in a cross-sectional study: joint effects of radiation and hepatitis-C virus on chronic liver disease. J Radiol Prot 2004;24:131–45. 10.1088/0952-4746/24/2/003 [DOI] [PubMed] [Google Scholar]
- 26. Gail MH, Wieand S, Piantadosi S. Biased estimates of treatment effect in randomized experiments with nonlinear regressions and omitted covariates. Biometrika 1984;71:431–44. 10.2307/2336553 [DOI] [Google Scholar]
- 27. Sjölander A, Dahlqwist E, Zetterqvist J. A note on the noncollapsibility of rate differences and rate ratios. Epidemiology 2016;27:356–9. 10.1097/EDE.0000000000000433 [DOI] [PubMed] [Google Scholar]
- 28. Cologne J, Furukawa K, Grant EJ et al. Effects of omitting non-confounding predictors from general relative-risk models for binary outcomes. J Epidemiol 2019;29:116–22. 10.2188/jea.JE20170226 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29. Pierce DA, Mendelsohn ML. A model for radiation-related cancer suggested by atomic bomb survivor data. Radiat Res 1999;152:642–54. 10.2307/3580260 [DOI] [PubMed] [Google Scholar]
- 30. Pierce DA. Mechanistic models for radiation carcinogenesis and the atomic bomb survivor data. Radiat Res 2003;160:718–23. 10.1667/rr3086 [DOI] [PubMed] [Google Scholar]
- 31. Pierce DA, Vaeth M. Age–time patterns of cancer to be anticipated from exposure to general mutagens. Biostatistics 2003;4:231–48. 10.1093/biostatistics/4.2.231 [DOI] [PubMed] [Google Scholar]
- 32. Little MP, Wright EG. A stochastic carcinogenesis model incorporating genomic instability fitted to colon cancer data. Math Biosci 2003;183:111–34. 10.1016/S0025-5564(03)00040-3 [DOI] [PubMed] [Google Scholar]
- 33. Little MP. Cancer models, genomic instability and somatic cellular Darwinian evolution. Biol Direct 2010;5:19. 10.1186/1745-6150-5-19 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34. Castelletti N, Kaiser JC, Simonetto C et al. Risk of lung adenocarcinoma from smoking and radiation arises in distinct molecular pathways. Carcinogenesis 2019;40:1240–50. 10.1093/carcin/bgz036 [DOI] [PubMed] [Google Scholar]
- 35. Kaiser JC, Misumi M, Furukawa K. Biologically-based modeling of radiation risk and biomarker prevalence for papillary thyroid cancer in Japanese a-bomb survivors 1958–2005. Int J Radiat Biol 2021;97:19–30. 10.1080/09553002.2020.1784488 [DOI] [PubMed] [Google Scholar]
- 36. Grant EJ, Cologne JB, Sharp GB et al. Bioavailable serum estradiol may alter radiation risk of postmenopausal breast cancer: a nested case-control study. Int J Radiat Biol 2018;94:97–105. 10.1080/09553002.2018.1419303 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37. Ohishi W, Fujiwara S, Cologne JB et al. Impact of radiation and hepatitis virus infection on risk of hepatocellular carcinoma. Hepatology 53:1237–1245. 10.1002/hep.24207 [DOI] [PubMed] [Google Scholar]
- 38. Fujiwara S, Sharp GB, Cologne JB et al. Prevalence of hepatitis B virus infection among atomic bomb survivors. Radiat Res 2003;159:780–6. 10.1667/0033-7587(2003)159[0780:pohbvi]2.0.co;2 [DOI] [PubMed] [Google Scholar]
- 39. Kim YM, Cologne JB, Jang E et al. Causal mediation analysis in nested case-control studies using conditional logistic regression. Biom J 2020;62:1939–59. 10.1002/bimj.201900120 [DOI] [PubMed] [Google Scholar]
- 40. Muthén BO, Muthén LK, Asparouhov T. Regression and Mediation Analysis Using Mplus. Los Angeles: Muthén and Muthén, 2016. [Google Scholar]
- 41. Nakamizo T, Takahashi I, Ohishi W et al. Study of arteriosclerosis in the Adult Health Study population (Part 2. Analysis of the cytokine network regulating differentiation of mesenchymal stem cells in artery). Research Protocol RP-2-11. Hiroshima, Japan: Radiation Effects Research Foundation, 2011.
- 42. Brenner AV, Preston DL, Sakata R et al. Comparison of all solid cancer mortality and incidence dose-response in the Life Span Study of atomic bomb survivors, 1958-2009. Radiat Res 2022;197:491–508. 10.1667/RADE-21-00059.1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43. Li CI, Nishi N, McDougall JA et al. Relationship between radiation exposure and risk of second primary cancers among atomic bomb survivors. Cancer Res 2010;70:7187–98. 10.1158/0008-5472.CAN-10-0276 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44. Tatsukawa Y, Cologne JB, Hsu W-L et al. Radiation risk of individual multifactorial diseases in offspring of the atomic-bomb survivors: a clinical health study. J Radiol Prot 2013;33:281–93. 10.1088/0952-4746/33/2/281 [DOI] [PubMed] [Google Scholar]
- 45. Pierce DA, Preston DL. Joint analysis of site-specific cancer risks for the atomic bomb survivors. Radiat Res 1993;134:134–42. 10.2307/3578452 [DOI] [PubMed] [Google Scholar]
- 46. Beyersmann J, Scheike TH. Classical regression models for competing risks. In: Klein JP, van Houwelingen HC, Ibrahim JG, Scheike TH (eds.), Handbook of Survival Analysis. Boca Raton, FL: CRC Press, 2014, 157–77. [Google Scholar]
- 47. Sposto R, Misumi M, Cologne J. A note on potential gains in precision of radiation risk estimates from joint analysis. Sci Rep 2024;14:26750. 10.1038/s41598-024-76920-x [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48. Utada M, Brenner AV, Preston DL et al. Radiation risk of ovarian cancer in atomic bomb survivors: 1958–2009. Radiat Res 2021;195:60–5. 10.1667/RADE-20-00170.1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49. Sakata R, Preston DL, Brenner AV et al. Radiation-related risk of cancers of the upper digestive tract among Japanese atomic bomb survivors. Radiat Res 2019;192:331–44. 10.1667/RR15386.1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50. Carlin BP, Louis TA. Bayesian Methods for Data Analysis. 3rd edn. Boca Raton, FL: Chapman & Hall/CRC, 2009. [Google Scholar]
- 51. Beyersmann J, Allignol A, Schumacher M. Competing Risks and Multistate Models with R. New York: Springer, 2012. [Google Scholar]
- 52. Skrondal A, Rabe-Hesketh S. Generalized Latent Variable Modeling: Multilevel, Longitudinal, and Structural Equation Models. Boca Raton, FL: Chapman & Hall/CRC, 2004. [Google Scholar]
- 53. Nakamizo T, Cologne J, Cordova K et al. Radiation effects on atherosclerosis in atomic bomb survivors: a cross-sectional study using structural equation modeling. Eur J Epidemiol 2021;36:401–14. 10.1007/s10654-021-00731-x [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54. Kline RB. Principles and Practice of Structural Equation Modeling. 4th edn. New York: The Guilford Press, 2016. [Google Scholar]
- 55. Ye W, Yu M. Joint models of longitudinal and survival data. In: Klein JP, van Houwelingen HC, Ibrahim JG, Scheike TH (eds.), Handbook of Survival Analysis. Boca Raton, FL: CRC Press, 2014, 523–47. [Google Scholar]
- 56. Rizopoulos D. Joint Models for Longitudinal and Time-to-Event Data: With Applications in R. Boca Raton, FL: CRC Press, 2012. [Google Scholar]
- 57. Yoshida K, French B, Yoshida N et al. Radiation exposure and longitudinal changes in peripheral monocytes over 50 years: the Adult Health Study of atomic-bomb survivors. Br J Haematol 2019;185:107–15. 10.1111/bjh.15750 [DOI] [PubMed] [Google Scholar]
- 58. Yoshida K, Misumi M, Kusunoki Y et al. Longitudinal changes in red blood cell distribution width decades after radiation exposure in atomic-bomb survivors. Br J Haematol 2021;193:406–9. 10.1111/bjh.17296 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59. Furukawa K, Misumi M, Cologne JB et al. A Bayesian semiparametric model for radiation dose-response estimation. Risk Anal 2016;36:1211–23. 10.1111/risa.12513 [DOI] [PubMed] [Google Scholar]
- 60. Furukawa K, Misumi M. A semiparametric approach to evaluate the harm of low-dose exposures. J Radiol Prot 2018;38:286–98. 10.1088/1361-6498/aaa57c [DOI] [PubMed] [Google Scholar]
- 61. LeCun Y, Bengio Y, Hinton G. Deep learning. Nature 2015;521:436–44. 10.1038/nature14539 [DOI] [PubMed] [Google Scholar]
- 62. Pierce DA, Vaeth M, Cologne JB. Allowance for random dose estimation errors in atomic bomb survivor studies: a revision. Radiat Res 2008;170:118–26. 10.1667/RR1059.1 [DOI] [PubMed] [Google Scholar]
- 63. Tekwe CD, Carter RL, Cullings HM et al. Multiple indicators, multiple causes measurement error models. Stat Med 2014;33:4469–81. 10.1002/sim.6243 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64. Misumi M, Furukawa K, Cologne JB et al. Simulation-extrapolation for bias correction with exposure uncertainty in radiation risk analysis utilizing grouped data. Appl Statist 2018;67:275–89. 10.1111/rssc.12225 [DOI] [Google Scholar]
- 65. Nakamizo T, Cologne J, Kishi T et al. Reliability, stability during long-term storage, and intra-individual variation of circulating levels of osteopontin, osteoprotegerin, vascular endothelial growth factor-A, and interleukin-17A. Eur J Med Res 2024;29:133. 10.1186/s40001-024-01722-w [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66. Greenland S, Finkle WD. A critical look at methods for handling missing covariates in epidemiologic regression analyses. Am J Epidemiol 1995;142:1255–64. 10.1093/oxfordjournals.aje.a117592 [DOI] [PubMed] [Google Scholar]
- 67. Furukawa K, Preston DL, Misumi M et al. Handling incomplete smoking history data in survival analysis. Stat Methods Med Res 2017;26:707–23. 10.1177/0962280214556794 [DOI] [PubMed] [Google Scholar]
- 68. Cologne J, Loo L, Shvetsov YB et al. Stepwise approach to SNP-set analysis illustrated with the Metabochip and colorectal cancer in Japanese Americans of the Multiethnic Cohort. BMC Genomics 2018;19:524. 10.1186/s12864-018-4910-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69. Wu MC, Lee S, Cai T et al. Rare-variant association testing for sequencing data with the sequence kernel association test. Am J Hum Genet 2011;89:82–93. 10.1016/j.ajhg.2011.05.029 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 70. Lin X, Lee S, Wu MC et al. Test for rare variants by environment interactions in sequencing association studies. Biometrics 2016;72:156–64. 10.1111/biom.12368 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 71. González JR, Cáceres A. Omic Association Studies with R and Bioconductor. Boca Raton, FL: CRC Press, 2019. [Google Scholar]
- 72. Okimoto G, Zeinalzadeh A, Wenska T et al. Joint analysis of multiple high-dimensional data types using sparse matrix approximations of rank-1 with applications to ovarian and liver cancer. BioData Min 2016;9:24. 10.1186/s13040-016-0103-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 73. Cologne J, Yoshida K, Taga M et al. Preliminary project to establish capability for integrated analyses of multiple molecular (omic) endpoints. Research Protocol RP-P1-17. Hiroshima, Japan: Radiation Effects Research Foundation, 2017.
- 74. Cologne JB, Preston DL. Longevity of atomic-bomb survivors. Lancet 2000;356:303–7. 10.1016/S0140-6736(00)02506-X [DOI] [PubMed] [Google Scholar]
- 75. Raghu VK, Weiss J, Hoffmann U et al. Deep learning to estimate biological age from chest radiographs. JACC Cardiovasc Imaging 2021;14:2226–36. 10.1016/j.jcmg.2021.01.008 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 76. Nakamizo T, Liu Z, Ono S et al. Artificial intelligence-estimated “Chest X-ray Age” among atomic bomb survivors (Part1 Application to AHS and characterization of features). Research Protocol RP-2-23. Hiroshima, Japan: Radiation Effects Research Foundation, 2023.
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
Although no new data are presented in this article, the Radiation Effects Research Foundation provides some downloadable summary data on its website: https://www.rerf.or.jp/. Please see the “Library (Papers/Data)” link there.





