Skip to main content
npj Health Systems logoLink to npj Health Systems
. 2026 Sep 2;3:83. doi: 10.1038/s44401-026-00143-7

Structural requirements for intelligent clinical digital twins in feedback-driven care

Yanfei Wang 1, Glenn E Smith 2, Gloria P Lipori 3, Ramon C Sun 4, Elizabeth A Shenkman 1, Mei Liu 1, Yonghui Wu 1, Yi Guo 1,✉
PMCID: PMC13538474  PMID: 42686992

Abstract

Clinical digital twins are increasingly promoted for decision support, yet most are designed and validated for prediction under observed care rather than decision-support validity. We identify four structural requirements for intelligent clinical digital twins used in bidirectional clinical care. Using Alzheimer’s disease and related dementias as an example, we show how routine deployment can induce care-path-dependent epistemic drift despite stable predictive performance.

Subject terms: Computational biology and bioinformatics, Health care, Mathematics and computing, Medical research

Introduction

In recent years, digital twins have attracted growing attention in medicine as a framework for representing patients, diseases, and care processes over time1. The digital twin concept originated in engineering and industrial systems2–4, and its application to healthcare is relatively recent and remains under active development. Existing studies have proposed clinical digital twins for a range of use cases, including individualized prognostic modeling5–7, patient monitoring and early warning8,9, operational and system-level optimization10, and longitudinal disease management11,12. To date, however, most efforts conceptualize digital twins as data-driven models constructed and evaluated retrospectively or offline, with their primary use in risk estimation, trajectory characterization, or exploratory analysis13,14, rather than direct participation in ongoing clinical care.

The term “digital twin” is used broadly in the literature and spans multiple conceptual layers. At a minimum, a clinical digital twin refers to a data-driven representation of a patient or disease process that evolves over time15,16. Recent frameworks have expanded this concept to include patient-specific representation, longitudinal updating, multi-source data integration, simulation, prediction, interoperability, verification and validation, uncertainty quantification, governance, and clinical implementation13–20. Systems built on these ideas have been applied to monitoring, risk prediction, disease progression modeling, treatment-response simulation, and operational decision support5–10. These contributions establish important technical and operational foundations for clinical digital twins.

Authoritative guidance from the National Academies of Sciences, Engineering, and Medicine has articulated a vision for digital twins as persistent, patient-anchored systems that adapt over time through learning or reasoning and interact bidirectionally with the care process17. This perspective reframes digital twins from static analytical models toward continuously evolving systems that may influence clinical decisions and, in turn, the data used to update them. In such systems, information exchange is not unidirectional. Model outputs may influence clinical actions, follow-up schedules, and care pathways, such that what data are collected, when they are collected, and how outcomes are observed depend in part on prior decisions and on the digital twin’s own outputs. For example, a risk estimate may trigger intensified monitoring, which in turn alters subsequent observations, patient-care interactions, and clinical documentation21. As a result, digital twins may become embedded in closed-loop systems in which care decisions and data generation co-evolve over time.

This bidirectional feedback creates a gap between prediction under observed care and decision-relevant interpretation: existing frameworks do not fully address how digital twin outputs should be interpreted when used to support clinical decisions under routine deployment. Most clinical digital twins are designed and validated primarily for prediction, which estimates what is likely to be observed next if care continues as before. Decision-making, however, requires reasoning about what would occur under alternative sequences of actions that were not taken. Under feedback-driven care, predictive performance under observed data does not characterize outcomes under alternative strategies and therefore cannot, on its own, support valid comparisons between them. When validation and interpretation are misaligned, digital twin outputs may be used to justify comparisons between clinical strategies that the system was not designed or assessed to support. Such misinterpretation can accumulate over time and may reinforce existing care patterns rather than inform meaningful alternatives22–24.

This manuscript addresses this prediction-to-decision gap. Our contribution is not another general checklist of clinical digital twin components. Instead, we identify the structural conditions required when outputs from a clinical digital twin designed and validated under observed care are used to support decision-relevant claims. In this setting, predictive validity under the observed care pathway is not sufficient for decision-support validity. A system may remain well calibrated to the care pathway it helps produce while becoming structurally misaligned for comparing alternative clinical strategies.

In this manuscript, we use the term intelligent clinical digital twin (ICDT) to refer to a specific class of clinical digital twin systems rather than to all digital twins used in medicine. An ICDT is a repeatedly updated, patient-level digital representation embedded in a clinical workflow whose outputs may influence subsequent monitoring, testing, documentation, referral, treatment, or care pathways and may be interpreted as supporting clinical decisions or comparisons among alternative clinical strategies. This definition is based on deployment context and intended use, not on a specific algorithmic architecture, artificial intelligence method, or level of autonomy. These requirements are therefore not intended as universal standards for all clinical digital twins. Rather, they apply most directly to ICDTs used to support decision-relevant claims under feedback-driven deployment.

Our approach is both diagnostic and normative. It is diagnostic because it identifies structural sources of misalignment between prediction-focused digital twin design and decision-support interpretation. It is normative because it articulates the minimum structural conditions required for ICDTs to support interpretable decision-relevant use under bidirectional interaction with clinical care. These conditions are not procedural recommendations or algorithmic prescriptions. Rather, they define what a system must be able to represent for its outputs to remain interpretable when used to reason about clinical decisions over time.

We therefore identify four minimum structural requirements for ICDTs used in feedback-driven care: (1) representation of decision-dependent data generation, (2) explicit temporal representation of states, actions, and observations, (3) representation of alternative clinical strategies, and (4) alignment between validation and intended decision use. Using Alzheimer’s disease and related dementias (ADRD) as an illustrative example, we show how routine deployment can induce care-path-dependent epistemic drift, in which model outputs become increasingly shaped by prior care decisions and their effects on measurement and patient-care interactions rather than underlying disease progression alone.

Structural requirements for intelligent clinical digital twins

We use several system-level terms to specify the deployment context for ICDTs. An embedded system refers to a digital twin integrated into a clinical workflow (e.g., embedded within an electronic health record [EHR] system) rather than used only for offline retrospective analysis. A closed-loop system refers to a setting in which model outputs may influence clinical actions, those actions shape subsequent data generation, and the resulting data are then used to update the digital twin. A feedback-driven clinical environment refers to a care setting in which clinical decisions, observations, documentation, and model outputs co-evolve over time. An operational system refers to a digital twin functioning as part of routine clinical infrastructure, with repeated updates, outputs, and interactions with care processes, rather than as a one-time analytic model.

Once embedded in clinical workflows, an ICDT functions as part of such a closed-loop system. Clinical decisions influence subsequent data generation, incoming data update the ICDT’s representation of patient state, state representations generate model outputs, and those outputs may in turn inform future decisions (Fig. 1). In this setting, state refers to the patient’s underlying clinical and biological condition at a given point in time, whereas state representation refers to the ICDT’s internal encoding or estimate of that condition. These two are not equivalent. The patient state is unobserved and evolving, while the state representation is constructed from data shaped by prior clinical decisions, observation processes, documentation practices, and model-influenced care. As a result, interpretability depends not only on model performance, but also on whether the ICDT is structurally capable of representing the feedback loop through which data, decisions, and state estimates co-evolve.

Fig. 1. Conceptual architecture of an intelligent clinical digital twin embedded in feedback-driven care.

Fig. 1

Schematic illustrating the bidirectional coupling between a physical clinical system and an intelligent clinical digital twin (ICDT) system. Clinical decisions shape data generation through monitoring frequency, testing, visits, and patient or caregiver engagement, producing decision-dependent observations. These observations are mapped into an internal state representation that is constructed under the observed care pathway, updated over time as new data arrive, and used to generate outputs for prediction under observed care, decision support, or alternative-strategy evaluation. Those outputs are then interpreted to guide subsequent clinical actions, creating feedback between care decisions and future data generation. The figure highlights four minimum structural requirements for interpretable, decision-relevant use: representation of decision-dependent data generation, explicit temporal representation of states, actions, and observations, representation of alternative clinical strategies separate from observed-care state tracking, and alignment between validation and intended decision use. Created in BioRender. Yuan, H. (2026) https://BioRender.com/h48upm7.

As noted above, these requirements are not intended as universal standards for all clinical digital twins, nor as prescriptions for a single computational architecture; they define structural conditions needed for decision-relevant interpretation under feedback-driven deployment. We describe four requirements: (1) representation of decision-dependent data generation, (2) explicit temporal representation of states, actions, and observations, (3) representation of alternative clinical strategies, and (4) alignment between validation and intended decision use.

The four requirements are structural rather than algorithm-specific. Different computational approaches may satisfy the same requirement, and a single ICDT may combine statistical, causal, knowledge-representation, and machine-learning components. Table 1 maps each requirement to its design objective, possible methodological or computational approaches, required system functions, and operational diagnostics. These examples are not exhaustive and are not intended to define a required architecture. Instead, they indicate the types of computational and operational functions that an ICDT would need to preserve for decision-relevant interpretation.

Table 1.

Operationalizing the structural requirements for intelligent clinical digital twins

Structural requirement Design objective Possible approaches Required system functions Operational diagnostics
Req 1. Representation of decision-dependent data generation Separate disease change from observation and documentation change Observation-process models; encounter-intensity models; missingness models; process mining; workflow logs; provenance metadata; reporter-source metadata Visit timing; test ordering; referral history; documentation density; source lineage; reporter type; care-path context Visit-frequency shift; test-ordering shift; referral shift; documentation-density shift; source-composition shift; patient-to-proxy reporting shift
Req 2. Explicit temporal representation of states, actions, and observations Preserve clinical order and timing Event-history models; state-space models; dynamic Bayesian models; temporal knowledge graphs; time-aware sequence models; temporal causal models Event time; documentation time; decision time; model-update time; time-indexed covariates; action history; outcome timing Temporal leakage; action-observation ordering; lag diagnostics; missingness over time; event-time versus record-time mismatch
Req 3. Representation of alternative clinical strategies Separate observed-care tracking from strategy evaluation Target trial emulation; causal estimands; structural causal models; g-formula; marginal structural models; dynamic treatment-regime frameworks; off-policy evaluation; counterfactual simulation Eligibility; time zero; strategies; follow-up; outcomes; censoring; assumptions; empirical support; deviation tracking Strategy-definition audit; positivity and support diagnostics; deviation diagnostics; unsupported-contrast audit; observed-care representation reuse audit
Req 4. Alignment between validation and intended decision use Match validation evidence to intended decision claim Context-of-use specification; claim-specific validation; calibration; discrimination; uncertainty quantification; decision-curve analysis; impact evaluation; model cards; lifecycle governance Output-use label; permitted uses; non-permitted uses; uncertainty; assumptions; validation evidence; governance trigger Decision-use audit; validation-by-claim; calibration by care pathway; uncertainty review; drift monitoring; semantic-stability monitoring; repurposing review

Requirement 1: representation of decision-dependent data generation

Most proposed clinical digital twins are developed in a prediction-first paradigm, in which the care processes that generate EHR measurements, including visit patterns, clinician-driven testing, and documentation practices, are effectively treated as incidental noise in data analysis25–27. In real-world clinical practice, however, data are neither collected automatically nor uniformly over time28. What is measured, when it is measured, and how it is documented depend on clinical decisions, perceived risk, and ongoing interactions among patients, caregivers, and clinicians29,30. As a result, care processes can systematically shape, and at times redefine, the meaning of recorded signals.

These processes are routinely adjusted in response to earlier assessments and actions21,31. When an ICDT treats observed data as a direct reflection of disease stage without accounting for how those data were generated, it risks conflating changes in measurement with changes in the underlying condition29,32. For example, intensified monitoring following perceived decline produces denser measurements and more detailed documentation, which can create the appearance of accelerated deterioration even if the disease process itself has not changed26,33. In such cases, the internal state representation begins to encode care intensity and observation patterns alongside disease severity. Without explicit representation of these observation processes, an ICDT cannot distinguish between changes in disease and changes in how disease is measured.

This issue extends beyond individual measurements to the integration of heterogeneous data sources across systems. Clinical data are distributed across heterogeneous sources, including EHRs, laboratories, imaging systems, pharmacy databases, devices, and registries, each reflecting distinct workflows, documentation incentives, and timing constraints. Once an ICDT influences care, integration of these sources becomes structurally critical rather than incidental. Model outputs that affect monitoring intensity, testing, referrals, or follow-up directly shape which data are generated, when they are generated, and how they are recorded across systems34. Data availability and measurement are therefore no longer external inputs, but components of a feedback loop linking decisions, data generation, integration, and interpretation35. Without explicit representation of these processes and their lineage, an ICDT cannot reliably distinguish changes in patient state from artifacts introduced by observation, documentation, or data fusion. As a result, downstream outputs may remain predictive yet lack stable interpretation in longitudinal clinical use. Accordingly, the system must explicitly represent the observation and data-generating processes underlying recorded measurements.

Requirement 2: explicit temporal representation of states, actions, and observations

In many existing digital twins, clinical data extracted from EHRs are treated as direct representations of patient state. Diagnoses, treatments, laboratory values, or risk factors are often collapsed into static indicators, such as whether a patient ever received a medication. While such representations may be adequate for retrospective association studies or cross-sectional prediction, they are misaligned with how clinical care unfolds over time, where decisions, measurements, and outcomes occur in sequence and influence one another21. When temporally structured processes are collapsed into static variables, an ICDT loses information about when events occurred, in what order, and under what clinical context. Without time-indexed representations, the system cannot distinguish cause from consequence, anticipation from response, or baseline risk from treatment effect. Apparent associations may therefore reflect documentation timing rather than disease progression or intervention effects.

Without explicit representation of clinical time, alternative clinical strategies defined as sequences of decisions over time cannot be coherently defined or compared. ICDTs used for decision-relevant claims should therefore represent covariates, treatments, observations, model outputs, and outcomes as evolving processes rather than static summaries. Temporal structure is not only descriptive. It is necessary for preserving clinical order, decision context, and the distinction between observations generated before, during, or after a clinical action.

Static state representations may be appropriate for prediction-only, monitoring, or risk-stratification uses, particularly when uncertainty models account for irregular monitoring, missingness, or variation in observation frequency. For example, monitoring intervals may widen uncertainty bounds rather than directly change the estimated patient state. Such approaches can mitigate some forms of measurement irregularity when the output is used to forecast what is likely to be observed under the current care pathway.

However, uncertainty modeling does not by itself preserve the temporal ordering of clinical actions, observations, model outputs, and outcomes once these processes have been collapsed into a static representation. For ICDTs used to support sequential decisions or comparisons among alternative clinical strategies, this ordering is part of the decision problem. Without it, the system cannot distinguish whether a measurement preceded a decision, resulted from a decision, triggered a decision, or was generated because of a prior model output. Requirement 2, therefore, does not require a single full-trajectory architecture for all clinical digital twins. It requires sufficient representation of clinical time for the intended decision claim.

Requirement 3: representation of alternative clinical strategies

ICDTs typically maintain an internal representation of patient state13, such as disease severity, risk, or functional status, that is updated as new clinical data become available. This state representation summarizes what is known about the patient under the care pathway actually followed and may support prediction under continued observed care36. However, it should not be treated as a neutral basis for evaluating alternative clinical strategies. The representation is constructed from data generated under a specific course of care, and visit timing, monitoring intensity, diagnostic workup, documentation density, and assessment type may themselves be consequences of prior clinical actions.

The core issue is that a state representation constructed under one care pathway may already encode the effects of that pathway29. When the same representation is reused to compare hypothetical alternative courses of action, the ICDT may implicitly treat the representation as independent of the care history that produced it. This assumption is rarely justified in feedback-driven clinical environments. Apparent differences between alternative strategies may therefore reflect differences in how patient state was observed, documented, and updated, rather than differences in how the underlying disease process would respond under those strategies.

For ICDTs used for decision-relevant claims, observed-care state tracking should therefore be distinguished from alternative-strategy evaluation. This does not require the system to perfectly reconstruct the true counterfactual state of the patient under every intervention that was not taken. Rather, it requires explicit representation of the alternative clinical strategies being compared, including the relevant decision point, eligible population, action sequence, follow-up period, outcomes, assumptions, empirical support, and uncertainty. The purpose is to ensure that comparisons reflect differences between specified strategies, rather than artifacts of reusing an observed-care representation as if it were independent of the care pathway that generated it.

This requirement is closely related to defining causal or interventional estimands. However, estimand definition alone is not sufficient if the ICDT does not structurally represent how care processes shape the data used to update state representations and evaluate strategy-specific outcomes. Requirement 3 therefore links decision-relevant interpretation to explicit strategy representation, support assessment, and separation between prediction under the realized care pathway and evaluation of alternative clinical strategies.

Requirement 4: alignment between validation and decision use

Digital twins are commonly evaluated using predictive accuracy and calibration under observed care. These criteria are appropriate when the system is used for forecasting, which estimates what is likely to be observed next if care continues as before. Even if alternative strategies are properly represented (Requirement 3), a separate challenge remains in whether the system’s validation supports the decision claims being made. However, once an ICDT is used to inform or compare clinical decisions, the meaning of its outputs changes. Outputs are no longer interpreted as expected observations, but as support for choices among alternative actions23,24. Strong predictive performance under observed care is no longer sufficient to justify the decision claims being made22–24.

The core risk is not lack of validation, but misalignment between what is validated and how the outputs are used37. An ICDT may be rigorously validated for short-term prediction yet implicitly relied upon to support long-term strategy comparisons, policy decisions, or care optimization. For ICDTs intended to support decision-making, validation evidence and the associated uncertainty characterization must be aligned with the specific decision claims the system is used to support. This requires making explicit which decisions an output is intended to inform, what assumptions are required for those claims to be meaningful, and whether the available validation evidence actually addresses those assumptions. Treating predictive accuracy as a general-purpose guarantee undermines interpretability once decisions are at stake. Accordingly, validation must be aligned with the specific decision claims the system is intended to support.

Alzheimer’s disease and related dementias as an illustrative example of care-path-induced epistemic drift

We use Alzheimer’s disease and related dementias (ADRD) as an illustrative example to show how an ICDT can become embedded in routine longitudinal care and how the meaning of its outputs may change through sustained use (Fig. 2). ADRD care is characterized by long time horizons, recurrent clinical decisions, adaptive monitoring, caregiver involvement, and evolving documentation sources. These features make ADRD a useful setting for examining bidirectional interaction between digital twin outputs and clinical care pathways. The purpose of this example is not to propose a disease-specific modeling architecture, but to illustrate how the four structural requirements become operational in a care pathway where monitoring, referral, reporting source, and documentation intensity evolve over time.

Fig. 2. Operational architecture of an intelligent clinical digital twin for Alzheimer’s disease and related dementias care.

Fig. 2

The figure distinguishes the physical Alzheimer’s disease and related dementias (ADRD) care system from the intelligent clinical digital twin (ICDT) system. The physical care system includes patients, caregivers or proxies, clinicians, clinical encounters, cognitive testing, imaging, referrals, documentation, and care transitions. The ICDT system includes data generation and integration, state representation, state update, memory, output, and a decision interface. The memory layer is not a separate structural requirement, but an operational component that preserves prior observations, model outputs, care actions, reporting sources, and care-context history and supports Requirements 2 and 3. Outputs from the ICDT may influence follow-up, referral, testing, documentation, and caregiver involvement, thereby reshaping the data stream used for future updates. Created in BioRender. Yuan, H. (2026) https://BioRender.com/5dg0ba6.

An ADRD-focused ICDT as an operational system

An ADRD-focused ICDT can be conceptualized as an operational system comprising data generation and integration, state representation, state update, memory, and an output layer with a decision interface (Fig. 2). The data generation and integration layer includes the clinical and related processes that produce measurements and patient-care interactions, drawing from EHRs, laboratory results, imaging, cognitive testing, medication records, referrals, social work, and ancillary services. An ADRD-focused ICDT would not replace core EHR platforms, such as Epic, a widely used commercial EHR system. Instead, it would operate alongside such systems by ingesting longitudinal clinical and nonclinical data, updating patient-state representations, and generating outputs that may inform care decisions. This operation assumes sustained access to, and longitudinal linkage across, clinical and nonclinical data sources, an integration that is often incomplete in practice.

The memory layer in Fig. 2 is not a fifth structural requirement. It is an ADRD-specific operational component that supports the general requirements, particularly Requirement 2 and Requirement 3. It preserves prior state representations, model outputs, clinical actions, observation intensity, reporting sources, and care-context history. In ADRD care, this function is important because the current state representation may be shaped by prior monitoring intensity, specialist referral, diagnostic workup, caregiver involvement, and documentation practices. The memory layer, therefore, records the care pathway through which the current representation was produced.

Embedding the ICDT in routine ADRD care

In routine ADRD care, data used to update the ICDT reflect not only underlying disease processes, but also clinical workflows. Decisions about follow-up frequency, cognitive testing, neuroimaging, medication review, specialist referral, caregiver engagement, social support, and care transitions determine which data are produced, when they are produced, and how they are documented. For example, perceived cognitive decline may prompt more frequent visits, repeated assessments, neuroimaging, or referral to a memory clinic. These actions increase the volume and density of available data while also encoding prior clinical decisions into the observed data stream.

Information relevant to ADRD management also arrives asynchronously and from multiple sources. Cognitive test results, imaging, medication records, clinician notes, caregiver reports, social work documentation, and functional assessments may be recorded at different temporal resolutions and with different documentation practices. As monitoring intensifies or care pathways change, the composition of the data stream evolves. Data integration in an ADRD-focused ICDT is therefore not a one-time build step, but a continuously operating process that co-evolves with care delivery. This illustrates Requirement 1, because data generation and measurement processes are shaped by prior clinical decisions. It also illustrates Requirement 2, because meaningful interpretation depends on preserving temporal structure across evolving data streams.

Illustrative trajectory—how care-path-induced epistemic drift may arise

Consider a patient with mild cognitive impairment who initially has sparse cognitive documentation and primarily patient-reported functional information. An ICDT generates an elevated risk estimate, which prompts closer follow-up, neuropsychological testing, neuroimaging, medication review, and referral to a memory clinic. Subsequent visits increasingly involve caregiver input, and functional decline is documented in greater detail through proxy reports. The ICDT now receives denser cognitive and functional data, but these data reflect not only disease progression, but also intensified observation, referral, diagnostic workup, and a shift from patient-reported to proxy-reported information.

If the updated state representation is later used as a neutral baseline for comparing routine monitoring with memory-clinic referral, the comparison may be distorted because the representation already encodes the pathway triggered by the prior risk estimate. Apparent worsening may partly reflect closer observation, more detailed documentation, and changed reporting source rather than accelerated biological decline. Conversely, apparent stability may reflect structured follow-up and better ascertainment rather than absence of disease progression. This illustrates Requirement 3: state representations constructed under observed care should not be reused as neutral baselines for evaluating alternative clinical strategies.

When validation succeeds but meaning drifts

Care-path-induced epistemic drift can occur even when the ICDT functions as designed. The system may continue to ingest data, update state representations, generate outputs, and remain well calibrated for observed outcomes. The failure mode is not necessarily a conventional prediction error. Instead, the meaning of the internal state representation changes. What was initially intended to represent underlying disease burden becomes, over time, a hybrid representation shaped by disease dynamics, prior alerts, follow-up decisions, diagnostic workup, documentation practices, and reporting regimes.

This failure mode is cumulative because repeated alerts, follow-up decisions, referrals, tests, and documentation practices can progressively change the data stream used to update the ICDT. It is difficult to observe because predictive performance may remain stable while the meaning of the state representation shifts. In ADRD, measurable indicators include increased visit frequency, new memory-clinic referral, increased neuropsychological testing or neuroimaging, greater documentation density, increased caregiver or proxy involvement, and a shift from patient-reported to proxy-reported functional information. Potential transition points include model-generated alerts, intensified follow-up, specialist referral, initiation of diagnostic workup, transition to caregiver-mediated reporting, and entry into home-health or institutional care.

Methodological approaches to prevent or mitigate such drift include provenance-aware state updating, observation-process modeling, source-stratified monitoring, process mining of care pathways, change-point detection in longitudinal data streams, reporter-type metadata, and strategy-specific validation. These approaches do not eliminate feedback between care and data generation. Rather, they make the feedback visible and prevent the ICDT from treating care-path-induced changes in observation or documentation as if they were purely changes in underlying disease.

This reflects Requirement 4. Validation under observed care may remain successful while becoming misaligned with decision uses that compare alternative strategies. In bidirectional clinical settings, validation can become endogenous to the care pathway that the ICDT helps produce. An ADRD-focused ICDT may therefore remain predictively accurate and operationally useful while the interpretability of its decision-relevant outputs gradually degrades.

Discussion

Many existing clinical digital twins are described as supporting clinical decisions in ways that may exceed what their underlying structure has been designed or validated to support, particularly in longitudinal care settings with decision-dependent data generation20,36,38. Rather than proposing new algorithms or modeling techniques, this manuscript identifies minimum structural requirements for ICDTs whose outputs are interpreted as supporting decision-relevant claims over time.

A central implication is that interpretability failures in ICDTs are often structural rather than algorithmic. They arise from misalignment between how data are generated, how patient state is represented, and how model outputs are used to justify decisions. Methodological sophistication and predictive accuracy are therefore neither necessary nor sufficient for reliable decision support. A system may integrate multiple data streams, update continuously, and predict well, yet still fall short for decision use if it cannot distinguish prediction under the realized care pathway from evaluation of alternative clinical strategies. Improving predictive performance alone does not resolve this misalignment, and may even obscure it when predictive success is treated as sufficient evidence for decision use. Future systems can better align with decision-relevant use by making the intended decision claim explicit, preserving observation provenance and timing, separating observed-care state tracking from alternative-strategy evaluation, and validating outputs in relation to the decision claim rather than prediction alone.

Although this manuscript is intentionally diagnostic rather than procedural, the four requirements have direct implications for design, deployment, and governance. We outline three operational safeguards below, organized by the requirement each most directly addresses; Requirement 2’s temporal-representation needs are largely embedded within the implementation of Safeguards 1 and 3, rather than requiring a separate mechanism.

First, consistent with the need to represent decision-dependent data generation (Requirement 1), deployment should include diagnostics that monitor semantic stability, rather than predictive error alone. Semantic stability monitoring should assess whether the meaning of the data stream changes after deployment. Operational diagnostics can be organized at three levels. Workflow diagnostics, including process mining, event-log analysis, and care-pathway sequence analysis, can detect changes in follow-up frequency, referral patterns, test ordering, imaging use, medication review, social work involvement, or other care actions after model deployment. Data-stream diagnostics can monitor longitudinal shifts in measurement density, missingness, documentation frequency, assessment timing, source composition, and structured versus unstructured data contribution, using drift detection, change-point analysis, and source-stratified monitoring. Source-provenance diagnostics can track who or what generated the information used to update the state representation, including reporter type, note source, assessment source, and caregiver or proxy involvement. The ADRD example above illustrates several of these diagnostics in practice, including monitoring for a shift from patient-reported to proxy-reported information and tracking testing or documentation changes following model-generated alerts.

Second, consistent with the need to represent alternative clinical strategies (Requirement 3), ICDTs used for decision-relevant claims should avoid using the observed-care state representation as the sole basis for evaluating alternative clinical strategies. This does not mean that an ICDT must perfectly infer the true counterfactual state of a patient under every intervention that was not taken. Such a state is not directly observable and cannot be recovered from routine clinical data without assumptions. The requirement is instead one of structural role separation. The state representation constructed under the observed care pathway may support tracking and prediction under continued observed care, but it should not serve as a neutral starting point for evaluating alternative strategies.

This separation could be implemented through a modular architecture in which an observed-state tracking module maintains the patient representation under actual care, while a separate strategy-evaluation layer defines the clinical decision, eligible population, time zero, alternative strategies, outcomes, assumptions, empirical support, and uncertainty. This layer could be implemented using target trial emulation39,40, structural causal models41, the g-formula42, marginal structural models43, dynamic treatment-regime frameworks44,45, off-policy evaluation46, or counterfactual simulation branches47,48. Alternatively, action-conditioned state-space or sequence models could update patient state, observation processes, and predicted outcomes conditional on specified actions while retaining provenance information about which observations were generated under the observed care pathway. Required technical capabilities include time-indexed records of actions, observations, model outputs, and subsequent care changes; explicit encoding of alternative strategies as action sequences; provenance tracking for observations; empirical support diagnostics for strategy comparisons; and uncertainty propagation when evaluation depends on extrapolation beyond observed care.

Third, in line with the need to align validation with intended use (Requirement 4), governance mechanisms should specify the permitted decision claims for each ICDT output, monitor whether outputs are used beyond those claims, and require additional validation when outputs are repurposed, extended to new clinical contexts, or used to support comparisons among alternative clinical strategies. Governance for ICDTs can draw on lifecycle approaches developed for AI-enabled medical devices. Regulatory lifecycle frameworks, including FDA guidance on AI-enabled device software functions, predetermined change control plans, and Good Machine Learning Practice principles, provide useful examples for governing intended use, model modification, risk management, performance monitoring, and post-deployment updates49–53.

These frameworks are relevant because ICDTs, like other AI-enabled clinical systems, may face changes in patient populations, clinical practice, data availability, treatment patterns, and documentation processes that can weaken the relevance of prior validation evidence. However, ICDTs introduce an additional governance problem. For many AI-enabled medical devices, the central lifecycle challenge is that the external environment changes around the model. For an ICDT embedded in routine care, the model may also change the environment that generates its future data. A risk estimate, alert, or strategy recommendation can alter follow-up intensity, testing, referral, documentation, treatment decisions, and patient or caregiver engagement. These changes then shape the observations used to update the ICDT and the data against which it is later evaluated. As the ADRD example illustrates, validation can therefore become endogenous to the care pathway that the system itself helps produce, allowing performance to remain stable while the interpretability of decision-relevant outputs quietly degrades.

Governance should therefore track not only whether model performance remains acceptable, but also whether each output is being used only for claims supported by validation evidence. Practical mechanisms include claim-specific validation, explicit use constraints, decision-use audits, monitoring of post-deployment data-generating processes, and review when outputs are repurposed, extended to new clinical contexts, or used to compare alternative clinical strategies. A system validated for forecasting under observed care should not be repurposed to justify alternative clinical strategies unless the corresponding decision claim, assumptions, support, and uncertainty have been explicitly evaluated.

While this manuscript assumes the availability and linkability of multi-source data for illustration, in practice, such data are often fragmented across clinical and community settings, captured with inconsistent identifiers and documentation practices, and subject to access, governance, and data quality constraints. These practical constraints further complicate longitudinal integration and interpretation in settings such as ADRD care.

Recent advances in generative artificial intelligence may help accelerate progress toward ICDTs, but only when used to strengthen, rather than bypass, the structural requirements articulated here. Generative AI may support components such as organizing and summarizing multi-source clinical information, preserving clinically meaningful temporality in internal representations, and improving communication of outputs and uncertainty through decision interfaces54,55. However, these capabilities do not replace the need for explicit representation of decision-dependent data generation, structural separation between state tracking under observed care and evaluation of alternative decision sequences, or alignment between validation, uncertainty characterization, and the specific decision claims a system is used to support.

Limitations and future directions

Several limitations should be acknowledged. First, this manuscript provides a conceptual and diagnostic framework rather than an implemented ICDT or engineering blueprint. Second, the ADRD example is illustrative and is not an evaluation of a deployed system. Third, the four requirements will require different operational implementations across clinical domains, because data-generating processes, decision timing, measurement sources, and care pathways vary across diseases and health systems. Fourth, the framework does not prescribe a single computational architecture. The same structural requirement may be addressed through different combinations of observation-process modeling, temporal state representation, causal strategy-evaluation modules, provenance tracking, validation procedures, and lifecycle governance. Finally, decision-relevant interpretation remains dependent on assumptions, empirical support, uncertainty characterization, and validation aligned with the intended use.

Future research should develop and evaluate methods for implementing these requirements in operational systems. Important methodological directions include observation-process-aware ICDT design, provenance-aware temporal state representation, source-sensitive state updating, methods for monitoring semantic stability and detecting semantic drift after deployment, explicit representation of alternative clinical strategies, validation methods aligned with decision claims rather than prediction alone, and lifecycle governance for adaptive systems whose outputs may reshape future data generation. Empirical studies are also needed to evaluate whether these requirements improve interpretability, clinical usefulness, safety, and robustness across care settings.

Conclusion

The central contribution of this manuscript is to distinguish predictive validity from decision-support validity in clinical digital twins embedded in feedback-driven care. A clinical digital twin should not be considered decision-ready merely because it updates over time, integrates multiple data streams, or predicts accurately under observed care. For decision-relevant use, developers and evaluators should ask whether the system represents how observations are generated, preserves the temporal ordering of states, actions, and observations, distinguishes observed-care state tracking from alternative-strategy evaluation, and validates outputs for the specific decision claims they are used to support.

The framework can be used as a diagnostic checklist when designing, evaluating, or governing clinical digital twins. If a system is intended only for prediction or monitoring under observed care, not all requirements may be necessary. If outputs are used to support clinical decisions or comparisons among alternative strategies, however, the four requirements define minimum structural conditions for interpretable decision-relevant use.

Acknowledgements

Y.G. is partially supported by National Institutes of Health (NIH) grants 5R01CA284646-03, 5R01AG080624-03, and 1R01CA305812-01, and Centers for Disease Control and Prevention (CDC) grant 2U54OH011230. The authors wish to thank the UF Health Cancer Center Research Development Team for their assistance with graphics.

Author contributions

Y.Wang and Y.G. drafted the initial version of the manuscript and finalized it for submission. G.E.S. provided clinical input and contributed to the interpretation of the clinical context. G.P.L., R.C.S., E.A.S., M.L., and Y.Wu. critically reviewed the manuscript and provided substantive edits to improve accuracy and clarity. All authors reviewed and approved the final manuscript.

Data availability

No datasets were generated or analysed during the current study.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Drummond, D. & Gonsard, A. Definitions and characteristics of patient digital twins being developed for clinical use: scoping review. J. Med Internet Res.26, e58504 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Grieves, M. & Vickers, J. in Transdisciplinary Perspectives on Complex Systems: New Findings and Approaches (eds Franz-Josef Kahlen, Shannon Flumerfelt, & Anabela Alves) 85–113 (Springer International Publishing, 2017).
  • 3.Tuegel, E. J., Ingraffea, A. R., Eason, T. G. & Spottswood, S. M. Reengineering aircraft structural life prediction using a digital twin. Int. J. Aerosp. Eng.2011, 154798 (2011). [Google Scholar]
  • 4.Negri, E., Fumagalli, L. & Macchi, M. A review of the roles of digital twin in CPS-based production systems. Procedia Manuf.11, 939–948 (2017). [Google Scholar]
  • 5.Gu, F. et al. Identification of digital twins to guide interpretable AI for diagnosis and prognosis in heart failure. npj Digit. Med.8, 110 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Wu, C. et al. MRI-based digital twins to improve treatment response of breast cancer by optimizing neoadjuvant chemotherapy regimens. npj Digit. Med.8, 195 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Allen, A. et al. A digital twins machine learning model for forecasting disease progression in stroke patients. Appl. Sci.11, 5576 (2021). [Google Scholar]
  • 8.Hou, G. Y. et al. Informing intensive care unit digital twins: dynamic assessment of cardiorespiratory failure trajectories in patients with sepsis. Shock63, 573–578 (2025). [DOI] [PubMed] [Google Scholar]
  • 9.Geddes, J. R. et al. Digital twins for noninvasively measuring predictive markers of right heart failure. npj Digit. Med.8, 545 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Silva-Aravena, F., Morales, J. & Jayabalan, M. e-Health strategy for surgical prioritization: a methodology based on digital twins and reinforcement learning. Bioengineering12, 605 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Sarani Rad, F., Hendawi, R., Yang, X. & Li, J. Personalized diabetes management with digital twins: a patient-centric knowledge graph approach. J. Pers. Med.14 10.3390/jpm14040359 (2024). [DOI] [PMC free article] [PubMed]
  • 12.Hwang, T. et al. Clinical usefulness of digital twin guided virtual amiodarone test in patients with atrial fibrillation ablation. npj Digit. Med.7, 297 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Katsoulakis, E. et al. Digital twins for health: a scoping review. NPJ Digit Med7, 77 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Tortora, M. et al. Medical digital twin: a review on technical principles and clinical applications. J. Clin. Med.14 10.3390/jcm14020324 (2025). [DOI] [PMC free article] [PubMed]
  • 15.Sadée, C. et al. Medical digital twins: enabling precision medicine and medical artificial intelligence. Lancet Digit Health7, 100864 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Committee on Foundational Research, G., Future Directions for Digital, T., National Academies of Sciences, E. & Medicine. Foundational Research Gaps and Future Directions for Digital Twins (National Academies Press, 2024). [PubMed]
  • 17.National Academies of Sciences, E., and Medicine. Foundational Research Gaps and Future Directions for Digital Twins (National Academies Press, 2024). [PubMed]
  • 18.Zhang, K. et al. Concepts and applications of digital twins in healthcare and medicine. Patterns5, 101028 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Mulder, S. T. et al. Dynamic digital twin: diagnosis, treatment, prediction, and prevention of disease during the life course. J. Med Internet Res.24, e35675 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Ringeval, M., Etindele Sosso, F. A., Cousineau, M. & Paré, G. Advancing health care with digital twins: meta-review of applications and implementation challenges. J. Med Internet Res.27, e69544 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Chiolero, A., Santschi, V. & Paccaud, F. Public health surveillance with electronic medical records: at risk of surveillance bias and overdiagnosis. Eur. J. Public Health23, 350–351 (2013). [DOI] [PubMed] [Google Scholar]
  • 22.Kappen, T. H. et al. Evaluating the impact of prediction models: lessons learned, challenges, and recommendations. Diagnost. Prognost. Res.2, 11 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.van Geloven, N. et al. Prediction meets causal inference: the role of treatment in clinical prediction models. Eur. J. Epidemiol.35, 619–630 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Dickerman, B. A. & Hernán, M. A. Counterfactual prediction is not only for causal inference. Eur. J. Epidemiol.35, 615–617 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Agniel, D., Kohane, I. S. & Weber, G. M. Biases in electronic health record data due to processes within the healthcare system: retrospective observational study. BMJ361, k1479 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Phelan, M., Bhavsar, N. A. & Goldstein, B. A. Illustrating informed presence bias in electronic health records data. How Patient Interact. a Health Syst. Can. Impact Inference EGEMS (Wash. DC)5, 22 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Weber, G. M. et al. Biases introduced by filtering electronic health records for patients with “complete data. J. Am. Med Inf. Assoc.24, 1134–1141 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Sisk, R. et al. Informative presence and observation in routine health data: A review of methodology for clinical risk prediction. J. Am. Med Inf. Assoc.28, 155–166 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Sisk, R. et al. Informative presence and observation in routine health data: A review of methodology for clinical risk prediction. J. Am. Med. Inform. Assoc.28, 155–166 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Luo, L., Small, D., Stewart, W. F. & Roy, J. A. Methods for estimating kidney disease stage transition probabilities using electronic medical records. EGEMS (Wash. DC)1, 1040 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Howe, C. J., Cole, S. R., Lau, B., Napravnik, S. & Eron, J. J. Jr Selection bias due to loss to follow up in cohort studies. Epidemiology27, 91–97 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Bower, J. K., Patel, S., Rudy, J. E. & Felix, A. S. Addressing bias in electronic health record-based surveillance of cardiovascular disease risk: finding the signal through the noise. Curr. Epidemiol. Rep.4, 346–352 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Goldstein, B. A., Bhavsar, N. A., Phelan, M. & Pencina, M. J. Controlling for informed presence bias due to the number of health encounters in an electronic health record. Am. J. Epidemiol.184, 847–855 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Kim, G. Y. E., Corbin, C. K., Grolleau, F., Baiocchi, M. & Chen, J. H. Monitoring strategies for continuous evaluation of deployed clinical prediction models. J. Biomed. Inf.168, 104854 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.van Amsterdam, W. A. C., van Geloven, N., Krijthe, J. H., Ranganath, R. & Cinà, G. When accurate prediction models yield harmful self-fulfilling prophecies. Patterns (N. Y)6, 101229 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Khoshfekr Rudsari, H. et al. Digital twins in healthcare: a comprehensive review and future directions. Front Digit Health7, 1633539 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Dziadkowiec, O., Durbin, J., Muralidharan, V. J., Novak, M. & Cornett, B. Improving the quality and design of retrospective clinical outcome studies that utilize electronic health records. HCA Health. J. Med1, 131–138 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Tudor, B. H. et al. A scoping review of human digital twins in healthcare applications and usage patterns. npj Digit. Med.8, 587 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Wang, Y., Li, Y., Lin, T. & Guo, Y. An operational target trial emulation framework for causal inference using electronic health record data. npj Digit. Med.9, 424 (2026). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Hernán, M. A. How to estimate the effect of treatment duration on survival outcomes using observational data. Bmj360, k182 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Pearl, J. An introduction to causal inference. Int. J. Biostat.6, Article 7 10.2202/1557-4679.1203 (2010). [DOI] [PMC free article] [PubMed]
  • 42.Young, J. G., Cain, L. E., Robins, J. M., O’Reilly, E. J. & Hernán, M. A. Comparative effectiveness of dynamic treatment regimes: an application of the parametric g-formula. Stat. Biosci.3, 119–143 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Hernán, M. A., Brumback, B. & Robins, J. M. Marginal structural models to estimate the causal effect of zidovudine on the survival of HIV-positive men. Epidemiology11, 561–570 (2000). [DOI] [PubMed] [Google Scholar]
  • 44.Robins, J. M. in Proceedings of the Second Seattle Symposium in Biostatistics: Analysis of Correlated Data (Lin, D. Y. & Heagerty, P. J. eds) 189–326 (Springer New York, 2004).
  • 45.Murphy, S. A. Optimal dynamic treatment regimes. J. R. Stat. Soc. Ser. B Stat. Methodol.65, 331–355 (2003). [Google Scholar]
  • 46.Thomas, P. S. & Brunskill, E. in Proceedings of the 33rd International Conference on International Conference on Machine Learning - Volume 48 2139–2148 (JMLR.org, New York, NY, USA, 2016).
  • 47.Sutton, R. S. Dyna, an integrated architecture for learning, planning, and reacting. SIGART Bull.2, 160–163 (1990). [Google Scholar]
  • 48.Ha, D. & Schmidhuber, J. World models arXiv Prepr. arXiv:1803. 101222, 440 (2018). [Google Scholar]
  • 49.Bodnari, A. & Travis, J. Scaling enterprise AI in healthcare: the role of governance in risk mitigation frameworks. NPJ Digit Med. 8, 272 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Reddy, S., Allan, S., Coghlan, S. & Cooper, P. A governance model for the application of AI in health care. J. Am. Med Inf. Assoc.27, 491–497 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Food, U. S. & Drug, A. Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence-Enabled Device Software Functions: Guidance for Industry and Food and Drug Administration Staff (U.S. Food and Drug Administration, Silver Spring, MD, 2025).
  • 52.Food, U. S. & Drug, A. Artificial Intelligence-Enabled Device Software Functions: Lifecycle Management and Marketing Submission Recommendations: Draft Guidance for Industry and Food and Drug Administration Staff (U.S. Food and Drug Administration, Silver Spring, MD, 2025).
  • 53.Food, U. S., Drug, A., Health, C., Medicines & Healthcare products Regulatory, A. Good Machine Learning Practice for Medical Device Development: Guiding Principles (U.S. Food and Drug Administration, Silver Spring, MD, 2021).
  • 54.Singhal, K. et al. Large language models encode clinical knowledge. Nature620, 172–180 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Abbas, Q., Jeong, W. & Lee, S. W. Explainable AI in clinical decision support systems: a meta-analysis of methods, applications, and usability challenges. Healthcare (Basel)13 10.3390/healthcare13172154 (2025). [DOI] [PMC free article] [PubMed]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

No datasets were generated or analysed during the current study.


Articles from npj Health Systems are provided here courtesy of Nature Publishing Group

RESOURCES