Skip to main content
Scientific Reports logoLink to Scientific Reports
. 2026 May 12;16:21765. doi: 10.1038/s41598-026-52579-4

Enhancing randomized controlled trials through smartwatch-guided participant matching for infectious disease outcomes

Edan Shahmoon 1, Matan Yechezkel 2, Shachar Snir 1, Marco V Perez 3, Margaret L Brandeau 4, Dan Yamin 1,4,✉
PMCID: PMC13358071  PMID: 42120660

Abstract

Randomized controlled trials (RCTs) aim to maximize statistical power while minimizing cost and recruitment burden. In practice, randomization is often stratified or restricted using demographic variables such as age and sex, while physiological heterogeneity that may influence treatment response is rarely incorporated. Consumer smartwatches are now widely used and provide continuous, real-world measurements of cardiovascular physiology and daily activity patterns, including resting heart rate, heart rate variability, sleep timing and regularity, and physical activity, capturing stable individual-level characteristics outside clinical settings. Leveraging these data, we developed Smartwatch-Informed Matching (SIM), a pre-randomization framework that groups physiologically similar participants and applies constrained randomization to assign participants to intervention and control arms. Using a prospective cohort of 4,795 individuals, we compared SIM with conventional age- and sex-based stratification. SIM improved covariate balance and increased similarity in symptom severity (Spearman ρ = 0.176 vs. 0.012) and physiological response profiles (Pearson r = 0.245 vs. 0.112). Power analyses showed that SIM reduced the sample size required to maintain statistical power by 9–18% across a range of effect sizes. These findings demonstrate that incorporating smartwatch-derived physiological similarity into pre-randomization design can enhance the efficiency and precision of randomized clinical trials. The SIM framework is also readily applicable to retrospective matched analyses that aim to reduce confounding.

Supplementary Information

The online version contains supplementary material available at 10.1038/s41598-026-52579-4.

Keywords: Wearable devices, Randomized controlled trials, Participant matching, Smartwatches, Trial design optimization

Subject terms: Cardiology, Health care, Medical research, Physiology

Introduction

Randomized controlled trials (RCTs) are widely regarded as the gold standard for assessing the safety and efficacy of medical interventions.1 A key challenge in designing RCTs is maximizing statistical power to detect treatment efficacy and identify rare adverse events while minimizing resource use.2,3 Larger sample sizes enhance statistical power but require greater participant recruitment and incentive resources. Participants are often the most costly component of RCTs, with recruitment, retention, and compensation accounting for a significant portion of trial budgets.

In 2020, a COVID-19-driven vaccine revolution occurred. Just 11 months after development, the Pfizer-BioNTech vaccine received regulatory approval, supported by large-scale clinical trials involving over 43,000 participants demonstrating its safety and efficacy.4 However, post-approval vaccine safety monitoring systems identified a causal link between mRNA COVID-19 vaccines and myocarditis5– an undetected association during the clinical trials. This emphasizes the critical need for enhancing clinical trial designs to capture rare but significant adverse events.

Contemporary clinical trials face two significant challenges. First, their stratified randomization processes often rely on relatively simple criteria, such as sex, age, comorbidities, and (where applicable) disease progression.6 These factors are likely insufficient to capture the complexity and variability of patient populations. While randomization ensures unbiased estimates in expectation, these factors may be insufficient to capture the full complexity and variability of patient populations. As a result, important prognostic factors, including behavioral and lifestyle characteristics7, may remain unaccounted for, potentially increasing outcome variability and reducing statistical efficiency. Second, contemporary trials predominantly depend on self-reported symptoms to assess outcomes. These self-reports are inherently subjective and susceptible to variability in individual symptom perception.8

A substantial methodological literature has developed trial allocation procedures that explicitly leverage baseline covariates to improve prognostic balance and statistical efficiency, including stratified randomization, covariate-adaptive (minimization-type) schemes, and matched-pair designs.9–12 In practice, these approaches most often operationalize balance using a limited set of prespecified, low-dimensional demographic or clinical factors, which may not fully reflect baseline physiological heterogeneity relevant to outcomes.10,11 Wearable sensors such as smartwatches offer a powerful and scalable tool for capturing rich, continuous physiological data.11 Prior research has demonstrated the potential of these devices to monitor infection, vaccination, or treatment efficacy and side effects.13–21 Such devices capture subtle physiological changes linked to infection progression, severity, and recovery, including fluctuations in heart rate, heart rate variability (HRV)22, sleep patterns, activity levels, and skin temperature.23 Because smartwatches capture key indicators of an individual’s health, activity and rest patterns, as well as social and lifestyle factors, these data can potentially be leveraged to identify similarities between individuals, thereby enabling a more precise and effective matching process in RCTs.

Here, we present a novel framework for the design of RCTs that integrates traditional recruitment methods with smartwatch-derived data to optimize both participant matching and outcome assessment. We refer to our method as smartwatch-informed matching. Participants are matched into intervention and control groups based on baseline physiological data collected via smartwatches during a fixed pre-trial period (e.g., one month). Subsequently, individual responses—to infection, vaccination, or treatment—are assessed using a combination of standard measures (e.g., self-reported symptoms) and deviations from each participant’s smartwatch-derived baseline. To evaluate our approach, we analyzed data from a prospective cohort of 4,795 individuals in Israel, identifying those who became infected with COVID-19 and calculating pre-infection smartwatch-based metrics. We demonstrate that smartwatch-based matching can improve similarity between matched pairs and can significantly reduce the number of participants required in an RCT while maintaining high statistical power.

Methods

Overview

First, we introduce the Smartwatch-Informed Matching process (SIM), which integrates demographic variables (age and sex) with physiological data derived from smartwatches. Next, we evaluate the performance of this matching approach using data from a two-year prospective observational study involving 4,795 participants (Appendix A), comparing its effectiveness against traditional covariate matching methods, such as those used in the Pfizer-BioNTech COVID-19 vaccine Phase 3 Trial.4 We then perform a series of power analysis simulations to assess how SIM influences statistical power. Specifically, we quantify the reduction in required sample size relative to standard matching methods to assess the potential of SIM to improve trial efficiency.

Smartwatch-Informed Matching

We developed SIM as a framework for pairing individuals based on physiological and behavioral similarity, enabling the construction of well-matched intervention and control groups for clinical trials (Fig. 1).

Fig. 1.

Fig. 1

Overview of the Smartwatch-Informed Matching (SIM) framework. Smartwatch-derived data (top panel) are used to compute physiological similarity between individuals (middle panel). These similarity metrics guide a matching algorithm (bottom panel) that pairs participants into intervention and control groups based on their baseline characteristics.

Smartwatches contain various biosensors—such as accelerometers, temperature monitors, and optical heart rate sensors—that can capture a wide range of physiological and behavioral traits relevant to individual health. While many of these signals may offer value for covariate matching in clinical trials, incorporating too many variables introduces the risk of overfitting, multiple testing concerns, and reduced interpretability. To address these challenges, we focused exclusively on heart rate and HRV—core physiological measures that are widely available across devices and reflect the activity of the cardiovascular and autonomic nervous systems. Heart rate is a fundamental vital sign commonly used to monitor general physiological status, including responses to infection and inflammation. HRV, while not a traditional vital sign, is a well-established digital biomarker of autonomic regulation and has been linked to stress, recovery, immune function, and early signs of physiological disruption. HRV-based signals have been shown to reliably capture meaningful physiological responses, including reactions to vaccination and early signs of infection – often preceding self-reported symptoms.13,24,25 Building on this evidence, and following prior work24, we focused on daily heart rate and HRV distributions during sedentary periods; these provide a stable, interpretable signal for comparing individuals’ physiological baselines.

The SIM procedure is detailed in Appendix B. Briefly, our matching process begins by stratifying the population into subgroups based on age and sex, as done in standard clinical trial designs. Within each subgroup, we calculate the distribution of heart rate and HRV for each individual during a predefined baseline period. To assess similarity between individuals, we compare their heart rate distributions using the Earth Mover’s Distance (EMD), a metric that quantifies how much one distribution would need to shift to match another. Intuitively, the lower the EMD, the more similar the heart rate patterns. We apply the same process to compare HRV distributions. We then combine these two distance measures using a simple Euclidean distance to compute an overall dissimilarity score between each pair of individuals. To avoid a fully deterministic implementation of the matching procedure, we introduced a stochastic component by perturbing the pairwise dissimilarity matrix prior to matching. Specifically, we added a noise matrix with values drawn from a distribution scaled to at most 15% of the empirical standard deviation of the dissimilarity scores. Matching was then performed on the perturbed matrix rather than the original scores. Using this perturbed matrix, we apply a nearest-neighbor greedy matching algorithm to pair individuals within each subgroup. This results in matched intervention and control groups, structured similarly to how randomization is done in clinical trials, but guided by physiological similarity.

Data-driven evaluation of SIM

PerMed study

To evaluate our approach, we analyzed the data from the PerMed prospective observational study.24,26,27 The study enrolled 4,795 participants aged 18 and older, who were recruited between November 16, 2020, and May 11, 2023, from various locations across Israel (Appendix A). Recruitment was facilitated through social media advertisements and word-of-mouth. All participants provided informed consent after receiving a thorough explanation of the study from a professional survey company. Eligibility for the PerMed study required participants to be members of Maccabi Healthcare Services (Maccabi) for at least two years prior to enrollment, be smartphone users, and capable of independently providing written informed consent. All participants received both oral and written information regarding the study, and participation was contingent upon written informed consent. This study was approved by the Institutional Review Board (IRB) of Tel Aviv University headed by Prof. Meir Lahav (protocol number 0122-20-TAU). We hereby declare that all methods were performed in accordance with the IRB guidelines and regulations.

Participants completed a one-time enrollment questionnaire, providing personal and health-related information. Each participant was provided with a Garmin Vivosmart 4 smartwatch and asked to install two mobile applications: (1) the PerMed application, which collects daily self-reported symptom and sleep data, and (2) a passive data collection app that records smartwatch measurements. All participants, regardless of operating system, downloaded the same applications. They were encouraged to wear their smartwatches consistently, and a survey company monitored compliance through a dedicated dashboard. This ensured that participants completed their questionnaires at least twice a week, kept their devices charged and worn, and resolved any technical issues with the apps or smartwatches.

Participants completed a daily questionnaire to report symptoms related to COVID-19, with the option to include additional symptoms in free text (Appendix A). The exact date and hour of testing for each positive diagnosis was recorded in the individual’s medical record if they sought care. If participants conducted a rapid test, they were instructed to report the testing time in the PerMed app. For participants with more than one positive record, only the earliest recorded testing time was used for each disease.

The smartwatches continuously recorded heart rate approximately every 60 s and HRV-based stress levels every 180 s. HRV-based stress is a Garmin-computed measure: specifically, the device calculates the interval between heartbeats using heart rate data, where lower variability during sedentary periods corresponds to higher stress levels, and greater variability indicates lower stress.28–31

To minimize participant attrition and ensure data quality, several strategies were employed. Participants who did not complete their daily questionnaire by 19:00 received reminders through the PerMed app. Additionally, the survey company used the dashboard to identify participants who neglected their tasks, prompting follow-up via text or phone call. To increase engagement, a weekly summary report was generated, and a monthly newsletter with study updates and smartwatch tips was sent to participants. Data from the smartphones and Garmin devices were securely collected and stored at Tel Aviv University.

Baseline period

We defined a baseline period – analogous to a run-in phase in clinical trials – for each participant as the five-week window between seven and two weeks prior to their COVID-19 diagnosis. This timeframe was chosen to avoid the influence of early symptoms or infection-related physiological changes, as the typical incubation period for SARS-CoV-2 is 4 to 6 days.32 We included participants who had at least 25 days with a minimum of 16 hours of smartwatch data per day, and missing data were imputed linearly. To minimize potential confounding effects from recent vaccination, we analyzed data from individuals who had received a COVID-19 vaccine at least three months prior to infection, based on evidence that the most pronounced physiological effects, particularly the presence of neutralizing antibodies, occur within the first three months after vaccination.33

Infection period

We defined the infection period as the four-week period spanning from one week before to three weeks after the testing date (days − 7 to + 21 relative to testing). For each participant, we identified the first new symptom reported during this phase. Only symptoms reported during this period and absent one week prior (days − 14 to −7 relative to testing) were included in the analysis. For example, if a participant reported a headache one day before testing and also nine days prior, it was not included in the analysis. Participants who did not complete the questionnaire during days − 13 to −7 were excluded from the analysis due to the inability to determine whether their symptoms existed before diagnosis.

Outcome measures

We included two primary outcome measures representing participants’ responses to COVID-19 infection: a self-reported symptom score and a smartwatch-derived heart rate reaction score.

We collected daily symptom reports through the structured questionnaires. Using data from the US Centers for Disease Control and Prevention and the Pfizer clinical trial, we categorized participant-reported symptoms by severity as follows:

  • No symptoms.

  • Mild symptoms: abdominal pain, back or neck pain, feeling cold, muscle pain, weakness, headache, dizziness, vomiting, sore throat, diarrhea, cough, leg pain, ear pain, loss of taste and smell, swelling of the lymph nodes.

  • Moderate-to-severe symptoms: a measured temperature over 37.5 °C, corresponding to at least a low-grade fever or reported fever (“fever”), fast heartbeat, hypertension, chest pain, dyspnea (shortness of breath), confusion, and chills.

We categorized symptom severity on an ordinal scale: 0 = no symptoms, 1 = mild symptoms, and 2 = moderate-to-severe symptoms. This score captured the subjective experience of illness and was used to assess similarity between matched participants, as well as to simulate vaccine effectiveness in reducing symptoms.

We used continuous heart rate data from smartwatches to compute an objective physiological response to infection. For each participant, we calculated the squared differences between hourly heart rate values during the post-infection period and the final week of baseline. We smoothed these values using a 72-hour rolling average and defined the maximum value of this curve as the participant’s HR reaction score. This measure served as a continuous outcome for evaluating physiological similarity between matched individuals and modeling the impact of vaccination on reducing physiological burden.

We selected the 72-hour window based on both statistical and clinical reasoning. This window is long enough to smooth out momentary fluctuations due to stress, activity, or measurement noise, while still capturing meaningful physiological deviations. This choice is also consistent with findings from a previous study, which reported that the most severe reactions typically last between 24 and 72 h.19 From the smoothed curve, we extracted the maximum value to represent each participant’s peak physiological reaction to infection. This peak reflects the most intense deviation from the individual’s baseline, analogous to clinical markers such as peak fever or peak inflammatory response. Reducing each time series to a single interpretable metric enables direct comparison of physiological responses across individuals and between matched pairs in both our similarity evaluation and power analysis simulations.

Statistical analysis

To evaluate the effectiveness of SIM, we performed a series of analyses using baseline and infection-period data from the PerMed study. Our goal was to assess whether SIM improved covariate balance and better paired individuals with similar characteristics and responses compared to standard age- and sex-based matching. We first stratified participants into four demographic strata based on sex (male, female) and age (18–55, 56–85), following categories used in previous clinical trials.34 Within each stratum, we applied nearest-neighbor matching using either (a) physiological similarity derived from heart rate and HRV distributions during the baseline period (SIM) or (b) random matching based on age and sex.

To assess covariate balance, we calculated the correlation between matched pairs for a set of individual-level characteristics known to influence infection risk. These attributes included lifestyle and social context (e.g., household composition, smoking and alcohol habits), general health (e.g., body mass index (BMI), medical history), and activity/sleep patterns (e.g., step count, activity time, sleep duration). We derived these measures from three sources: enrollment questionnaires, daily symptom and behavior questionnaires, and baseline smartwatch data. For each matching method, we computed the correlation across matched pairs for each covariate. To quantify the uncertainty around correlation estimates for the standard matching approach, we used bootstrapping (1,000 repetitions) and reported the 2.5th and 97.5th percentiles as confidence bounds.

To assess predictive validity, we examined the extent to which matched participants exhibited similar responses to COVID-19 infection. For this analysis, we used two outcome types: (1) self-reported symptom severity, recorded through daily questionnaires and categorized on an ordinal scale (no symptoms, mild, or moderate-to-severe), and (2) objective cardiovascular response, calculated from smartwatch data as the mean absolute change in heart rate and HRV-derived stress between baseline and infection periods. For each of these outcomes, we calculated the correlation between matched intervention and control individuals. To reflect the dependency structure observed in real data, we used Spearman’s rank correlation (ρ) for the ordinal symptom score and Pearson’s correlation (r) for the continuous cardiovascular response score. To examine potential differences in symptom severity across demographic groups, we performed a chi-square test of independence and conducted post-hoc pairwise comparisons with Bonferroni correction to control for multiple comparisons.

Simulation study

We describe the simulation study using the ADEMP Framework35; further details are in Appendix C.

Aims: The primary objective of the simulations was to quantitatively evaluate the extent to which the SIM procedure reduces the required sample size in hypothetical vaccine efficacy trials compared to traditional matching strategies based solely on demographic characteristics (age and sex), while maintaining predefined statistical power.

Data-generating mechanisms: We simulated 100,000 hypothetical clinical trials to ensure comprehensive coverage of realistic epidemiological conditions. In each trial, matched participant pairs (one intervention, one control) were generated. The baseline joint probability of becoming infected and developing symptomatic illness (π) was set to 0.003 (low-risk), 0.030 (moderate-risk), or 0.100 (high-risk) for the control group. For matched intervention participants, this probability was reduced proportionally according to the specified vaccine efficacy (VE), evaluated at 40%, 50%, 60%, and 90%.

To capture the degree of physiological and symptomatic similarity between matched participants, we incorporated outcome dependencies within pairs. These dependencies were modeled using empirical correlations observed in the PerMed dataset: Spearman’s ρ for self-reported symptom severity and Pearson’s r for smartwatch-derived physiological measures. In addition, we conducted a sensitivity analysis with correlation structures ranging from 5% to 40%.

Estimands: The target estimand was the treatment effect (vaccine efficacy) on the probability of developing symptomatic infection over the trial period.

Methods: We compared two trial design strategies: SIM-based matching, which incorporates age, sex, and physiological similarity, and conventional matching based solely on age and sex. For each scenario, simulated trials were analyzed to detect a statistically significant vaccine effect using a fixed significance level of α = 0.05. Analyses were conducted separately for self-reported symptom-based outcomes and smartwatch-derived physiological outcomes.

Performance measures: The primary performance measure was the minimum sample size (number of matched pairs) required to achieve statistical power of 0.80 and 0.95. Relative performance was assessed by comparing required sample sizes between SIM and conventional matching across all parameter combinations, thereby quantifying improvements in trial efficiency.

Results

Validation of SIM before COVID-19 infection

A total of 1,015 participants from the PerMed study were included in the analysis (Fig. 2). Of these, 533 (52.5%) were female and 482 (47.5%) were male. Participants ranged in age from 19 to 85 years, with a median age of 46 years.

Fig. 2.

Fig. 2

Trial profile for the study cohort.

We first validated that SIM captures individual-level factors known to influence infection risk—such as general health, physical activity, and social and lifestyle characteristics. Using baseline (pre-infection) data, we found that SIM-matched individuals demonstrated stronger correlations across these factors compared to individuals matched only by age and sex (Table 1). For instance, the correlation in daily step count between SIM-matched pairs was 16.9% (95% CI: 10.9%–29.1%), substantially higher than the 2.4% correlation (95% CI: − 6.4%–11.1%) observed with age- and sex-based matching. Similarly, perceived general health showed a correlation of 9.5% (95% CI: 0.2%–18.3%) with SIM, compared to 2.8% (95% CI: − 6.6%–11.3%) with traditional matching. The number of children in the household correlated at 20.0% (95% CI: 11.1%–28.7%) for SIM-matched pairs, versus 13.4% (95% CI: 5.5%–20.9%) for age- and sex-based pairs.

Table 1.

Activity and rest, general health, and social and lifestyle correlations among matched pairs.

Measure N (%) or [IQR] Total N Age- and sex-based covariate matching
Mean correlation %, (95% CI)
Age-, sex- and smartwatch-based covariate matching
Mean correlation %, (95% CI)
Activity and rest patterns, median [IQR]
 Daily step count* 6965 [5259, 8957] 1015 2.4 (− 6.4–11.1)† 16.9 (10.9–29.1)†
 Active time per day* 8833 [6721, 11018] 1015 1.8 (− 6.7–11.0)† 12.6 (3.7–22.1)†
 Sleep duration* 6.847 [6.277, 7.440] 960 4.0 (− 5.0–13.1)† 7.6 (− 0.7–18.9)†
 Sleep quality (1–5 scale) 0.630 [0.083, 1.0] 961 0.2 (− 9.0–9.5)† 4.9 (− 5.7–14.0)†
General health
 BMI ≥ 30 203 (20.0%) 1015 0.4 (− 8.8–9.5) 11.0 (0.8–21.5)
 Hypertension 95 (9.5%) 1002 12.1 (1.2–24.0) 22.2 (7.4–36.3)
 With background diseases 389 (38.3%) 1015 5.6 (− 2.2–14.3) 10.3 (1.0–19.6)
 With cardiovascular disease history 182 (18.2%) 1002 16.3 (6.0–26.7) 20.6 (8.3–33.5)
 Perceived general health (1–5 scale), Median [IQR] 4 [4, 5] 1012 2.8 (− 6.6–11.3) 9.5 (0.2–18.3)
Social and lifestyle factors
 Alcohol
Less than 4 times per month 601 (59.2%) 1015 2.6 (− 6.0–11.4) 8.1 (− 1.3–17.1)
4 or more times per month 414 (40.8%)
 Smoking
Yes 97 (9.6%) 1013 0.0 (− 8.3–9.7) –0.1 (− 9.0–9.9)
No 918 (90.4%)
 Number of children in household
None 264 (26.0%) 1013 13.4 (5.5–20.9) 20.0 (11.1–28.7)
1 83 (8.2%)
2 262 (25.8%)
3 277 (27.3%)
4+ 127 (12.5%)
Unknown 2 (0.2%)
 Number of residents in household
1 78 (7.7%) 1005 9.1 (1.5–16.7) 5.2 (− 3.7–13.9)
2 289 (28.5%)
3 187 (18.4%)
4 206 (20.3%)
5+ 245 (24.1%)
Unknown 10 (1%)
 Self-reported stress, Median [IQR] − 0.7 [− 1.0, − 0.2] 961 1.0 (− 8.6–11.0)† 5.6 (− 5.6–14.2)†

*Measure recorded by smartwatch; other measures are self-reported.

†Pearson correlation coefficient; all other correlation coefficients are Spearman’s rank correlation.

Predictive validity of SIM during COVID-19 infection

During COVID-19 infection, 834 participants completed sufficient questionnaires for our analysis. Among them, 222 (26.6%) reported no symptoms, 295 (35.4%) reported mild symptoms, and 317 (38.0%) reported moderate-to-severe symptoms (Table 2).

Table 2.

Distribution of self-reported symptom severity levels and smartwatch-based reaction severity levels within each stratum.

Strata Self-reported symptoms Smartwatch-based reaction severity
None Mild Moderate to severe Total Bottom third Middle third Top
third
Total
Female, 18–55 60 (20%) 104 (35%) 131 (45%) 295 52 (26%) 70 (35%) 77 (39%) 199
Male, 18–55 73 (26%) 104 (38%) 99 (36%) 276 39 (19%) 67 (33%) 97 (48%) 203
Female, 56–85 45 (30%) 52 (35%) 53 (35%) 150 85 (59%) 45 (31%) 15 (10%) 145
Male, 56–85 44 (40%) 35 (30%) 34 (30%) 113 41 (39%) 35 (34%) 28 (27%) 104
Total 222 (27%) 295 (35%) 317 (38%) 834 217 (33%) 217 (33%) 217 (33%) 651

We evaluated the ability of SIM matching to identify similarities between individuals’ severity of COVID-19 infection based on reported symptoms. Correlation of self-reported infection severity between matched pairs was r = 0.176 (95% CI: 0.075–0.282) for SIM, versus r = 0.012 (95% CI: 0.010–0.015) for age- and sex-based matching. We also compared SIM to a one-nearest-neighbor matching based on sex and exact age. This age-based matching yielded a correlation of r = 0.011 for self-reported infection severity between matched pairs, similar to that for age- and sex-based covariate matching.

We assessed the ability of SIM to identify similarities between individuals based on objective physiological responses recorded by the smartwatch rather than self-reported symptoms. Specifically, for each participant, we calculated the mean absolute change in heart rate and HRV-derived stress between the baseline and the COVID-19 infection period. Using the magnitude of these changes as an outcome measure, we found that SIM yielded a correlation of r = 0.245 (95% CI: 0.137–0.345) between matched individuals, compared to r = 0.112 (95% CI: 0.101–0.116) using age- and sex-based matching. A one-nearest-neighbor approach based on exact age and sex produced a lower correlation of r = 0.104 for the same smartwatch-measured cardiovascular response.

We conducted a post-hoc analysis to validate that differences in vaccination timing between matched individuals did not bias our findings. This analysis showed that approximately 80% of matched pairs were vaccinated within 20 days of each other, confirming that vaccination timing was well-balanced and unlikely to influence our results (Appendix D).

Power analysis of clinical trials using SIM

We simulated clinical trial scenarios to evaluate the impact of SIM on the required sample size when using either self-reported symptoms or smartwatch-based cardiovascular responses as the primary outcome. Across all combinations of vaccine efficacy (VE) and the probability of becoming infected and developing symptoms (π), SIM consistently led to reductions in sample size compared to standard age- and sex-based matching (Table 3).

Table 3.

Required sample sizes with standard stratified matching (n0) and smartwatch-based matching (n) for different levels of statistical power (0.95, 0.80) and different outcome metrics (smartwatch-based reaction score, self-reported symptom severity).

Metric π VE Power = 0.95 Power = 0.80
n0 n n0 - n % Reduction in sample size n0 n n0 - n % Reduction in sample size
Smartwatch-based reaction score 0.003 0.40 57,454 48,155 6,373 16.2 31,983 26,238 5,745 18.0
0.50 33,374 28,653 4,720 14.1 18,193 15,945 2,247 12.4
0.60 22,080 18,894 3,185 14.4 11,919 10,527 1,391 11.7
0.90 7,712 7,490 221 2.9 4,209 4,011 198 4.7
0.030 0.40 5,496 4,458 1,038 18.9 3,080 2,520 560 18.2
0.50 3,220 2,755 465 14.4 1,692 1,426 266 15.8
0.60 2,081 1,744 336 16.2 1,171 1,042 128 11.0
0.90 772 747 25 3.3 429 413 16 3.8
0.100 0.40 1,398 1,182 216 15.4 944 778 165 17.6
0.50 981 823 157 16.1 545 453 91 16.7
0.60 611 534 77 12.6 346 301 45 13.0
0.90 203 195 8 3.9 120 116 4 3.3
Self-reported symptom severity 0.003 0.40 57,454 51,081 6,373 11.1 31,983 26,775 5,208 16.3
0.50 33,374 30,231 3,142 9.4 18,193 16,276 1916 10.5
0.60 22,080 20,093 1,987 9.0 11,919 10,614 1305 11.0
0.90 7,712 7,548 164 2.1 4,209 4,078 130 3.0
0.030 0.40 5,496 5,144 352 6.4 3,080 2,817 263 8.6
0.50 3,220 2,882 337 10.5 1,692 1,495 197 11.6
0.60 2,081 1,826 254 12.2 1,171 1,076 94 8.1
0.90 772 743 29 3.8 429 419 10 2.5
0.100 0.40 1,398 1,244 154 11.0 944 846 98 10.4
0.50 981 907 73 7.5 545 482 62 11.5
0.60 611 559 52 8.5 346 312 34 9.9
0.90 203 200 3 1.5 120 118 2 1.7

When using self-reported symptoms, SIM improved matching efficiency and lowered the number of participants needed to achieve the desired statistical power. For example, under trial conditions with π = 0.003 and VE = 0.50, a trial targeting 95% power required 33,374 participants with age- and sex-based matching, compared to 30,231 with SIM, a reduction of 3,142 participants (9.4%). For 80% power under the same conditions, the required sample size dropped from 18,193 to 16,276 (10.5% reduction). The magnitude of reduction varied with VE and π but remained consistent across settings.

When using objective cardiovascular responses from smartwatches as an outcome measure, SIM produced even greater reductions in required sample size. In the same scenario (π = 0.003, VE = 0.50), the sample size required for 95% power dropped from 33,374 to 28,653, representing a reduction of 4,720 participants (14.1%) compared to age- and sex-based matching. At 80% power, the sample size decreased from 18,193 to 15,945 (12.4% reduction). The reductions in required sample size were particularly notable in settings where the probability of becoming infected and developing symptoms (π) was low and vaccine efficacy (VE) was modest, conditions that typically require the largest trials.

To evaluate the impact of varying prognostic associations between the smartwatch-derived metrics and trial outcomes (r), we conducted a sensitivity analysis assuming the correlation strength of the SIM approach ranged from r = 0.05 to r = 0.4. Our results indicated that a strong prognostic association (r = 0.4) could lead to a dramatic reduction in the required sample size, ranging from 24.8% to 31.0%. In contrast, a weak association (r = 0.05) yields no improvement or modest reductions in sample size, of up to 8.4%, compared to standard age- and sex-based stratification alone.

Discussion

Our findings demonstrate that wearable sensors can enhance clinical trial design by enabling more precise matching of participants in intervention and control groups. By incorporating smartwatch-based physiological data into the matching process, our framework enhances the statistical sensitivity to detect intervention-related changes and increases the precision of outcome assessments. These improvements can enhance the timeliness of clinical trials by accelerating the accumulation of meaningful data, thereby supporting faster evaluation and delivery of effective medical interventions.

In simulations across a range of realistic scenarios, we found that physiological matching consistently reduced the required sample size compared to standard stratification by age and sex. Using data from participants infected with COVID-19, the reduction ranged from 9% to 18% in settings where vaccine efficacy is moderate (40–60%). Given that the cost of enrolling a participant in a clinical trial ranges from 31,000$ to 82,000$36 and that participant costs constitute more than one-third of total trial costs37, these reductions in required sample size represent significant potential savings.

Smartwatch data provides an objective measure of physiological responses, reducing the bias of self-reported symptoms and offering valuable insights into treatment effects. Notably, we found that younger individuals, particularly males aged 18–55, had more pronounced physiological responses during infection than older individuals, despite older individuals being at higher risk for severe COVID-19 complications. This may reflect heightened immune activation and autonomic reactivity in younger participants, leading to more detectable changes in metrics like heart rate and heart rate variability. These findings highlight the value of wearable technology in clinical trials for precise, continuous monitoring and a deeper understanding of demographic differences, enabling more tailored treatments.

We used participants’ daily heart rate distributions and HRV-derived stress metrics for our matching process, as they correlate with key physical and mental health factors, as well as lifestyle metrics such as activity, steps, and sleep patterns. Because these lifestyle and baseline health metrics are strong prognostic factors in many therapeutic areas, we believe this matching method may be generalizable to other types of clinical trials.

Indeed, previous studies have found that heart rate metrics derived from smartwatches are associated with a wide range of health outcomes, extending well beyond cardiovascular conditions. These include mood regulation38,39, sleep quality40,41 and disorders42,43, chronic pain44, cognitive decline45, and even subtle health indicators such as low-grade fever and changes in resting heart rate46and HRV-derived stress.47,48 Previous studies have also shown that variations in heart rate, HRV, and resting heart rate can predict and detect responses to infections such as influenza, group A streptococcus, and COVID-19.24,49–51 By accounting for individuals’ overall health, lifestyle baselines and their physiological responses to diverse stimuli (including, but not limited to, infectious diseases), our approach enables a more efficient trial design across various clinical domains while potentially reducing the number of participants needed to achieve robust outcomes.

As an immediate practical application, our framework can be implemented into retrospective analyses of vaccination and healthcare technology assessments that rely on matching approaches. Current matching methods in retrospective research typically rely on demographics and static health measures, overlooking dynamic physiological variation at the individual level.52–54 This limitation can leave meaningful confounding unaddressed. SIM introduces physiologic matching that captures individual-level variation. When continuous physiological data are available, SIM can substantially improve covariate balance in retrospective cohorts and strengthen causal inference, particularly in comparative effectiveness research, where physiologic state may influence both treatment selection and outcomes.

Due to the limited number of participants in our study, we focused on heart measures. However, given the rich, diverse, and continuous data generated by various biosensors in smartwatches, we believe our study represents just the tip of the iceberg in what can be achieved for more accurate matching of individuals in clinical trials. Specifically, in clinical trials focused on infectious diseases, smartwatch biosensors can provide insights into lifestyle factors such as step counts, proximity to other Bluetooth devices, and social interaction patterns, which may help predict individuals who are more prone to becoming infected. Other smartwatch biosensors, such as those that measure respiratory rate, oxygen level, blood pressure, and sleeping patterns, could be valuable in identifying individuals at greater risk for complications if infected.55,56 Such information could be used to create even more accurate participant matching than we have considered.

Our study has several limitations. Participants were recruited via social media and word-of-mouth, resulting in a convenience sample that may not fully represent the general population. The demanding requirements – wearing a smartwatch and completing daily questionnaires for two years – could have discouraged certain individuals from participating. As highlighted by recent literature, device-based studies are susceptible to self-selection biases, as not all demographic groups are equally willing or able to adopt smartwatches.57 Notably, approximately one in four U.S. adults (25%) owns a smartwatch, with adoption extending across diverse demographic and socioeconomic groups.58,59 Furthermore, even among willing participants, daily wear time varies.60 While we mitigated missing data by requiring a baseline wear-time threshold and applying linear imputation, excluding highly non-compliant individuals may bias the cohort toward those with higher baseline health engagement.60 Additionally, COVID-19 is known to be associated with heart-related issues, which may have contributed to the effectiveness of our smartwatch matching approach based on heart rate and HRV. Furthermore, the Garmin smartwatches worn by trial participants are not medical-grade devices and may not fully capture all relevant physiological changes. Moreover, the current implementation assumes simultaneous availability of all participants, whereas clinical trials typically enroll individuals sequentially. In prospective settings, SIM could be applied in batches of newly recruited participants or integrated into “matching-on-the-fly” frameworks, such as Sequential Matched Randomization.61 Finally, while this analysis demonstrated the utility of SIM in the context of communicable diseases, the generalizability of these findings to other types of clinical trials remains to be established.

In conclusion, our study demonstrates the potential of smartwatch-derived physiological data to improve clinical trial design by shifting from traditional static matching methods to a dynamic, data-driven approach. By integrating continuous metrics such as heart rate and HRV, our method can reduce sample size requirements while maintaining statistical power, compared to traditional age- and sex-based matching methods. This marks a potentially transformative step in clinical research, where wearable technology can enable improved participant matching in clinical trials, as well as real-time, objective monitoring of health and disease progression, enhancing the efficiency, accuracy, and personalization of clinical trials.

Supplementary Information

Below is the link to the electronic supplementary material.

Supplementary Material 1 (772.9KB, docx)

Author contributions

Conception and design: DY and ES. Collection and assembly of data: DY, MY, SS, and ES had access to the raw data and were responsible for verifying the data. Analysis and interpretation of the data: ES, SS, MY, DY, MP and MLB. Statistical analysis: ES and DY. Drafting the article: ES, MY, SS, DY, and MLB. Critical revision of the article for important intellectual content: all authors. Final approval of the article: All authors. Obtaining funding: DY.

Funding

This work was supported by the European Research Council, project #949850, and the Israel Science Foundation (ISF), grant No. 3409/19, within the Israel Precision Medicine Partnership program. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.

Data availability

According to this study’s Tel Aviv University IRB guidelines, no patient-level data is to be shared outside the permitted researchers. The statistical analysis code along with an aggregated version of the data sufficient to reproduce the study’s results is publicly available at https://github.com/eeddaann/Smartwatch-Guided-Participant-Matching.

Declarations

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Hariton, E. & Locascio, J. J. Randomised controlled trials - The gold standard for effectiveness research. BJOG125, 1716 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Gul, R. B. & Ali, P. A. Clinical trials: The challenge of recruitment and retention of participants. J. Clin. Nurs.19, 227–233 (2010). [DOI] [PubMed] [Google Scholar]
  • 3.Speich, B. et al. Systematic review on costs and resource use of randomized clinical trials shows a lack of transparent and comprehensive data. J. Clin. Epidemiol.96, 1–11 (2018). [DOI] [PubMed] [Google Scholar]
  • 4.Polack, F. P. et al. Safety and efficacy of the BNT162b2 mRNA Covid-19 vaccine. N Engl. J. Med.383, 2603–2615 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Wise, J. Covid-19: European countries suspend use of Oxford-AstraZeneca vaccine after reports of blood clots. BMJ372, n699 (2021). [DOI] [PubMed] [Google Scholar]
  • 6.Wallach, J. D., Sullivan, P. G., Trepanowski, J. F., Steyerberg, E. W. & Ioannidis, J. P. Sex based subgroup differences in randomized controlled trials: Empirical evidence from Cochrane meta-analyses. BMJ355, i5826 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Ioannidis, J. P. & Adami, H. O. Nested randomized trials in large cohorts and biobanks: Studying the health effects of lifestyle factors. Epidemiology19, 75–82 (2008). [DOI] [PubMed] [Google Scholar]
  • 8.Bauhoff, S. Systematic self-report bias in health data: Impact on estimating cross-sectional and treatment effects. Health Serv. Outcomes Res. Methodol.11, 44–53 (2011). [Google Scholar]
  • 9.Greevy, R., Lu, B., Silber, J. H. & Rosenbaum, P. Optimal multivariate matching before randomization. Biostatistics5, 263–275 (2004). [DOI] [PubMed] [Google Scholar]
  • 10.Kernan, W. N., Viscoli, C. M., Makuch, R. W., Brass, L. M. & Horwitz, R. I. Stratified randomization for clinical trials. J. Clin. Epidemiol.52, 19–26 (1999). [DOI] [PubMed] [Google Scholar]
  • 11.Lin, Y., Zhu, M. & Su, Z. The pursuit of balance: An overview of covariate-adaptive randomization techniques in clinical trials. Contemp. Clin. Trials. 45, 21–25 (2015). [DOI] [PubMed] [Google Scholar]
  • 12.Pocock, S. J. & Simon, R. Sequential treatment assignment with balancing for prognostic factors in the controlled clinical trial. Biometrics31103. (1975). [PubMed] [Google Scholar]
  • 13.Guan, G. et al. Higher sensitivity monitoring of reactions to COVID-19 vaccination using smartwatches. npj Digit. Med.5, 140 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Guk, K. et al. Evolution of wearable devices with real-time disease monitoring for personalized healthcare. Nanomaterials (Basel)10.3390/nano9060813 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Gupta, A. S., Patel, S., Premasiri, A. & Vieira, F. At-home wearables and machine learning sensitively capture disease progression in amyotrophic lateral sclerosis. Nat. Commun.14, 5080 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Perez, M. V. et al. Large-scale assessment of a smartwatch to identify atrial fibrillation. N Engl. J. Med.381, 1909–1917 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Radin, J. M. et al. Assessment of prolonged physiological and behavioral changes associated with COVID-19 infection. JAMA Netw. Open.4, e2115959 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Sotirakis, C. et al. Identification of motor progression in Parkinson’s disease using wearable sensors and machine learning. npj Parkinsons Dis.9, 142 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Straus, L. D. et al. Utility of wrist-wearable data for assessing pain, sleep, and anxiety outcomes after traumatic stress exposure. JAMA Psychiatry. 80, 220–229 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Wells, C. I. et al. Wearable devices to monitor recovery after abdominal surgery: Scoping review. BJS Open6, (2022). [DOI] [PMC free article] [PubMed]
  • 21.Levi, Y., Brandeau, M. L., Shmueli, E. & Yamin, D. Prediction and detection of side effects severity following COVID-19 and influenza vaccinations: Utilizing smartwatches and smartphones. Sci. Rep.14, 6012 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Hasty, F. et al. Heart rate variability as a possible predictive marker for acute inflammatory response in COVID-19 patients. Mil. Med.186, e34–e38 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Holt, S. G. et al. Monitoring skin temperature at the wrist in hospitalised patients may assist in the detection of infection. Intern. Med. J.50, 685–690 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Snir, S. et al. Changes in behavior and biomarkers during the diagnostic decision period for COVID-19, influenza, and group A streptococcus (GAS): A two-year prospective cohort study in Israel. Lancet. Reg. Health42, 100934 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Yechezkel, M. et al. Comparison of physiological and clinical reactions to COVID-19 and influenza vaccination. Commun. Med.4, 169 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Mofaz, M. et al. Self-reported and physiological reactions to the third BNT162b2 mRNA COVID-19 (booster) vaccine dose. Emerg. Infect. Dis.28, 1375–1383 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Yechezkel, M. et al. Safety of the fourth COVID-19 BNT162b2 mRNA (second booster) vaccine: A prospective and retrospective cohort study. Lancet. Respir. Med.11, 139–150 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Firstbeat Technologies Ltd. Stress and recovery analysis method based on 24-hour heart rate variability, https://assets.firstbeat.com/firstbeat/uploads/2015/11/Stress-and-recovery_whitepaper_20145.pdf (2014).
  • 29.Garmin. Vivosmart owner’s manual, https://www8.garmin.com/manuals/webhelp/vivosmart/EN-US/GUID-E0C28652-FD9C-4FF4-B859-0265F3FB696A.html (2018).
  • 30.Garmin Support Center. What is the stress level feature on my Garmin watch?https://support.garmin.com/en-US/?faq=WT9BmhjacO4ZpxbCc0EKn9 (2021).
  • 31.Google Patents. Procedure for detection of stress by segmentation and analyzing a heart beat signal. Patent number US20050256414A1, https://patents.google.com/patent/US20050256414A1/en (2005).
  • 32.Lauer, S. A. et al. The incubation period of coronavirus disease 2019 (COVID-19) from publicly reported confirmed cases: estimation and application. Ann. Intern. Med.172, 577–582 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Xu, Q. Y. et al. Response and duration of serum anti-SARS-CoV-2 antibodies after inactivated vaccination within 160 days. Front. Immunol.12, 786554 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.BioNTech, S. E. Study to describe the safety, tolerability, immunogenicity, and efficacy of RNA vaccine candidates against COVID-19 in healthy individuals. https://clinicaltrials.gov/study/NCT04368728 (2023).
  • 35.Morris, T. P., White, I. R. & Crowther, M. J. Using simulation studies to evaluate statistical methods. Stat. Med.38, 2074–2102 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Moore, T. J., Zhang, H., Anderson, G. & Alexander, G. C. Estimated costs of pivotal trials for novel therapeutic agents approved by the US Food and Drug Administration, 2015–2016. JAMA Intern. Med.178, 1451–1457 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Deloitte Centre for Health Solutions. Intelligent Clinical Trials: Transforming through AI-Enabled Engagement, https://www2.deloitte.com/content/dam/insights/us/articles/22934_intelligent-clinical-trials/DI_Intelligent-clinical-trials.pdf (2020).
  • 38.Bachmann, A. et al. Leveraging smartwatches for unobtrusive mobile ambulatory mood assessment. In 2015 ACM International Joint Conference on Pervasive and Ubiquitous Computing and 2015 ACM International Symposium on Wearable Computers. 1057–1062.
  • 39.Quiroz, J. C., Geangu, E. & Yong, M. H. Emotion recognition using smart watch sensor data: mixed-design study. JMIR Ment Health. 5, e10153 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Sano, A. et al. Recognizing academic performance, sleep quality, stress level, and mental health using personality traits, wearable sensors and mobile phones. In 2015 IEEE 12th International Conference on Wearable and Implantable Body Sensor Networks (BSN). 1–6. [DOI] [PMC free article] [PubMed]
  • 41.Sathyanarayana, A. et al. Sleep quality prediction from wearable data using deep learning. JMIR Mhealth Uhealth. 4, e125 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Penzel, T., Schobel, C. & Fietze, I. New technology to assess sleep apnea: wearables, smartphones, and accessories. F1000Res7, 413 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Xu, S. et al. A review of automated sleep disorder detection. Comput. Biol. Med.150, 106100 (2022). [DOI] [PubMed] [Google Scholar]
  • 44.Zheng, N. S. et al. Sleep patterns and risk of chronic disease as measured by long-term monitoring with commercial wearable devices in the All of Us Research Program. Nat. Med.30, 2648–2656 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Boletsis, C., McCallum, S. & Landmark, B. F. Human Aspects of IT for the Aged Population. Design for Everyday Life. (eds Zhou, J. & Salvendy, G.) (Springer, 2015).
  • 46.Quer, G., Gouda, P., Galarnyk, M., Topol, E. J. & Steinhubl, S. R. Inter- and intraindividual variability in daily resting heart rate and its associations with age, sex, sleep, BMI, and time of year: Retrospective, longitudinal cohort study of 92,457 adults. PLoS One15, e0227709 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Grosicki, G. J. et al. Self-recorded heart rate variability profiles are associated with health and lifestyle markers in young adults. Clin. Auton. Res.32, 507–518 (2022). [DOI] [PubMed] [Google Scholar]
  • 48.Santa-Rosa, F. A. et al. Impact of an active lifestyle on heart rate variability and oxidative stress markers in offspring of hypertensives. Sci. Rep.10, 12439 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Levi, Y. et al. Smartwatch-derived versus self-reported outcomes of physiological recovery after COVID-19, influenza, and group A streptococcus: A 2-year prospective cohort study. Lancet Digit. Health8, 100956 (2026). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Mishra, T. et al. Pre-symptomatic detection of COVID-19 from smartwatch data. Nat. Biomed. Eng.4, 1208–1220 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Radin, J. M. et al. Long-term changes in wearable sensor data in people with and without Long Covid. npj Digit. Med.7, 246 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Dagan, N. et al. BNT162b2 mRNA Covid-19 vaccine in a nationwide mass vaccination setting. N. Engl. J. Med.384, 1412–1423 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Tsanani, S. E. et al. Effectiveness of influenza vaccination in preventing severe COPD exacerbations and pneumonia before, during, and after the COVID-19 pandemic: A retrospective cohort study. Lancet Reg. Health53, 101307 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Yaron, S. et al. Incremental benefit of high dose compared to standard dose influenza vaccine in reducing hospitalizations. npj Vaccines10, 3 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Angelucci, A., Greco, M., Cecconi, M. & Aliverti, A. Wearable devices for patient monitoring in the intensive care unit. Intensive Care Med. Exp.13, 26 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Hanly, P. J. et al. The effect of oxygen on respiration and sleep in patients with congestive heart failure. Ann. Intern. Med.111, 777–782 (1989). [DOI] [PubMed] [Google Scholar]
  • 57.Tackney, M. S. et al. Digital endpoints in clinical trials: Emerging themes from a multi-stakeholder knowledge exchange event. Trials25, 521 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Pawar, P. Smartwatch Statistics By Brands, Revenue and Facts. https://electroiq.com/stats/smartwatch-statistics/ (2025).
  • 59.Slotta, D. Wearables - Statistics and Facts. https://www.statista.com/topics/1556/wearable-technology/?srsltid=AfmBOopEoRZdDg3FMzgirltSN2_N7Hi2xQzkMRMYk-AU1cYcF7lRfOdh (2025).
  • 60.Di, J. et al. Considerations to address missing data when deriving clinical trial endpoints from digital health technologies. Contemp. Clin. Trials113, 106661 (2022). [DOI] [PubMed] [Google Scholar]
  • 61.Chipman, J. J., Mayberry, L. & Greevy Jr., R. A. Rematching on-the-fly: Sequential matched randomization and a case for covariate-adjusted randomization. Stat. Med.42, 3981–3995 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material 1 (772.9KB, docx)

Data Availability Statement

According to this study’s Tel Aviv University IRB guidelines, no patient-level data is to be shared outside the permitted researchers. The statistical analysis code along with an aggregated version of the data sufficient to reproduce the study’s results is publicly available at https://github.com/eeddaann/Smartwatch-Guided-Participant-Matching.


Articles from Scientific Reports are provided here courtesy of Nature Publishing Group

RESOURCES