ABSTRACT
Early‐phase clinical trials focus heavily on demonstrating safety and tolerability of drug candidates. Clinical laboratory test results outside of reference intervals (RIs) are often interpreted as possible indicators of harm. However, values outside RIs do not necessarily constitute drug‐induced adverse effects. This retrospective analysis aims to position results outside RIs as common phenomena during clinical pharmacology trials. For this, a retrospective analysis was performed on 37 commonly used blood‐based hematology, chemistry and coagulation markers of 11.535 screened participants and patients at screening, and 888 placebo‐randomized healthy participants from 88 trials performed at the Centre for Human Drug Research between 2005 to 2025. On average, participants with RI excursions during trial participation already show more extreme values at screening and baseline. Also, after having been medically screened, values outside RIs are still commonly observed in placebo‐randomized participants. Results appear centered around the mean of individual participants. The frequency and extent of RI excursions is measurand‐specific but generally within 0.7–1.3 times the RI. More significant excursions over 0.5–2.0 times the RI are observed, including for liver transferases, but comprise only 1.5% of all excursions. Diurnal variation impacts some hematological measurands resulting in RI excursion especially around through times. To conclude, an individual predisposition exists towards more extreme values challenging the applicability of RIs in a longitudinal setting. Results during trial conduct can best be interpreted compared to individual baselines. RI excursions remain common in placebo‐randomized participants and are therefore poorly indicative of drug‐induced adverse events when not considered in their full clinical context.
Keywords: clinical trials, healthy participants, laboratory values, placebo
Study Highlights
-
What is the current knowledge on the topic?
Healthy participants for early‐phase clinical trials are selected, in part, based on acceptable laboratory values in line with reference intervals. During trials, these values are used to monitor participants for drug‐induced adverse effects.
-
What question did this study address?
This study aims to provide guidance on the interpretation and implementation of common laboratory results during clinical trials in healthy participants.
-
What does this study add to our knowledge?
Despite having been screened, placebo‐randomized participants commonly show values exceeding reference ranges throughout their participation. A predisposition exists towards more extreme values within participants. The occurrence of these events in participants receiving placebo indicates deviations are not necessarily drug‐induced and may result from within‐participant variation or external factors.
-
How might this change clinical pharmacology or translational science?
This retrospective analysis indicates that reference intervals applied to a longitudinal setting in healthy participants should be interpreted with less stringency. By offering retrospect and guidance on the use of reference intervals and the timing of measurements in trials, future trials can be conducted more effectively, resulting in less confusion and more reliable outcomes.
1. Introduction
Early‐phase clinical trials are strongly focused on the safety and tolerability of novel drugs in humans. Based on preclinical models and safety factors, appropriate dose levels are selected and thereafter evaluated in small groups of healthy participants pending further evaluation in patients when safe and meaningful doses have been identified [1]. However, drug attrition is high as only a fraction of all candidates tested in man become available for patients [2]. Accompanied with the disproportionate increase in development costs throughout successive stages of clinical development, a high‐risk environment is created where the overinterpretation of benign and insignificant events due to individual or laboratory variability could lead to unnecessary pauses and inefficient decision making during trial conduct, especially during first‐in‐human studies and dose escalation where a strong focus on true safety signals is warranted [3, 4].
Decisions on whether to continue the development of a drug candidate in any stage are heavily based on the emergence of adverse events and other safety signals at each evaluated dose level [5]. In this, standard laboratory tests offer a direct and essential reflection on many physiological processes [6]. In clinical practice, these tests are routinely used to help diagnose or treat patients with apparent medical conditions [7]. In contrast, laboratory testing in the context of early‐phase clinical trials is employed to exclude pre‐existing sub‐clinical conditions and to affirm eligibility of participants prior to enrollment into a study [8, 9]. Thereafter, laboratory testing is used repeatedly to monitor any possible effect when the study drug is administered. This enables investigators to readily observe deviating results that can be precursors of more serious injury. In the event of such safety signals, the investigator can implement mitigating strategies such as additional follow‐up tests, study drug interruption or adjustment, or pausing the trial to minimize the risk of harm to current and future participants.
Laboratory readouts span several domains such as cell counts, coagulation, basic homeostasis, renal function and liver function. In trials, liver function is particularly closely monitored as drug‐induced liver injury (DILI) is a relatively rare but important and dangerous adverse effect that can cause considerable harm to participants and can be responsible for drug failure during clinical development and even for withdrawal after market approval [10]. Besides safety, laboratory measurements can also offer an easy and accessible way to measure pharmacodynamic effects in trial participants or the effect of established drugs in patients [11, 12, 13]. Even when primarily used for pharmacodynamic purposes, the safety aspect remains strongly represented in monitoring off‐target effects or preventing cases of exaggerated pharmacology in which drug effects become too potent [14, 15]. This makes the definition of the acceptable range of measurand values, and therefore also the safe range, especially important.
Generally, laboratory results are flagged for clinical evaluation when they exceed predefined Lower and Upper Limits of Normal (LLN and ULN, respectively). These limits often represent both extremes of the 95% reference interval (RI), which is based on measurements obtained from an external healthy reference population [16]. With this, an estimated 5% of measurements will logically fall outside this RI that, when interpreted as the absolute limits, may lead to the wrongful interpretation that a participant is not considered healthy [17]. Additionally, RIs are commonly established using cross‐sectional data, which do not account for intra‐individual variation, even though this may be considered more relevant to a clinical trial setting where participants are tested repeatedly [18]. These factors should be considered when interpreting deviating laboratory results from healthy participants during clinical trials and classifying them as clinically significant.
Instead of focusing on the suitability of RIs for selecting healthy participants, we focus on their application in a longitudinal setting during early‐phase clinical studies after medical screenings have already been performed. A retrospective pooled analysis of placebo‐controlled trials is performed on data obtained over 20 years at a clinical pharmacology unit in the Netherlands in order to provide trialists with guidance and reference. By providing this retrospective overview of common results at screening and during placebo treatment in healthy participants, an exemplification of the usual longitudinal characteristics of laboratory results from trial participants can be established. This may position laboratory values outside of RIs as common phenomena of early‐phase clinical trials and steer discussion on when findings can be deemed clinically significant based on laboratory results alone.
2. Methods
2.1. Source Data
This retrospective analysis was performed on trial data gathered in the period of March 2005 to Januari 2025 at the Centre for Human Drug Research, Leiden, the Netherlands. Studies were performed conform the principles of the Declaration of Helsinki and after ethical approval. For the screening dataset, both healthy participants and patient groups are included. Only placebo‐controlled studies in healthy participants were included in the longitudinal dataset. All participants gave written informed consent prior to any study related procedures and underwent a medical screening prior to enrollment.
2.2. Measurements and Data Curation
Blood samples for hematology, chemistry and coagulation measurements were obtained by trained staff through standardized phlebotomy protocols with date and clocktime recorded. Measurements were performed at the clinical chemistry lab of the Leiden University Medical Centre (LUMC), Leiden, the Netherlands. Data was extracted from the Promasys database (Anju Software, Leiden, the Netherlands). Non‐fasted RIs were used as fasting state could not be confirmed. RIs were retrieved from the 10‐year historical limits of the LUMC KCL Laboratory, taking the widest range in the case of multiple RIs in that period, and subsequently imposed on the dataset. RIs were additionally verified by performing a review of data superimposed with the 2.5‐ and 97.5‐percentiles of all measurements obtained at screenings. The RI for the International Normalized Ratio was retrieved from literature [19]. The applied RIs are shown in Table S1.
The 37 most frequently reported measurands were selected for analysis, combined into one dataset and a blind data review was subsequently performed per measurand to establish outliers and consistency of analysis. Only participants between 18 to 64 years, inclusive, were considered for analysis.
For screening, results from participants were included when the same measurand was measured both during the screening and when first reassessed up to maximally 100 days thereafter during participation in a trial. Positive selection was applied to remove any measurands with RI excursions at screening.
For longitudinal analysis, only participants randomized to placebo with at least 3 measurements after screening were included. Outliers were identified with a blind data review of all results per measurand, see supplemental methods. Trials were removed from the dataset in its entirety if outliers could be traced back to patient populations. Studies identified during this step were also excluded from the screening dataset. All final included study protocols were investigated to further exclude systemic challenge studies and patients.
Results of unscheduled samples are not included in any calculations and only graphed, if applicable (Figures 1, 2 and S2). Diurnal rhythm of measurands is determined using a cosinor‐fit using the clock time irrespective of sampling date.
FIGURE 1.

Differences in measurand value at screening and when re‐examined at enrollment in a trial. Measurands and participants are stratified in Cases (result within reference interval at screening, outside reference interval at start of the trial) and Normals (within reference interval at both screening and start of the trial). Average Z‐scores are shown in a radar plot per measurand (A). Additionally, the amount of time in days between the last measurement at screening and first measurement during the trial is shown (B). ALT: alanine aminotransferase, AST: aspartate aminotransferase, Gamma‐GT: gamma‐glutamyl transferase, HDL: high‐density lipoprotein, LDL: low‐density lipoprotein.
FIGURE 2.

Longitudinal overview of all hematology, chemistry and coagulation results in participants randomized to placebo for the first 8 days of the occasion. Graphs show the overall variation within measurands using the associated Z‐score (A) and the variation within a participant for each measurand (B). Values outside the reference intervals are indicated when above (high) or below (low) the reference intervals. Unscheduled events are indicated with triangles.
2.3. Data Analysis
Visualization and analysis were performed in R (version 4.5.0, R Core Team 2024). Normality and equal variance of the data were assessed using a Shapiro–Wilk and Levene's test, respectively. Statistical significance was tested using Mann–Whitney tests with Bonferroni correction.
3. Results
A total of 11.535 healthy participants and patients that were reassessed at the research unit within 100 days of screening were included in the screening dataset with a total of 303.981 measurands from 26.863 samples (11.343 chemistry, 11.295 hematology and 4.225 coagulation blood tubes) obtained during screening. Patients were excluded from the longitudinal dataset, resulting in a total of 888 placebo‐randomized healthy participants from 88 different trials with 22.896 measurands obtained ≥ 3 times from 9.606 total samples (3.840 chemistry, 3.905 hematology and 1.861 coagulation tubes), of which 219 tubes were obtained as unscheduled events. Demographics are presented in Table 1.
TABLE 1.
Basic demographics of the study participants in the dataset. Demographics are provided for the group of participants in the screening dataset with at least one measurand within reference interval, the screening dataset with participants without any RI excursions and the longitudinal placebo‐randomized dataset.
| Screening dataset | Screening dataset* | Longitudinal dataset | |
|---|---|---|---|
| Participants (n) | 11,535 | 2089 | 888 |
| Sex (n) | |||
| Male | 8335 | 1206 | 657 |
| Female | 3196 | 883 | 231 |
| Age (±sd, range) (years) | 29.97 (±11.97, 18–64) | 29.35 (±11.20, 18–64) | 30.12 (±11.30, 18–64) |
| Body Mass Index (±sd, range) (kg/m²) | 23.57 (±3.08, 16.0–44.5) | 23.32 (±2.85, 16.8–36.9) | 23.71 (±3.00, 17.7–34.6) |
Demographics corresponding to Figure 1B.
3.1. Screened Participants Show Reference Interval Excursions Upon Reexamination Irrespective of the Interval Between Screening and Trial Start
A set of 37 frequently measured measurands were compared at screening and when reassessed at trial start. Measurands are defined as Normal when within RI at both screening and re‐examination. Cases represent measurands that convert from within RIs at screening to outside RIs upon re‐examination. On a per‐measurand basis, only alkaline phosphatase, neutrophils, and lymphocytes were not significantly different between Case‐measurands and Normal‐measurands at screening (Figure 1A). Alkaline phosphatase, bicarbonate, and lymphocytes were not significantly different in Case‐measurands compared to Normal‐measurands at the occasion. Of note, no measurements for basophils or high‐density lipoprotein (HDL) were outside RIs. The proportion of cases per measurand remained below 10%, except for creatine phosphokinase (10.5%) (see also Table S2).
Strict limits are often imposed on markers for liver function. Liver function test that convert from within RI at screening to outside RI at reassessment constitute up to 4.5% of all measured results (alkaline phosphatase; 2.1%, ALT [alanine transaminase]; 2.7%, AST [aspartate aminotransferase]; 4.5%, total bilirubin; 3,8%, gamma‐glutamyl transferase; 0.7%).
Generalizing by using the absolute Z‐score as a measure of deviation from the mean, the average Z‐score of individual measurands is marginally higher at screening for Case‐measurands (0.60 ± 0.58) compared to Normal‐measurands (0.52 ± 0.47, p < 0.005) but, as a direct consequence of the positive selection, evidently higher at enrollment (Case‐measurands: 1.54 ± 1.64, Normal‐measurands: 0.53 ± 0.48, p < 0.005) (Figure 1B).
When focusing on whole participants instead of single measurands, only 2.089 of 11.535 participants do not show any flagged measurands during screening (Figure S1). Defining participants in this selected population as Normals or Cases when they have none or ≥ 1 flagged result at reassessment, respectively, the averaged Z‐score of all measurements within a participant remains significantly higher for Cases at both screening (Cases: 0.52 ± 0.12, Normals: 0.51 ± 0.14, p = 0.011) and at re‐examination (Cases: 0.62 ± 0.19, Normals: 0.53 ± 0.14, p < 0.005). However, the difference in averaged Z‐score at screening is negligible and insufficient to discriminate Normals and Cases in a practical setting.
Lastly, the number of days between both measurements was not significantly different between Cases (18.94 ± 13.15 days) and Normals (18.96 ± 13.68 days, p = 0.79) indicating that the incidence of results outside RIs does not increase with time (Figure 1B).
Of note, results remain comparable when limiting the dataset to studies included in the longitudinal analysis (Figure S2) with respect to individual measurands at screening and reassessment. However, when the Z‐score of all measurands is first averaged within participants of this group of 803 participants without RI excursions at screening, the difference is no longer significantly different between the 345 Normals (0.64 ± 0.18) and 458 Cases (0.65 ± 0.14, p = 0.22) at screening.
3.2. Flagged Values Are a Result of Consistent Increases Within Participants
During trial conduct, measurements are performed repeatedly in an already medically screened population. Therefore, the incidence of flagged results cannot readily be extrapolated from screening to trial participation. Visualizing their occurrence in placebo‐randomized participants is therefore most appropriate. RI excursions are frequently observed in healthy participants during the first seven days, with a roughly 24‐h interval reflecting the common sampling scheme of trials (Figure 2A). Flagged values typically deviate with a magnitude of up to 2.5 and 5 standard deviations from the mean for the LLN and ULN, respectively. However, excursions between 5 and 10 standard deviations are observed on occasion. More sporadic occurrences above 10 standard deviations are also present. Of note, high standard deviations associated with some measurands result in flagged values close to zero (Table S3).
After taking into consideration intra‐participant variation using the individual participant's mean instead of the population mean, the range of Z‐scores decreases significantly to within 5 standard deviations (Figure 2B). Flagged measurements with low Z‐scores, or even on opposite sides of the x‐axis, indicate that participants show results consistently around the RI limits. This also indicates that outliers are relatively rare and excursions of the RIs are often part of a pattern. Although high individual variation may result in high standard deviations and subsequently in relatively small Z‐scores which can therefore obscure outliers, coefficients of variation (CV%) highlight a smaller individual CV% compared to the population CV% for all measurands by a factor of 1.4 (for potassium) up to 7.7 (for Gamma‐glutamyl transferase) (Table S3).
Flagged results do not consistently originate from the same series of measurements. From all individual measurands followed over time, 3174 (13.9%) show at least one result outside RIs (Figure 3A). For 1612 of these, excursions were observed only once throughout follow‐up (643 below LLN, 970 above ULN), indicating RI excursions are often sporadic. Conversely, 141 measurands were always below LLN (including 34 hemoglobin, 32 erythrocytes, 18 hematocrit, 17 leucocytes and 14 creatinine) and 293 always above ULN (including 62 cholesterol, 38 albumin, 29 total bilirubin, 22 ALT and 20 LDL). Only one phosphate measurand showed both a result below ULN and above ULN. When focusing on the 1700 (7.4%) measurands within RI at baseline but with at least one low‐ or high‐flagged result during follow‐up, their mean Z‐score at baseline is already significantly lower when exceeding LLN (n = 803, −0.62 ± 0.77, p < 0.001) or higher when exceeding ULN (n = 896, 0.48 ± 0.82, p < 0.001) than measurands without any RI excursions at all (0.06 ± 0.83) (Figure 3B). However, distributions are strongly overlapping.
FIGURE 3.

Trajectories of measurands with at least one result outside of reference intervals (RI) when measurements are above upper limit of normal (ULN), within RI or below lower limit of normal (LLN) at a certain sample (A). Percentages indicate the fraction of samples included at each sample, as samples without any RI excursions are not included. Lines colors are maintained based on the flag at the first sample and maintained throughout. A distribution showing the baseline value for measurements that are within RI at baseline, and show at least one RI excursion below LLN, at least one RI excursion above ULN or without any RI excursions is presented (b).
Besides high‐density lipoprotein (HDL) and basophil counts, all measurands show results outside RIs with up to 40% for cholesterol levels on day 6 (Figure 4). Over the first 8 days together, 8 measurands show a percentage of flagged measurements of 10% or higher: cholesterol (27.4%), LDL (22.5%), albumin (22.2%), erythrocytes (15.3%), bicarbonate (14.2%), hemoglobin (12.7%), creatine phosphokinase (11.6%), and triglycerides (10.4%). Ultimately, 6.1% of all results are outside RIs (Table S5). Of note, differences in the frequency that measurands are assessed may skew results at some days for C‐reactive protein, LDL, fibrinogen, and cholesterol (Figure S4).
FIGURE 4.

Heat map showing the fraction of measurands that are outside of the reference interval as proportion of the total measurements performed for each 24‐h time block from zero point. The absolute number of measurements per day is shown in Figure S4.
3.3. Liver Function Tests in Placebo‐Randomized Participants
Flagged alanine aminotransferase (ALT) and aspartate aminotransferase (AST) measurements remain relatively low in the first 4 days, corresponding with the usual timeframe of confinement in early‐phase trials, and total up to 5.4% and 2.8% after 8 days, respectively (Figure 4). Overall, ALT and AST remained below 3 times ULN and total bilirubin below 2 times ULN in the first 7 days post‐dose precluding any Hy's Law cases (Figure S3). Trajectories of ALT, AST and bilirubin with at least one result exceeding ULN constituted 69 (8.7%), 93 (11.8%), 126 (16.2%) trajectories, respectively. Of which, 40 ALT, 36 AST and 126 bilirubin trajectories exceeded RI only once and 21 ALT, 20 AST and 35 bilirubin trajectories did so twice. Conversely, consistent RI excursions above the ULN were observed in 2 ALT, 22 AST and 29 bilirubin trajectories. Gamma‐glutamyl transferase exceeded ULN only in 28 (3.9%) of trajectories, of which 8 single excursions and 12 consistent excursions (data not shown).
3.4. The Incidence of Flagged Results Is Measurand Specific
The risk that screened, placebo‐randomized participants obtain a flagged result increases with time. Flagged results are observed for all measurands from the first pre‐dose measurement onward, except for HDL and basophil counts (Figure 5A). An initial steep increase in cumulative incidence of flagged values is observed throughout chemistry and coagulation markers, which may be partly attributed to sampling of participants with inherently more extreme values resulting in flagged results already at baseline. This increase plateaus after 14 days. Bicarbonate, cholesterol, LDL, and albumin exceed a cumulative incidence of 0.3 at 7 days. This is also observed for phosphate, triglycerides, and creatin phosphokinase after 31 days. On the other hand, excursions for lactate dehydrogenase, gamma‐glutamyl transferase, and the International Normalized Ratio (INR) remain low.
FIGURE 5.

Cumulative incidence plot of separate measurands with values outside the reference intervals over the first 31 days for chemistry and coagulation measurands (A) and hematology samples (B). A diurnal rhythm is shown for white blood cell differential counts (C) and the distribution of values below the lower limit of normal, within the reference interval, and above the upper limit of normal is show on a 24‐h timescale (D).
For hematology, the incidence of flagged results for erythrocyte counts, leukocyte counts, hemoglobin and hematocrit shows a similar incidence of 0.30–0.45 (Figure 5B). However, thrombocytes and differential blood counts seem less prone to deviations outside RIs. Substantial diurnal variation is observed in differential blood counts (Figure 5C). For leukocytes, a difference of 1.73 × 109/L is observed between 10:15 AM and 22:15 PM (Table S4). Neutrophils and lymphocytes show similar rhythms. In practice, this results in a decrease in counts when comparing evening and morning samples. Indeed, the distribution of LLN flags in leukocyte counts is skewed towards the hours between 08:00 AM and 11:00 AM which coincides with the lowest diurnal point (Figure 5D). In contrast, almost no LLN flags are observed after 16:00 PM while results above ULN are evenly distributed throughout the day. Correcting for diurnal variation by using the mesor, or average fluctuation during the day, the amount of LLN flags is reduced from 294 to 5 for leukocytes, from 52 to 1 for lymphocytes and from 78 to 45 for neutrophils. Conversely, ULN flags increase for leucocytes (46 to 120), lymphocytes (50 to 135) and monocytes [20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48] but remain similar for neutrophils [24 to 24]. The effect of diurnal correction on flagged basophils and eosinophils count was small to non‐existent.
3.5. Most Excursions Are Within 0.7–1.3 Times the Reference Interval
In respect to the participant instead of individual measurand, the incidence of participants with at least one measurement outside RIs is considerable already pre‐dose (Figure 6A). Within the first week, the cumulative incidence increases from 0.73 at zeropoint to 0.91 at 7 days, further increasing up to 0.96 after 31 days. The occurrence of participants experiencing at least one flagged result is effectively decreased when a range of up to 0.8–1.2 is used (7 days; 0.42, 31 days; 0.59) and plateaus at 0.5–1.5 (7 days; 0.20, 31 days; 0.30). While further extending this range symmetrically results in a dysfunctional LLN, extending the range to 0.5–2.0 shows a final reduced incidence of 0.06 and 0.13 at 7 and 31 days, respectively. The frequency and extent of RI excursions is, as with their cumulative incidence, measurand specific (Table S5). RI excursions outside 0.7–1.3 account for 10.3% of total RI excursions and are not observed for albumin, aPTT, basophils, bicarbonate, calcium, chloride, creatinine, erythrocytes, HDL, hematocrit, hemoglobin, INR, phosphate, potassium, PT, sodium and uric acid. Focusing on the number of flags per participant, the cumulative incidence is above 0.5 at day 7 for up to 5 flags and increases to at least 0.79 after 31 days (Figure 6B). The final incidence at 31 days remains high and comparable up to 5 flags per participant, but plateaus at an increasingly later moment. This indicates that not all 5 flags are obtained at the same time but spread out over different measurements. The cumulative incidence is considerably decreased from > 15 flagged measurements with 0.09 and 0.32 after 7 and 31 days, respectively. Ultimately, only 97 of the 888 participants show no flagged measurands after 31 days (Figure 6C). A total of 3 flags per participant is most commonly observed and the average number of flags per participant is 5.50 ± 4.97 and 6.80 ± 6.47 flags after 7 and 31 days, respectively.
FIGURE 6.

The cumulative incidence of participants with at least one reference interval excursion is shown during the first 31 days for the standard reference range (1.0) and extended up to 0.5 times the lower limit of normal and 2.0 times the upper limit of normal (0.5–2.0) (A). Additionally, the risk of participants obtaining 1 up to 30 flagged measurements (B) and the number of participants with that number of flags.
4. Discussion
Blood laboratory results outside RIs are regular events during early‐phase clinical trials. In this retrospective analysis spanning 20 years of clinical chemistry data obtained in the same laboratory, we demonstrate that these prevalently occur in a group of 888 placebo‐treated participants throughout almost all commonly measured measurands, including hepatic enzymes, regardless of medical screening and with hematological measurands impacted by individual predisposition and diurnal variation. With this observation, these treatment unrelated events should be considered just as likely to occur in the active arm of clinical trials. Indeed, these trials frequently start with the evaluation of low or subtherapeutic doses due to the high safety factors applied where no pharmacological effect is expected [15, 20].
Already at screening, results exceeding the RI are apparent which is expected as this range comprises only 95% of normal laboratory results [17]. In trials with stringent criteria surrounding laboratory values, these excursions may prevent a participant from enrolling [8, 21, 22]. Strictly adhering to RIs may therefore result in the wrongful exclusion of participants. This is supported by the presence of results outside RIs for medically screened participants at the start of a trial even when those were within RIs during screening, possibly due to intra‐participant variation [23, 24]. This is especially relevant for liver function tests that are often tightly controlled in early‐phase trials. Using the standard RI as absolute limit would result in the exclusion of up to 4.5% of participants based on the prevalence of AST results that are outside RI upon reassessment, significantly impacting trial enrollment. However, the European Medicines Agency explicitly states that justified deviations should be possible [25]. Having committed to a medical screening and subsequent trial participation may be sufficient justification when deviations are not clinically significant and assure more straightforward enrollment of participants that exhibit flagged values when reassessed. Unfortunately, shortening the screening window does not appear to limit this as the interval between participants with values within and outside RIs does not significantly differ. Indeed, more time between repeated measurements does not necessarily increase variability [26]. Of note, patients are also included in the screening dataset. However, participants with RI excursions at screening were excluded which effectively limits the dataset to participants, including patients, with normal laboratory values. Additionally, comparable results are obtained when solely participants from studies included in the longitudinal analysis are evaluated which further indicates the appropriateness of the chosen analysis population.
Even amongst successfully screened and enrolled healthy participants, flagged results remain prevalent during the first 8 days of participation. In this period after dosing, maximum drug plasma concentrations are frequently achieved which are inherently connected to the likelihood of adverse events [27]. Though values far outside RIs are rare, in small cohort sizes they may appear significant and cause for alarm. Elevations of hepatic enzymes in placebo‐randomized participants have been observed in phase 1 trials [28, 29]. Besides factors such as possible alcohol intake, exercise or viral infections, these can be partly attributed to intra‐subject variation as baseline values appear important. This is especially probable for subjects that already demonstrate slight deviations from the mean at earlier timepoints [30, 31]. Additionally, participants with persistent mild elevations of bilirubin as a result of Gilbert's syndrome may be enrolled as they comprise a substantial proportion of the healthy population [32]. Indeed, variation significantly decreased when measurements are averaged within the individual participant. Focusing on intra‐participant variation might therefore be more applicable for trials where serial testing is performed as it can reveal clinically significant changes from a prior result that remain within RIs and would go unnoticed when only focusing on RI excursions [33]. However, the feasibility of this approach on an ongoing basis during trials is limited as solely screening and baseline measurements are present at the start of a study. Herein, adaptation of a Bayesian approach where personalized RIs are estimated from historical or external data and finetuned as participant data becomes available may prove effective [24, 34].
Similarly, diurnal fluctuation can contribute to variation between timepoints at different times of the day. Although ubiquitous throughout many laboratory measurands, these are especially evident for leucocyte counts [35, 36]. Even though few measurements are obtained during nighttime in this population, which may confound actual peak and trough times, the diurnal fluctuation in this study is in good agreement when purposely determined within participants over 48 h [35]. In practice, an apparent change in leucocyte count when comparing samples taken at different times of day may be incorrectly attributed to a drug effect. This is especially relevant for immunomodulatory compounds where (off‐target) effects on cell counts may be expected based on the drug's mechanism of action [37]. Collecting samples at consistent times of day is advisable to minimize the impact of diurnal variation, with additional stratification between active and placebo arms during interim reporting to preclude concomitant drug‐induced effects. Doing so may additionally help reveal drug effects that are otherwise harder to identify [38].
Ultimately, excursions of the RI are observed for many measurands and can be drastically reduced by extending the RI with a factor 0.7–1.3. Conversely, more stringent criteria are often used for first‐in‐human trials for relevant measurands such as liver transferases [39]. Although slightly elevated transferases at baseline are not necessarily associated with a higher susceptibility of liver injury, especially in healthy participants demonstrating prior results within RIs, it may mask drug‐induced transferase elevations [40]. However, commonly accepted predictors of DILI such as Hy's Law are defined by severe elevations of hepatic enzymes in the absence of alternate causality instead of slight RI excursions and, together with follow‐up of individual cases and their overall incidence in drug‐exposed participants, constitute the main signal for the assessment of DILI [41]. No cases of Hy's law were observed in this analysis.
The still relatively high prevalence of results outside the 0.7–1.3 RI in placebo‐randomized participants, also for hepatic enzymes, indicates that false positives regularly occur. However, this is heavily dependent on individual measurands as the prevalence of INR and basophil counts outside RIs is extremely rare and therewith more cause for alarm when observed. Though INR was therefore considered a robust indicator of coagulopathy, it may also indicate a high degree of insensitivity [32]. Additionally, extended RIs can only be implemented when both extremes remain within safe ranges as this approach would, for example, push the upper limit of potassium into severe hyperkalemia of 6.8 mmol/L [42]. Ultimately, the difference in extent of RI excursions varies between measurands, which necessitates that their significance is always placed in the context of typically observed RI excursions of the individual measurand taking any physiological implications into consideration.
By focusing solely on laboratory results, an incomplete overview of the health status of a participant is obtained. In practice, deviating laboratory results are supplemented by context and other clinical assessments prior to interpretation by medical staff [43]. This holistic clinical judgment should outweigh that of solely laboratory results. Though mild RI excursions may not be reason for alarm, RIs serve as important thresholds and excursions deserve attention by the investigator irrespective of the severity of these excursions. To support operational feasibility, broader clinical‐discretion RIs could be defined within which excursions should not trigger automatic discontinuation lest there is sufficient accompanying rationale.
Many external factors may contribute to flagged results and the frequency thereof. On top of possible non‐compliance [44], lifestyle restrictions differ in this retrospective analysis with not all studies barring physical exercise, requiring dietary restrictions or imposing multiple‐day confinement during participation. These constitute important preanalytical variables [45, 46, 47]. In this retrospective analysis, samples were obtained from the same Dutch population at the same research unit and analyzed by the same analytical laboratory that defined the RIs, both solidifying results but also limiting extrapolation to other populations that differ in demographics or laboratories that may have less consistent assays. A blind data review was performed to exclude any inconsistencies in this dataset and to exclude outliers due to common preanalytical errors such as undue clotting and hemolysis, if still reported [48]. This monocenter experience, though typical for early‐phase studies, may vary from multicenter settings as aspects such as sample transport can negatively affect results and may further impact measurement validity [6].
To conclude, laboratory values outside RIs commonly occur in early‐phase clinical trials. Their prevalence in medically screened and placebo‐treated subjects indicates they can occur independently of active treatment, which is also observed for adverse events in general [29, 49, 50]. This should be taken into consideration during the conduct of trials. Ultimately, the complete clinical context of individual participants and those of other participants in the same trial, rather than individual laboratory results, should remain central in dose escalation and trial continuity.
Author Contributions
Jannik Rousel, Tessa Niemeyer‐van der Kolk, Matthijs Moerland, Robert Rissmann, Naomi B. Klarenbeek and Jacobus Burggraaf wrote the manuscript. Jannik Rousel designed the research. Jannik Rousel performed the research. Jannik Rousel analyzed the data.
Funding
The authors have nothing to report.
Conflicts of Interest
The authors declare no conflicts of interest.
Supporting information
Table S1: Reference intervals for the measurands included in analysis. Unfasted reference ranges are applied. Relevant sex is indicated if reference intervals differ between man and women. An asterisk (*) indicates that the reference range is extended to comprise the widest range in effect.
Table S2: Overview of the results at screening and occasion for normals (within reference interval at screening and occasion) and cases (within reference interval at screening, outside reference interval at occasion). Statistical significance between cases and normals are determined at screening and occasion separately. Units are included in Table S2.
Figure S1: Average Z‐scores are shown when all measurands from Figure 1A are averaged together for Normal‐ and Case‐measurands (A). Alternatively, the average of all measurands obtained at screening or the start of the trial is determined for participants in which cases are defined as individuals with at least one measurand result outside the reference interval (B).
Figure S2: Differences in measurand value at screening and when re‐examined at enrollment in a trial with only for studies included in the longitudinal analysis. Measurands (n = 126.006) and participants (n = 4.472) are stratified in Cases (result within reference interval at screening, outside reference interval at start of the trial) and Normals (within reference interval at both screening and start of the trial). Average Z‐scores are shown in a radar plot per measurand (A). Significance between Normal‐ and Case‐measurands at screening and first re‐assessment during trial participation is indicated below the measurand name (significance as screening/significance at re‐assessment) with p > 0.05; n.s., p ≤ 0.05; *, p < 0.01; **, p < 0.001; *** and na; not applicable. Additionally, the average z‐score for all Normal and Case measurands at screening (Normals: 0.64 ± 0.50, Cases: 0.79 ± 0.61, p < 0.001) or enrollment (Normals: 0.66 ± 0.51, Cases: 1.99 ± 1.72, p < 0.001) are shown (B). Alternatively, the average of all measurands obtained at screening or the start of the trial is determined for the 803 participants in which Cases are defined as individuals with at least one measurand result outside the reference interval at the occasion, demonstrating significantly increased z‐scores for the 458 Cases compared to the 345 Normals at enrollment (Normals: 0.66 ± 0.17, Cases: 0.76 ± 0.20, p < 0.001) but not at screening (Normals: 0.64 ± 0.18, Cases: 0.65 ± 0.14, p = 0.220) (C). Lastly, the amount of time in days between the last measurement at screening and first measurement during the trial is shown which does not significantly differ between Normals and Cases (16.65 ± 9.66 compared to 17.17 ± 10.54 days, respectively, p = 0.62) (D). ALT; alanine aminotransferase, AST; aspartate aminotransferase, Gamma‐GT; gamma‐glutamyl transferase, HDL; high‐density lipoprotein, LDL; low‐density lipoprotein.
Table S3: the mean, standard deviation (SD), coefficient of variation (CV, represented by the SD divided by the mean) of the entire population (CVpop) or when the CV is determined within a participant and subsequently averaged throughout the population (CVind).
Figure S3: Overview of alanine aminotransferase (ALT, panel A), aspartate aminotransferase (AST, panel B) and bilirubin (panel C) over time and their trajectories. Values exceeding the reference intervals are highlighted in red. The 1.5 and 3.0 times upper limit of normal (ULN) is shown, stratified for male and females if applicable. Both increases of ALT were related to participants with possible viral infections. Results from unscheduled samples are only included for the results over time (left) and are indicated with a triangle. Unscheduled samples are not included for the trajectories (right).
Figure S4: Heatplot showing the proportion of results that are outside the reference interval per measurand per day. The number of results outside of reference intervals and total results obtained within 24‐h blocks from zeropoint are shown.
Table S4: Parameters for the diurnal rhythm of white blood cell counts. Time is given in decimals.
Table S5: Listing of total (flagged) measurements of all parameters from day −1 to day 7, corresponding to Figures 3 and S3. Percentages are given as the part of flagged measurements from the total amount of measurements obtained.
Acknowledgments
The authors are grateful for all that participated in the studies at the Centre for Human Drug Research.
References
- 1. Wright B., “Clinical Trial Phases,” in A Comprehensive and Practical Guide to Clinical Trials, vol. 11–15 (Elsevier, 2017), 10.1016/B978-0-12-804729-3.00002-X. [DOI] [Google Scholar]
- 2. Waring M. J., Arrowsmith J., Leach A. R., et al., “An Analysis of the Attrition of Drug Candidates From Four Major Pharmaceutical Companies,” Nature Reviews. Drug Discovery 14 (2015): 475–486. [DOI] [PubMed] [Google Scholar]
- 3. Williams R. J., Tse T., DiPiazza K., and Zarin D. A., “Terminated Trials in the ClinicalTrials.Gov Results Database: Evaluation of Availability of Primary Outcome Data and Reasons for Termination,” PLoS One 10 (2015): e0127242. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4. de Visser S. J. and Cohen A. F., “Question‐Based Drug Development and the Value of Novel Treatments,” Expert Review of Pharmacoeconomics & Outcomes Research 25 (2024): 275–278. [DOI] [PubMed] [Google Scholar]
- 5. Cutler N. R., Sramek J. J., Greenblatt D. J., et al., “Defining the Maximum Tolerated Dose: Investigator, Academic, Industry and Regulatory Perspectives,” Journal of Clinical Pharmacology 37 (1997): 767–783. [DOI] [PubMed] [Google Scholar]
- 6. Lippi G., Simundic A. M., Rodriguez‐Manas L., Bossuyt P., and Banfi G., “Standardizing In Vitro Diagnostics Tasks in Clinical Trials: A Call for Action,” Annals of Translational Medicine 4 (2016): 181. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. Lippi G., Guidi G. C., and Plebani M., “One Hundred Years of Laboratory Testing and Patient Safety,” Clinical Chemistry and Laboratory Medicine 45 (2007): 797–798. [DOI] [PubMed] [Google Scholar]
- 8. Breithaupt‐Groegler K., Coch C., Coenen M., et al., “Who Is a ‘Healthy Subject’?‐Consensus Results on Pivotal Eligibility Criteria for Clinical Trials,” European Journal of Clinical Pharmacology 73 (2017): 409–416. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9. Devine E. G., Waters M. E., Putnam M., et al., “Concealment and Fabrication by Experienced Research Subjects,” Clinical Trials 10 (2013): 935–948. [DOI] [PubMed] [Google Scholar]
- 10. Hunt C. M., Papay J. I., Edwards R. I., et al., “Monitoring Liver Safety in Drug Development: The GSK Experience,” Regulatory Toxicology and Pharmacology 49 (2007): 90–100. [DOI] [PubMed] [Google Scholar]
- 11. Hallworth M. J., Epner P. L., Ebert C., et al., “Current Evidence and Future Perspectives on the Effective Practice of Patient‐Centered Laboratory Medicine,” Clinical Chemistry 61 (2015): 589–599. [DOI] [PubMed] [Google Scholar]
- 12. Cai Y., Chai D., Falagas M. E., et al., “Immediate Hematological Toxicity of Linezolid in Healthy Volunteers With Different Body Weight: A Phase I Clinical Trial,” Journal of Antibiotics 65, no. 4 (2012): 175–178. [DOI] [PubMed] [Google Scholar]
- 13. De Kam P. J., El Galta R., Kruithof A. C., et al., “No Clinically Relevant Interaction Between Sugammadex and Aspirin on Platelet Aggregation and Coagulation Parameters,” International Journal of Clinical Pharmacology and Therapeutics 51 (2013): 976–985. [DOI] [PubMed] [Google Scholar]
- 14. Ferreira G. S., Dijkstra F. M., Veening‐Griffioen D. H., et al., “Translatability of Preclinical to Early Clinical Tolerable and Pharmacologically Active Dose Ranges for Central Nervous System Active Drugs,” Translational Psychiatry 13, no. 1 (2023): 1–12. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15. Shen J., Swift B., Mamelok R., Pine S., Sinclair J., and Attar M., “Design and Conduct Considerations for First‐In‐Human Trials,” Clinical and Translational Science 12 (2019): 6–19. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. Ozarda Y., “Reference Intervals: Current Status, Recent Developments and Future Considerations,” Croatian Society for Medical Biochemistry and Laboratory Medicine 26, no. 5 (2016): 5–16. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17. Den Elzen W. P. J., van Gerven J., Schenk P. W., et al., “How to Define Reference Intervals to Rule in Healthy Individuals for Clinical Trials?,” Clinical Chemistry and Laboratory Medicine 55 (2017): e59–e61. [DOI] [PubMed] [Google Scholar]
- 18. Brombacher P. J., Wersch J. W. J., and Bas B. M., “Short‐Term and Long‐Term Intra‐Individual Variations and Critical Differences of Haematological Laboratory Parameters,” Clinical Chemistry and Laboratory Medicine 23 (1985): 69–76. [PubMed] [Google Scholar]
- 19. Chen H., Ahmad G., Colasurdo M., et al., “Mildly Elevated INR Is Associated With Worse Outcomes Following Mechanical Thrombectomy for Acute Ischemic Stroke,” Journal of NeuroInterventional Surgery 15 (2022): e117. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20. Baldrick P., “Getting a Molecule Into the Clinic: Nonclinical Testing and Starting Dose Considerations,” Regulatory Toxicology and Pharmacology 89 (2017): 95–100. [DOI] [PubMed] [Google Scholar]
- 21. Li H., Liu Y., He Y., et al., “A Retrospective Study: Screening Failure Analysis of 1,058 Healthy Volunteers in Phase I Clinical Trials,” Annals of Palliative Medicine 11 (2022): 2464–2477. [DOI] [PubMed] [Google Scholar]
- 22. Chalasani N., Hayashi P. H., Luffer‐Atlas D., Regev A., and Watkins P. B., “Assessment of Liver Injury Potential of Investigational Medicines in Drug Development,” Hepatology (2025), 10.1097/HEP.0000000000001281. [DOI] [PubMed] [Google Scholar]
- 23. Ricós C., Cava F., García‐Lario J. V., et al., “The Reference Change Value: A Proposal to Interpret Laboratory Reports in Serial Testing Based on Biological Variation,” Scandinavian Journal of Clinical and Laboratory Investigation 64 (2004): 175–184. [DOI] [PubMed] [Google Scholar]
- 24. Clayton G. L., Schachter A. D., Magnusson B., Li Y., and Colin L., “How Often Do Safety Signals Occur by Chance in First‐In‐Human Trials?,” Clinical and Translational Science 11 (2018): 471–476. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25. European Medicines Agency , “C. for M. P. for H. U. (CHMP). Guideline on Strategies to Identify And Mitigate Risks for First‐In‐Human and Early Clinical Trials With Investigational Medicinal Products,” (2017), https://www.ema.europa.eu/en/about‐us/contacts‐european‐medicines‐agency/send‐question‐european‐medicines‐agency. [DOI] [PMC free article] [PubMed]
- 26. Ricós C., Alvarez V., Cava F., et al., “Integration of Data Derived From Biological Variation Into the Quality Management System,” Clinica Chimica Acta 346 (2004): 13–18. [DOI] [PubMed] [Google Scholar]
- 27. van Gerven J. and Cohen A., “Integrating Data From the Investigational Medicinal Product Dossier/Investigator's Brochure. A New Tool for Translational Integration of Preclinical Effects,” British Journal of Clinical Pharmacology 84 (2018): 1457–1466. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28. Rosenzweig P., Miget N., and Brohier S., “Transaminase Elevation on Placebo During Phase I Trials: Prevalence and Significance,” British Journal of Clinical Pharmacology 48 (1999): 19–23. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29. Jung D., Braun I. V., and Wensing G., “Safety of Healthy Subjects in First‐In‐Human Multiple‐Dose Studies: A Pooled Analysis,” International Journal of Clinical Pharmacology and Therapeutics 59 (2020): 1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30. Cai Z., Christianson A. M., Ståhle L., and Keisu M., “Reexamining Transaminase Elevation in Phase I Clinical Trials: The Importance of Baseline and Change From Baseline,” European Journal of Clinical Pharmacology 65 (2009): 1025–1035. [DOI] [PubMed] [Google Scholar]
- 31. Sibille M., Deigat N., Durieu I., et al., “Laboratory Data in Healthy Volunteers: Reference Values, Reference Changes, Screening and Laboratory Adverse Event Limits in Phase I Clinical Trials,” European Journal of Clinical Pharmacology 55 (1999): 13–19. [DOI] [PubMed] [Google Scholar]
- 32. Deiteren A., Coenen E., Lenders S., Verwilst P., Mannaert E., and Rasschaert F., “Data Driven Evaluation of Healthy Volunteer Characteristics at Screening for Phase I Clinical Trials to Inform on Study Design and Optimize Screening Processes,” Clinical and Translational Science 14 (2021): 2450–2460. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33. De MacEdo D. V., Nunes L. A. S., and Brenzikofer R., “Reference Change Values of Blood Analytes From Physically Active Subjects,” European Journal of Applied Physiology 110 (2010): 191–198. [DOI] [PubMed] [Google Scholar]
- 34. Badrick T., “Biological Variation: Understanding Why It Is So Important?,” Practical Laboratory Medicine 23 (2021): e00199. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35. Ackermann K., Revell V. L., Lao O., Rombouts E. J., Skene D. J., and Kayser M., “Diurnal Rhythms in Blood Cell Populations and the Effect of Acute Sleep Deprivation in Healthy Young Men,” Sleep 35 (2012): 933–940. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36. Pocock S. J., Ashby D., Shaper A. G., Walker M., and Broughton P. M. G., “Diurnal Variations in Serum Biochemical and Haematological Measurements,” Journal of Clinical Pathology 42 (1989): 172–179. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37. Radanovic I., Klarenbeek N., Rissmann R., et al., “Integration of Healthy Volunteers in Early Phase Clinical Trials With Immuno‐Oncological Compounds,” Frontiers in Oncology 12 (2022): 954806. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38. Hassing G. J., van Esdonk M. J., van Westen G. J. P., Cohen A. F., Burggraaf J., and Gal P., “Dose Escalations in Phase I Studies: Feasibility of Interpreting Blinded Pharmacodynamic Data,” British Journal of Clinical Pharmacology 88 (2022): 5412–5419. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39. Sibille M., Patat A., Caplain H., and Donazzolo Y., “A Safety Grading Scale to Support Dose Escalation and Define Stopping Rules for Healthy Subject First‐Entry‐Into‐Man Studies: Some Points to Consider From the French Club Phase I Working Group,” British Journal of Clinical Pharmacology 70 (2010): 736–748. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40. Chalasani N. and Regev A., “Drug‐Induced Liver Injury in Patients With Preexisting Chronic Liver Disease in Drug Development: How to Identify and Manage?,” Gastroenterology 151 (2016): 1046–1051. [DOI] [PubMed] [Google Scholar]
- 41. Lewis J. H., “‘Hy's Law,’ the ‘Rezulin Rule,’ and Other Predictors of Severe Drug‐Induced Hepatotoxicity: Putting Risk‐Benefit Into Perspective,” Pharmacoepidemiology and Drug Safety 15 (2006): 221–229. [DOI] [PubMed] [Google Scholar]
- 42. An J. N., Lee J. P., Jeon H. J., et al., “Severe Hyperkalemia Requiring Hospitalization: Predictors of Mortality,” Critical Care 16 (2012): R225. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43. Charney A. N. and Dourmashkin J. T., “Interpreting Clinical and Laboratory Tests: Importance and Implications of Context,” Diagnosis (Berl) 8 (2019): 33–36. [DOI] [PubMed] [Google Scholar]
- 44. Dresser R., “Subversive Subjects: Rule‐Breaking and Deception in Clinical Trials,” Journal of Law, Medicine & Ethics 41 (2013): 829–840. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45. Plumelle D., Lombard E., Nicolay A., and Portugal H., “Influence of Diet and Sample Collection Time on 77 Laboratory Tests on Healthy Adults,” Clinical Biochemistry 47 (2014): 31–37. [DOI] [PubMed] [Google Scholar]
- 46. Sanchis‐Gomar F. and Lippi G., “Physical Activity ‐ an Important Preanalytical Variable,” Biochem. Med. (Zagreb) 24 (2014): 68. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47. Danielsson J., Kangastupa P., Laatikainen T., Aalto M., and Niemelä O., “Impacts of Common Factors of Life Style on Serum Liver Enzymes,” World Journal of Gastroenterology : WJG 20 (2014): 11743–11752. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48. Giavarina D. and Lippi G., “Blood Venous Sample Collection: Recommendations Overview and a Checklist to Improve Quality,” Clinical Biochemistry 50 (2017): 568–573. [DOI] [PubMed] [Google Scholar]
- 49. Emanuel E. J., Bedarida G., Macci K., Gabler N. B., Rid A., and Wendler D., “Quantifying the Risks of Non‐Oncology Phase I Research in Healthy Volunteers: Retrospective Analysis of Phase I Studies,” BMJ 350 (2015): h3271. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50. Lutfullin A., Kuhlmann J., and Wensing G., “Adverse Events in Volunteers Participating in Phase I Clinical Trials: A Single‐Center Five‐Year Survey in 1,559 Subjects,” International Journal of Clinical Pharmacology and Therapeutics 43 (2005): 217–226. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Table S1: Reference intervals for the measurands included in analysis. Unfasted reference ranges are applied. Relevant sex is indicated if reference intervals differ between man and women. An asterisk (*) indicates that the reference range is extended to comprise the widest range in effect.
Table S2: Overview of the results at screening and occasion for normals (within reference interval at screening and occasion) and cases (within reference interval at screening, outside reference interval at occasion). Statistical significance between cases and normals are determined at screening and occasion separately. Units are included in Table S2.
Figure S1: Average Z‐scores are shown when all measurands from Figure 1A are averaged together for Normal‐ and Case‐measurands (A). Alternatively, the average of all measurands obtained at screening or the start of the trial is determined for participants in which cases are defined as individuals with at least one measurand result outside the reference interval (B).
Figure S2: Differences in measurand value at screening and when re‐examined at enrollment in a trial with only for studies included in the longitudinal analysis. Measurands (n = 126.006) and participants (n = 4.472) are stratified in Cases (result within reference interval at screening, outside reference interval at start of the trial) and Normals (within reference interval at both screening and start of the trial). Average Z‐scores are shown in a radar plot per measurand (A). Significance between Normal‐ and Case‐measurands at screening and first re‐assessment during trial participation is indicated below the measurand name (significance as screening/significance at re‐assessment) with p > 0.05; n.s., p ≤ 0.05; *, p < 0.01; **, p < 0.001; *** and na; not applicable. Additionally, the average z‐score for all Normal and Case measurands at screening (Normals: 0.64 ± 0.50, Cases: 0.79 ± 0.61, p < 0.001) or enrollment (Normals: 0.66 ± 0.51, Cases: 1.99 ± 1.72, p < 0.001) are shown (B). Alternatively, the average of all measurands obtained at screening or the start of the trial is determined for the 803 participants in which Cases are defined as individuals with at least one measurand result outside the reference interval at the occasion, demonstrating significantly increased z‐scores for the 458 Cases compared to the 345 Normals at enrollment (Normals: 0.66 ± 0.17, Cases: 0.76 ± 0.20, p < 0.001) but not at screening (Normals: 0.64 ± 0.18, Cases: 0.65 ± 0.14, p = 0.220) (C). Lastly, the amount of time in days between the last measurement at screening and first measurement during the trial is shown which does not significantly differ between Normals and Cases (16.65 ± 9.66 compared to 17.17 ± 10.54 days, respectively, p = 0.62) (D). ALT; alanine aminotransferase, AST; aspartate aminotransferase, Gamma‐GT; gamma‐glutamyl transferase, HDL; high‐density lipoprotein, LDL; low‐density lipoprotein.
Table S3: the mean, standard deviation (SD), coefficient of variation (CV, represented by the SD divided by the mean) of the entire population (CVpop) or when the CV is determined within a participant and subsequently averaged throughout the population (CVind).
Figure S3: Overview of alanine aminotransferase (ALT, panel A), aspartate aminotransferase (AST, panel B) and bilirubin (panel C) over time and their trajectories. Values exceeding the reference intervals are highlighted in red. The 1.5 and 3.0 times upper limit of normal (ULN) is shown, stratified for male and females if applicable. Both increases of ALT were related to participants with possible viral infections. Results from unscheduled samples are only included for the results over time (left) and are indicated with a triangle. Unscheduled samples are not included for the trajectories (right).
Figure S4: Heatplot showing the proportion of results that are outside the reference interval per measurand per day. The number of results outside of reference intervals and total results obtained within 24‐h blocks from zeropoint are shown.
Table S4: Parameters for the diurnal rhythm of white blood cell counts. Time is given in decimals.
Table S5: Listing of total (flagged) measurements of all parameters from day −1 to day 7, corresponding to Figures 3 and S3. Percentages are given as the part of flagged measurements from the total amount of measurements obtained.
