Skip to main content
mBio logoLink to mBio
. 2025 Sep 9;16(10):e01881-25. doi: 10.1128/mbio.01881-25

Bayesian estimation of HIV acquisition dates for prevention trials

Raabya Rossenkhan 1, Elena E Giorgi 1, Danica Shao 1, James Ludwig 1, Phillip Labuschagne 1,2, Craig A Magaret 1, Thumbi Ndung'u 3,4,5,6, Daniel Muema 3,4, Kamini Gounder 3,4, Krista L Dong 6, Bruce D Walker 3,6,7, Morgane Rolland 8,9, Merlin L Robb 9, Leigh Anne Eller 8,9, Fredrick Sawe 10,11, Sorachai Nitayaphan 12, Eduard Grebe 13,14,15, Michael P Busch 14,15, Kevin P Delaney 16, Shelley Facente 15,17,18, Lindsay N Carpp 1, Allan C deCamp 1, Yunda Huang 1,19, Bette Korber 20,21, Michal Juraska 1, Erika Rudnicki 1, Ewelina Kosmider 1, Daniel B Reeves 1,19, Bryan T Mayer 1, John Hural 1, Wenjie Deng 22, Dylan H Westfall 22, Anna Yssel 23, David Matten 23, Tanmoy Bhattacharya 20,24, Lawrence Corey 1,25,26,27, Peter B Gilbert 1,25,28, Carolyn Williamson 23,29, James I Mullins 19,22,26, Paul T Edlefsen 1,
Editor: Suresh Mahalingam30
PMCID: PMC12505913  PMID: 40923785

ABSTRACT

Accurate timing estimates of when participants acquire HIV in HIV prevention trials are necessary for determining antibody levels at acquisition. The Antibody-Mediated Prevention (AMP) Studies showed that a passively administered broadly neutralizing antibody can prevent the acquisition of HIV from a neutralization-sensitive virus. We developed a pipeline for estimating the date of detectable HIV acquisition (DDA) in AMP Study participants using diagnostic and viral sequence data. Using a Bayesian strategy that combines three streams of data (REN [rev/vpu/env/Δnef] sequence, GP [gag/Δpol] sequence, and diagnostic) where their 95% credible intervals overlap based on pre-specified criteria and decision rules. We evaluated the performance of our AMP pipeline using PacBio viral sequence data from 41 participants across two prospective acute HIV acquisition cohort studies, FRESH and RV217, with twice-weekly sampling. These cohort studies enrolled young women in South Africa and men and women in Kenya and Thailand, respectively, with a high likelihood of HIV acquisition. In evaluating performance, “true DDA” was the center of bounds between last-negative and first-positive RNA diagnostic tests (median time 4 days, range 2–7 days); bias was the mean difference between estimated and true DDA. Using diagnostic data alone yielded timing estimates with a bias of 2.4 days and root mean square error (RMSE) of 7.9 days. These results were improved using sequence + diagnostic data (bias 1.5 days, RMSE 6.9 days), as well as by restricting sequence-based estimation to samples from ≤5 weeks post-DDA (bias 0.2 days, RMSE 7.8 days).

IMPORTANCE

In HIV prevention trials, accurate timing estimates of when individual participants acquire HIV can be used to estimate antibody levels at the time of acquisition, which is useful for projecting antibody levels needed for prevention. The results we report here suggest that if sequence-based estimation of acquisition timing is used in future clinical trials of combination broadly neutralizing antibody (bnAb) regimens or multispecific bnAbs for HIV prevention, a sampling frequency of at least monthly is needed. Moreover, in the samples analyzed here, we observed less bias in sequence-based timing estimation for samples taken <5 weeks post-DDA. This observation is consistent with the timing of immune-driven selective pressures that may negatively impact the power to detect acquisition sieve effects.

KEYWORDS: acute acquisition cohort, Antibody-Mediated Prevention (AMP) Studies, Bayesian posterior distribution, date of detectable acquisition, FRESH, HIV, RV217

INTRODUCTION

The Antibody-Mediated Prevention (AMP) Studies, HVTN 704/HPTN 085 (NCT02716675) and HVTN 703/HPTN 081 (NCT02568215), were the first clinical trials to provide proof of concept that a broadly neutralizing antibody (bnAb) could prevent HIV acquisition. The two trials were conducted in distinct cohorts and had harmonized protocols wherein participants were randomized to low-dose (10 mg/kg) VRC01 (a CD4-binding site bnAb), high-dose (30 mg/kg) VRC01, or saline placebo, with 10 intravenous infusions given at 8-week intervals (1). While overall prevention efficacy of VRC01 (both dose groups pooled vs. placebo) against the primary endpoint of documented HIV diagnosis by the week 80 study visit was not significant, prespecified secondary analyses showed high prevention efficacy (75.4% [45.5%, 88.9%]) of the two VRC01 dose groups pooled vs. placebo against VRC01-sensitive viruses (defined as an 80% in vitro inhibitory concentration [IC80 ] < 1 µg/mL) (1). These estimates were much lower for more resistant viruses: 4.2% (−108.7%, 56.0%) against viruses with IC80 1–3 µg/mL and 3.3% (−48.0%, 36.8%) against viruses with IC80 >3 µg/mL. These results, along with neutralization sieve results, showed that prevention efficacy varied by virus susceptibility to VRC01-mediated neutralization, with high prevention efficacy achieved against sensitive viruses.

The AMP results underscore the importance of testing combination bnAb regimens or multispecific bnAbs to improve coverage against resistant viruses and facilitate the discovery of important biomarkers and pharmacokinetics (25) as benchmarks to inform future trials. To accurately estimate the efficacy of bnAbs for prevention, it is critical to obtain the best possible estimates of the time of HIV exposure leading to acquisition, in addition to understanding virus population characteristics at that time. In practice, the date of detectable acquisition (DDA; previously date of detectable infection [DDI] and modified here for person-first language) is estimated, which is about ~1 week after the exposure event due to the eclipse phase (6) (i.e., the time window after which HIV exposure has occurred, but before HIV is detectable in the blood). Current strategies for estimating the time of HIV acquisition include diagnostic-based methods (7) and subsequent generalizations/extensions (810), or, alternatively, sequence and molecular clock-based approaches (11, 12). Both methods have distinct strengths and weaknesses. For example, there are substantial cost implications of the relatively frequent sampling required for accurate diagnostic timing in clinical trials, while sequencing approaches (11, 12) rely on assumptions that do not always hold, depending on when the viral population is sampled. Examples of such assumptions include growth in the absence of immune-driven selection pressure and, for the Poisson Fitter tool (12), a homogeneous infection and constant mutation rate. A combined approach of using data streams with both diagnostic and sequence information can narrow the uncertainty window and ameliorate potential weaknesses contributed by these approaches individually, toward a better estimate of the DDA.

We have previously published HIV acquisition timing and multiplicity founder estimation methods that combine the strengths of both diagnostic and sequence-based strategies (13). These methods were based on maximum likelihood estimators that shift and scale (calibrate) estimates by fitting approximately true detectable acquisition times and founder multiplicities (from the RV217 [14] and CAPRISA 002 [15] acute HIV acquisition cohort studies) to a linear regression model with independent variables defined by data on HIV sequences, viral loads, and HIV diagnostics. As part of that work, we developed a cross-platform software pipeline distributed as a Docker container (16) (https://hub.docker.com/r/philliplab/hiv-founder-id) to provide reproducible estimation of detectable acquisition times for participants in HIV Vaccine Trials Network (HVTN) efficacy trials (13). That pipeline can also be used for generating reports and for sharing data.

We have previously published approaches and a software pipeline for estimating dates of HIV acquisition in AMP Study participants (4). In the present work, we describe the origins of the methodology used in the AMP studies in Gilbert et al. by developing a Bayesian analog to the Poisson Fitter software, as well as adding quality checks for iterative blinded analysis using Poisson Fitter within the Los Alamos National Laboratory (LANL) and HVTN pipelines. We apply our pipeline to estimate the DDA for participants in two prospective cohort studies, FRESH (17, 18) and RV217 (14), each of which notably used twice-weekly fingerstick blood collection for HIV RNA testing. FRESH enrolled 945 women 18–23 years of age in KwaZulu-Natal, South Africa, with a high likelihood of HIV acquisition; 42 of these women were diagnosed with acute HIV between December 1, 2012, and June 30, 2016 (17, 18). RV217 enrolled 2276 men and women 18–50 years of age in Kenya and Thailand with a high likelihood of HIV acquisition into a surveillance phase; 50 of these participants were diagnosed with acute HIV and met criteria for inclusion in the primary analysis (14). The present analysis included sequence and diagnostic data from 13 FRESH participants and 28 RV217 participants who were diagnosed with acute HIV, selected based on the following criteria: (i) sufficient sample volume to allow one stored tube to be shipped to the Mullins or Williamson laboratories for viral PacBio sequencing and (ii) the last blood sample with a negative HIV RNA test result (no detectable HIV RNA) and the first blood sample with a positive HIV RNA test result (HIV RNA above the test’s limit of detection) were separated by a maximum of 7 days (median 4 days; range 2–7 days). For the purposes of assessing the performance of our estimators, the DDA (taken as ground truth) was calculated as the center of bounds (COB) between these last-negative and first-positive RNA diagnostic tests. (iii) We also only included samples from participants not on antiretroviral therapy during AMP-like timing selection scenarios. Using sampling timepoints chosen to best-match the monthly samplings in the AMP Studies, GP (2.5 kb gag and Δpol, where Δ = only part of the gene is included) and REN (3.0 kb, rev-vpu-env-Δnef) were sequenced using the PacBio SMRT-UMI (Single Molecule Real-Time platform using amplicons derived from cDNA templates tagged with Unique Molecular Identifiers) protocol (19, 20). Our main objective was to evaluate the bias and accuracy of multiple estimators, using these acute acquisition data sets to have the best possible assessment of the DDA as measured in the AMP Studies. The methods we developed and evaluated to address this objective apply to prospective studies with regular and frequent participant testing and are less applicable to observational cohort or surveillance studies with infrequent testing schedules.

We find that by using the median of the Bayesian posterior distribution, computed by combining sequence-based methods with the diagnostic curves as described in Grebe et al. (9), we obtained accurate estimates of the time of HIV RNA conversion (i.e., DDA).

MATERIALS AND METHODS

Participants

Plasma samples for HIV sequencing were selected from FRESH participants (n = 13) and RV217 participants (n = 28). Sequencing was performed on one to three samples per participant, taken at consecutive time points with intervals to match sampling scenarios in the AMP studies.

HIV testing

Antibody-positive (Ab+) and antibody-negative (Ab−) serostatus for the FRESH and RV217 participants was determined according to the diagnostic methods described previously (14, 17, 18).

Viral sequencing and processing

Long-read PacBio sequences were obtained using the optimized SMRT-UMI sequencing and PORPIDpipeline methods described in Westfall et al. (19). Lineages were identified using phylogenies, Highlighter, and Matches plots (21) as implemented in Phylobook (22). Screening for recombinant sequences and APOBEC enrichment was also done using the tools RAPR and Hypermut available from the LANL HIV database (https://www.hiv.lanl.gov/). Two regions of the viral genome were evaluated for each participant sample—the GP region, spanning bases 790–3,297 relative to the HXB2 reference, including the entire gag gene and first 1 kb of the pol gene, and the REN region, spanning bases 5,970–9,012 and including both exons of rev, the second exon of tat, all of vpu and env and the first 215 bp of nef. For the data analyzed here, a median and range of 252 (4–1,220) GP and 208 (2–698) REN sequences were obtained from the RV217 and FRESH samples at each time point. A preliminary data set of sequences was used for the current study, whereas a more heavily curated data set will be provided in Mullins et al. (in preparation). Maximum likelihood trees using the GTR model of evolution were generated using PhyML in DIVEIN (23) and displayed using Figtree (https://github.com/rambaut/figtree).

Sequence processing and founder lineage identification

Sequence data were further processed as in Supplementary Note 2 of Gilbert et al. (4), except for lineage determination. Analyses were conducted in R v4.3.2 and python 2. Standard linear regression and Wilcoxon testing were used as noted. Unadjusted P-values ≤ 0.05 were considered significant. Our sequence-based strategy to estimate DDA is based on the assumption that early in the infection, intra-host sequence diversity within a participant will conform to a Poisson distribution whose mean is proportional to the number of generation cycles the virus has undergone since the beginning of infection. We measure this intra-host sequence diversity by calculating all pairwise Hamming distances (HD, defined as the number of mutations between all sequence pairs within a single time point sample) for each intra-host lineage observed, and then fitting a Poisson distribution to the HD frequency counts, as previously described (21, 24). For these distance distributions to fit a Poisson, the mutations across all sequence pairs have to be randomly or close to randomly distributed. This is no longer true when multiple genetically distinct lineages (independently of whether they originated at transmission or post-transmission through early selection) are present, as all inter-lineage pairs will share the same set of mutations. Therefore, when multiple lineages are detected, after excluding hypermutated sequences and recombinants, we subset the sequences into these groups and calculate within-lineage Hamming distances. Because some of these lineages arise post-acquisition and do not always coincide with the transmitted founder lineages, we refer to these groupings as “bottleneck lineages” to distinguish them from the transmitted founder lineages defined elsewhere (Mullins lab, unpublished data).

Screening for hypermutation and recombination

APOBEC (apolipoprotein B mRNA editing enzyme, catalytic polypeptide-like) enzymes are polynucleotide cytidine deaminases, with APOBEC3 proteins being potent inhibitors of HIV replication due to the resulting guanine-to-adenine (G-to-A) hypermutations that are introduced into the viral genome at distinct sequence motifs/hotspots and occur at a higher rate than the baseline viral mutation rate (2527). Because of this, APOBEC-induced hypermutation can cause the Poisson model to either fail or yield an artificially inflated estimated time of infection due to the additional mutations. To prevent this, we screened viral GP and REN sequences from FRESH and RV217 participants for overall enrichment for hypermutation (defined as a higher than expected rate of G-to-A mutations in APOBEC context scattered across many sequences in the alignment) as well as for sequences individually enriched for such mutations prior to fitting our models, as described in Gilbert et al. (4) using the LANL tool Hypermut (28). Hypermutation was detected in the FRESH and RV217 samples, especially at earlier time points, consistent with previous findings (29). When detected, to minimize the potential impact of hypermutation on our timing estimates, we removed all positions in the alignment where APOBEC-induced hypermutation may occur, that is, positions in an “APOBEC context”—which we define (as in Giorgi et al. [29]) as positions in the genome where the putative lineage founder has a guanine followed by either a thymine or an adenine. We refer to this removal as “APOBEC masking.”

Recombination can also artificially skew the mutation rate, creating sequences with higher than expected diversity when the parental strains are divergent or lower than expected when the parental strains are very homogeneous (30). Therefore, we also screened for recombination using the LANL tool RAPR (https://www.hiv.lanl.gov/content/sequence/RAP2017/rap.html) and, when detected, statistically significant recombinants were excluded from time estimate analyses.

Data employed in diagnostic (DX)-based timing estimation

The data available from the diagnostic assays were used in an adapted version of the Infection Dating Tool (IDT; https://zenodo.org/records/1972956) (9). The IDT was originally developed by the Consortium for the Evaluation and Performance of HIV Incidence Assays (CEPHIA) using information from multiple previous studies (9) and expands on the “Fiebig staging” approach, allowing for flexible use of different diagnostic assays with specific window periods determined for each assay. Our adapted method employs the window period distributions of the last-negative and first-positive test results to construct a Bayesian posterior distribution for the time of acquisition. Our implementation is called “acquisition dating tool in R” (adtr). Our local R implementation of the open-source web-based method is available at https://github.com/HVTN-SDMC/AMP_FRESH_RV217_AcquisitionTimingPaper.

Creation of a hypothetical AMP-like data set

To evaluate the prospective performance of our methods, we constructed a data set of sequences as well as diagnostic test histories that mimicked the AMP Studies data and sampling timeline. For each sample, we used the date of sampling as a hypothetical first-positive visit (FP), reporting all negative and positive tests performed on that date with that participant in the original trial. We also constructed a hypothetical last-negative visit (LN) that would be consistent with both the true date of HIV DDA for that sample (known to within ~3 days in FRESH and RV217), as well as the monthly sampling timeline employed in AMP. To that end, we set the LN to be 30 days before FP in cases where the true interval between LN and FP was 30 days or smaller, or 60 days before FP in cases where the true interval was between 30 and 60 days (to simulate a missed visit scenario in the AMP Studies).

Pipeline

Figure 1 shows the refined analysis pipeline for estimating the time of HIV DDA, based on the pipeline reported in Rossenkhan et al. (13). Using a Bayesian strategy, this pipeline combines the three streams of data (sequence REN, sequence GP, and diagnostic data [DX]) where their 95% credible intervals (CIs) overlap based on the pre-specified criterion and as described above. This generates a point estimate and CI. Congruence was observed between these methods and their Bayesian implementations. Decisions on which streams of data to use were based on the rule employed in the AMP Studies due to visits being scheduled monthly. The AMP Studies’ rule allows for SEQ and DX estimates to be merged by using the median of their combined posterior probability distributions only when the 95% credible interval (CI) of the posterior distributions overlap, and when the sampling date is within 5 weeks of the last-negative visit.

Fig 1.

HIV acquisition timing methodology integrating diagnostic (DX) and sequence data (REN/GP = SEQ) using Bayesian analysis to produce posterior probabilities and credible intervals for infection dates, with decision paths for data reliability assessment.

(A) Schematic overview of the types of data input, decision-making steps, data combining, and resultant output of the HIV acquisition timing methods used here. (B) Decision tree diagram employed for the AMP Studies and a guide to subsequent results shown from FRESH and RV217 data analysis to match AMP.

Bayesian implementation of existing methods

We developed software that implements procedures for computing Bayesian posterior distributions of detectable acquisition time for (i) individual sequence alignments and (ii) last-negative and first-positive diagnostic test results using the Grebe method (9) as employed in the IDT. In the Grebe method, data sets consisting of information on individual-level participant diagnostic test dates, specific test information (manufacturer/assay type), and result (positive or negative) are used along with test-specific diagnostic delay parameter estimates to output point estimates for each participant’s DDA and their earliest plausible and latest plausible DDAs (95% posterior credible interval) (9). In our R language implementation of this method, we also output the posterior probability by day for downstream Bayesian integration. The test-specific diagnostic delay parameter estimates used in this procedure are summary measures of population-level sensitivities of each particular diagnostic test, expressed as parameters of 3-parameter Weibull models describing inter-test intervals. Summaries of these parameters are provided in an open-source database in the IDT source code (9, 31), and the derived Weibull parameters are provided in an analogous database within our open-source software (https://github.com/HVTN-SDMC/AMP_FRESH_RV217_AcquisitionTimingPaper).

To output the full daily posterior distribution, we reimplemented the Grebe procedure in R. We employed these estimators to yield posterior distributions of DDA for simulated AMP participants in the manner employed for the AMP Studies, given only sequence data and diagnostic tests collected from one sample at one time point. This was accomplished by a Bayesian combination of posterior distributions across input sources, including last-negative diagnosis date, first-positive diagnostic test results and date, and sequence alignments representing potentially multiple lineages per region for each of the two sequenced regions of the HIV genome (GP and REN).

We identified three distinct sources of independent information (statistical definition) pertinent for our analysis: (i) diagnostic test data, (ii) GP sequence data, and (iii) REN sequence data. However, we anticipated potential discrepancies among these three sources, particularly given the limitations of sequence-based methods. These limitations derive from immune selection-driven bottlenecks (32, 33) and demographic processes (biological mechanisms that restrict transmission and early viral dynamics) (34), the latter of which are particularly relevant in the first few weeks post-transmission. Moreover, considering our expectation of heightened selection pressure on Env due to the presence of VRC01, we sequenced the GP region to potentially augment the accuracy of acquisition time estimates in the AMP Studies in case we had evidence of a discrepancy. Mutations due to selection would violate the assumptions of a constant molecular clock and likely bias estimators toward the time of the selection bottleneck, away from the true DDA. However, there is a caveat that selection pressure on Env in the form of an advantageous mutation or genotype that outgrows the originally transmitted lineage—if such an event were to occur—might also affect the GP region, although these disparate regions of the viral genome appear to evolve nearly independently due to rampant recombination (35). In addition, in the structurally more constrained Gag-Pol proteins, negative selection pressure would be more pronounced compared to that in Env and would likely result in a population more homogeneous in GP than in REN. Here, for simplicity, we ignore these effects, though we acknowledge that doing so may result in an underestimate of the DDA based on GP data. Our prespecified strategy was to combine all independent estimators of the DDA unless the two posterior 95% CIs did not overlap, in which case we concluded that one or both estimators were not estimating the DDA well. In this case, in the absence of any further flags, we prespecified a plan to prioritize based on an assumed order of impacts of selection effects: if GP and REN estimates disagreed (defined as non-overlapping 95% CIs), we prespecified to use the GP estimate, and if the SEQ estimate (defined as the combination of the GP and REN estimates if they agreed, or the GP estimate alone if they did not) did not agree with the DX estimate (based solely on diagnostic data), we used only the DX for our final estimator (see Fig. 1B).

To combine the sequence- and diagnostic-based estimators, we first needed to ensure that the outputs from each method were reasonable to combine for all possible dates of detectable acquisition. To output posterior probabilities by date for the sequence-based estimators, we created custom scripts based on methods implemented by Giorgi et al. (12). For the diagnostic-based estimators, we created a new R package (adtr) based on the previously published and publicly accessible online Infection Dating Tool (IDT), which expands upon widely accepted Fiebig staging methods (7). Next, we multiplied posterior probabilities to create combined distributions, from which we generated point estimates and CIs for the target value, as well as DDA, to compare them to “ground truth” data (see “Statistical Analysis,” below).

Definitions of acquisition timing estimators

We defined the sequence-based estimator of DDA (SEQ) as the median of the combined posterior probability distributions of the REN- and GP-based estimators when the 95% CIs overlapped. We initially planned to prefer GP over REN estimators when the 95% CIs did not overlap; however, while still blinded, the prespecified plan was revised to prefer REN over GP such that the SEQ estimate was identical to the REN-based estimate if the 95% CIs did not overlap. For the combined estimator, selected sequence data (REN and GP from up to 5 weeks post-COB) were combined with DX data where they overlap. After 5 weeks, the combined estimator excluded the SEQ data (DX only), and all diagnostic data were used.

Statistical analysis

We did not formally compare estimators across time or across serostatus. However, to account for within-participant correlation across time in analyses evaluating accuracy and precision within categories of serostatus (Ab− and Ab+) and time since COB (1–14 days, 15–28 days, or 29+ days, with bins selected to achieve approximately equal sample sizes), we subsetted the data within categories. Using this approach, in the case of multiple samples from the same participant falling in a category (e.g., Ab+), only one sample was used. We arbitrarily decided a priori on using the first sample chronologically. When we evaluated other prespecified subsampling schemes in an available early subset of the data, we found empirically that the choice would have minimal potential downstream impact as it resulted in none or only negligible differences in selected samples.

For each participant, we computed our “ground truth” estimator for DDA (“true DDA”) as the COB between their true last-negative (LN) and first-positive (FP) diagnostic tests (using the actual test dates in the FRESH and RV217 studies). To assess the accuracy of DDA point estimates based on later samples, we computed the error (number of days between known COB and estimated DDA, with negative error representing estimates that are too early), bias (mean error), and RMSE of each estimator overall, as well as within categories of serostatus (Ab+ vs. Ab−) and time since COB (1–14 days vs. 15–28 days vs. 29+ days). To assess the uncertainty intervals of these estimators, we also reported the mean width of the 95% posterior CI, as well as a frequentist coverage probability assessment (the proportion of 95% posterior CIs that included the known COB). All analyses were performed in R v4.3.2. (36).

Calibration of the model through cross-validated machine learning

To assess how to improve the performance of the Combined estimator and the DX DDA estimators, we first generated a scatterplot of the estimated number of days that the sequenced sample was collected after the COB against the estimated number of days that the sequenced sample was collected since the DDA, calculated using either the Combined or DX estimator. We then performed linear regression with 10-fold cross-validation, with the intercept fixed at the origin, and the original estimates were calibrated via this cross-validated model. We then generated a scatterplot of the original estimate from each estimator against the estimate from the cross-validation-calibrated estimator.

P values for testing estimator improvement (i.e., a reduction in RMSE) via cross-validation calibration were obtained through a one-sample t-test of the differences of squared errors between the original estimator and the cross-validation-calibrated estimator.

RESULTS

Sequence data for evaluation

Sequences from 41 participants (13 from FRESH and 28 from RV217) were evaluated. Table 1 provides descriptive statistics for the FRESH and RV217 samples and sequence data by genomic region. APOBEC mutations were found in all participants in both cohorts in at least one lineage and at least one time point.

TABLE 1.

Numbers of viral sequences from participants in the FRESH (n = 13) and RV217 (n = 28) studiesa

Parameter GP REN Total
Samples across all time points, both studies, and all participants 139 139 278
Total number of viral templates sequenced 38,150 33,509 71,659
Mean 274 240
Median 252 208
Range 4–1,220 2–698
a

GP (2.5 kb gag and Δpol) and REN (3.0 kb, rev-vpu-env-Δnef); Δ = only part of these genes were included.

AMP simulation results

While developing the implementation of this methodology for the AMP Studies, and still blinded to treatment status, we employed the software on these data to emulate AMP participants using data from the FRESH and RV217 cohort studies. This approach worked well with no optimization apart from two changes to earlier decisions, based on initial impressions from partial analyses of these data. (i) We swapped the original specified order, which prioritized GP over REN when they disagree, because of the absence of VRC01, there seemed to be modestly better agreement with REN than GP. This likely indicates that early purifying selection impacts GP more than REN in natural acquisition scenarios. (ii) We prioritized the use of diagnostics (DX) for later samples, which we define as samples later than five weeks post COB, which in the AMP Studies would indicate a missed test or infusion visit just prior to diagnosis. We chose this priority even when posterior intervals overlap, as the posterior CIs for the diagnostic method are wide, making overlap of posterior CIs overly frequent. In addition, we observed that the sequence-based estimators appear to be downward biased after about a month, possibly reflecting a sequence bottleneck occurring in the 4- to 6-week time frame that biases estimates from both GP and REN regions. We found that sequence-based estimators perform well and are unbiased when applied to samples collected within the first 4 weeks, even among participants who are already seropositive.

Acquisition timing methods using FRESH and RV217 samples matched to AMP monthly sampling

To address timing uncertainty in the AMP Study scenario of monthly testing using these FRESH and RV217 samples, we explored two key variables: (i) serostatus, categorizing samples as either Ab− or Ab+, and (ii) time since HIV RNA detection, categorized into intervals (time bins) of 2 weeks (1–14 days), 4 weeks (15–28 days), and 29+ days. Data were filtered by these criteria to prevent duplication of participant counts within each category.

We employed six different types or combinations of input data to estimate HIV acquisition timing. Of these six, three showed the best performance: (i) using sequence data alone, restricting to samples taken ≤5 weeks from the COB (SEQ [first 5 weeks]), (ii) using diagnostic data only (DX), and (iii) using diagnostic and all sequence data (“Combined”) (Fig. 2). Across these three methods, overall bias ranged from 0.2 to 2.4 days (Fig. 2A) and RMSE from 6.9 to 7.9 days (Fig. 2B). The combined GP and REN sequence-based estimator (SEQ) had relatively high negative bias (−15.1 days) when the sample was taken 29 or more days since COB (Fig. 2A). Therefore, for the evaluation of the hypothetical AMP-like data set, we only used sequence data from <5 weeks, which mitigates this bias by combining SEQ and DX, using solely DX when the sample was taken more than 5 weeks after COB (Fig. 2E).

Fig 2.

HIV timing comparison: combined approaches show lowest bias/error while sequence-only methods underestimate after 29+ days. Early samples yield better accuracy. Scatter plots confirm combined methods align closer to “true” (COB) HIV acquisition dates.

Performance of the six different HIV acquisition timing estimation approaches for the 41 selected FRESH and RV217 participants and the effect of sampling time and HIV serostatus on performance. (A, B) Heatmaps showing (A) bias in days (average error) and (B) RMSE (root mean square error) for six different approaches of estimating date of detectable acquisition (DDA) for FRESH and RV217 participants. In (A), a negative value means an underestimation of the DDA, that is, the given estimator estimated the number of days that the sample was taken from the DDA to be fewer than those yielded by the ground truth COB estimate, and a positive value means an overestimation of the DDA, that is, the given estimator estimated the number of days that the sample was taken from the DDA to be greater than those yielded by the ground truth COB estimate. Rows denote the method used as follows (bottom to top): GP: sequence-based estimation using the GP region; REN: sequence-based estimation using the REN region; SEQ (All): combined sequence-based estimate using both regions, with no restriction on the length of time after which the sample was taken post-COB; SEQ (first 5 weeks): combined sequence-based estimate (GP and REN regions) using only data from the first 5 weeks after COB; DX: diagnostic-based estimation alone; Combined: combined diagnostic and sequence-based estimation. Stacked panels denote (A) the bias or (B) RMSE of various subsets of the data as follows (left to right): (i) overall; (ii) sampled 1–14 days after COB; (iii) sampled 15–28 days after COB; (iv) sampled more than 28 days after COB; (v) HIV antibody-negative (Ab−) sampling periods; and (vi) HIV antibody-positive (Ab+) sampling periods. Note that the 1–14, 15–28, and 29+ bins pool across Ab− and Ab +participants. (C, D, E) Scatterplots of (C) SEQ estimates, (D) DX estimates, and (E) combined estimates of DDA. Approach used in AMP vs actual known COB between last-negative and first-positive HIV test result (most accurate estimate of DDA available from actual acute acquisition study data). In (C) and (D), darker points depict estimates using data sampled more than 5 weeks from COB. In (E), the combined estimate is shown as orange circles when only the DX estimate is used, and as purple crosses when the DX and SEQ estimates are combined. In (C–E), the dotted line is a line of equality and the solid gray line is a regression line.

SEQ is an accurate estimator of DDA for antibody-negative samples and samples collected up to 1 month post-DDA

The SEQ estimator of DDA yielded reasonable accuracy (bias) and precision (RMSE) in Ab− samples across all sampling times post-DDA, as estimated by COB (bias 4.4 days [Fig. 2A], RMSE 7.5 days [Fig. 2B]). In samples collected within the initial 5 weeks post-DDA as defined by COB (pooling across Ab− and Ab +participants), the SEQ estimator performed well (bias –0.1 days for the 15–28 day bin and 5.1 days for the 1–14 day bin [Fig. 2A], RMSE 5.8 for the 15–28 day bin and 8.0 days for the 1–14 day bin [Fig. 2B)]). However, bias and RMSE were much larger in samples collected beyond 5 weeks post-DDA as defined by COB (bias –15.1 days, RMSE 19.6 days). Consequently, we decided to restrict the incorporation of sequence data to the initial 5 weeks in AMP. The SEQ (first 5 weeks) estimator also performed acceptably well on samples from Ab− participants: bias 4.4 days (versus −0.9 days on Ab+ participants) and RMSE 7.5 days (versus 7.9 days in Ab+ participants). This is the first evaluation of the Bayesian implementation of the Poisson Fitter method on human samples, and it shows that the method works well for samples collected up to 5 weeks after the COB and that SEQ is also a good estimator of DDA for samples from Ab− participants.

DX is a comparably accurate estimator of DDA for samples from 1 to 14 days vs. 15 to 28 days post-COB

Across both serostatuses and across all sampling times, the DX estimator had a bias of 2.4 days and a RMSE of 7.9 days (Fig. 2A and B). Breaking this down into sampling time groups, the DX estimator had a bias of 2.7 days and a RMSE of 6.8 days for the 1–14 day group; a bias of 1.3 days and a RMSE of 6.9 days for the 15–28 day group; and a bias of 4 days and a RMSE of 10.2 days for the 29+ day group (Fig. 2A and B). These findings show that the DX estimator performs comparably well for samples from 1 to 14 days vs. 15 to 28 days post-COB and prompted us to evaluate the utility of combining sequence and diagnostic data to further reduce bias and error in the DDA estimates.

Improved performance of the combined estimator compared to the DX estimator across both serostatus and across all sampling times

Across both serostatuses and across all sampling times, the Combined estimator had a bias of 1.5 days and an RMSE of 6.9 days (Fig. 2A and B). These values were slightly lower, and thus of moderately better performance, compared to the DX results discussed above. This trend generally held across all sampling time groups and across both serostatuses, with some exceptions, including slightly lower bias in the DX estimator in samples from Ab− participants (1.2 and 2.1 days bias for DX and Combined, respectively).

Establishing decision rules for the AMP studies: analyzing bias and consistency in GP-based and REN-based acquisition timing estimates of FRESH and RV217 samples

Estimation error for the GP and REN regions is shown in Fig. 3A, with GP tending to underestimate the number of days since DDA to a greater extent than REN. Given the potential likelihood of VRC01-induced selection (3739) on REN in the AMP Studies, we had planned to prioritize GP over REN in sequence-based estimates if solely GP-based and solely REN-based timing estimates were largely concordant. As shown in Fig. 3B, this is broadly true, although the GP and REN estimates for the FRESH and RV217 samples tended to be less concordant for estimating acquisition times more distant from the sample date. Compared to the COB-estimated DDA, both sequence-based methods tended to underestimate acquisition times more distant from the sample date (Fig. 3C and D). Overall, this bias was noticeably greater in the GP-based estimates (Fig. 2A and B: GP bias –8.4 days, RMSE 14.9 days; REN bias –2.1 days, RMSE 12.5 days). Thus, we prioritized REN while still utilizing sequence information from both regions when appropriate (see Materials and Methods for decision rules and definition of acquisition timing estimators).

Fig 3.

GP method shows less estimation error than REN for HIV timing in FRESH/RV217 participants. Scatterplots reveal correlation between methods, with both showing variable accuracy versus COB estimates. GP typically underestimates days since DDA.

Comparison of GP vs. REN sequence-based timing estimates for the 41 selected FRESH and RV217 participants. (A) Boxplots, overlaid with individual data of estimation error in days obtained using REN (top, red) vs. GP (bottom, blue) sequence data. Gray lines between dots connect the two sequence-based results from the same sample. For all estimators of DDA, we defined the estimation error as the difference in days between the estimator and the ground truth estimate using the COB between the actual LN and FP dates in RV217 and FRESH cohort data. Positive values indicate overestimation of the number of days between DDA and FP (i.e., the estimate of DDA is too early compared to COB), and negative values indicate underestimation (i.e., the estimate of DDA is too late). (B) Concordance between GP and REN acquisition timing estimates of days since DDA. (C, D) Scatterplots of (C) REN-based estimates of DDA and (D) GP-based estimates of DDA versus the COB DDA estimate. In panels B–D, the dotted line is a line of equality, and the solid gray line is a regression line.

Relatively uniform distributions of participant-level error obtained by the SEQ estimator across sampling-time groups and serostatus

Figure 4 shows participant-level estimation error for SEQ for participants in each group of sampling interval post-COB (1–14 days, 15–28 days, and 29+ days) (Fig. 4A) and for participants stratified by Ab+ or Ab− HIV serostatus (Fig. 4B). The individual-level participant distribution of SEQ error was more similar between the 1–14 day group and the 15–28 day group compared to the 29+ day group, which seemed to have one outlier participant with an overestimation of nearly 20 days. There also seemed to be a little difference between the Ab− and Ab+ participant distributions of SEQ estimation error, except that the Ab− participants seemed to have a somewhat smaller range of error distribution.

Fig 4.

SEQ-based DDA estimation errors in FRESH/RV217 HIV cohorts show earlier sampling (1-14 days) yields more accurate results with narrower distributions. Antibody-negative participants demonstrate smaller errors than antibody-positive participants.

Effects of sampling time and HIV serostatus on SEQ-based DDA estimates for the 41 selected FRESH and RV217 participants. Boxplots, overlaid with estimation error in days obtained using (A) For both serostatuses pooled (Ab− and Ab+), results are stratified by sampling interval post-COB (1–14 days, 15–28 days, and 29+ days); (B) for all sampling times post-COB pooled, results are stratified by HIV serostatus (Ab− vs. Ab+). The dotted vertical line represents perfect estimation of DDA in terms of the ground truth (COB). In interior box plots, left and right vertical edges show the 25th and 75th percentiles of the SEQ estimation error, the middle line indicates the median, and vertical bars extend to the outermost data points within 1.5 times the IQR from these quartiles. For all estimators of DDA, we defined the estimation error as the difference in days between the estimator and the ground truth estimate using the COB between the actual LN and FP dates in RV217 and FRESH cohort data. Positive values indicate overestimation of the number of days between DDA and FP (i.e., the estimate of DDA is too early compared to COB), and negative values indicate underestimation (i.e., the estimate of DDA is too late).

A wider range of participant-level errors obtained by the DX estimator

The DX estimator generally crosses the true estimate, although the interval widths of uncertainty are wider. Figure 5 similarly shows participant-level estimation error for DX. Compared to the SEQ estimator, there seems to be a wider range of estimation errors in the 29+ day group, and the ranges of the estimation errors for participants in the 1–14 day, 15–28 day, and 29+ day groups largely overlapped (Fig. 5A). There also seemed to be little difference in the participant-level distributions of estimation error for the Ab− and Ab +group (Fig. 5B).

Fig 5.

Boxplots of HIV DX-based DDA estimation error in FRESH/RV217 cohorts. Error spread increases with sampling interval post-COB (1-14 days least, 29+ most). Ab+ shows wider distribution than Ab-. All groups trend toward slight overestimation.

Effects of sampling time and HIV serostatus on DX-based DDA estimates for the 41 selected FRESH and RV217 participants. Boxplots, overlaid with estimation error in days obtained using (A) For both serostatuses pooled (Ab− and Ab+), results are stratified by sampling interval post-COB (1–14 days, 15–28 days, and 29+ days); (B) for all sampling times post-COB pooled, results are stratified by HIV serostatus (Ab− vs. B+). The dotted vertical line represents perfect estimation of DDA in terms of the ground truth (COB). In interior box plots, left and right vertical edges show the 25th and 75th percentiles of the DX estimation error, the middle line indicates the median, and vertical bars extend to the outermost data points within 1.5 times the IQR from these quartiles. For all estimators of DDA, we defined the estimation error as the difference in days between the estimator and the ground truth estimate using the COB between the actual last negative (LN) and first positive (FP) dates in RV217 and FRESH cohort data. Positive values indicate overestimation of the number of days between DDA and FP (i.e., the estimate of DDA is too early compared to COB), and negative values indicate underestimation (i.e., the estimate of DDA is too late).

Integration of diagnostic and sequence data from overlapping 95% CI (Combined)

Figure 6 similarly shows participant-level estimation error for the Combined estimator. These results were more similar to those from the DX estimator in that the ranges of the estimation errors for participants in the 1–14 day, 15–28 day, and 29+ day groups largely overlapped (Fig. 6A). There also seemed to be little difference in the participant-level distributions of estimation error for the Ab− and Ab+ group (Fig. 6B).

Fig 6.

Boxplots comparing Combined-based DDA estimation errors in HIV cohorts. Early sampling shows tighter error distribution than later intervals. Antibody-negative samples show narrower error ranges than antibody-positive ones. Most medians near zero.

Effects of sampling time and HIV serostatus on combined-based DDA estimates for the 41 selected FRESH and RV217 participants. Boxplots, overlaid with estimation error in days obtained using (A) For both serostatuses pooled (Ab−- and Ab+), results are stratified by sampling interval post-COB (1–14 days, 15–28 days, 29+ days); (B) for all sampling times post-COB pooled, results are stratified by HIV serostatus (Ab− vs. Ab+). The dotted vertical line represents perfect estimation of DDA in terms of the ground truth (COB). In interior box plots, left and right vertical edges show the 25th and 75th percentiles of the Combined estimation error, the middle line indicates the median, and vertical bars extend to the outermost data points within 1.5 times the IQR from these quartiles. For all estimators of DDA, we defined the estimation error as the difference in days between the estimator and the ground truth estimate using the COB between the actual LN and FP dates in RV217 and FRESH cohort data. Positive values indicate overestimation of the number of days between DDA and FP (i.e., the estimate of DDA is too early compared to COB), and negative values indicate underestimation (i.e., the estimate of DDA is too late).

Overall (across all sampling times post-COB and both serostatuses), the 95% CI was smaller for the Combined estimator (median width 6 days, IQR 4–21.5 days) than for the DX estimator (median width 15 days, IQR 11–23 days) (Table 2). However, the overall coverage of the COB was lower for the Combined than the DX estimator (0.515 vs. 0.835, respectively). The differences between the Combined and DX estimators were especially notable in Ab− samples (median 95% CI width [IQR] = 3 days [3–5 days] versus 12 days [7–14 days], respectively; coverage 0.31 versus 0.69, respectively) and samples collected up to 28 days post-COB (e.g., for the 1–14 day group, median 95% CI width [IQR] = 3 days [3–4.3 days] versus 10.5 days [7–13 days] and coverage 0.25 versus 0.63, respectively).

TABLE 2.

Coverage of COB by the 95% credible interval and median 95% credible interval width of the DX and of the Combined Bayesian posterior distributions

Parameter n 95% CI coveragea Median (IQR) 95% CI width (days)
SEQ DX Combined SEQ DX Combined
Overallb 103 20.3% 78.6% 51.5% 5 (3.25–7) 17 (14–41) 8 (4–41)
1–14 days 24 20.8% 58.3% 20.8% 3.5 (3–5.25) 14 (9–15.5) 3.5 (3–5.25)
15–28 days 32 25.0% 68.8% 34.4% 6 (3.75–10.25) 16 (14–17.25) 6 (3.75–10.25)
29+ days 32 0.1% 96.9% 84.4% 41 (41–41.25) 41 (41–42) 41 (41–41.25)
Ab− 29 24.1% 62.1% 27.6% 4 (3–6) 14 (10–15) 4 (3–6)
Ab+ 39 26.7% 82.1% 53.8% 6 (4–8) 20 (16-41) 10 (6-41)
a

Defined as the proportion of 95% CIs that include the ground truth estimate (COB).

b

Includes all sampling times post-COB and both HIV serostatuses.

The coverage of an estimation interval refers to how often it contains the true value. What is shown here are credible intervals, which are analogous to confidence intervals; however, unlike frequentist methods, Bayesian methods are not optimized for coverage. The wider CIs of DX vs SEQ need to be considered when assessing coverage information. Still, the results presented in Table 2 and the individual-level plots (Fig. 7 and 8) can provide insights with respect to our approaches.

Fig 7.

Eight plots of DDA estimators for recent HIV acquisition (12-27 days from COB). GP estimator shows narrower confidence intervals than other methods. Ab+ participants demonstrate wider intervals than Ab- participants with varied estimator agreement.

Individual-level plots showing results of 5 of the DDA estimators for eight participants whose COB was less than 5 weeks from the post-acquisition sampling date. Plots show point estimates and 95% CIs for the GP (blue), REN (red), SEQ (green), DX (yellow), and Combined (purple) DDA estimators. For each sampled observation, the COB is shown as a vertical dotted line, and the number of days that the COB is distant from the sampling date is provided at the top of each plot. HIV serostatus is designated as follows: square, Ab+; diamond, Ab−.

Fig 8.

HIV timing estimates from five DDA methods for six antibody-positive RV217 participants. DDA occurred 35-51 days pre-sampling. GP and SEQ estimators show earlier DDA with narrower confidence intervals than DX/Combined methods.

Individual-level plots showing results of five of the DDA estimators for 6 of the studied RV217 participants whose COB was 5 weeks or further from the post-acquisition sampling date. All participants shown had antibody-positive HIV serostatus. Plots show point estimates and 95% CIs for the GP (blue), REN (red), SEQ (green), DX (yellow), and Combined (purple) DDA estimators. For each sampled observation, the COB is shown as a vertical dotted line, and the number of days that the COB is distant from the sampling date is provided at the top of each plot.

Figure 7 shows eight examples of participants with COB less than 5 weeks before the sampling collection date, whereas Fig. 8 shows six examples of participants with COB 5 weeks or further from the sampling collection date. Both figures present point estimates and 95% CIs for five of the estimators (GP, REN, SEQ, DX, and Combined), as well as the COB, for 14 selected participants (2 FRESH, 12 RV217). Analogous plots for all FRESH and RV217 participant data evaluated in this study are shown in Fig. S1. As we see in the individual plots, the addition of overlapping streams of data gives us an approach that is generally closer to the ground truth (COB).

Panels A–D in Fig. 7 provide examples where the GP, REN, and SEQ estimators had narrower 95% CIs compared to the DX estimator; moreover, the Combined estimator had the tightest 95% CIs compared to any of the individual estimators. Panels F and H in Fig. 7 provide examples where the GP estimate was more distant from the COB than any of the other individual estimators.

In all panels of Fig. 8, which includes data from participants whose COB was >5 weeks from the post-acquisition sampling date, the DX and Combined estimators were nearly identical in terms of point estimates and 95% CIs. Generally, the point estimates of the DX and Combined estimators were much closer to the COB than the point estimates of the GP, REN, and SEQ estimators, which tended to underestimate the DDA by a large amount.

Evaluation of the performance of the combined and DX estimators using machine learning

To assess the performance of the Combined and DX estimators, we used 10-fold cross-validation to calibrate each estimator (see Materials and Methods and Fig. S2 and S3 for details). For the CV-calibrated Combined estimator, the adjusted R2 was 0.879 when days since COB were plotted against days since DDA, with RMSE of 5.8 days for the CV-calibrated Combined estimator compared to RMSE of 6.1 days for the original Combined estimator (P = 0.075 for improvement with calibration) (Fig. 9A). Results were similar for the DX estimator, with a reduction in the RMSE from 5.6 days using the original DX estimator to 5.4 days using the CV-calibrated DX estimator (P = 0.191 for improvement with calibration) (Fig. 9C). Based on these results, we conclude that 10-fold cross-validation calibration resulted in modest improvement of the Combined and the DX estimators.

Fig 9.

Scatter plots comparing timing estimators. Combined and DX methods show correlation between original and CV-calibrated DDA estimates relative to COB. Calibration improves accuracy with reduced RMSE as data points align closer to reference lines.

Machine learning cross-validation (CV) evaluation of the ability of the Combined and DX estimators to estimate the DDA. Plots show original vs. cross-validation calibrated estimates of the DDA obtained using (A, B) the Combined estimator and (C, D) the DX estimator. In panels A and C, the estimated number of days that the sequenced sample was collected after the COB is plotted against the estimated number of days that the sequenced sample was collected since the DDA. Gray open circles represent original estimates for each estimator; black cross symbols represent CV-calibrated estimates for each estimator. In panels B (Combined estimator) and D (DX estimator), the estimated number of days that the sequenced sample was collected after the DDA (original estimate) is plotted against the estimated number of days that the sequenced sample was collected after the DDA (calibrated estimate). CV was done by linear regression with the intercept fixed at the origin (see Materials and Methods). COB, center of bounds between last-negative and first-positive RNA diagnostic tests (ground truth); DDA, date of detectable acquisition; RMSE, root mean square error.

DISCUSSION

Here we leveraged a first-of-its-kind data set, distinguished by its highly frequent (twice-weekly) blood sample collection and longitudinal follow-up (up to 60 days post-acquisition), and generated rich HIV viral deep sequence data to evaluate the performance of different timing estimators of HIV acquisition. We employed the AMP Studies’ sequence and diagnostic test results analysis pipeline for estimating the date of detectable acquisition (DDA) using these sequence data and diagnostic data, which depends on Poisson Fitter and Infection Dating Tool (IDT) via new Bayesian implementations. We then compared these estimates to the “ground truth” DDAs (the COB) for each of 41 evaluated participants (13 from FRESH, 28 from RV217), using 1–3 time points for each participant. We optimized the sampling design with a range of days since diagnosis and HIV serostatus to match the sampling observed in the AMP Studies with their monthly study visits. Overall, the best DDA estimators (in terms of yielding lowest bias and root mean square error [RMSE]) were sequence-based estimation when restricting to samples from ≤5 weeks post-DDA, 0.2 days and 7.8 days, respectively; sequence + diagnostic data, 1.5 days and 6.9 days, respectively; diagnostic data alone, 2.4 days and 7.9 days, respectively.

Excellent work (reviewed in [40]) has already been conducted using HIV sequence data (such as HIV sequence diversity estimated by next-generation sequencing) and/or serological/immunological biomarker data (such as CD4+ T-cell count, immunoassays for detection of recent HIV acquisition) from observational or surveillance studies to estimate timing of HIV acquisition (4145). Given that these observational and surveillance studies did not employ as highly frequent participant sampling, the estimated times since acquisition are generally longer, typically on the order of weeks or months, with measurements of uncertainty also on the order of weeks or months. Such results are well-suited to applications such as estimating annual HIV incidence, but have insufficient accuracy and precision for our application, which is the estimation of HIV acquisition timing in clinical efficacy trials for use in immune correlates and sieve analyses.

Consistent with previous application of Poisson Fitter to data from non-human primate studies (29), bias in the DDA estimate was reduced when Poisson Fitter was used on sequence data from samples collected within 5 weeks of the COB, or when evaluating only HIV-seronegative samples. For samples collected more than 5 weeks after COB, we found that using sequence data alone led to underestimation of the DDA, consistent with selection induced by adaptive immunity (21, 29) or changes in viral fitness or, alternatively, potentially explained by a shift to non-exponential population growth, for example, due to target cell limitations. Observational (34) and recent modeling work (46) characterizing early post-acquisition driving forces behind HIV evolution has suggested that adaptive immunity may not play a major role in driving viral evolution, highlighting the complicated interplay between population dynamics and evolution.

We found that these results were largely consistent across the two sequenced regions (GP [gag/Δpol] and REN [rev/vpu/env/Δnef]) of the HIV genome, which are located at opposite extremes, despite representing potentially different selection pressure patterns, albeit another potential explanation is a shift in population dynamics to non-exponential growth.

We found that both the GP and REN estimators had good performance, with most bias (average difference between estimated DDA and true DDA) limited to ≤1 week. While the DDA point estimates differed slightly when using the GP vs. REN regions, there seemed to be modestly better agreement with REN than GP. However, it should be noted that participants were not receiving VRC01 as in the AMP Studies, and therefore these results are more translatable to the placebo group. In future clinical trials of combination bnAb regimens or multispecific bnAbs for HIV prevention, it will be important to determine the timing of HIV acquisition in the presence of bnAbs, with bnAb concentration being a major consideration. We are planning to interrogate and modify our methods to accommodate this consideration as we look more closely at data from the AMP trials. In previous work, we have utilized VRC01 pharmacokinetic and pharmacodynamic data, along with in vitro virus isolate-specific neutralization data, to generate a mathematical model of viral load and assess whether/how infused VRC01 influences viral load (47). This model also estimated acquisition timing, where the predicted serum neutralization 80% inhibitory dilution titers at acquisition estimated by the viral dynamics-based model (47) were similar to those obtained using our previous approach (13), incorporating diagnostic and sequence-based information.

While we are planning future work to quantify whether and how the GP and REN acquisition timing estimates differed in the AMP Studies, the present results suggest that there would be a difference between the GP vs. REN estimates in the control arms, especially with samples collected within 5 weeks of the true DDA. The collection of data from different sources (GP, REN, and DX) provides us with a valuable way to study the potential effects of VRC01 and the ability to compare results here with what we find in the placebo, high-, and low-dose VRC01 arms of AMP.

We found that, for Ab− samples and for samples collected within about 1 month of RNA conversion (up to ~5 weeks post-acquisition, accounting for the eclipse phase), the estimators had both low bias and low error variation. Moreover, for samples collected at times more distant from the time of RNA conversion, the estimator bias verged toward predicting the DDA closer to the time of sample collection. In other words, all estimators tended to underestimate the time that had passed from RNA conversion to sample collection. The diagnostic-based estimators also yielded similar bias but generally wider 95% CIs. In future applications, we may conduct a calibration process for the diagnostic test results (see the slope of the line in Fig. 2D), as we previously conducted (13), to narrow these 95% CIs.

In addition to selection-induced reductions in diversity, our processing pipeline addresses the complexity in sequence diversity of these later samples using techniques that tend to further reduce diversity, including masking regions identified as under targeted selection as described above, and APOBEC and recombinant filtering. The need for techniques that can correct estimates that are biased downward by these processes is a limitation of this work. Moreover, as the strength of the different forms of selection (diversifying, purifying) is unknown, the rate of evolution is unknown for a given individual, which poses a challenge to the general approach here. Future work may include the development and application of techniques designed to correct these later-time-point estimates and to evaluate the extent to which diagnostic test data can help improve them. Noting that the Ab+ cases have the least informative diagnostic-based timing information, such techniques may be most relevant for Ab− cases.

Findings of this study may impact future projects on multiple levels. First, as we found iterative analyses to be critical for optimizing sequence pre-processing steps, our work underscores the importance of such iteration. Second, our findings suggest that monthly, or more frequent, sampling will be important in potential future clinical trials of HIV prevention modalities for which unbiased sequence-based estimation of elapsed time since DDA or acquisition is important to the analysis plan, for example, in planned trials of combination bnAb regimens or multi-specific bnAbs for HIV prevention. However, the combined predictor we describe here maintains the good qualities of the diagnostic-based predictors beyond this period. For studies in which some participants are Ab+ at first sample collection, it may be beneficial for analyses of sequence data to focus on the Ab− or first-month sample cases, because sequence-based timing for Ab+ or later cases may be biased downward. Third, some of the most time-consuming aspects of the sequence processing work (e.g., iterative lineage classification for good Poisson fit) can be obviated, specifically for cases that are Ab+ or were sampled more than 4–5 weeks since RNA conversion. Determining this has saved us modest time and energy on this project but will greatly reduce the burden in future studies.

We note that these RV217 and FRESH participants were not on PrEP nor had received any bnAb infusions. As such, modalities may impact timing estimation; a limitation of this study is that its findings may not be applicable to individuals on PrEP until further research is performed to understand whether and how receipt of PrEP impacts timing estimation, although delays in diagnostic indicators have been reported previously (48). Moreover, PrEP can be delivered by multiple methods using one of several different drugs (49), with some variance in effectiveness across PrEP regimens (50)—raising the possibility that timing estimators may need to vary by delivery method used by the participant. Nevertheless, our results provide an important step toward understanding the contributions of sequence and diagnostic information when estimating DDA. Another limitation of our study is that our target of inference was the DDA, which occurs approximately 1 week after actual virus acquisition (6), with the latter target being most relevant in terms of HIV prevention. However, there are insufficient available data sets with actual times of acquisition, and thus our study comes as close to this goal as is possible.

The present article reports on the development and application of the method used for estimating HIV estimated DDAs in the AMP Studies. In the context of blinded selection of parameters for application to AMP, our pre-specified statistical analysis plan constrained our approach, and there was no opportunity for optimization of the procedure beyond tweaking choices of how to resolve inconsistent sources of information. Thus, before unblinded analysis of AMP data but after evaluating the FRESH and RV217 data presented here, we were able to alter final algorithm choices based on the evidence presented here of a violation of our prior assumption that observed evolutionary rates across these two regions of the genome would be concordant in the absence of antibody-induced selective pressure. This led to us altering the choice of which sequence-based estimate to use when the two estimates are not concordant. Otherwise, we were bound to use the procedure as described, which simply implemented existing procedures and combined them with Bayesian logic. It remains for future work to optimize the parameters of this procedure and its input procedures. Indeed, these are precious data for calibrating both the sequence-based estimators as well as the diagnostic testing-based estimators that are components of our final estimates, as these sequences are the first human data available for calibrating the former and these diagnostic test results offer a rare new opportunity to increase and diversify the available data for calibrating the inter-test reactivity intervals that the IDT estimators, and ours are currently relying on, which were estimated from American blood bank donors as described in Delaney et al. (31). It is astonishing how well-calibrated the sequence-based estimators are without such optimization, considering that the evolutionary rate is assumed fixed across the whole genome. Here we show with a cursory calibration of the final estimators that improvement is possible, and that, for example, the diagnostic-based estimators may be independently calibrated to improve their accuracy, suggesting that further test-specific calibration in a cross-validated analysis is likely to yield further improvements. We leave this important next step to future work.

In summary, we conducted a sequencing study to employ the same PacBio sequencing processes and sequence analysis pipeline that had been applied in the AMP Studies to a set of data for which the true earliest DDA is known with high accuracy. We conducted sequencing and analysis of 41 people who were experiencing acute HIV infection at the time of their sample collection, for whom the true date of conversion to a state of having detectable HIV RNA in the blood was known to within only a few days (time between last-negative and first-positive HIV RNA detection dates was median 4 days; range, 2–7 days). We show that our timing methods perform well on human samples, demonstrating that sequence-based estimation of the DDA in humans using a variant of the Poisson Fitter methodology yields reasonably accurate and precise estimates for both antibody-negative samples (bias 4.4 days, RMSE 7.5 days) and samples collected up to 5 weeks post-acquisition (bias –0.1 to 5.1 days, RMSE 5.8–8.0 days), but not samples collected after 5 weeks post-acquisition (bias −15.1 days, RMSE 19.6 days). When the sample is antibody negative, these robust error characteristics hold, which might be helpful for application in cases in which the true number of days since RNA seroconversion cannot be limited to an earlier sampling time window. We demonstrate that our HIV time estimation approach, integrating sequence-based estimates with diagnostic-based estimates, yields unbiased approximations of the DDA with about 5–6 days of expected error as approximated by the RMSE. Further error introduced by uncertainty regarding the eclipse phase duration limits the possible accuracy of the final downstream estimates of exposure times, which we computed in the AMP Studies by convolution of these DDA estimators with a distribution reflecting variation in the duration of the eclipse phase. This has important implications for the power of analyses to evaluate correlates of HIV acquisition risk that depend on estimating antibody concentration at the time of exposure (13).

ACKNOWLEDGMENTS

We thank all study participants and staff from the RV217 and FRESH cohorts, without whom this work would not have been possible. We also thank the HVTN 703/HPTN 081 and HVTN 704/HPTN 085 trial teams and participants. We thank William Hahn, Ann C. Duerr, and Gail Broder for careful review of the manuscript.

This work was supported, in whole or in part, by the Bill and Melinda Gates Foundation (Investment Record ID INV-016189 to J.I.M.). Under the grant conditions of the Foundation, a Creative Commons Attribution 4.0 Generic License has already been assigned to the Author Accepted Manuscript version that might arise from this submission. This work was also supported in part by the National Institute of Allergy and Infectious Diseases, National Institutes of Health awards UM1AI068635 (to PBG), R37AI054165 (to P.B.G.), and K25 AI155224 (to D.B.R.). M.R., M.L.R., and L.A.E. received support from a cooperative agreement (W81XWH-18-2-0040) between the Henry M. Jackson Foundation for the Advancement of Military Medicine and the U.S. Department of Defense.

The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health. Material has been reviewed by the Walter Reed Army Institute of Research. There is no objection to its presentation and/or publication. The opinions or assertions contained herein are the private views of the author and are not to be construed as official or as reflecting true views of the Department of the Army, the Department of Defense, or HJF (M.R., M.L.R., L.A.E.). This paper is the result of funding in whole or in part by the National Institutes of Health (NIH). It is subject to the NIH Public Access Policy. Through acceptance of this federal funding, NIH has been given a right to make the manuscript publicly available in PubMed Central upon the official date of publication, as defined by NIH.

Contributor Information

Paul T. Edlefsen, Email: pedlefse@fredhutch.org.

Suresh Mahalingam, Emory University Vaccine Center, Gold Coast, Queensland, Australia.

ETHICS APPROVAL

All work described here complied with all relevant ethical regulations. For the AMP Studies, central- and site-specific institutional review boards and ethics committees reviewed and approved the initial protocol and each subsequent version. All participants provided written informed consent, and new consent was obtained for each version of the protocol. The FRESH study was approved by the biomedical research ethics committee of the University of KwaZulu-Natal and the institutional review board of Massachusetts General Hospital. All participants provided written consent to enrollment. The RV217 protocol was approved by the local ethics review boards and the Walter Reed Army Institute of Research. The investigators have adhered to the policies for protection of human participants as prescribed in AR 70–25. Written informed consent was obtained from all participants.

DATA AVAILABILITY

A preliminary data set of sequences was used for the current study. A more heavily curated data set is available at GenBank via accession numbers PV963961 to PV968620 (FRESH GP), PV968621 to PV972048 (FRESH REN), PX001706 to PX010774 (RV217 GP), and PX010775 to PX020890 (RV217 REN). Code is available at https://github.com/HVTN-SDMC/AMP_FRESH_RV217_AcquisitionTimingPaper.

SUPPLEMENTAL MATERIAL

The following material is available online at https://doi.org/10.1128/mbio.01881-25.

Figure S1. mbio.01881-25-s0001.pdf.

Results of 5 of the DDA estimators for all studied participants in the RV217 and FRESH cohorts.

mbio.01881-25-s0001.pdf (88.1KB, pdf)
DOI: 10.1128/mbio.01881-25.SuF1
Figures S2 and S3. mbio.01881-25-s0002.docx.

Additional results related to Fig. 9.

DOI: 10.1128/mbio.01881-25.SuF2

ASM does not own the copyrights to Supplemental Material that may be linked to, or accessed through, an article. The authors have granted ASM a non-exclusive, world-wide license to publish the Supplemental Material files. Please contact the corresponding author directly for reuse.

REFERENCES

  • 1. Corey L, Gilbert PB, Juraska M, Montefiori DC, Morris L, Karuna ST, Edupuganti S, Mgodi NM, deCamp AC, Rudnicki E, et al. 2021. Two randomized trials of neutralizing antibodies to prevent HIV-1 acquisition. N Engl J Med 384:1003–1014. doi: 10.1056/NEJMoa2031738 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2. Huang Y, Naidoo L, Zhang L, Carpp LN, Rudnicki E, Randhawa A, Gonzales P, McDermott A, Ledgerwood J, Lorenzo MMG, Burns D, DeCamp A, Juraska M, Mascola J, Edupuganti S, Mgodi N, Cohen M, Corey L, Andrew P, Karuna S, Gilbert PB, Mngadi K, Lazarus E. 2021. Pharmacokinetics and predicted neutralisation coverage of VRC01 in HIV-uninfected participants of the Antibody Mediated Prevention (AMP) trials. EBioMedicine 64:103203. doi: 10.1016/j.ebiom.2020.103203 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3. Huang Y, Zhang L, Eaton A, Mkhize NN, Carpp LN, Rudnicki E, DeCamp A, Juraska M, Randhawa A, McDermott A, Ledgerwood J, Andrew P, Karuna S, Edupuganti S, Mgodi N, Cohen M, Corey L, Mascola J, Gilbert PB, Morris L, Montefiori DC. 2022. Prediction of serum HIV-1 neutralization titers of VRC01 in HIV-uninfected antibody mediated prevention (AMP) trial participants. Hum Vaccin Immunother 18:1908030. doi: 10.1080/21645515.2021.1908030 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4. Gilbert PB, Huang Y, deCamp AC, Karuna S, Zhang Y, Magaret CA, Giorgi EE, Korber B, Edlefsen PT, Rossenkhan R, et al. 2022. Neutralization titer biomarker for antibody-mediated prevention of HIV-1 acquisition. Nat Med 28:1924–1932. doi: 10.1038/s41591-022-01953-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5. Seaton KE, Huang Y, Karuna S, Heptinstall JR, Brackett C, Chiong K, Zhang L, Yates NL, Sampson M, Rudnicki E, et al. 2023. Pharmacokinetic serum concentrations of VRC01 correlate with prevention of HIV-1 acquisition. EBioMedicine 93:104590. doi: 10.1016/j.ebiom.2023.104590 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6. Rolland M, Tovanabutra S, Dearlove B, Li Y, Owen CL, Lewitus E, Sanders-Buell E, Bose M, O’Sullivan A, Rossenkhan R, et al. 2020. Molecular dating and viral load growth rates suggested that the eclipse phase lasted about a week in HIV-1 infected adults in East Africa and Thailand. PLoS Pathog 16:e1008179. doi: 10.1371/journal.ppat.1008179 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. Fiebig EW, Wright DJ, Rawal BD, Garrett PE, Schumacher RT, Peddada L, Heldebrant C, Smith R, Conrad A, Kleinman SH, Busch MP. 2003. Dynamics of HIV viremia and antibody seroconversion in plasma donors: implications for diagnosis and staging of primary HIV infection. AIDS 17:1871–1879. doi: 10.1097/00002030-200309050-00005 [DOI] [PubMed] [Google Scholar]
  • 8. Pilcher CD, Porco TC, Facente SN, Grebe E, Delaney KP, Masciotra S, Kassanjee R, Busch MP, Murphy G, Owen SM, Welte A, Consortium for the Evaluation and Performance of HIV Incidence Assays (CEPHIA) . 2019. A generalizable method for estimating duration of HIV infections using clinical testing history and HIV test results. AIDS 33:1231–1240. doi: 10.1097/QAD.0000000000002190 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9. Grebe E, Facente SN, Bingham J, Pilcher CD, Powrie A, Gerber J, Priede G, Chibawara T, Busch MP, Murphy G, Kassanjee R, Welte A, Consortium for the Evaluation and Performance of HIV Incidence Assays (CEPHIA) . 2019. Interpreting HIV diagnostic histories into infection time estimates: analytical framework and online tool. BMC Infect Dis 19:894. doi: 10.1186/s12879-019-4543-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10. Facente SN, Grebe E, Pilcher CD, Busch MP, Murphy G, Welte A. 2020. Estimated dates of detectable infection (EDDIs) as an improvement upon Fiebig staging for HIV infection dating. Epidemiol Infect 148:e53. doi: 10.1017/S0950268820000503 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. Bouckaert R, Vaughan TG, Barido-Sottani J, Duchêne S, Fourment M, Gavryushkina A, Heled J, Jones G, Kühnert D, De Maio N, Matschiner M, Mendes FK, Müller NF, Ogilvie HA, du Plessis L, Popinga A, Rambaut A, Rasmussen D, Siveroni I, Suchard MA, Wu C-H, Xie D, Zhang C, Stadler T, Drummond AJ. 2019. BEAST 2.5: an advanced software platform for Bayesian evolutionary analysis. PLoS Comput Biol 15:e1006650. doi: 10.1371/journal.pcbi.1006650 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12. Giorgi EE, Funkhouser B, Athreya G, Perelson AS, Korber BT, Bhattacharya T. 2010. Estimating time since infection in early homogeneous HIV-1 samples using a poisson model. BMC Bioinformatics 11:532. doi: 10.1186/1471-2105-11-532 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. Rossenkhan R, Rolland M, Labuschagne JPL, Ferreira R-C, Magaret CA, Carpp LN, Matsen Iv FA, Huang Y, Rudnicki EE, Zhang Y, et al. 2019. Combining viral genetics and statistical modeling to improve HIV-1 time-of-infection estimation towards enhanced vaccine efficacy assessment. Viruses 11:607. doi: 10.3390/v11070607 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Robb ML, Eller LA, Kibuuka H, Rono K, Maganga L, Nitayaphan S, Kroon E, Sawe FK, Sinei S, Sriplienchan S, et al. 2016. Prospective study of acute HIV-1 infection in adults in East Africa and Thailand. N Engl J Med 374:2120–2130. doi: 10.1056/NEJMoa1508952 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15. Mlisana K, Werner L, Garrett NJ, McKinnon LR, van Loggerenberg F, Passmore J-AS, Gray CM, Morris L, Williamson C, Abdool Karim SS, Centre for the AIDS Programme of Research in South Africa (CAPRISA) 002 Study Team . 2014. Rapid disease progression in HIV-1 subtype C-infected South African women. Clin Infect Dis 59:1322–1331. doi: 10.1093/cid/ciu573 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. Alser M, Lawlor B, Abdill RJ, Waymost S, Ayyala R, Rajkumar N, LaPierre N, Brito J, Ribeiro-Dos-Santos AM, Almadhoun N, Sarwal V, Firtina C, Osinski T, Eskin E, Hu Q, Strong D, Kim B-DBD, Abedalthagafi MS, Mutlu O, Mangul S. 2024. Packaging and containerization of computational methods. Nat Protoc 19:2529–2539. doi: 10.1038/s41596-024-00986-0 [DOI] [PubMed] [Google Scholar]
  • 17. Dong KL, Moodley A, Kwon DS, Ghebremichael MS, Dong M, Ismail N, Ndhlovu ZM, Mabuka JM, Muema DM, Pretorius K, Lin N, Walker BD, Ndung’u T. 2018. Detection and treatment of Fiebig stage I HIV-1 infection in young at-risk women in South Africa: a prospective cohort study. Lancet HIV 5:e35–e44. doi: 10.1016/S2352-3018(17)30146-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18. Ndung’u T, Dong KL, Kwon DS, Walker BD. 2018. A FRESH approach: combining basic science and social good. Sci Immunol 3:eaau2798. doi: 10.1126/sciimmunol.aau2798 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. Westfall DH, Deng W, Pankow A, Murrell H, Chen L, Zhao H, Williamson C, Rolland M, Murrell B, Mullins JI. 2024. Optimized SMRT-UMI protocol produces highly accurate sequence datasets from diverse populations-Application to HIV-1 quasispecies. Virus Evol 10:veae019. doi: 10.1093/ve/veae019 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20. Anonymous . 2024. Correction to: optimized SMRT-UMI protocol produces highly accurate sequence datasets from diverse populations—Application to HIV-1 quasispecies. Virus Evol 10:veae038. doi: 10.1093/ve/veae038 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21. Keele BF, Giorgi EE, Salazar-Gonzalez JF, Decker JM, Pham KT, Salazar MG, Sun C, Grayson T, Wang S, Li H, et al. 2008. Identification and characterization of transmitted and early founder virus envelopes in primary HIV-1 infection. Proc Natl Acad Sci U S A 105:7552–7557. doi: 10.1073/pnas.0802203105 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. Furlong JC, Darley PD, Deng W, Mullins JI, Bumgarner RE. 2024. Phylobook: a tool for display, clade annotation and extraction of sequences from molecular phylogenies. Biotechniques 76:263–274. doi: 10.2144/btn-2023-0056 [DOI] [PubMed] [Google Scholar]
  • 23. Deng W, Maust BS, Nickle DC, Learn GH, Liu Y, Heath L, Kosakovsky Pond SL, Mullins JI. 2010. DIVEIN: a web server to analyze phylogenies, sequence divergence, diversity, and informative sites. Biotechniques 48:405–408. doi: 10.2144/000113370 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Lee HY, Giorgi EE, Keele BF, Gaschen B, Athreya GS, Salazar-Gonzalez JF, Pham KT, Goepfert PA, Kilby JM, Saag MS, Delwart EL, Busch MP, Hahn BH, Shaw GM, Korber BT, Bhattacharya T, Perelson AS. 2009. Modeling sequence evolution in acute HIV-1 infection. J Theor Biol 261:341–360. doi: 10.1016/j.jtbi.2009.07.038 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25. Chen J, MacCarthy T. 2017. The preferred nucleotide contexts of the AID/APOBEC cytidine deaminases have differential effects when mutating retrotransposon and virus sequences compared to host genes. PLoS Comput Biol 13:e1005471. doi: 10.1371/journal.pcbi.1005471 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Malim MH. 2009. APOBEC proteins and intrinsic resistance to HIV-1 infection. Philos Trans R Soc Lond B Biol Sci 364:675–687. doi: 10.1098/rstb.2008.0185 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27. Simon V, Bloch N, Landau NR. 2015. Intrinsic host restrictions to HIV-1 and mechanisms of viral escape. Nat Immunol 16:546–553. doi: 10.1038/ni.3156 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28. Lapp Z, Yoon H, Foley B, Leitner T. 2024. Hypermut 3: Identifying specific mutational patterns in a defined nucleotide context that allows multistate characters. bioRxiv:2024.10.24.620069. doi: 10.1101/2024.10.24.620069 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29. Giorgi EE, Li H, Bhattacharya T, Shaw GM, Korber B. 2020. Estimating the timing of early simian-human immunodeficiency virus infections: a comparison between poisson fitter and BEAST. mBio 11:e00324-20. doi: 10.1128/mBio.00324-20 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30. Giorgi EE, Korber BT, Perelson AS, Bhattacharya T. 2013. Modeling sequence evolution in HIV-1 infection with recombination. J Theor Biol 329:82–93. doi: 10.1016/j.jtbi.2013.03.026 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31. Delaney KP, Hanson DL, Masciotra S, Ethridge SF, Wesolowski L, Owen SM. 2017. Time until emergence of HIV test reactivity following infection with HIV-1: implications for interpreting test results and retesting after exposure. Clin Infect Dis 64:53–59. doi: 10.1093/cid/ciw666 [DOI] [PubMed] [Google Scholar]
  • 32. Joseph SB, Swanstrom R, Kashuba ADM, Cohen MS. 2015. Bottlenecks in HIV-1 transmission: insights from the study of founder viruses. Nat Rev Microbiol 13:414–425. doi: 10.1038/nrmicro3471 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33. Tully DC, Ogilvie CB, Batorsky RE, Bean DJ, Power KA, Ghebremichael M, Bedard HE, Gladden AD, Seese AM, Amero MA, et al. 2016. Differences in the selection bottleneck between modes of sexual transmission influence the genetic composition of the HIV-1 founder virus. PLoS Pathog 12:e1005619. doi: 10.1371/journal.ppat.1005619 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34. Herbeck JT, Rolland M, Liu Y, McLaughlin S, McNevin J, Zhao H, Wong K, Stoddard JN, Raugi D, Sorensen S, Genowati I, Birditt B, McKay A, Diem K, Maust BS, Deng W, Collier AC, Stekler JD, McElrath MJ, Mullins JI. 2011. Demographic processes affect HIV-1 evolution in primary infection before the onset of selective processes. J Virol 85:7523–7534. doi: 10.1128/JVI.02697-10 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35. Shriner D, Rodrigo AG, Nickle DC, Mullins JI. 2004. Pervasive genomic recombination of HIV-1 in vivo. Genetics 167:1573–1583. doi: 10.1534/genetics.103.023382 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36. Team RC. 2021. R: a language and environment for statistical computing. R Foundation for Statistical Computing, Vienna, Austria. Available from: https://www.R-project.org [Google Scholar]
  • 37. Lynch RM, Boritz E, Coates EE, DeZure A, Madden P, Costner P, Enama ME, Plummer S, Holman L, Hendel CS, et al. 2015. Virologic effects of broadly neutralizing antibody VRC01 administration during chronic HIV-1 infection. Sci Transl Med 7:319ra206. doi: 10.1126/scitranslmed.aad5752 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38. Bar KJ, Sneller MC, Harrison LJ, Justement JS, Overton ET, Petrone ME, Salantes DB, Seamon CA, Scheinfeld B, Kwan RW, et al. 2016. Effect of HIV antibody VRC01 on viral rebound after treatment interruption. N Engl J Med 375:2037–2050. doi: 10.1056/NEJMoa1608243 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39. Cale EM, Bai H, Bose M, Messina MA, Colby DJ, Sanders-Buell E, Dearlove B, Li Y, Engeman E, Silas D, et al. 2020. Neutralizing antibody VRC01 failed to select for HIV-1 mutations upon viral rebound. J Clin Invest 130:3299–3304. doi: 10.1172/JCI134395 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40. Moyo S, Wilkinson E, Novitsky V, Vandormael A, Gaseitsiwe S, Essex M, Engelbrecht S, de Oliveira T. 2015. Identifying recent HIV infections: from serological assays to genomics. Viruses 7:5508–5524. doi: 10.3390/v7102887 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41. Puller V, Neher R, Albert J. 2017. Estimating time of HIV-1 infection from next-generation sequence diversity. PLoS Comput Biol 13:e1005775. doi: 10.1371/journal.pcbi.1005775 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42. Sommen C, Commenges D, Vu SL, Meyer L, Alioum A. 2011. Estimation of the distribution of infection times using longitudinal serological markers of HIV: implications for the estimation of HIV incidence. Biometrics 67:467–475. doi: 10.1111/j.1541-0420.2010.01473.x [DOI] [PubMed] [Google Scholar]
  • 43. Romero-Severson EO, Lee Petrie C, Ionides E, Albert J, Leitner T. 2015. Trends of HIV-1 incidence with credible intervals in Sweden 2002-09 reconstructed using a dynamic model of within-patient IgG growth. Int J Epidemiol 44:998–1006. doi: 10.1093/ije/dyv034 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44. Giardina F, Romero-Severson EO, Axelsson M, Svedhem V, Leitner T, Britton T, Albert J. 2019. Getting more from heterogeneous HIV-1 surveillance data in a high immigration country: estimation of incidence and undiagnosed population size using multiple biomarkers. Int J Epidemiol 48:1795–1803. doi: 10.1093/ije/dyz100 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45. Lundgren E, Romero-Severson E, Albert J, Leitner T. 2022. Combining biomarker and virus phylogenetic models improves HIV-1 epidemiological source identification. PLoS Comput Biol 18:e1009741. doi: 10.1371/journal.pcbi.1009741 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46. Swan DA, Rolland M, Herbeck JT, Schiffer JT, Reeves DB. 2022. Evolution during primary HIV infection does not require adaptive immune selection. Proc Natl Acad Sci U S A 119:e2109172119. doi: 10.1073/pnas.2109172119 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47. Reeves DB, Mayer BT, deCamp AC, Huang Y, Zhang B, Carpp LN, Magaret CA, Juraska M, Gilbert PB, Montefiori DC, et al. 2023. High monoclonal neutralization titers reduced breakthrough HIV-1 viral loads in the Antibody Mediated Prevention trials. Nat Commun 14:8299. doi: 10.1038/s41467-023-43384-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48. Donnell D, Ramos E, Celum C, Baeten J, Dragavon J, Tappero J, Lingappa JR, Ronald A, Fife K, Coombs RW. 2017. The effect of oral preexposure prophylaxis on the progression of HIV-1 seroconversion. AIDS 31:2007–2016. doi: 10.1097/QAD.0000000000001577 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49. Rivera CG, Zeuli JD, Smith BL, Johnson TM, Bhatia R, Otto AO, Temesgen Z. 2023. HIV pre-exposure prophylaxis: new and upcoming drugs to address the HIV epidemic. Drugs (Abingdon Engl) 83:1677–1698. doi: 10.1007/s40265-023-01963-9 [DOI] [PubMed] [Google Scholar]
  • 50. Chou R, Spencer H, Bougatsos C, Blazina I, Ahmed A, Selph S. 2023. Prophylaxis for the prevention of HIV: updated evidence report and systematic review for the US preventive services task force. JAMA 330:746–763. doi: 10.1001/jama.2023.9865 [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Figure S1. mbio.01881-25-s0001.pdf.

Results of 5 of the DDA estimators for all studied participants in the RV217 and FRESH cohorts.

mbio.01881-25-s0001.pdf (88.1KB, pdf)
DOI: 10.1128/mbio.01881-25.SuF1
Figures S2 and S3. mbio.01881-25-s0002.docx.

Additional results related to Fig. 9.

DOI: 10.1128/mbio.01881-25.SuF2

Data Availability Statement

A preliminary data set of sequences was used for the current study. A more heavily curated data set is available at GenBank via accession numbers PV963961 to PV968620 (FRESH GP), PV968621 to PV972048 (FRESH REN), PX001706 to PX010774 (RV217 GP), and PX010775 to PX020890 (RV217 REN). Code is available at https://github.com/HVTN-SDMC/AMP_FRESH_RV217_AcquisitionTimingPaper.


Articles from mBio are provided here courtesy of American Society for Microbiology (ASM)

RESOURCES