Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2016 Apr 30.
Published in final edited form as: Stat Med. 2015 Nov 15;35(9):1471–1487. doi: 10.1002/sim.6795

A Bayesian hierarchical model with novel prior specifications for estimating HIV testing rates

Qian An a, Jian Kang b,*,, Ruiguang Song a, H Irene Hall a
PMCID: PMC4845103  NIHMSID: NIHMS778831  PMID: 26567891

Abstract

Human immunodeficiency virus (HIV) infection is a severe infectious disease actively spreading globally, and acquired immunodeficiency syndrome (AIDS) is an advanced stage of HIV infection. The HIV testing rate, that is, the probability that an AIDS-free HIV infected person seeks a test for HIV during a particular time interval, given no previous positive test has been obtained prior to the start of the time, is an important parameter for public health. In this paper, we propose a Bayesian hierarchical model with two levels of hierarchy to estimate the HIV testing rate using annual AIDS and AIDS-free HIV diagnoses data. At level one, we model the latent number of HIV infections for each year using a Poisson distribution with the intensity parameter representing the HIV incidence rate. At level two, the annual numbers of AIDS and AIDS-free HIV diagnosed cases and all undiagnosed cases stratified by the HIV infections at different years are modeled using a multinomial distribution with parameters including the HIV testing rate. We propose a new class of priors for the HIV incidence rate and HIV testing rate taking into account the temporal dependence of these parameters to improve the estimation accuracy. We develop an efficient posterior computation algorithm based on the adaptive rejection metropolis sampling technique. We demonstrate our model using simulation studies and the analysis of the national HIV surveillance data in the USA.

Keywords: HIV testing rate, Bayesian hierarchical model, temporal dependence, adaptive rejection metropolis sampling

1. Introduction

Human immunodeficiency virus (HIV) infection is a severe infectious disease actively spreading globally. Acquired immunodeficiency syndrome (AIDS) is an advanced stage of HIV infection with CD4 T-lymphocyte count less than 200/uL or opportunistic illnesses. The diagnosis of HIV infection by HIV testing is one of the most important tools for HIV prevention and treatment. It can foster early detection of HIV infection and is the first essential step for entry to clinical care to reduce morbidity and mortality. In addition, persons aware of their infections and on effective treatment are less likely to transmit the virus to others. Since HIV testing first became available in 1985, the importance of HIV testing was soon recognized, emphasized, and promoted [1]. However, a significant proportion of individuals infected with HIV still remain untested, undiagnosed, and unaware of their infection. As of December 2010, more than 1.1 million people were living with HIV infection in the USA and about 1 in 6 were unaware of their infections [2]. The HIV testing rate is defined as the probability that an AIDS-free HIV infected person seeks a test for HIV during a particular time interval, given no previous positive test has been obtained prior to the start of the time. In this study, we assume that the HIV test has perfect sensitivity and specificity, and an HIV positive person who gets tested is thereafter diagnosed. The accurate and timely estimation of the HIV testing rate is crucial for public health in that (1) it is a crucial component in the modeling framework that estimates the HIV incidence and reducing HIV incidence is the ultimate goal of HIV prevention; (2) it can be used to assess the effectiveness of HIV testing programs and provide guidance for future public health responses; and (3) it provides additional information on the average length of time between HIV infection and the first positive HIV test, informing how soon an HIV infection gets diagnosed.

It is challenging to obtain an accurate estimate of the HIV testing rate in that modeling the probability of HIV testing often involves the complex time lag from HIV infection to HIV diagnosis because the time of HIV infection is rarely observable [3, 4]. The HIV testing rate was initially introduced by Marschner [5] and Farewell [6], independently. In both of their work, HIV testing rate was simply assumed to be a constant, implying constant testing hazard for all individuals across calendar years. Following their work, various parametric models have been proposed to characterize individual's HIV testing behavior based on the time since HIV infection to diagnosis using the frequentist approach. Among them, one work models the HIV testing process using a two-parameter exponential distribution incorporating dependence between the time since infection and the calendar time [7]. Other methods include the heterogeneous mixed exponential model, leading to a Pareto distribution with a decreasing hazard function for the duration between HIV infection and HIV diagnosis [810], and the additive hazards model, which partitions the HIV testing process to the constant routine HIV testing and the symptoms-driven testing characterized by an exponential distribution [4, 11, 12]. Another framework of modeling the HIV disease and diagnosis process is through a multi-state formulation, describing the progression through various disease stages from infection to AIDS. Often various HIV disease stages are characterized by laboratory markers such as the CD4 T-cell count [13]. In recent years, Sweeting and Birrell proposed to use a Bayesian formulation of a similar multi-state model for estimation of HIV incidence, in which the natural disease progression probabilities and HIV testing probabilities collectively define the transition probabilities from one state to another. The HIV testing probabilities are estimated from external HIV diagnoses data [14, 15]. In addition, observed longitudinal CD4 data are used to model the date of HIV infection using joint linear mixed models [16]. However, the inclusion of additional information in the model increases the complexity of the model. This potentially introduces bias when the observed laboratory data are not representative of all the new diagnoses.

Although the HIV testing rate has been explored and modeled in various frameworks, it does not serve as a primary parameter of interest. It is an important parameter to the so-called back-calculation or back-projection model, which mainly focuses on reconstructing the past pattern of HIV infections in many countries using the annual numbers of diagnosed AIDS and AIDS-free HIV cases [9, 12, 1719]. In those models, HIV testing rates are assumed to have certain parametric forms or to be a constant for a few years. These models are not flexible and do not adequately characterize changes over time. To the best of our knowledge, there is no existing statistical framework to systematically model and estimate the annual HIV testing rates over time. To fill this gap, this paper focuses on the Bayesian modeling of annual HIV testing rates over the calendar years using the annual observed numbers of persons diagnosed with HIV, with or without AIDS at HIV diagnosis. Numerous countries have established national HIV registries or surveillance systems to collect the numbers of AIDS and AIDS-free HIV diagnoses. For example, the dataset used for this work comes from the national HIV surveillance system in the USA that started collecting data on the number of persons diagnosed with AIDS since 1981 and the number of HIV diagnoses since 1994. Over time, additional states and local jurisdictions started reporting HIV diagnoses to national HIV surveillance system [19].

In this work, we assume that the HIV testing rate is only dependent on the calendar year, and these annual probabilities can be considered as the discrete-time analogue of the HIV testing intensity. To introduce the smoothness on the HIV testing rates and the HIV incidence rates over years, we resort to a Bayesian shrinkage approach by developing a new class of priors. In the past decade, the Bayesian shrinkage methods using various priors, such as Laplace priors, have been successful in many applications [2024]. These methods are mainly developed for variable selection under a regression framework. In particular, the fused lasso priors [24, 25], that is, extended Laplace priors, are used to impose smoothness between the model parameters. In a similar fashion, for our problem, we develop a new class of priors that smooth the annual HIV testing rates and the HIV incidence rates.

Compared with existing models involving the HIV testing rate, our proposed approach has the following remarkable features: (1) our model is among the very first to propose a structured and systematic model to estimate the annual HIV testing rate in the USA; (2) unlike other back-calculation models [19], our model does not impose the constraints that HIV testing rates remain constant within certain years, nor does it make assumptions on the parametric form of the HIV testing rates; (3) our model has an ability to characterize the temporal dependence and smoothness between the annual HIV testing rates, which substantially improve the model fitting and parameter inference accuracy; and (4) our model is widely applicable in that it only needs the annual numbers of AIDS-free HIV and AIDS diagnoses as data in contrast to other approaches, such as the multi-state formulation that requires good representative laboratory data such as the CD4 data.

The paper is organized as follows: in Section 2, we present the detailed formation of our Bayesian hierarchical model for estimating annual HIV testing rates, where two new Laplace-type prior models are introduced. The corresponding properties are discussed in Section 2.3 and the posterior computation strategy are developed in Section 2.4. We demonstrate the superiority of our proposed method via simulation studies in Section 3 and illustrate our method via analysis of the US national HIV surveillance data in Section 4. We conclude our paper with a discussion of future work in Section 5.

2. The model

2.1. Data and notation

First, we outline the problem and the observed data. We use January 1, 1977 as the time origin in this analysis because the earliest diagnosis dates of AIDS cases in the USA were in 1977. The time unit used in this analysis is calendar year.

We consider HIV infection from year 1 to year T in this paper. When an individual becomes infected with HIV in year i, he or she could be (1) diagnosed with AIDS-free HIV in year t, where itT; or (2) diagnosed with AIDS in year t, where itT; or (3) remain undiagnosed as of the most recent year T. The numbers of AIDS-free HIV and AIDS diagnoses in a calendar year include persons infected with HIV any time up to and including that year. Note that there does not exist a diagnosis test to determine when the individual became infected.

Denote by Ait the number of persons infected in year i and diagnosed with AIDS in year t, by Hit the number of persons infected in year i and diagnosed with AIDS-free HIV in year t and by UiT as the number of persons infected in year i but remain undiagnosed at the end of year T. Let Ni be the total number of new HIV infections in year i. We have

Ni=t=iT(Ait+Hit)+UiT. (1)

We observe the total number of cases diagnosed with AIDS and the total number of cases diagnosed with AIDS-free HIV in year t, denoted by At and Ht, respectively. We have

At=i=1tAitandHt=i=1tHit. (2)

Table I illustrates the relationship between Ni, Ait, Hit, UiT, At, and Ht, where columns characterize the number of persons diagnosed with AIDS and AIDS-free HIV, and those undiagnosed, and rows represent the number of infections at different years.

Table I.

Illustration of the relationship between observed data and latent quantities in the model.

Year of infection (i) Year of AIDS or HIV diagnosis (t) Undiagnosed by T

1 2 T Incidence
1 A 11 H 11 A 12 H 12 A 1T H 1T U 1T N 1
2 A 22 H 22 A 2T H 2T U 2T N 2
T ATT HTT UTT NT
Observed A 1 H 1 A 2 H 2 AT HT UT

The column totals represent the number of persons diagnosed with AIDS and AIDS-free HIV in each year and those undiagnosed by the end of the most recent year. Row totals represent the number of new infections in each year. In each cell, Ait represents the number of persons infected in year i and diagnosed with AIDS in year t and Hit represents the number of persons infected in year i and diagnosed with AIDS-free HIV in year t. Only column totals {At}t=1T and {Ht}t=1T are observed, and all other quantities are latent.

2.2. A hierarchical model

The primary interest of this study is to estimate the annual HIV testing rates and re-construct the trend in the HIV testing process over the years. We propose a Bayesian hierarchical model with two levels of hierarchy. To begin with the top level, we model the latent total numbers of HIV infections Ni, for i = 1, …, T, which are assumed to be mutually independent and follow Poisson distributions with intensity λi, that is,

[Niλi]Poisson(λi), (3)

where Poisson(μ) denotes a Poisson distribution with mean μ.

At level 2, given Ni, we model the annual numbers of AIDS, AIDS-free HIV cases, and all the undiagnosed cases stratified by the HIV infections at different years. Write NiAHU=(Aii,Ai,i+1,,AiT,Hii,Hi,i+1,,HiT,UiT) for i = 1, …, T. It represents a collection of the numbers of persons diagnosed with AIDS, AIDS-free HIV, and undiagnosed persons, infected with HIV in year i and diagnosed in different years. Let multinomial(p, n) represent a multinomial distribution with event probability p and number of trials n, for i = 1, …, T, then we assume

[NiAHUqiAHU,Ni]multinomial(qiAHU,Ni), (4)

where qiAHU(qiiA,qi,i+1A,,qiTA,qiiH,qi,i+1H,,qiTH,qiTU). Given that a person is infected with HIV in year i, qitA and qitH represent the probability of being diagnosed with AIDS and AIDS-free HIV in year t, respectively, for 1 ⩽ itT, and qiTU is the probability of remaining undiagnosed at the end of year T. Note that for all i, t=iT(qitA+qitH)+qiTU=1 and qitA, qitH, qiTU0.

To characterize qiAHU, we introduce two types of conditional probabilities: the annual AIDS diagnosis rate denoted by ptiA and the annual HIV testing rate denoted by ptH, which is the primary interest of this study. The annual AIDS diagnosis rate ptiA is the probability that a person infected with HIV in year i gets diagnosed with AIDS in year t, ti given no previous positive test has been obtained prior to the start of year t. We assume that all persons newly diagnosed with AIDS did not receive HIV treatment before diagnosis. The treatment-free AIDS diagnosis rate ptiA can be generated from the known AIDS incubation period, which has been studied and modeled by a gamma distribution with the shape parameter of 2 and the scale parameter of 4 [26, 27]. The AIDS incubation period is only determined by the time interval from HIV infection to AIDS diagnosis, that is, the value of (ti). This means that without treatment, the AIDS diagnosis rate ptiA only depends on how long a person has been infected with HIV.

The annual HIV testing rate ptH is the probability that an AIDS-free HIV positive person seeks an HIV test in year t given no previous positive test has been obtained prior to the start of year t. We assume that persons diagnosed with HIV in the same year have the same HIV testing rate, regardless of when they became infected. That is, ptH is only dependent on the calendar year t and is independent of infection time i. In the first year of HIV infection, assuming HIV infection happens uniformly in the calendar year, HIV testing can only happen after HIV infection and before the calendar year-end. Therefore, the probability of being diagnosed with AIDS-free HIV in the year of HIV infection is proportional to the time interval from HIV infection to the calendar year-end. On average, the HIV testing rate in the year of HIV infection is half of the annual HIV testing rate. Thus, for 1 ⩽ itT, we can represent qitH using ptiA and ptH, which is given by

qitH={12piH×(1p0A)t=ipiH(112piH)k=i+1t1(1pkH)×k=it(1pkiA)ti+1} (5)

where we define Πk=i+1i(1pkH)=1. The term 12piH represents the conditional probability that a person gets HIV infected and tested in year i (the year of infection) given no AIDS diagnosis in the same year. The term (1p0A) represents the probability that a person is not diagnosed with AIDS in the same year of HIV infection. The term ptH(112piH)Πk=i+1t1(1pkH) represents the conditional probability that a person gets infected with HIV in year i but is not tested until year t given no AIDS diagnosis between year i and year t. The term Πk=it(1pkiA) represents the probability that a person is not diagnosed with AIDS from year i to year t. Similarly, we can represent qitA as

qitA={p0At=iptiAk=it1(1pkiA)×(112piH)k=i+1t1(1pkH)ti+1.} (6)

This further implies that qiTU can be written as a function of pkiA and piH using qiTU=1Σk=iT(qikA+qikH).

All the quantities in models (3)–(6) are latent. The only observed data are A={At}t=1T and H={Ht}t=1T in (2). The estimated values of parameter pA={ptiA}ti can be obtained from published literatures [26, 27]. Our primary interest is in making inference on the HIV testing rates pH={ptH}t=1T.

Another quantity of great interest in the public health community related to the HIV testing rates pH is the expected time-since-infection. Let ξt denote the time from infection to AIDS-free HIV or AIDS diagnosis for individuals diagnosed during year t. According to the definition of qitA and qitH, we can represent the probability mass function for ξt as

Pr(ξt=tiqitA,qitH)(qitA+qitH)λi,fori=1,,t.

This implies that the expected time-since-infection, denoted ηt, can be represented in terms of {qiAHU}i=1t, which are functions of pH and λ, that is,

ηt=E(ξtpH,λ)=i=1t(ti)(qitA+qitH)λii=1t(qitA+qitH)λi,fort=1,,T.

This informs on average how long individuals diagnosed with AIDS-free HIV or AIDS in a certain year have been infected. Because ηt is a deterministic function of pH and λ, the posterior inference on ηt can be obtained straightforwardly through the posterior inference on pH, for which we introduce a class of new priors in Section 2.3.

2.3. Prior specifications

In this section, we discuss the prior specifications for our proposed hierarchical model. We start from the most important parameter in our analysis, that is, the HIV testing rates pH. It is believed that the HIV testing process usually does not change dramatically between two succeeding years. Thus, it is reasonable to assume that the HIV testing rate in the current year is associated with the rate in the previous year. To characterize such an association, we propose a new family of probability distributions defined on [0, 1], based on which we construct a prior for pH taking into account the temporal dependence between HIV testing rates over the years. Specifically, we introduce the following definition of the new distribution.

Definition 2.1 (Laplace-beta distribution)

Let μ ∈ [0, 1], a, b, cR+. For x ∈ [0, 1], let

π(x;μ,a,b,c)=1L(μ,a,b,c)xa1(1x)b1exp(cxμ), (7)

where L(μ,a,b,c)=01xa1(1x)b1exp(cxμ)dx. We refer to π(x; μ, a, b, c) as the probability density function of a Laplace-beta distribution with the location parameter μ, the rate parameter c and two shape parameters a and b. A random variable X following this distribution is denoted as X ~ Laplace-beta(μ, a, b, c).

Remark

1) when c = 0, the Laplace-beta distribution reduces to a beta distribution with shape parameters a and b. 2) When a = b = 1, the Laplace-beta distribution becomes a truncated Laplace distribution with location parameter μ ∈ [0, 1] and rate parameter c. Also, the properties of the Laplace-beta distribution are summarized in the following proposition.

Proposition 2.1

Let X ~ Laplace-beta(μ, a, b, c). Then

  • (1)

    1 − X ~ Laplace-beta(1 − μ, b, a, c).

  • (2)
    The nth moment of X is given by
    E[Xn]=L(μ,a+n,b,c)L(μ,a,b,c),forn=1,2,, (8)
    where L(μ, a, b, c) is defined in Definition 2.1.
  • (3)
    Given a, bR+ and μ ∈ [0, 1], we have
    limc+E[X]=μandlimc0E[X]=aa+b. (9)

See the proof in the Appendix. Proposition 2.1(3) implies that the parameters c and μ control how the mean of Laplace-beta(μ, a, b, c) deviates from the mean of beta(a, b). It approximately equals to μ when c is sufficiently large, and it gets close to the mean of beta(a, b) for a small c.

Now, we assign the priors for the annual HIV testing rates using Laplace-beta distributions. Specifically, we have

p1H~beta(aH,bH)and[pt+1HptH]~Laplace-beta(ptH,aH,bH,cH), (10)

for t = 1, …, T−1. This implies as a priori, the distribution of pt+1H borrows information from ptH. According to the property of the Laplace-beta distribution, the parameter cH controls the difference between E(pt+1HptH) and ptH, which reflects the average change of the HIV testing rate from year t to t + 1. The larger cH is, the closer E(pt+1HptH) gets to ptH. In other words, the choice of cH controls the overall smoothness of the HIV testing rates over the years. A range of the cH can be specified according to the coefficient of variation of pt+1H given ptH. Several values of cH can be picked within this range and an optimal value can be chosen by maximizing the Bayes factor [28, 29]. We will discuss more details in Section 2.4. In light of the published literature, it is believed that the annual HIV testing rate in the past three decades should range from 0 to 0.4 with a mean of 0.2 and lower testing rates in the earlier years [30]. For hyperparameters aH and bH, we choose the values such that E(p1H)=0.05.

Similar to the HIV testing rates, the annual HIV incidence rates that are reflected by parameters λ={λi}i=1T in (3) is also dependent on time measured in years. To characterize such temporal dependence, we introduce another new family of probability distributions to specify the priors for λ.

Definition 2.2 (Laplace-gamma distribution)

Let a, b, c, μR+. For xR+, let

π(x;μ,a,b,c)=1K(μ,a,b,c)xa1exp(bxcxμ),

where K(μ,a,b,c)=0xa1exp(bxcxμ)dx. We refer to π(x; μ, a, b, c) as the probability density function of a Laplace-gamma distribution with the location parameter μ, the shape parameter a and two rate parameters b and c. A random variable X following this distribution is denoted as X ~ Laplace-gamma(μ, a, b, c).

Remark

1) when c = 0, the Laplace-gamma distribution reduces to a gamma distribution with shape a and rate b. 2) when a = 1 and b = 0, the Laplace-gamma distribution becomes a truncated Laplace distribution with location μ and rate c. The properties of the Laplace-gamma distribution are summarized in the following proposition.

Proposition 2.2

Let X ~ Laplace-gamma(μ, a, b, c). Then

  • (1)

    For τR+, X/τ ~ Laplace-gamma(μ/τ, a, τb, τc).

  • (2)
    The nth moment of X is given by
    E[Xn]=K(μ,a+n,b,c)K(μ,a,b,c),forn=1,2,, (11)
    where K(μ, a, b, c) is defined in Definition 2.2.
  • (3)
    Given a, b, μR+, we have
    limc+E[X]=μandlimc0E[X]=ab. (12)

The proof is straightforward and similar to proposition 2.1. Proposition (2.2)(3) implies that the parameters c and μ reflect how different the mean of Laplace-gamma(μ, a, b, c) is from the mean of gamma(a, b). It gets close to μ when c is sufficiently large and it approaches to the mean of gamma(a, b) when c is very small. Based on this property, we assign the following priors for λ:

λ1~gamma(aλ,bλ)and[λt+1λt]~Laplace-gamma(λt,aλ,bλ,cλ), (13)

for t = 1, …, T − 1. As a priori, λt+1 is assumed to follow a distribution characterized by λt, where the difference between E[λt+1 | λt] and λt is controlled by cλ > 0. This prior specification implies that the number of HIV infections in the current year can be dependent on that of the previous year to a certain extent. This further assists better posterior inference on the HIV testing rate. We demonstrate this in simulation studies. For hyperparameters in (13), we choose aλ and bλ such that E(λ1) is close to the mean of the annual numbers of total diagnoses. Similar to cH, the range of cλ is determined by the coefficient of variation of λt+1 given λt and the choice of the cλ can be determined via Bayes factors. More details will be discussed in Section 2.4.

The proposed Laplace-beta and Laplace-gamma distributions provide an alternative way to other methods that incorporate serial dependence. The notable features of these distributions include: (1) model parameters are easily interpreted and can be straightforwardly specified and estimated; (2) the prior specifications lead to computationally feasible posterior inference and we developed an MCMC algorithm based on the adaptive rejection sampling along with a random walk updating scheme; (3) the special case of the proposed model can be considered as the Bayesian counterpart of the penalized likelihood approach with the fused lasso penalty, which has a good theoretical foundation and been widely used in many applications; and (4) compared with the fussed lasso approach, the Bayesian approach automatically achieves the post-selection inference that takes into account all source of uncertainty and produces more reliable results.

2.4. Posterior inference

To simplify the posterior computation, we consider an equivalent model representation by integrating out the latent quantities NiAHU and Ni in models (3) and (4).

Because Ni follows a Poisson distribution with mean λi, it is straightforward to show that Ait and Hit follow Poisson distributions with means qitAλi and qitHλi, respectively, that is,

[Aitλi,qitA]~Poisson(qitAλi)and[Hitλi,qitH]~Poisson(qitHλi). (14)

See the proof in the Appendix. Note that Ait and Hit are mutually independent given λ, qitA and qitH. Given pH and λ, the observed total numbers of cases diagnosed with AIDS (At=i=1tAit) and HIV (Ht=i=1tHit) follow Poisson distributions with means i=1tqitAλi and i=1tqitHλi, respectively, that is,

[AtpH,λ]~Poisson(i=1tqitAλi)and[HtpH,λ]~Poisson(i=1tqitHλi), (15)

where both qitA and qitH are functions of pH according to models (6) and (5). The joint posterior distribution of pH and λ given data A and H and hyperparameters (cH, cλ) is given by

π(pH,λA,H,cH,cλ)t=1Tπ(At,HtpH,λ)×π(p1H)t=1T1π(pt+1HptH,cH)×π(λ1)t=1T1π(λt+1λt,cλ).

The posterior distributions of parameters given the data are complicated and have no closed form solutions. Thus, to sample from the posterior distributions for given cλ and cH, we resort to the adaptive rejection metropolis sampling within Gibbs sampling [31]. Details of the full conditional distributions of ptH and λt are provided in the Appendix. To achieve a better mixing of the simulated Markov chain, we suggest a random walk step to jointly update all the elements in pH and λ given all other parameters [32]. The proposal distributions are independent normal distributions with mean zero, and the proposal variances can be determined by adjusting the acceptance rate to be around 30%. This updating scheme is helpful to produce reliable credible intervals of parameters, especially when there exists high posterior correlation among parameters.

For the choice of hyperparameters (cH, cλ), we first determine the ranges of cH and cλ based on the ranges of the coefficient of variation (CV) for pt+1H given ptH and λt+1 given λt, respectively. The ranges of the CV were chosen based on the desired variances, that is, the standard deviation not to exceed a certain percentage of the mean. Thus, the desired variance drives the choice of the range for hyperparameters. Given all other parameters, we have

CV(pt+1HptH)L(E[p1H],aH+2,bH,cH)L(E[p1H],aH,bH,cH)L(E[p1H],aH+1,bH,cH)2L(E[p1H],aH+1,bH,cH),

where E[p1H]=aHaH+bH, and

CV(λt+1λt)K(E[λ1],aλ+2,bλ,cλ)K(E[λ1],aλ,bλ,cλ)K(E[λ1],aλ+1,bλ,cλ)2K(E[λ1],aλ+1,bλ,cλ),

where E[λ1]=aλbλ. Based on our experience, it is reasonable to assume that CV ranges from 0.1 to 0.3 for both pt+1H given ptH and λt+1 given λt. We, thus, determine the ranges and specify a set of values for cH and cλ, respectively. Then, on a set of pre-specified values {(cH(k),cλ(k))}k=1K, we maximize the approximate Bayes factors, where the reference model is the case when (cH, cλ) = (0, 0), that is, we choose

(cH^,cλ^)=(cH(k^),cλ(k^))withk^=argmaxk[s=1Sπ1(A,HpH(k,s),λ(k,s))]1,

where {pH(k,s),λ(k,s)}s=1S are the simulated samples from the posterior distribution π(pH, λ | A, H, cH(k), cλ(k)).

3. Simulation study

To demonstrate the performance of the proposed model, we conducted simulation studies to estimate the HIV testing rates and the expected time-since-infection. We specify the AIDS diagnosis rate (pA(x)), shown in Table II, as the conditional probability P(x < A < x + 1|A > x) based on the distribution of A, AIDS incubation period, given by gamma(2,4). To simulate the observed numbers of AIDS-free HIV and AIDS diagnoses over the years (A and H), we specify the true values for the mean numbers of new HIV infections over the years (λ) and the HIV testing rates (pH) (Table A.1 in Appendix). In particular, we have two scenarios for estimating testing rates. In Scenario 1, we consider a 34-year period with a gradual increasing trend in the HIV testing rate. In Scenario 2, we consider a 20-year period with an increasing trend followed by a decreasing trend in the HIV testing rate.

Table II.

Values of AIDS diagnosis rate generated from the AIDS incubation time period given by a gamma distribution with shape parameter of 2 and scale parameter of 4.

t − i
ptiA
t − i
ptiA
t − i
ptiA
t − i
ptiA
t − i
ptiA
0 0.00934 7 0.14701 14 0.17670 21 0.18941 28 0.19648
1 0.04761 8 0.15346 15 0.17910 22 0.19066 29 0.19724
2 0.07934 9 0.15889 16 0.18126 23 0.19181 30 0.19795
3 0.10124 10 0.16351 17 0.18321 24 0.19288 31 0.19863
4 0.11727 11 0.16749 18 0.18498 25 0.19387 32 0.19926
5 0.12952 12 0.17095 19 0.18659 26 0.19480 33 0.19986
6 0.13919 13 0.17400 20 0.18806 27 0.19567 34 0.20043

t is year of AIDS diagnosis and i is year of infection.

Given a set of values of λ, pH and pA, the HIV infections and diagnoses data are simulated from a process that mimics data generation and collection in real-life; a person becomes infected with HIV in year i, and he or she gets diagnosed in a later year ti or remains undiagnosed as of the most recent year T. At the time of diagnosis, he or she can be diagnosed with AIDS or AIDS-free HIV. The diagnosis date and disease status are determined and reported to a national surveillance registry. Annual numbers of AIDS-free HIV and AIDS diagnoses are thus summarized. This process is simulated as follows:

  • Step 1: For each year i, i = 1, …, T, the number of new HIV infections Ni is generated through a Poisson distribution based on a mean λi.

  • Step 2: For each case of infection in year i,
    • (1)
      Simulate the time of HIV infection in the year from uniform(0,1), assuming an HIV infection happens uniformly throughout the year.
    • (2)
      Simulate the time interval from HIV infection to AIDS diagnosis (i.e., the AIDS incubation period) using a gamma(2, 4) distribution.
    • (3)
      Determine the year of diagnosis and categorize the case as either `AIDS-free HIV' or `AIDS' at the time of diagnosis.
      • (a)
        If the AIDS incubation period is smaller than one (i.e., AIDS diagnosis happens in the same year of HIV infection), the case is categorized as `AIDS' and the year of diagnosis is the year of HIV infection.
      • (b)
        If the AIDS incubation period is greater than one (i.e., AIDS is diagnosed in years after the year of HIV infection), determine whether the case had an HIV test in a year before the year of AIDS diagnosis based on the HIV testing rates in each year before AIDS diagnosis.
        • If a case has an HIV test before AIDS diagnosis, it is categorized as `AIDS-free HIV' and the year of diagnosis is the year of HIV test;
        • Else if the year of AIDS diagnosis is earlier than the most recent year, the case is an `AIDS' case, and the year of diagnosis is the year of AIDS diagnosis.
        • Otherwise, the case remains undiagnosed as of the most recent year.
  • Step 3: After looping through each infection and each year, summarize the diagnosed AIDS-free HIV cases (H) and AIDS cases (A) over years.

Given the simulated A and H, we set the initial values for λ as a half of the annual observed numbers of AIDS-free HIV and AIDS diagnoses and set pH as random values between 0 and 1, respectively. We choose the hyperparameters (cH, cλ) by maximizing the Bayes factors as discussed in Section 2.4 on a set of pre-specified values. The selected values for cH are 30, 45, and 60, which imply the coefficients of variation of pt+1H given ptH are 0.30, 0.17, and 0.12 respectively. The values for cλ are 0.001, 0.002, and 0.004, leading to the coefficients of variation of λt+1 given λt of 0.31, 0.23, and 0.14, respectively. We run the proposed posterior simulation algorithm 2,000 iterations with 200 burn-in for each combination of (cH, cλ) in both scenarios. We check the convergence of the simulated Markov chains using the Gelman and Rubin diagnostic [33] by running five additional Markov chains with different initial values. The potential scale reduction factors of the log-likelihood for Scenarios 1 and 2 are respectively 0.99 and 1.06, which are both close to 1, indicating the convergence of the posterior simulations.

The selected value of (cH, cλ) is (60, 0.002) for both Scenario 1 and Scenario 2. The posterior mean and 95% credible interval of ptH and ηt are shown in Figure 1. For Scenarios 1 and 2, the estimated posterior means of ptH and ηt, for t = 1, …, T, are quite close to the true values of the HIV testing rates, and the expected time-since-infection and the associated 95% credible intervals cover the true values for all years. The results show that our model can provide accurate estimates under different trends and different periods of time. In Table III, for both scenarios, we compare the model fitting results for different choices of hyperparameters (cH, cλ). For both ptH and ηt, Table III summarizes the average mean square error (AMSE) over years, the average length of 95% credible intervals (ACI) over years, and the approximate Bayes factors (BF). A special case (shown in Figure 1) is when (cH, cλ) = (0, 0), corresponding to a regular choice of independent beta priors for ptH and independent gamma priors for λt, compared with which our method has a better model fitting (larger BF) and provides more accurate estimates and inference on the HIV testing rate and the time-since-infection (smaller AMSE and ACI).

Figure 1.

Figure 1

Estimated posterior mean and 95% credible intervals for HIV testing rates and expected time-since-infection with different choices of (cH, cλ) in simulation Scenarios 1 and 2.

Table III.

Simulation model fitting results for different choices of hyperparameters (cH, cλ).

Scenario 1
Scenario 2
pH
η
pH
η
(cH, cλ) BF AMSE ACI AMSE ACI BF AMSE ACI AMSE ACI
(0,0) 1.00 3.1e–05 0.032 0.008 0.30 1.00 2.8e–05 0.036 0.009 0.32
(30,0.001) 61.49 2.7e–05 0.027 0.006 0.29 3.51 2.8e–05 0.030 0.006 0.29
(30,0.002) 36.33 2.8e–05 0.026 0.006 0.27 2.58 2.8e–05 0.031 0.008 0.29
(30,0.004) 36.05 3.0e–05 0.026 0.007 0.25 2.53 2.8e–05 0.031 0.007 0.29
(45,0.001) 62.15 2.8e–05 0.025 0.006 0.29 3.74 3.1e–05 0.028 0.008 0.30
(45,0.002) 122.01 2.6e–05 0.025 0.005 0.26 10.05 2.5e–05 0.030 0.007 0.28
(45,0.004) 99.69 2.7e–05 0.024 0.007 0.24 5.49 2.5e–05 0.031 0.007 0.28
(60,0.001) 116.94 2.2e–05 0.025 0.006 0.27 13.91 4.4e–05 0.028 0.008 0.29
(60,0.002) 648.10 2.3e–05 0.024 0.005 0.26 18.14 2.3e–05 0.027 0.008 0.27
(60,0.004) 169.48 2.9e–05 0.024 0.006 0.24 2.45 3.1e–05 0.029 0.007 0.28

BF is the approximate Bayes factor, AMSE is the average mean square error, and ACI is the average length of 95% credible intervals.

4. Application

In this section, we apply the proposed Bayesian hierarchical model to the data from national HIV surveillance in USA.

4.1. Analysis of the US HIV surveillance data

Since 1982, all 50 states and the District of Columbia have reported AIDS cases to the Centers for Disease Control and Prevention (CDC) using a standardized case report form. In 1994, the CDC implemented data management for national reporting of HIV integrated with AIDS case reporting, at which time 25 states with confidential name-based HIV surveillance started submitting case reports to the CDC. Over time, additional states implemented name-based HIV surveillance, and all states had implemented such surveillance in 2008.

In this study, we use HIV and AIDS data reported to the CDC through June 2012. The data are adjusted for incomplete reporting, reporting delay, detection and elimination of duplicate reports, and misclassification of the diagnosis dates [3436]. The resulting estimated annual numbers of AIDS-free HIV and AIDS diagnosed cases each year from 1977 (the beginning of the HIV epidemic) to 2010 are presented in Table IV.

Table IV.

The estimated annual numbers of HIV (Ht) and AIDS (At) diagnoses from 1977 to 2010 from the National HIV Surveillance System in the USA.

Year At Ht Year At Ht Year At Ht
1977 2 3 1989 23225 53650 2001 20407 40962
1978 8 10 1990 25785 61233 2002 20024 39652
1979 13 17 1991 26013 58726 2003 18094 36028
1980 75 162 1992 28297 62485 2004 17585 36036
1981 307 144 1993 32276 51410 2005 16600 35568
1982 1144 725 1994 29813 46438 2006 15418 35260
1983 3039 1387 1995 30122 43850 2007 15360 36986
1984 6122 3211 1996 27312 40258 2008 14940 37295
1985 11515 30242 1997 23890 40692 2009 13946 33250
1986 14994 40777 1998 21582 39066 2010 13686 32356
1987 17593 46323 1999 19846 38098
1988 21698 50522 2000 22126 44360

The initial values for HIV testing rates are drawn from uniform distribution on [0, 1]. The annual numbers of new HIV infections are initially assigned a half of the observed numbers of AIDS-free HIV and AIDS diagnoses. We choose the hyperparameters (cH, cλ) as (45, 0.0005), which has the maximal value (9.1) of the Bayes factor among a set of pre-specified values. We allow a wide range for the coefficients of variation and hyperparameters to ensure a less chance to miss the best hyperparameters. The selected values for cH are 15, 30, 45, 60, corresponding to the coefficients of variation of 0.69, 0.30, 0.17, 0.12 and the values for cλ are 0.0001, 0.0002, 0.0003, 0.0005, and 0.0008, corresponding to the coefficients of variation of 0.58, 0.18, 0.12, 0.07, and 0.04, respectively. We run the posterior simulation algorithm 2,000 iterations with 200 burn-in. The potential scale reduction factors of the loglikelihood is 1.04 from running five additional chains, indicating the convergence of the posterior simulations.

Figure 2 presents the posterior means and 95% credible intervals for the HIV testing rate and the expected time-since-infection (in years) from 1985 when the first HIV test became available in the USA, to 2010. In the first few years after HIV test became available, HIV testing was widely adopted and it continued to increase until 1990. The testing rate went down in early 1990s. This is likely caused by the change in the CDC AIDS definition during that time resulting in a high proportion of simultaneous AIDS diagnoses, which could indicate low HIV testing rate. Another possible reason could be a sudden increase in the number of new HIV infections in early 1990s resulting in a high number of undiagnosed HIV infections and consequently low HIV testing rate. After that, the HIV testing rate gradually increased and sustained the increasing trend ever since. In the most recent years since 2007, the annual HIV testing rate has been stable around 0.22. As for the expected time interval (in years) since HIV infection to HIV diagnosis among individuals diagnosed in a specific calendar year, there was an increasing trend followed by a decreasing trend. The expected time-since-infection was short among those diagnosed in the early years of the epidemic, which is likely because the early infections were mainly concentrated among men who have sex with men and the targeted HIV testing among this population could result in shorter time interval from infection to diagnoses. As time went by, with the HIV epidemic spread to a more general population and the treatment improving, the expected time-since-infection gradually increased and reached the peak of 4.5 years among cases diagnosed in 1997. Since 1997, because of the increased HIV testing rate, the estimated expected time-since-infection decreased and was 3.3 years in 2010.

Figure 2.

Figure 2

Estimated posterior mean and 95% credible intervals for the annual HIV testing rates and expected time-since-infection from 1985 to 2010 in United States.

The estimated HIV testing rates and expected time-since-infection reflect the impact of important public health initiatives and recommendations on HIV testing. As shown in the results, the HIV testing rate increased when the first HIV test became available in 1985. In 1987, the United States Public Health Service issued guidelines making HIV counseling and testing a priority as a prevention strategy for people with high risk behaviors. [37] As a result, the HIV testing rate increased in the late 1980s. Although HIV testing went down during 1991 to 1993, the HIV testing rate started increasing since 1993 when CDC updated the recommendations regarding HIV counseling and testing of patients in acute-care hospital settings. [38] In 1995, when National HIV Testing Day was observed, the HIV testing rate sustained the upward trend until 2000. Throughout the first decade of 2000, a few important recommendations were published in 2001 [39], 2003 [40], and 2006 [1] respectively to emphasize routine HIV testing as an important HIV prevention tool. The HIV testing rate increased after each recommendation, and it maintained an increasing trend since 2001. As a result of the increased HIV testing, the expected time-since-infection decreased for cases diagnosed since 2001. These results indicate that public health recommendations on HIV testing have a consistently positive impact on people's HIV testing awareness and testing behavior.

4.2. Model assessment

We conduct the posterior predictive model assessment using two approaches. First, we calculate the χ2 discrepancy, which is a summary statistic for the sum of squares of standardized residuals of the data with respect to their expectations under the model [41]. For our model, the χ2 discrepancy is defined as:

χ2(H;pH,λ)=t=134(HtE(HtpH,λ))2Var(HtpH,λ).

We calculate the posterior predictive p-value based on χ2 as p = P2(Hrep; pH, λ) ⩾ χ2(H; pH, λ)), where Hrep represents the predictive replication and H represents the observed data. The p-value of the predictive versus realized χ2 discrepancies is 0.48. This implies that the model fits the data pretty well.

We also assess the validity of the 95% credible intervals for the annual HIV testing rates and expected time-since-infection from 1985 to 2010. We generate 100 simulated datasets given the estimated HIV testing rates and expected time-since-infection. For each simulated dataset, we refit the model and obtain the estimated HIV testing rates and expected time-since-infection. Then we count the number of times that the newly estimated HIV testing rates and expected time-since-infection in the estimated 95% credible intervals for each year. We found that for both HIV testing rates and the expected time-since-infection, the newly estimated HIV testing rates and the expected time-since-infection fall in the estimated 95% credible intervals more than 95 times for 24 of the 26 years. This demonstrates the 95% credible intervals are plausible.

5. Discussion

In this paper, we develop a Bayesian hierarchical model to estimate the intensities of HIV testing from 1985 to 2010 using annual numbers of AIDS-free HIV and AIDS diagnosed cases collected through national HIV surveillance in the USA. Our model takes the most general form and makes no parametric assumptions for HIV testing rates. We assume that the HIV testing rate is only dependent on the calendar year, and these annual probabilities can be considered the discrete-time analogue of the HIV testing intensity. We propose Laplace-beta and Laplace-gamma priors to characterize the temporal dependence for annual HIV testing rates and annual HIV incidence rates, respectively, which greatly improve the estimation accuracy and the model fitting. The simulation studies show that our model can make much more accurate inference on HIV testing rates of different trends for either long or short periods of time compared with a regular choice of priors.

The proposed Laplace-beta, completely different from the beta-Laplace distribution by Cordeiro and Lemonte [42], and Laplace-gamma distributions are motivated by the need of incorporating the prior temporal dependence among HIV testing rates and incidence rates in the model; however, they are not the only way in which temporal dependence could have been incorporated. The proposed priors are general and can have different applications. Further applications will help to determine the utility of this approach compared with other more traditional approaches. For example, in spatial statistics, these two prior distributions can impose the smoothness over the spatial parameters, which are alternatives to the use of Gaussian random fields requiring the normality assumption that are not valid in certain cases. In the spatial modeling of disease mapping, the disease infection probabilities over space can be assigned with Laplace-beta priors, and the intensity of the disease clusters can be modeled by Laplace-gamma priors. Also, in the analysis of functional neuro-imaging data, the voxel-wise probabilities of activation can be modeled with Laplace-beta priors and the Laplace-gamma model should be a good choice for the intensity of peak activation locations.

There are several possible future directions that extend our current work. One extension is to include covariates such as sex, race/ethnicity, or transmission category in the model, which can adjust for the confounding factors that might affect the HIV testing rates. Also, in contrast to the current modeling of the whole population in the USA, this extension can provide stratified estimates for sub-populations, for example, HIV testing rates for demographic groups or different regions in the country, which are of particular interest to public health officials and program evaluation. Another direction is that we can develop alternative strategies to choose the hyperparameters (cH, cλ) that are strongly related to the performance of model fitting and parameter estimation. A fully Bayesian approach can be considered by jointly updating (cH, cλ) in the posterior simulation and using the Bayesian model averaging to make the inference. The key steps are to choose appropriate priors for (cH, cλ) and to develop an efficient posterior sampling strategy that is worthy of investigation in that this approach would take into account more sources of variation in the model and potentially produce better posterior inference on the model parameters.

Appendix A

Derivations for marginal distributions for Ait and Hit

Noting that if (X1, X2, X3) ~ multinormial(p1, p2, p3, N) and N ~ Poisson(λ), then the joint distribution of (X1, X2, X3) and N is given by

P(X1,X2,X3,N)=N!X1!X2!X3!p1X1p2X2p3X3λNN!eλ.

Integrating out N from the previous joint probability density, we have

P(X1,X2)=N=X1+X2p1X1p2X2(1p1p2)NX1X2X1!X2!(NX1X2)!λNX1X2eλλX1+X2=N=X1+X2[(1p1p2)λ]NX1X2(NX1X2)!eλ(1p1p2)e(p1+p2)λp1X1p2X2X1!X2!λX1+X2=(p1λ)X1X1!ep1λ(p2λ)X2X2!ep2λ

This implies that X1 and X2 follow independent Poisson distributions. In a similar fashion of the derivation, it is straightforward to show that if

(Aii,,AiT,Hii,,Hit,UiT)~multinormial(qiiA,,qiTA,qiiH,,qiTH,qiTU,Ni)

and

Ni~Poisson(λi),

then

Ait~Poisson(qitAλi)andHit~Poisson(qitHλi),t=i,,T,

where Ait and Hit are mutually independent.

Full conditional distributions

  • Full conditional distribution of pH
    π(pHλ,A,H)t=1Tei=1tqitAλiAt!(i=1tqitAλi)Atei=1tqitHλiHt!(i=1tqitHλi)Ht×t=1T(ptH)a21(1ptH)b21B(a2,b2)×t=2Texp{cHptHpt1H}
  • Full conditional distribution of λ
    π(λpH,A,H)t=1Tei=1tqitAλiAt!(i=1tqitAλi)Atei=1tqitHλiHt!(i=1tqitHλi)Htb0a0Γ(a0)λta01exp(b0λt)×t=2Texp{cλλtλt1}

Proofs for Proposition 2.1

Note that X ~ Laplace-beta(μ, a, b, c).

  • (1)
    The density function of Y = 1 − X is
    f(y)=π(1x;μ,a,b,c)(1x)a1xb1exp(c1xμ),
    which implies that 1 − X ~ Laplace-beta(1 − μ, b, a, c).
  • (2)
    E(Xn)=011L(μ,a,b,c)xn+a1(1x)b1exp(cxμ)dx=L(μ,a+n,b,c)L(μ,a,b,c)
  • (3)

    Given a, b > 0, when c → ∞, π(x; a, b, c) → I[x = μ], thus, E(X) → μ.

Table A.I.

Values for the parameters used in the simulation studies.

Scenario 1
Scenario 2
year λ pH A H A pH A H
1 24 0.060 0 1 24 0.060 0 1
2 86 0.060 2 4 86 0.060 2 4
3 86 0.060 6 8 86 0.060 6 8
4 244 0.060 14 17 244 0.060 14 17
5 244 0.096 28 46 244 0.096 28 46
6 862 0.096 49 91 862 0.096 49 91
7 862 0.096 92 156 862 0.096 92 156
8 2521 0.120 161 361 2521 0.120 161 361
9 2521 0.120 285 586 2521 0.120 285 586
10 3337 0.120 439 815 3337 0.120 439 815
11 3337 0.156 618 1356 3337 0.156 618 1356
12 4675 0.156 780 1649 4675 0.156 780 1649
13 4675 0.156 965 1970 4675 0.156 965 1970
14 5305 0.156 1151 2262 5305 0.156 1151 2262
15 5305 0.156 1335 2529 5305 0.156 1335 2529
16 5305 0.156 1500 2728 5305 0.144 1500 2518
17 3429 0.156 1619 2730 3429 0.144 1640 2547
18 3429 0.156 1647 2582 3429 0.144 1687 2431
19 3429 0.180 1621 2840 3429 0.120 1678 1944
20 2000 0.180 1522 2542 2000 0.120 1674 1835
21 2000 0.180 1379 2196
22 1500 0.180 1227 1895
23 1500 0.192 1074 1739
24 2346 0.192 934 1596
25 2346 0.192 852 1576
26 2346 0.192 808 1569
27 2346 0.192 784 1568
28 2642 0.192 775 1597
29 2642 0.204 783 1750
30 2642 0.204 786 1772
31 2400 0.204 789 1763
32 2400 0.204 785 1733
33 2400 0.204 777 1711
34 2400 0.204 769 1694

λ is the mean annual number of new HIV infections, pH is the annual HIV testing rate, A is the observed annual number of AIDS diagnoses and H is the observed annual number of HIV diagnoses. Scenario 1 is for a 34-year period with a gradual increasing trend in pH and Scenario 2 is for a 20-year period with an increasing trend followed by a decreasing trend in pH.

References

  • 1.Bernard BM, Handsfield HH, Lampe AM, Janssen SR, Taylor WA, Lyss BS, Clark EJ. Revised recommendations for HIV testing of adults, adolescents, and pregnant women in health-care settings. Mortality and Morbidity Weekly Report. 2006;55(RR-14):1–17. [PubMed] [Google Scholar]
  • 2.Centers for Disease Control and Prevention Monitoring selected national HIV prevention and care objectives by using HIV surveillance data – United States and 6 US dependent areas 2010. HIV Surveillance Supplemental Report. 2012;17(3):22. [Google Scholar]
  • 3.Blaxhult A, Svensson A. Assessing the extent of the HIV epidemic in Sweden, using information on the extent to which people who develop AIDS are already known to be HIV infected. International Journal of Epidemiology. 1992;21(4):784–791. doi: 10.1093/ije/21.4.784. [DOI] [PubMed] [Google Scholar]
  • 4.Becker GN, Lewis CJ, Li Z, McDonald A. Age-specific back-projection of HIV diagnosis data. Statistics in Medicine. 2003;22(13):2177–2190. doi: 10.1002/sim.1406. [DOI] [PubMed] [Google Scholar]
  • 5.Marschner CI. Using time of first positive HIV test and other auxiliary data in back-projection of AIDS incidence. Statistics in Medicine. 1994;13(19–20):1959–1974. doi: 10.1002/sim.4780131908. [DOI] [PubMed] [Google Scholar]
  • 6.Farewell VT, Aalen OO, De Angelis D, Mrc ND. Estimation of the rate of diagnosis of HIV infection in HIV infected individuals. Biometrika. 1994;81(2):287–294. [Google Scholar]
  • 7.Bellocco R, Marschner IC. Joint analysis of HIV and AIDS surveillance data in back-calculation. Statistics in Medicine. 2000;19(3):297–311. doi: 10.1002/(sici)1097-0258(20000215)19:3<297::aid-sim340>3.0.co;2-6. [DOI] [PubMed] [Google Scholar]
  • 8.Wand H, Wilson D, Yan P, Gonnermann A, McDonald A, Kaldor J, Law M. Characterizing trends in HIV infection among men who have sex with men in Australia by birth cohorts: results from a modified back-projection method. Journal of the International AIDS Society. 2009;12(1):1–8. doi: 10.1186/1758-2652-12-19. Article 19. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Wand H, Yan P, Wilson D, McDonald A, Middleton M, Kaldor J, Law M. Increasing HIV transmission through male homosexual and heterosexual contact in Australia: results from an extended back-projection approach. HIV Medicine. 2010;11(6):395–403. doi: 10.1111/j.1468-1293.2009.00804.x. [DOI] [PubMed] [Google Scholar]
  • 10.Yan P, Zhang F, Wand H. Using HIV diagnostic data to estimate HIV incidence: method and simulation. Statistical Communications in Infectious Diseases. 2011;3(1):1–28. Article 6. [Google Scholar]
  • 11.Cui J, Becker GN. Estimating HIV incidence using dates of both HIV and AIDS diagnoses. Statistics in Medicine. 2000;19(9):1165–1177. doi: 10.1002/(sici)1097-0258(20000515)19:9<1165::aid-sim419>3.0.co;2-7. [DOI] [PubMed] [Google Scholar]
  • 12.Chau PH, Yip SFP, Cui JS. Reconstructing the incidence of human immunodeficiency virus in Hong Kong by using data from HIV positive tests and diagnoses of acquired immune deficiency syndrome. Journal of the Royal Statistical Society: Series C (Applied Statistics) 2003;52(2):237–248. [Google Scholar]
  • 13.Aalen OO, Farewell TV, De Angelis D, Day EN, Gill O. A Markov model for HIV disease progression including the effect of HIV diagnosis and treatment: application to AIDS prediction in England and Wales. Statistics in Medicine. 1997;16(19):2191–2210. doi: 10.1002/(sici)1097-0258(19971015)16:19<2191::aid-sim645>3.0.co;2-5. [DOI] [PubMed] [Google Scholar]
  • 14.Sweeting JM, Angelis DD, Aalen OO. Bayesian back-calculation using a multi-state model with application with HIV. Statistics in Medicine. 2005;24(24):3991–4007. doi: 10.1002/sim.2432. [DOI] [PubMed] [Google Scholar]
  • 15.Birrell JP, Chadborn RT, Gill NO, Delpech CV, De Angelis D. Estimating trends in incidence, time-to-diagnosis and undiagnosed prevalence using a CD4-based Bayesian back-calculation. Statistical Communications in Infectious Diseases. 2012;4(1):1–29. Article 6. [Google Scholar]
  • 16.Taffe P, May M. A joint back calculation model for the imputation of the date of HIV infection in a prevalent cohort. Statistics in Medicine. 2008;27(23):4835–4853. doi: 10.1002/sim.3294. [DOI] [PubMed] [Google Scholar]
  • 17.Punyacharoensin N, Viwatwongkasem C. Trends in three decades of HIV/AIDS epidemic in Thailand by nonparametric back calculation method. AIDS. 2009;23(9):1143–1152. doi: 10.1097/QAD.0b013e32832baa1c. [DOI] [PubMed] [Google Scholar]
  • 18.Mallitt K, Wilson PD, McDonald A, Wand H. HIV incidence trends vary between jurisdictions in Australia: an extended back-projection analysis of men who have sex with men. Sexual Health. 2012;9(2):138–143. doi: 10.1071/SH10141. [DOI] [PubMed] [Google Scholar]
  • 19.Hall HI, Song R, Rhodes P, Prejean J, An Q, Lee ML, Karon J, Brookmeyer R, Kaplan HE, Mckenna TM. Estimation of HIV incidence in the United States. The Journal of the American Medical Association. 2008;300(5):520–529. doi: 10.1001/jama.300.5.520. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Figueiredo AM, Nowak DR. An EM algorithm for wavelet-based image restoration. IEEE Transactions on Image Processing. 2003;12(8):906–916. doi: 10.1109/TIP.2003.814255. [DOI] [PubMed] [Google Scholar]
  • 21.Bae K, Mallick KB. Gene selection using a two-level hierarchical Bayesian model. Bioinformatics. 2004;20(18):3423–3430. doi: 10.1093/bioinformatics/bth419. [DOI] [PubMed] [Google Scholar]
  • 22.Genkin A, Lewis DD, Madigan D. Large-scale Bayesian logistic regression for text categorization. Technometrics. 2007;49(3):291–304. [Google Scholar]
  • 23.Park T, Casella G. The bayesian lasso. Journal of the American Statistical Association. 2008;103(482):681–686. [Google Scholar]
  • 24.Kyung M, Gill J, Ghosh M, Casella G. Penalized regression, standard errors, and Bayesian lassos. Bayesian Analysis. 2010;5(2):369–411. [Google Scholar]
  • 25.Tibshirani R, Saunders M, Rosset S, Zhu J, Knight K. Sparsity and smoothness via the fused lasso. Journal of the Royal Statistical Society: Series B (Statistical Methodology) 2005;1(67):91–108. [Google Scholar]
  • 26.Longini IM, Clark SW, Byers HR, John W, Darrow WW, Lemp FG, Hethcote WH. Statistical analysis of the stages of HIV infection using a Markov model. Statistics in Medicine. 1989;8(7):831–843. doi: 10.1002/sim.4780080708. [DOI] [PubMed] [Google Scholar]
  • 27.Longini IM, Clark SW, Gardner IL, Brundage FJ. The dynamics of CD4+ T-lymphocyte decline in HIV infected individual: a Markov modeling approach. Journal of Acquired Immune Deficiency Syndromes. 1991;4(11):1141–1147. [PubMed] [Google Scholar]
  • 28.Aitkin M. Posterior bayes factors. Journal of the Royal Statistical Society. Series B (Methodological) 1991;53(1):111–142. [Google Scholar]
  • 29.Kass ER, Raftery EA. Bayes factors. Journal of the American Statistical Association. 1995;90(430):773–795. [Google Scholar]
  • 30.Karon MJ, Song R, Brookmeyer R, Kaplan HE, Hall HI. Estimating HIV incidence in the United States from HIV/AIDS surveillance data and biomarker HIV test results. Statistics in Medicine. 2008;27(23):4617–4633. doi: 10.1002/sim.3144. [DOI] [PubMed] [Google Scholar]
  • 31.Gilks RW, Best NG, Tan KKC. Adaptive rejection Metropolis sampling within Gibbs sampling. Applied Statistics. 1995;44(4):455–472. [Google Scholar]
  • 32.Roberts GO, Rosenthal JS. Optimal scaling of discrete approximations to Langevin diffusions. Journal of the Royal Statistical Society: Series B (Statistical Methodology) 1998;60(1):255–268. [Google Scholar]
  • 33.Gelman A, Rubin BD. Inference from iterative simulation using multiple sequences. Statistical Science. 1992;7(4):453–464. [Google Scholar]
  • 34.Song R, Green T. An improved approach to accounting for reporting delay in case surveillance systems. JP Journal of Biostatistics. 2012;7(1):1–14. [Google Scholar]
  • 35.Song R, Hall HI, Frey R. Uncertainties associated with incidence estimates of HIV/AIDS diagnoses adjusted for reporting delay and risk redistribution. Statistics in Medicine. 2005;24(3):453–464. doi: 10.1002/sim.1935. [DOI] [PubMed] [Google Scholar]
  • 36.Xiao Y, Song R, Chen M, Hall HI. Direct and unbiased multiple imputation methods for missing values of categorical variables. Journal of Data Science. 2012;10:465–481. [Google Scholar]
  • 37.CDC Public Health Service guidelines for counseling and antibody testing to prevent HIV infection and AIDS. Morbidity and Mortality Weekly Report. 1987;36(31):509–515. [PubMed] [Google Scholar]
  • 38.Ward JW, Janssen SR, Jaffe HW. Recommendations for HIV testing services for inpatients and outpatients in acute-care hospital settings. Morbidity and Mortality Weekly Report. 1993;42(RR-2):1–6. [PubMed] [Google Scholar]
  • 39.Allen D, Ammann A, Bailey H, Arms A, Baker C, Berger R, Birkhead G, Boland M, Colman C, Permanente K. Revised recommendations for HIV screening of pregnant women. Morbidity and Mortality Weekly Report. 2001;50(RR-19):59–86. [PubMed] [Google Scholar]
  • 40.Janssen SR, Onorato MI, Valdiserri OR, Durham MT, Nichols PW, Seiler ME, Jaffe WH. Advancing HIV prevention: new strategies for a changing epidemic – United States, 2003. Morbidity and Mortality Weekly Report. 2003;52(15):329. [PubMed] [Google Scholar]
  • 41.Gelman A, Meng X, Stern H. Posterior predictive assessment of model fitness via realized discrepancies. Statistica Sinica. 1996;6(4):733–760. [Google Scholar]
  • 42.Cordeiro MG, Lemonte JA. The beta Laplace distribution. Statistics & Probability Letters. 2011;81(8):973–982. [Google Scholar]

RESOURCES