Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Mar 1.
Published in final edited form as: Stat Probab Lett. 2024 Nov 22;218:110295. doi: 10.1016/j.spl.2024.110295

On exact Bayesian credible sets for discrete parameters

Chaegeun Song a,*, Bing Li a
PMCID: PMC11722005  NIHMSID: NIHMS2038506  PMID: 39803594

Abstract

We introduce a generalized Bayesian credible set that can achieve any preassigned credible level, addressing a limitation of the current credible sets. This is achieved by exploiting a connection between the highest posterior density set and the Neyman-Pearson lemma.

Keywords: Bayesian classification, Highest posterior density set, Neyman-Pearson lemma, Pattern recognition

1. Introduction

The credible set is one of the most useful tools in Bayesian inference to quantify the uncertainty in Bayesian estimation. See, for example, [1], and [2]. It plays a similar role as the confidence set in the frequentist setting. The most commonly used credible set is the highest posterior density credible set, which has the shortest length for a given credible level. Approximate Bayesian credible sets can be constructed using the Bernstein-von Mises theorem: see, for example, [3] and [4].

However, the classical definition of the highest posterior density credible set [1] is useful mostly for continuous parameters—the parameter supported on an interval or a set with a nonempty and connected interior. In the discrete case where the parameter is supported on a finite or a countable set, the highest posterior density credible set may not exist for a prescribed credible level. This problem is particularly acute for Bayesian classification where the highest posterior density credible set almost never exists for a given credible level. To our knowledge, there has been no systematic method to construct the shortest credible set for such problems.

Bayesian classification has been widely used in many contemporary scientific studies such as cancer research, genomics, and healthcare applications. See, for example, [5], [6],[7],[8], [9], and [10]. Thus, the need for precise uncertainty quantification for Bayesian classification is prevalent. It is high time to facilitate Bayesian classification with exact inference tools and bring it up to par with Bayesian inference for continuous parameters, where exact credible sets have long been available and widely used. The current paper is one step towards satisfying this need.

Our extension is based on the following insight: the process of identifying the highest posterior density credible set is rather similar to that of constructing the most powerful test based on the Neyman-Pearson lemma, and the latter can accommodate any prescribed significance level—even for discrete random variables—by strategically placing probability masses at some transitioning points. This consideration leads us to define the generalized highest posterior density credible set not as a set in the parameter space, but as a mapping from the parameter space to the closed interval [0, 1]. By this adjustment, there always exists a unique shortest credible set for any prescribed credible level, regardless of the nature of the parameter. Accompanying this methodology, we also introduce (in the Supplementary Material) a visualization tool for classification called the Steering Wheel Plot.

2. Generalized credible set

Definition 2.1.

For an α(0,1), we call any measurable function ϕ:ΩΘ×ΩX[0,1] such that

E{ϕ(ΘX)X}=1-αa.s.PX

a(1-α)-level generalized credible set for θ.

This definition is motivated by the randomized test in classical hypothesis testing theory (see, for example, [11]). Here, the value of ϕ(ΘX) can be interpreted as follows: conditioning on the data X, at each value of Θ, we perform a Bernoulli trial with success probability ϕ(ΘX); if the trial turns up head, we classify Θ as belonging to the credible set; otherwise, we classify Θ as not belonging to the credible set. Alternatively, we can also interpret ϕ(ΘX) as the membership function in fuzzy logic (see, for example, [12], [13], and [14]), where a larger value ϕ(ΘX) represents a stronger tendency of belonging to the credible set.

Under this new framework, the classical credible set, which is a subset of ΩΘ, can be viewed as a mapping ϕ from ΩΘ×ΩX to the set {0, 1}; that is, ϕ is the indicator function of the credible set. Since {0,1}[0,1], the generalized credible set indeed generalizes the credible set.

3. Generalized highest posterior density credible set

As mentioned earlier, the highest posterior density credible set is the smallest credible set in the sense that it has the smallest μΘ-measure among all credible sets of the same level. More precisely, if Cκα is the highest posterior density credible set of level 1α and C is any other credible set of level 1α, then μΘ{C(κα)}μΘ(C). In this section, we extend this optimal property to generalized credible sets. We first generalize the highest posterior density credible set.

Definition 3.1.

Let γ(θx) be a value in the range [0, 1]. If there is a constant κα(x)>0 such that the generalized credible set

ϕ*(θx)=1ifπΘX(θx)>κα(x)γ(θx)ifπΘX(θx)=κα(x)0ifπΘX(θx)<κα(x) (1)

satisfies E{ϕ*(ΘX)X}=1-α, then ϕ* is called the level 1-α generalized highest posterior density credible set.

If we take γ(θx) to be constantly 1, the above function is nothing but the indicator function of the set {θ:πΘX(θx)κα}, which is the definition of the highest posterior density credible set. The difference is that for the highest posterior density credible set we can only guarantee P[Θ-1{C(κα)}X]1-α, where the equality may not be achieved. In comparison, as we will show in Theorem 3.3, the exact level 1α can always be achieved by the generalized highest posterior density credible set.

Now that the measure μΘ(C) is replaced by the integral ϕ(θx)dμΘ(θ), the size of a generalized credible set has a new meaning: we simply use ϕ(θx)dμΘ(θ) to define the size of a generalized credible set with ϕ(θx). The next two theorems show that the generalized highest posterior density credible set is the smallest generalized credible set, that it exists, and that it is unique. The proofs of the theorems echo the proof of the Neyman-Pearson lemma.

Theorem 3.2.

If ϕ*:ΩΘ×ΩX[0, 1] is the level 1α generalized highest posterior density credible set, then it is the smallest level 1α generalized credible set; that is, for any ϕ:ΩΘ×ΩX[0, 1] satisfying E{ϕ(ΘX)X}=1-α, we have

ϕ*(θx)dμΘ(θ)ϕ(θx)dμΘ(θ).

We next show the existence of the generalized highest posterior density credible set and the uniqueness of the smallest generalized credible set up to the choice of γ(θx). The specific form of γ(θx) is established in Theorem 4.2.

Theorem 3.3.

For any 0<α<1 and xΩX, there exists a (1α)-level generalized highest posterior density credible set of the form ϕ* in (1) with γ(θx)=γ(x). Furthermore, if ϕ is any smallest generalized credible set of level 1α, then it has the form (1) a.s. PX. That is,

PΘX{ϕ*(θx)ϕ(θx),πΘX(θx)κα(x)X=x}=0a.s.PX.

Let Define ρ(κx)=PΘX{πΘX(θx)>κx}. The next proposition shows that the generalized highest posterior density credible set reduces to the highest posterior density credible set when ρ(x) is continuous at κα.

Proposition 3.1.

Let κα(x)=sup{κ:P{πΘX(θx)κx}1-α}. If ρ is continuous at κα(x), then the generalized highest posterior density credible set ϕ* in (1) reduces to

ϕ*(θx)=1ifπΘX(θx)κα(x)0ifπΘX(θx)<κα(x)

so that {θ:ϕ*(θx)=1} is the (1α)-level highest posterior density credible set.

4. Fair generalized highest posterior density credible set

The generalized highest posterior density credible set ϕ* in (1) is not unique, unless PΘX{πΘX(θx)=κα(x)x}=0. When PΘX{πΘX(θx)=κα(x)x}>0, any γ(θx) satisfying

K(x)γ(θx)πΘX(θx)dμΘ(θ)=1-α-PΘX{πΘX(θx)>κα(x)x} (2)

is optimal in the sense of Theorem 3.2 where K(x)={θ:πΘX(θx)=κα(x)}. In other words, if we let

S={ϕ*:ϕ*is of the form (1) satisfying (2)}, (3)

then every member of S is the generalized highest posterior density credible set. For definiteness, we introduce the fair generalized highest posterior density credible set as follows.

Definition 4.1.

Let S be the collection of all generalized highest posterior density credible set of level-(1α) as defined in (3). We call ϕf*S the fair generalized highest posterior density credible set if

var{ϕf*(θx)}var{ϕ*(θx)}for allϕ*(θx)S.

When there is more than one generalized highest posterior density credible set, it is desirable to make ϕ*(θx) as similar as possible to give each θ a fair treatment. That is, we should minimize the variation of ϕ*(θx) with respect to πΘX(θx).

The above definition is needed only when PΘX{πΘX(θx)=κα(x)}>0. When this probability is 0, it doesn’t matter how we define γ(θx). For example, we can take γ(θx)=0 on the set {θ:πΘX(θx)=κα(x)}. The following theorem characterizes the form of the fair generalized highest posterior density credible set.

Theorem 4.2.

Suppose PΘX{πΘX(θx)=κα(x)}>0. Then the fair generalized highest posterior density credible set ϕf* is the member of S with

γ(θx)=1-α-PΘX{πΘX(θx)>κα(x)x}PΘX{πΘX(θx)=κα(x)x}.

To compute the generalized highest posterior density credible set, imagine the posterior probabilities at θ1,θ2, as spikes whose heights correspond to the posterior probabilities. We simply collect the set of highest spikes so that adding another spike (or a few more spikes, if the next few spikes are equal in height) would make the total height exceed the credible level 1α. We then put a weight on the next spike (or distribute the weight evenly among the next few spikes if they are equal in height), so that the posterior probabilities and the weighted posterior probabilities add up to 1α.

5. Data application

We now apply the fair generalized highest posterior density credible set and the steering wheel plot, which can be found in the Supplementary Material, to speaker accent recognition data. It is available at the UCI Machine Learning Repository at https://archive.ics.uci.edu/ml/datasets/Speaker+Accent+Recognition#; see also [15]. There are 12 integer features and six different accents of English: Spanish, French, German, Italian, British, and American. The data has a sample size of 329. Speakers from six different countries read single English words. Twelve features for each subject are extracted from the soundtrack of no more than one second of reading a word using Mel-Frequency Cepstral Coefficients.

For better visual effect, we first perform sufficient dimension reduction to reduce the dimension of the predictors [16]. We apply sliced inverse regression [17] and directional regression [18] to find low-dimensional predictors with four dimensions. We then conduct quadratic discriminant analysis to classify the accents in the reduced space. The 95%-level fair generalized highest posterior density credible set and steering wheel plot (as described in the Supplementary Material) enable us to infer the classification uncertainty.

Figure 1 shows the distinction between the classification uncertainties of two different quadratic discriminant analysis classifiers. For example, considering the French accent, the spokes of the classifier with sliced inverse regression are French, American, Spanish, or British, whereas the spokes of the classifier with directional regression are French or American. We can infer that directional regression helps to reduce classification uncertainty when classifying French accents.

Figure 1.

Figure 1.

Speaker accent data: the steering wheel plots (as described in the Supplementary Material) of the first two sufficient predictors. Spanish is red, French is green, German is blue, Italian is black, British is orange, and American is purple. To enhance clarity, five samples are randomly selected for each classified accent, resulting in a total of 30 samples. (a) sliced inverse regression predictors; (b) directional regression predictors.

6. Discussion

In this paper, we generalize the classical Bayesian credible set to achieve exact credible levels for discrete parameters. This fills a gap in Bayesian uncertainty quantification when the random parameter is supported on a finite or discrete set, as in classification problems such as linear or quadratic discriminant analysis.

As a referee pointed out, our construction has a connection with fuzzy Bayesian inference (see, for example, [12], [13], and [14]), and, in particular, with the idea of the membership function. In fact, our generalized credible set ϕ(θ|x) can be viewed as a membership function whose greater value (close to 1) represents closer proximity to the credible set. However, there are some important differences in the motivation, formulation, and theory of our approach from fuzzy Bayesian inference.

The original motivation of our construction of the function ϕ(θ|x) is the “randomized test” in the classical theory of hypothesis testing. See, for example, [11]. The randomized test generalizes the critical region (rejection region), which is equivalent to an indicator function, by a function that takes values in the interval [0, 1] (as opposed to the set {0, 1}). This was precisely to take care of the scenarios where the significance level α cannot be precisely attained. We transplant this idea to the Bayesian setting and apply it to discrete parameters.

However, beyond the membership function, the overlap between our problem and that of fuzzy Bayesian inference is quite small: we do not consider any fuzzy data, where x is replaced by a fuzzy number, nor any fuzzy parameter, where θ is replaced by a fuzzy number. Our setting is completely that of classical Bayesian inference, except that we introduced the analogy of a randomized test, or membership function, to achieve the exact credible level.

Although [12] and [13] also developed the highest posterior density credible set, their purposes were to extend this concept to the context of fuzzy data and fuzzy parameters. However, our purpose in extending the highest posterior density credible set is simply to accommodate discrete parameters. More importantly, the main theoretical novelty of our paper is to establish the existence, optimality, and uniqueness of the generalized highest posterior density credible set by borrowing the argument of the Neyman-Pearson lemma. This key component, however, was absent in the mentioned earlier works.

Our optimal result in Theorem 3.2 can be stated through a Bayes rule in an appropriately formulated statistical decision, which will give further insight into our method. Consider the statistical decision problem with the following components.

  1. Parameter space ΩΘ: this is a discrete or finite set of parameter values (or class labels), say, ΩΘ=θ1,θ2,.

  2. Action space ΩA: this is the collection of all functions ϕ:ΩΘ[0, 1] such that
    E[ϕ(Θ)X]=1-α.

    That is, each member of ΩA is a generalized credible set of credible level 1 − α.

  3. Class of decision rules 𝒟: this is the set of all mappings from ΩX to ΩA; that is, a member of 𝒟 is a function xϕ(x), where ϕ(x) is a member of ΩA

  4. Loss function: given a parameter value θΩΘ, and an action ϕΩA, we define the loss function as
    L(θ,ϕ)=ΩΘϕ(θ)dμΘ(θ).
  5. Posterior Expected Loss: given a decision xϕ(x) in 𝒟, the posterior expected loss is
    ΩΘϕ(θX)dμΘ(θ).
  6. Bayes Rule: Our Theorem 3.2 shows that the generalized highest posterior density credible set defined in Definition 3.1 is, in fact, the Bayes rule with respect to the loss L(θ,ϕ), among the decision rules in 𝒟.

Supplementary Material

1

Acknowledgments

The authors would like to thank two anonymous referees, an Associate Editor and the Editor for their constructive comments that improved the quality of this paper. Specifically, the decision theoretical formulation in Section 6 and the connection with fuzzy Bayesian inference were suggested by a referee.

Funding

Bing Li was supported in part by the National Science Foundation (NSF) grant DMS-2210775 and the National Institutes of Health (NIH) grant (1 R01 GM152812-01).

Footnotes

Supplement

The supplementary material includes omitted proofs of the paper and additional results. The code is available at https://github.com/chaegeunsong/ExactBayesianCredibleSet.

Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.

References

  • [1].Berger JO, Statistical Decision Theory and Bayesian Analysis, Springer Science & Business Media, 1985. 1 [Google Scholar]
  • [2].Li B, Babu GJ, A graduate course on statistical inference, Springer, 2019. 1 [Google Scholar]
  • [3].Le Cam L, On some asymptotic properties of maximum likelihood estimates and related bayes estimates, Univ. Calif. Publ. Statist. 1 (1953) 277–330. 1 [Google Scholar]
  • [4].van der Vaart AW, Asymptotic statistics, Univ. Calif. Publ. Statist. 1 (1953) 277–330. 1 [Google Scholar]
  • [5].Mallick BK, Ghosh D, Ghosh M, Bayesian classification of tumours by using gene expression data, Journal of the Royal Statistical Society: Series B (Statistical Methodology) 67 (2) (2005) 219–234. 1 [Google Scholar]
  • [6].Sambo F, Trifoglio E, Di Camillo B, Toffolo GM, Cobelli C, Bag of naïve bayes: biomarker selection and classification from genome-wide snp data, BMC bioinformatics 13 (14) (2012) 1–10. 1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [7].Bhuvaneswari R, Kalaiselvi K, Naive bayesian classification approach in healthcare applications, International Journal of Computer Science and Telecommunications 3 (1) (2012) 106–112. 1 [Google Scholar]
  • [8].Knight JM, Ivanov I, Dougherty ER, Mcmc implementation of the optimal bayesian classifier for non-gaussian models: model-based rna-seq classification, BMC bioinformatics 15 (1) (2014) 1–13. 1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [9].Tavtigian SV, Greenblatt MS, Harrison SM, Nussbaum RL, Prabhu SA, Boucher KM, Biesecker LG, Group CSVIW, et al. , Modeling the acmg/amp variant classification guidelines as a bayesian classification framework, Genetics in Medicine 20 (9) (2018) 1054–1060. 1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [10].Ibeni WNLWH, Salikon MZM, Mustapha A, Daud SA, Salleh MNM, Comparative analysis on bayesian classification for breast cancer problem, Bulletin of Electrical Engineering and Informatics 8 (4) (2019) 1303–1311. 1 [Google Scholar]
  • [11].Lehmann EL, Testing statistical hypothesis, John Wiley & Sons, Inc., 1959. 2, 4 [Google Scholar]
  • [12].Viertl R, Hule H, On bayes’ theorem for fuzzy data, Statistical papers 32 (1) (1991) 115–122. 2, 4, 5 [Google Scholar]
  • [13].Frühwirth-Schnatter S, On fuzzy bayesian inference, Fuzzy sets and systems 60 (1) (1993) 41–58. 2, 4, 5 [Google Scholar]
  • [14].Heller KA, Williamson S, Ghahramani Z, Statistical models for partial membership, in: Proceedings of the 25th International Conference on Machine learning, 2008, pp. 392–399. 2, 4 [Google Scholar]
  • [15].Ma Z, Fokoué E, A comparison of classifiers in performing speaker accent recognition using mfccs, arXiv preprint arXiv:1501.07866. 4 [Google Scholar]
  • [16].Li B, Sufficient dimension reduction: Methods and applications with R, CRC Press, 2018. 4 [Google Scholar]
  • [17].Li K-C, Sliced inverse regression for dimension reduction, Journal of the American Statistical Association 86 (414) (1991) 316–327. 4 [Google Scholar]
  • [18].Li B, Wang S, On directional regression for dimension reduction, Journal of the American Statistical Association 102 (479) (2007) 997–1008. 4 [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

1

RESOURCES