Abstract
We introduce a generalized Bayesian credible set that can achieve any preassigned credible level, addressing a limitation of the current credible sets. This is achieved by exploiting a connection between the highest posterior density set and the Neyman-Pearson lemma.
Keywords: Bayesian classification, Highest posterior density set, Neyman-Pearson lemma, Pattern recognition
1. Introduction
The credible set is one of the most useful tools in Bayesian inference to quantify the uncertainty in Bayesian estimation. See, for example, [1], and [2]. It plays a similar role as the confidence set in the frequentist setting. The most commonly used credible set is the highest posterior density credible set, which has the shortest length for a given credible level. Approximate Bayesian credible sets can be constructed using the Bernstein-von Mises theorem: see, for example, [3] and [4].
However, the classical definition of the highest posterior density credible set [1] is useful mostly for continuous parameters—the parameter supported on an interval or a set with a nonempty and connected interior. In the discrete case where the parameter is supported on a finite or a countable set, the highest posterior density credible set may not exist for a prescribed credible level. This problem is particularly acute for Bayesian classification where the highest posterior density credible set almost never exists for a given credible level. To our knowledge, there has been no systematic method to construct the shortest credible set for such problems.
Bayesian classification has been widely used in many contemporary scientific studies such as cancer research, genomics, and healthcare applications. See, for example, [5], [6],[7],[8], [9], and [10]. Thus, the need for precise uncertainty quantification for Bayesian classification is prevalent. It is high time to facilitate Bayesian classification with exact inference tools and bring it up to par with Bayesian inference for continuous parameters, where exact credible sets have long been available and widely used. The current paper is one step towards satisfying this need.
Our extension is based on the following insight: the process of identifying the highest posterior density credible set is rather similar to that of constructing the most powerful test based on the Neyman-Pearson lemma, and the latter can accommodate any prescribed significance level—even for discrete random variables—by strategically placing probability masses at some transitioning points. This consideration leads us to define the generalized highest posterior density credible set not as a set in the parameter space, but as a mapping from the parameter space to the closed interval [0, 1]. By this adjustment, there always exists a unique shortest credible set for any prescribed credible level, regardless of the nature of the parameter. Accompanying this methodology, we also introduce (in the Supplementary Material) a visualization tool for classification called the Steering Wheel Plot.
2. Generalized credible set
Definition 2.1.
For an , we call any measurable function such that
-level generalized credible set for .
This definition is motivated by the randomized test in classical hypothesis testing theory (see, for example, [11]). Here, the value of can be interpreted as follows: conditioning on the data , at each value of , we perform a Bernoulli trial with success probability ; if the trial turns up head, we classify as belonging to the credible set; otherwise, we classify as not belonging to the credible set. Alternatively, we can also interpret as the membership function in fuzzy logic (see, for example, [12], [13], and [14]), where a larger value represents a stronger tendency of belonging to the credible set.
Under this new framework, the classical credible set, which is a subset of , can be viewed as a mapping from to the set {0, 1}; that is, is the indicator function of the credible set. Since , the generalized credible set indeed generalizes the credible set.
3. Generalized highest posterior density credible set
As mentioned earlier, the highest posterior density credible set is the smallest credible set in the sense that it has the smallest -measure among all credible sets of the same level. More precisely, if is the highest posterior density credible set of level and is any other credible set of level , then . In this section, we extend this optimal property to generalized credible sets. We first generalize the highest posterior density credible set.
Definition 3.1.
Let be a value in the range [0, 1]. If there is a constant such that the generalized credible set
| (1) |
satisfies , then is called the level generalized highest posterior density credible set.
If we take to be constantly 1, the above function is nothing but the indicator function of the set , which is the definition of the highest posterior density credible set. The difference is that for the highest posterior density credible set we can only guarantee , where the equality may not be achieved. In comparison, as we will show in Theorem 3.3, the exact level can always be achieved by the generalized highest posterior density credible set.
Now that the measure is replaced by the integral , the size of a generalized credible set has a new meaning: we simply use to define the size of a generalized credible set with . The next two theorems show that the generalized highest posterior density credible set is the smallest generalized credible set, that it exists, and that it is unique. The proofs of the theorems echo the proof of the Neyman-Pearson lemma.
Theorem 3.2.
If is the level generalized highest posterior density credible set, then it is the smallest level generalized credible set; that is, for any satisfying , we have
We next show the existence of the generalized highest posterior density credible set and the uniqueness of the smallest generalized credible set up to the choice of . The specific form of is established in Theorem 4.2.
Theorem 3.3.
For any and , there exists a ()-level generalized highest posterior density credible set of the form in (1) with . Furthermore, if is any smallest generalized credible set of level , then it has the form (1) a.s. . That is,
Let Define . The next proposition shows that the generalized highest posterior density credible set reduces to the highest posterior density credible set when is continuous at .
Proposition 3.1.
Let . If is continuous at , then the generalized highest posterior density credible set in (1) reduces to
so that is the ()-level highest posterior density credible set.
4. Fair generalized highest posterior density credible set
The generalized highest posterior density credible set in (1) is not unique, unless . When , any satisfying
| (2) |
is optimal in the sense of Theorem 3.2 where . In other words, if we let
| (3) |
then every member of is the generalized highest posterior density credible set. For definiteness, we introduce the fair generalized highest posterior density credible set as follows.
Definition 4.1.
Let be the collection of all generalized highest posterior density credible set of level-() as defined in (3). We call the fair generalized highest posterior density credible set if
When there is more than one generalized highest posterior density credible set, it is desirable to make as similar as possible to give each a fair treatment. That is, we should minimize the variation of with respect to .
The above definition is needed only when . When this probability is 0, it doesn’t matter how we define . For example, we can take on the set . The following theorem characterizes the form of the fair generalized highest posterior density credible set.
Theorem 4.2.
Suppose . Then the fair generalized highest posterior density credible set is the member of with
To compute the generalized highest posterior density credible set, imagine the posterior probabilities at as spikes whose heights correspond to the posterior probabilities. We simply collect the set of highest spikes so that adding another spike (or a few more spikes, if the next few spikes are equal in height) would make the total height exceed the credible level . We then put a weight on the next spike (or distribute the weight evenly among the next few spikes if they are equal in height), so that the posterior probabilities and the weighted posterior probabilities add up to .
5. Data application
We now apply the fair generalized highest posterior density credible set and the steering wheel plot, which can be found in the Supplementary Material, to speaker accent recognition data. It is available at the UCI Machine Learning Repository at https://archive.ics.uci.edu/ml/datasets/Speaker+Accent+Recognition#; see also [15]. There are 12 integer features and six different accents of English: Spanish, French, German, Italian, British, and American. The data has a sample size of 329. Speakers from six different countries read single English words. Twelve features for each subject are extracted from the soundtrack of no more than one second of reading a word using Mel-Frequency Cepstral Coefficients.
For better visual effect, we first perform sufficient dimension reduction to reduce the dimension of the predictors [16]. We apply sliced inverse regression [17] and directional regression [18] to find low-dimensional predictors with four dimensions. We then conduct quadratic discriminant analysis to classify the accents in the reduced space. The 95%-level fair generalized highest posterior density credible set and steering wheel plot (as described in the Supplementary Material) enable us to infer the classification uncertainty.
Figure 1 shows the distinction between the classification uncertainties of two different quadratic discriminant analysis classifiers. For example, considering the French accent, the spokes of the classifier with sliced inverse regression are French, American, Spanish, or British, whereas the spokes of the classifier with directional regression are French or American. We can infer that directional regression helps to reduce classification uncertainty when classifying French accents.
Figure 1.
Speaker accent data: the steering wheel plots (as described in the Supplementary Material) of the first two sufficient predictors. Spanish is red, French is green, German is blue, Italian is black, British is orange, and American is purple. To enhance clarity, five samples are randomly selected for each classified accent, resulting in a total of 30 samples. (a) sliced inverse regression predictors; (b) directional regression predictors.
6. Discussion
In this paper, we generalize the classical Bayesian credible set to achieve exact credible levels for discrete parameters. This fills a gap in Bayesian uncertainty quantification when the random parameter is supported on a finite or discrete set, as in classification problems such as linear or quadratic discriminant analysis.
As a referee pointed out, our construction has a connection with fuzzy Bayesian inference (see, for example, [12], [13], and [14]), and, in particular, with the idea of the membership function. In fact, our generalized credible set can be viewed as a membership function whose greater value (close to 1) represents closer proximity to the credible set. However, there are some important differences in the motivation, formulation, and theory of our approach from fuzzy Bayesian inference.
The original motivation of our construction of the function is the “randomized test” in the classical theory of hypothesis testing. See, for example, [11]. The randomized test generalizes the critical region (rejection region), which is equivalent to an indicator function, by a function that takes values in the interval [0, 1] (as opposed to the set {0, 1}). This was precisely to take care of the scenarios where the significance level cannot be precisely attained. We transplant this idea to the Bayesian setting and apply it to discrete parameters.
However, beyond the membership function, the overlap between our problem and that of fuzzy Bayesian inference is quite small: we do not consider any fuzzy data, where is replaced by a fuzzy number, nor any fuzzy parameter, where is replaced by a fuzzy number. Our setting is completely that of classical Bayesian inference, except that we introduced the analogy of a randomized test, or membership function, to achieve the exact credible level.
Although [12] and [13] also developed the highest posterior density credible set, their purposes were to extend this concept to the context of fuzzy data and fuzzy parameters. However, our purpose in extending the highest posterior density credible set is simply to accommodate discrete parameters. More importantly, the main theoretical novelty of our paper is to establish the existence, optimality, and uniqueness of the generalized highest posterior density credible set by borrowing the argument of the Neyman-Pearson lemma. This key component, however, was absent in the mentioned earlier works.
Our optimal result in Theorem 3.2 can be stated through a Bayes rule in an appropriately formulated statistical decision, which will give further insight into our method. Consider the statistical decision problem with the following components.
Parameter space : this is a discrete or finite set of parameter values (or class labels), say, .
-
Action space : this is the collection of all functions such that
That is, each member of is a generalized credible set of credible level 1 − .
Class of decision rules : this is the set of all mappings from to ; that is, a member of is a function , where is a member of
- Loss function: given a parameter value , and an action , we define the loss function as
- Posterior Expected Loss: given a decision in , the posterior expected loss is
Bayes Rule: Our Theorem 3.2 shows that the generalized highest posterior density credible set defined in Definition 3.1 is, in fact, the Bayes rule with respect to the loss , among the decision rules in .
Supplementary Material
Acknowledgments
The authors would like to thank two anonymous referees, an Associate Editor and the Editor for their constructive comments that improved the quality of this paper. Specifically, the decision theoretical formulation in Section 6 and the connection with fuzzy Bayesian inference were suggested by a referee.
Funding
Bing Li was supported in part by the National Science Foundation (NSF) grant DMS-2210775 and the National Institutes of Health (NIH) grant (1 R01 GM152812-01).
Footnotes
Supplement
The supplementary material includes omitted proofs of the paper and additional results. The code is available at https://github.com/chaegeunsong/ExactBayesianCredibleSet.
Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.
References
- [1].Berger JO, Statistical Decision Theory and Bayesian Analysis, Springer Science & Business Media, 1985. 1 [Google Scholar]
- [2].Li B, Babu GJ, A graduate course on statistical inference, Springer, 2019. 1 [Google Scholar]
- [3].Le Cam L, On some asymptotic properties of maximum likelihood estimates and related bayes estimates, Univ. Calif. Publ. Statist. 1 (1953) 277–330. 1 [Google Scholar]
- [4].van der Vaart AW, Asymptotic statistics, Univ. Calif. Publ. Statist. 1 (1953) 277–330. 1 [Google Scholar]
- [5].Mallick BK, Ghosh D, Ghosh M, Bayesian classification of tumours by using gene expression data, Journal of the Royal Statistical Society: Series B (Statistical Methodology) 67 (2) (2005) 219–234. 1 [Google Scholar]
- [6].Sambo F, Trifoglio E, Di Camillo B, Toffolo GM, Cobelli C, Bag of naïve bayes: biomarker selection and classification from genome-wide snp data, BMC bioinformatics 13 (14) (2012) 1–10. 1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [7].Bhuvaneswari R, Kalaiselvi K, Naive bayesian classification approach in healthcare applications, International Journal of Computer Science and Telecommunications 3 (1) (2012) 106–112. 1 [Google Scholar]
- [8].Knight JM, Ivanov I, Dougherty ER, Mcmc implementation of the optimal bayesian classifier for non-gaussian models: model-based rna-seq classification, BMC bioinformatics 15 (1) (2014) 1–13. 1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [9].Tavtigian SV, Greenblatt MS, Harrison SM, Nussbaum RL, Prabhu SA, Boucher KM, Biesecker LG, Group CSVIW, et al. , Modeling the acmg/amp variant classification guidelines as a bayesian classification framework, Genetics in Medicine 20 (9) (2018) 1054–1060. 1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [10].Ibeni WNLWH, Salikon MZM, Mustapha A, Daud SA, Salleh MNM, Comparative analysis on bayesian classification for breast cancer problem, Bulletin of Electrical Engineering and Informatics 8 (4) (2019) 1303–1311. 1 [Google Scholar]
- [11].Lehmann EL, Testing statistical hypothesis, John Wiley & Sons, Inc., 1959. 2, 4 [Google Scholar]
- [12].Viertl R, Hule H, On bayes’ theorem for fuzzy data, Statistical papers 32 (1) (1991) 115–122. 2, 4, 5 [Google Scholar]
- [13].Frühwirth-Schnatter S, On fuzzy bayesian inference, Fuzzy sets and systems 60 (1) (1993) 41–58. 2, 4, 5 [Google Scholar]
- [14].Heller KA, Williamson S, Ghahramani Z, Statistical models for partial membership, in: Proceedings of the 25th International Conference on Machine learning, 2008, pp. 392–399. 2, 4 [Google Scholar]
- [15].Ma Z, Fokoué E, A comparison of classifiers in performing speaker accent recognition using mfccs, arXiv preprint arXiv:1501.07866. 4 [Google Scholar]
- [16].Li B, Sufficient dimension reduction: Methods and applications with R, CRC Press, 2018. 4 [Google Scholar]
- [17].Li K-C, Sliced inverse regression for dimension reduction, Journal of the American Statistical Association 86 (414) (1991) 316–327. 4 [Google Scholar]
- [18].Li B, Wang S, On directional regression for dimension reduction, Journal of the American Statistical Association 102 (479) (2007) 997–1008. 4 [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.

