Abstract
Decision making is a fundamental subfield within neuroscience. While recent findings have yielded major advances in our understanding of decision making, confidence in such decisions remains poorly understood. In this paper, we present a confidence signal detection (CSD) model that combines a standard signal detection model yielding a noisy decision variable with a model of confidence. The CSD model requires quantitative measures of confidence obtained by recording confidence probability judgments. Specifically, we model confidence probability judgments for binary direction recognition (e.g., did I move left or right) decisions. We use our CSD model to study both confidence calibration (i.e., how does confidence compare with performance) and the distributions of confidence probability judgments. We evaluate two variants of our CSD model: a conventional model with two free parameters (CSD2) that assumes that confidence is well calibrated and our new model with three free parameters (CSD3) that includes an additional confidence scaling factor. On average, our CSD2 and CSD3 models explain 73 and 82%, respectively, of the variance found in our empirical data set. Furthermore, for our large data sets consisting of 3,600 trials per subject, correlation and residual analyses suggest that the CSD3 model better explains the predominant aspects of the empirical data than the CSD2 model, especially for subjects whose confidence is not well calibrated. Moreover, simulations show that asymmetric confidence distributions can lead traditional confidence calibration analyses to suggest “underconfidence” even when confidence is perfectly calibrated. These findings show that this CSD model can be used to help improve our understanding of confidence and decision making.
NEW & NOTEWORTHY We make life-or-death decisions each day; our actions depend on our “confidence.” Though confidence, accuracy, and response time are the three pillars of decision making, we know little about confidence. In a previous paper, we presented a new model — dependent on a single scaling parameter — that transforms decision variables to confidence. Here we show that this model explains the empirical human confidence distributions obtained during a vestibular direction recognition task better than standard signal detection models.
Keywords: confidence calibration, confidence rating, decision making, probability judgments, thresholds, vestibular
INTRODUCTION
In this paper, we analyze confidence probability judgments using a confidence signal detection (CSD, pronounced “kissed”) model that links a model of confidence with a standard signal detection model. As noted by a recent study, “Confidence judgments, self-assessments about the quality of a subject’s knowledge, are considered a central example of metacognition” (Kepecs and Mainen 2012). Until relatively recently, confidence has played a “Cinderella role . . . overlooked as an interesting variable in its own right” (Vickers 2001). But studies focused on confidence have recently gained increasing prominence, including a variety of confidence models (Fetsch et al. 2014; Grimaldi et al. 2015; Maniscalco and Lau 2012; Pleskac and Busemeyer 2010; Ratcliff and Starns 2013; Yu et al. 2015).
In a previous article (Yi and Merfeld 2016) that includes more introductory material, we used our CSD model to fit a psychometric function and a linked confidence function, which (as described in greater detail later) exactly matches the psychometric function representative of average accuracy when a subject is well calibrated, to individual data sets comprised of confidence probability judgments. Human confidence probability judgments were acquired from four human subjects, and extensive simulation data were also acquired and analyzed. All stimuli for that paper — both for human studies and simulations — were sampled via standard adaptive staircase procedures (i.e., a 3-down/1-up staircase). Adaptive procedures were used because these have been shown to yield efficient sampling when thresholds are not known in advance. The findings presented in that paper suggested that the use of confidence probability judgments have the potential to lead to a fivefold efficiency improvement when fitting psychometric functions, since ~20 trials utilizing confidence probability judgments delivered the same fit parameter precision (i.e., fit parameter error bars) for a direction recognition (sometimes called direction discrimination) threshold estimate as 100 trials using traditional methods.
To perform our CSD analysis (here and in the earlier paper) we assumed that perceptual decision making is probabilistic because physiological noise is present. Furthermore, like many earlier signal detection analyses (e.g., Green and Swets 1966), this noisy decision-making process for each individual trial is assumed to be represented by a decision variable sampled from a probabilistic (i.e., noisy) distribution, with each sampled decision variable (one per trial) leading to a deterministic decision for each trial depending on the location of the sampled decision variable relative to the subjective decision boundary. To link confidence to the psychometric function, we add three assumptions to the standard signal detection model: 1) that the same sampled decision variable used to make a decision is also used to determine confidence, 2) that the subject’s confidence is based on how far the sampled decision variable is from the decision boundary relative to the noise level, and 3) that confidence need not be well calibrated, allowing us to model data from overconfident or underconfident subjects.
It is crucial to note that our modeling builds on a number of earlier studies that include the idea that confidence depends on distance between the sampled decision variable and the decision boundary (e.g., Balakrishnan and Ratcliff 1996; Björkman et al. 1993; Ferrell 1995; Ferrell and McGoey 1980; Fetsch et al. 2014; Grimaldi et al. 2015; Kepecs and Mainen 2012; Maniscalco and Lau 2012; Massoni et al. 2014; Pleskac and Busemeyer 2010; Ratcliff and Starns 2013). In fact, our CSD2 model includes the predominant characteristics of these earlier models that used signal detection theory to help model confidence. (See more extended treatment in the discussion.) Despite these influential efforts, we are not aware of any previous efforts by others to use confidence to help estimate psychometric fit parameters nor any efforts to link underconfidence (or overconfidence) to the underlying decision variable via a model that scales confidence.
In the previous article, the CSD model was introduced analytically, which has the clear benefit of yielding analytic equations, but, for some readers, such analyses can obscure the underlying mechanisms. To provide a more mechanistic description of the model, we here introduce the CSD model using simulations. For some readers, this will help illustrate how the psychometric function and confidence are linked for each sampled decision variable.
We begin with a stimulus (i.e., a signal). Each stimulus is assumed well controlled relative to the physiological noise, so each specific stimulus is represented as a constant. The distribution for a specific well-controlled stimulus is represented by an impulsive probability density function (sometimes called a delta function) having infinite height and infinitesimal width (Fig. 1A). For the case shown, we arbitrarily assume a stimulus amplitude value of 2.0, which is plotted on the x-axis: .
Fig. 1.
Simulated confidence distribution under our signal detection confidence (CSD) model. A: the stimulus for this example is well controlled having an amplitude of +2.0 with little variation, so the objective probability density function is a delta function. B: a signal detection model assumes additive noise. For this example, Gaussian noise having zero-mean and a standard deviation of 1 was simulated ε = N(μ = 0, σ = 1); the distribution for 10,000 samples is shown. The dotted vertical line at zero represents a decision boundary. If a sampled decision variable (dj) on the jth trial falls to the right of the decision boundary, the subject decides positive. If the sampled decision variable falls to the left, the subject decides negative. For this example, 97.7% of the trials (i.e., decision variables) lead to the subject deciding positive. C: the asterisk, located at (2, 0.977), represents the example data point illustrated in the previous panel. When this process is repeated for a variety of different stimulus levels, it yields a psychometric function, Ψ(x) = φ(x, μ = 0, σ = 1) (black curve). D: the confidence function for a well-calibrated subject (k = 1) is the same as the psychometric function shown in C, χ(x) = Ψ(x). E: the confidence distribution that results from the 10,000 sampled decision variables represented in B. The distribution is calculated by taking each of the 10,000 decision variables represented in B and determining the confidence using the function in D, χ(j) = φ(dj, μ = 0, kσ = 1). For example, a decision variable of zero would convert to a confidence of 0.5, and decision variable of 2 would convert to a confidence of 0.977. Ordinate for E is truncated at 5 to help illustrate variations. F through J repeat the above process but for an underconfident subject having a confidence-scaling factor of two (k = 2). I shows this underconfidence confidence function, χ(j) = φ(dj, μ = 0, kσ = 2).
For this simulated demonstration, we assume that the physiologic noise (ε) is Gaussian with a standard deviation of 1 (σ = 1) and mean of zero (μ = 0), ε = N(0,1). Given this noise model, the simulated distribution of the decision variable across 10,000 trials for a stimulus amplitude of 2.0 is shown (Fig. 1B). Each of these sampled values represents the decision variable available to the nervous system for a single trial for the given noise distribution and the given stimulus level.
We assume that our signal detection subject has no decision biases — i.e., that the a priori probability for the direction of each stimulus (e.g., left or right) is equal and costs for all decisions are equal. Therefore, the ideal signal detector would set a decision boundary at zero, where a decision boundary represents the border that delineates whether the subject decides positive or negative. Placing the decision boundary at zero means that when an individual trial yields a decision variable that is positive, the subject reports positive (right); when a trial yields a negative decision variable, the subject reports negative (left). For the example shown, the predicted distribution falls above the decision boundary (right) 97.7% of the time and below the decision boundary (left) 2.3% of the time. This 97.7% data point is shown at the stimulus level of 2.0 in Fig. 1C.
When this process is repeated many times for many different stimulus levels controlled by the operator, a psychometric data set is generated, which is often quantified by fitting a psychometric function to these data. Such a psychometric function reflects expected average performance at each stimulus level. As described in the previous article, given the assumed Gaussian noise distribution, the pertinent theoretic psychometric curve is a cumulative Gaussian, which is shown in Fig. 1C. Using the nomenclature introduced in the previous article, a fitted Gaussian cumulative distribution function (ϕ) has two fit parameters () such that the psychophysical function1 can be written (see appendix a). But we emphasize that each point on this curve could be determined empirically by simply repeating the process described above and illustrated via Fig. 1, A and B for any stimulus level other than 2. With enough data, this empirically determined psychometric function is typically assumed to converge to a psychometric function representative of the subject’s underlying noise distribution, .
We next define the relationship between a confidence function and the psychometric function. For simplicity, we begin by defining a well-calibrated simulated subject, but we later relax this calibration assumption. For a symmetric task (having uniform priors) like ours (e.g., Merfeld 2011; Yi and Merfeld 2016), perfect confidence calibration would be defined to occur when subjective confidence matches objectively assessed accuracy (e.g., Fig. 1, D or I herein or Fig. 3A in Drugowitsch et al. 2014). Recall that average accuracy is represented by the psychometric function. Therefore, perfectly calibrated confidence would be reflected by a confidence function that equaled the psychometric function , as shown in Fig. 1D.
Specifically, imagine that on a given trial, a simulated well-calibrated subject has sampled a decision variable having an amplitude of 2.0 (i.e., one trial from the distribution shown in Fig. 1B) that he/she has converted to a confidence probability judgment. Assuming that the subject has an accurate estimate of the noise distribution (σ = 1 and μ = 0) and ignoring the actual neural processes employed, our simulated model of a well-calibrated subject calculates confidence by directly mapping the decision variable for each trial onto the “confidence function” to determine a confidence probability judgment, which for this trial yields 97.7%, ϕ(x = 2, μ = 0, σ = 1) = 0.977. When confidence is similarly calculated for each sampled decision variable (each of the trials represented in Fig. 1B), this process yields one confidence value for each sampled decision variable. For this simulation, these confidence values yield the normalized confidence histogram shown in Fig. 1E. The right column of Fig. 1 shows the same calculations as the left-hand column but for a simulated underconfident subject having a confidence scaling of two, which simply means that the confidence function has a slope parameter that is half that of the psychometric function. We specifically highlight that the confidence function for an underconfident subject (Fig. 1I) differs from that for a well-calibrated subject (Fig. 1D), which leads to different distributions of confidence (compare Fig. 1, E and J) even for identical stimuli (Fig. 1, A and F), physiologic noise (Fig. 1, B and G), and psychometric functions (Fig. 1, C and H).
While the previous article focused on fitting data (Yi and Merfeld 2016), herein we now apply this CSD model to analyze confidence probability judgments. These analyses will help us begin to evaluate the validity of the assumptions outlined above and in the previous paper. To perform this analysis, we empirically determined confidence distributions for four human subjects via extensive testing (circa 3,600 trials each).
METHODS
Since the general experimental and simulation methods match those described in more detail in the previous article (Yi and Merfeld 2016), we focus most of the methodological descriptions included herein on methods not described in the previous article, on methodological differences from the previous article, and on the more novel aspects of these quantitative confidence analyses.
Human Studies
Basic human experimental methods (e.g., forced-choice direction-recognition task performed in complete darkness on a Moog platform equipped with a human chair/restraint system, 1-Hz single-cycle sine yaw rotation motion stimuli, auditory white noise during motion, confidence analysis, and fitting methods) were identical to those reported in the previous article (Yi and Merfeld 2016), where these specific methods are described in more detail. The primary difference was that we chose to use a nonadaptive stimulus sampling procedure for the main study presented herein. Before beginning the main study, each subject was tested with a 3-down/1-up staircase procedure that was designed to determine an initial rough estimate of each subject’s threshold (), where i signifies that this is an initial estimate and j signifies the jth subject. For this adaptive test, 75 trials were performed.
After estimating each subject’s threshold () using unbiased psychometric fit methods (Chaudhuri and Merfeld 2013), a nonadaptive sampling procedure — having 12 different stimulus levels (±0.01, ± 0.1, ± 0.2, ± 0.3 ± 1.0, ± 2.0) — was designed for each subject. This allowed us to obtain repeated trials at different stimulus levels, which simplifies some analyses. Table 1 shows the actual stimulus levels used for each subject. Each stimulus was provided 10 times for each test, yielding 120 trials for each test. This nonadaptive test, each consisting of 120 trials, was repeated 30 times for each subject, yielding a total of 3,600 trials. This corresponds to 300 repeated trials at each of the 12 stimulus levels for each subject.
Table 1.
Nonadaptive stimulus levels used for each subject
| Subject | Experimental stimulus amplitude | |||||
|---|---|---|---|---|---|---|
| S1 | ± 0.03* | ± 0.10 | ± 0.20 | ± 0.29 | ± 0.98 | ± 1.95 |
| S2 | ± 0.03* | ± 0.06 | ± 0.12 | ± 0.19 | ± 0.62 | ± 1.24 |
| S3 | ± 0.03* | ± 0.05 | ± 0.06 | ± 0.09 | ± 0.29 | ± 0.59 |
| S4 | ± 0.03* | ± 0.15 | ± 0.30 | ± 0.45 | ± 1.51 | ± 3.02 |
Peak stimulus velocity is shown in degrees per second.
The smallest stimulus magnitude was set to be the minimum motion that could be reliably provided by our MOOG platform.
As for the previous study, both direction and confidence responses were reported using an iPad that was held by the subject using both hands. Once each trial began, the iPad was totally “dark,” meaning that all iPad backlighting was turned off. After each trial, while still in the dark, the subject tapped on the left side of the iPad to report perceived motion to the left and tapped on the right side to report perceived motion to the right. As soon as the iPad registered this binary decision, the iPad screen illuminated to display confidence-rating sliders that ranged from 50 to 100%. Subjects moved the selected slider up/down to indicate their confidence. The subject’s response — both direction (left/right) and confidence (50 to 100%) — were displayed on the screen. The subject could adjust their direction or confidence response (or both) until they were satisfied. When satisfied, they tapped a button labeled “Confirm,” and the iPad and subject prepared for the next trial. Subjects practiced before each session to be sure that they understood and/or remembered the task.
The confidence resolution available to the subject was 1% (e.g., 50, 51, 52, . . ., 100%), Confidence distributions presented in this manuscript were analyzed by binning the confidence judgments into bins having 10% resolution. To allow direct comparisons, the exact same binning was used for human data and simulations for all analyses.
Four healthy human subjects (2 male, 2 female, 26–34 yr old) were tested on 5 separate days. Written, informed consent was obtained from all subjects before participation in the study. The study was approved by the Massachusetts Eye and Ear Human Studies Committee and was performed in accordance with the ethical standards laid down in the 1964 Declaration of Helsinki.
All four subjects were the same subjects tested for the previous article; the testing for these two studies was separated by about a month. An author (Y. Yi) participated as one of the subjects. Since the computer randomly selected the stimuli, this subject did not have information to guide his binary reports or confidence judgments on each individual trial. More importantly, as noted in the results, this subject’s responses did not differ from the other subjects in any substantive manner.
Simulation
All simulations were performed in MATLAB R2015b (The MathWorks) on the Harvard Orchestra computation cluster using parallel IBM BladeCenter HS21 XMs with 3.16 GHz Xeon processors and 8 GB of RAM. Basic simulation methods were the same as described in previous article, except that we conducted simulations using the same nonadaptive sampling procedure described in the previous section for the human testing. The nonadaptive simulations were performed with n = 120 trials, using ± 0.01σ, ± 0.1σ, ± 0.2σ, ± 0.3σ, ± 1.0σ and ± 2.0σ as the stimulus levels, where σ = 1; 10,000 data sets were simulated for each condition.
Data Analysis
Psychometric function fits.
The exact same fitting procedures described in the previous article (Yi and Merfeld 2016) were used. Briefly, without providing all the details, we fit both a standard Gaussian psychometric function, , and a Gaussian confidence function, χ̂, to the data using maximum likelihood methods. To investigate the impact of the confidence-scaling factor (), we fit the data via two different ways. For one fit, we fixed the confidence-scaling factor, , equal to one (). We will refer to this as the CSD two-parameter (CSD2) model, since this model has two degrees of freedom ( and ) — no additional degrees of freedom beyond those inherent to the psychometric function. For this CSD2 model, the confidence function is defined to have the exact same parameters as the psychometric function — assuming perfect calibration. This model appears to include the predominant characteristics of previous signal detection models of human confidence (e.g., Björkman et al. 1993; Ferrell 1995; Ferrell and McGoey 1980; Maniscalco and Lau 2012; Massoni et al. 2014) modified as appropriate for our direction-recognition task. For the second fit, as in the previous article, the confidence-scaling factor provided a third free parameter. We refer to this fit as the CSD3 fit. As observed in Fig. 1, D and I, this CSD3 model allows different width parameters for the psychometric function and the confidence function to represent overconfidence () or underconfidence ().
Confidence distribution model.
We will utilize the general CSD model illustrated in Fig. 1 to predict confidence distributions for different stimulus levels and the two different CSD models — CSD2 (left column Fig. 1) and CSD3 (right column Fig. 1). To perform these calculations, we must provide a stimulus level (Sj) and a quantitative CSD model — both a psychometric function, , and a confidence function, , where μ, σ, and k are assumed known.
We illustrate with two examples. The first example uses the exact same CSD model as the right column of Fig. 1 (μ = 0, σ = 1, and k = 2). To illustrate how confidence distributions vary for different stimulus levels, for Fig. 2, we select five stimulus levels (s = −2, −1, 0, 1, or 2) and repeat the process shown in Fig. 1 for each of the different stimulus levels. In fact, Fig. 2E is identical to Fig. 1E except for different graphical aspect ratios. The top row of Fig. 2 shows these simulated confidence distributions for this well-calibrated subject for each of the five different stimulus amplitudes. Figure 2, A, B, and C shows confidence distributions for large negative stimuli, moderate negative stimuli, and no stimuli (i.e., stimulus amplitude equals 0), respectively. Figure 2, D and E shows the simulated confidence distributions for moderate positive stimuli and large positive stimuli, respectively. The second row (Fig. 2, G–L) shows the simulated confidence distributions for the exact same stimuli and the same CSD model except that these confidence data represent a simulated underconfident subject having a confidence-scaling factor of 2 (k = 2) — the same parameters as for the underconfident subject represented in the right column of Fig. 1.
Fig. 2.
Predicted confidence distributions for different models and different stimulus levels. Distributions are calculated using the process outlined in Fig. 1. From left to right, predicted confidence distributions are shown for stimuli equal to −2, −1, 0, 1, and 2, respectively. Top row shows predicted confidence distributions for the exact same noise model ε = N(μ = 0, σ = 1) and confidence function χ(x) = Ψ(x) = ϕ(x, μ = 0, kσ = 1) shown in the left column of Fig. 1. The bottom row shows predicted confidence distributions for the exact same noise model but for an underconfident subject having a confidence-scaling factor of 2 χ(x) = ϕ(x, μ = 0, kσ = 2).
Note that for stimuli with zero amplitude, the confidence distributions span broad confidence ranges (i.e., nearly 0% to 100%) (Fig. 2, C and I). In fact, for zero-amplitude stimuli and the well-calibrated simulated subject, a uniform distribution of confidence reports is predicted (Fig. 2C). This falls directly from the assumption of perfect confidence calibration.
We calculated standard information criteria as goodness of fit metrics to assess the CSD model fits. Deviance, which is a standard measure of maximized likelihood, was calculated for all maximum likelihood fits. Likelihood ratio was calculated to compare between both model fits.
Correlations were calculated between the observed/simulated confidence distributions as well as between the CSD2 and CSD3 models (Rosner et al. 2015). Mixed-effects models were used to compare CSD2 and CSD3 fits to the observed/simulated confidence distributions with method as the fixed effect and subjects as random effects for Table 4. Mixed-effects models were used to compare CSD2 and CSD3 fits to the observed/simulated confidence distributions with stimulus levels and confidence levels as fixed effects and subjects as random effects for Figs. 5 and 8. Both correlation analyses and mixed models take into account the correlations due to repeated measures from each subject. The analyses were performed using SAS Version 9.4.
Table 4.
Variance explained by CSD2 and CSD3
| Subject | Mean Model | CSD2 | CSD3 |
|---|---|---|---|
| S1 | 0.0141 | 0.0015 (89.4%) | 0.0011 (92.2%) |
| S2 | 0.0113 | 0.0028 (75.2%) | 0.0020 (82.3%) |
| S3 | 0.0154 | 0.0141 (8.44%) | 0.0115 (25.3%) |
| S4 | 0.0136 | 0.0078 (42.6%) | 0.0030 (77.9%) |
Values for three columns reflect the mean squared difference between the empirical confidence bin values and 1) the mean, 2) CSD2 model predictions, and 3) CSD3 model predictions, respectively. The mean model is the mean of the 10 bins, which is, by definition, 0.1. Values in parentheses show percent of overall variations explained by the two CSD models.
Fig. 5.
Confidence probability judgment distributions. Each row represents one of the four subjects (from top to bottom S1 through S4, ordered from subject with lowest confidence-scaling factor on top to highest on bottom). Histograms show empirical human data at each of five stimulus levels. Each column represents different stimulus levels; the largest stimulus magnitudes are represented by the 1st and 5th columns, the 2nd largest stimulus magnitudes are represented by the 2nd and 4th columns, and the remaining relatively small stimuli are represented by the middle (3rd) column. For S3, because the stimuli tested were smaller relative to the actual threshold than for the other three subjects, the 2nd and 4th column show the largest stimulus magnitudes and the center column shows the confidence for remaining subthreshold stimuli. See Table 1 for actual stimulus levels for each subject. Predicted confidence judgment distributions for the CSD2 (red +) and CSD3 (blue x) models using fitted parameters for each subject are overlapped for comparison.
Fig. 8.
Simulated confidence probability judgment distributions. Top row shows simulated subject S1. Bottom row shows simulated subject S4. Histograms show simulated data at each of five stimulus levels (−2, −1, 0, 1, 2). Predicted confidence judgment distributions for the CSD2 (red +) and CSD3 (blue ×) models using fitted parameters for each subject are overlapped for comparison.
RESULTS
Human Studies
Four human subjects recorded confidence for each of their binary decisions as described in the methods. To quantify whether we could find any evidence of learning, we analyzed the sequence of 30 tests (each consisting of 120 trials) for each subject specifically looking for a consistent trend across subjects. Table 2 shows fitted linear regression slopes and associated P values for each of 12 cases (4 subjects × 3 fit parameters each). The slopes of these sequential data sets for the four subjects were not consistent for any one of the three fit parameters (i.e., sign of slope for at least one of the subjects was different than for the other subjects). Furthermore, the lowest P value for any of the 12 cases evaluated was 0.03, which, given multiple comparisons, is not significant. Using a Bonferroni multiple comparison, the criterion boundary would be P < 0.0042 (0.05/12). These results demonstrated little or no evidence of systematic learning over the course of 3,600 trials for these subjects.
Figures 3 and 4 show the fitted psychometric function parameters (, ) and confidence-scaling parameter () as the trial number incrementally increased from 24 to 120 trials in increments of 12 (1 trial at each of the 12 stimulus levels); mean estimates and standard deviation across the 30 tests for each of the four subjects are shown. The plotted parameter estimates in Fig. 3 are determined using the CSD2 model that assumes that the subject is well calibrated (), and the estimates in Fig. 4 are determined using the CSD3 model, where the confidence-scaling factor () is a free parameter fit to the data.
Fig. 3.
Parameter fits for the CSD2 model that fits a psychometric function under a Gaussian noise model, Ψ(x) = ϕ(x, , ), with the additional assumption that confidence is perfectly calibrated χ(x) = Ψ(x) = ϕ(x, , ). Each column represents fitted parameters for one of the four subjects (from left to right S1 through S4) in the same order as the previous article (ordered from subject with lowest confidence-scaling factor on left to highest on right). Top row (A–D) shows fitted psychometric width parameter (). Bottom row (E–H) shows fitted psychometric function bias (). Thick black curves show average psychometric parameter estimates calculated using conventional forced-choice analyses. Thick red curves show average parameter estimates determined by fitting confidence probability judgment data. Errors bars (thin gray curves and thin red curves, respectively) represent standard deviation of parameter estimates. Data points at 24 (36, 48, . . . 120) trials represent the mean threshold value (across the 30 test sessions) obtained by analyzing data obtained during the first 24 (36, 48, . . . 120) trials for each of the 30 test sessions independently and separately.
Fig. 4.
Parameter fits for the CSD3 model that fits a psychometric function under a Gaussian noise model, Ψ(x) = ϕ(x, , ) and that simultaneously fits a Gaussian confidence function, χ(x) = ϕ(x, , ). Each column represents one of the four subjects in the same order as Fig. 3. Top row (A–D) shows fitted psychometric width parameter (). Middle row (E–H) shows fitted confidence-scaling factor (). Bottom row (I–L) shows fitted psychometric function bias (). Thick black curves show average psychometric parameter estimates calculated using conventional forced-choice analyses. Thick red curves show average parameter estimates determined by fitting confidence probability judgment data. Errors bars (thin gray curves and thin red curves, respectively) represent standard deviation of parameter estimates. Data points at 24 (36, 48, . . . 120) trials represent the mean threshold value (across the 30 test sessions) obtained by analyzing data obtained during the first 24 (36, 48, . . . 120) trials for each of the 30 test sessions independently and separately.
We focus first on the CSD2 fits (Fig. 3) and note that, aside from one notable exception, the parameters estimated using the CSD2 model (red) converged to be near the conventional parameter estimates. The notable exception was the average width parameter estimates () for the fourth subject (Fig. 3D); the CSD2 width parameter estimate was roughly twice the conventional estimate. This difference nearly disappears for the CSD3 fit (Fig. 4D).
Looking at the confidence-scaling factor (, Fig. 4H), we see that this subject’s fits suggest that this subject was underconfident — with the confidence-scaling factor converging to a stable value near 2. In comparison, note that one of the subjects seemed very well calibrated (Fig. 4E), with a confidence-scaling factor near 1, and the two other subjects seemed slightly underconfident (Fig. 4, F and G), with confidence-scaling factors converging to values just a little greater than 1.
Table 3 provides standard goodness of fit metrics for these maximum likelihood fits. Deviance was always less for the CSD3 model than for the CSD2 model; this reflects that the additional degree of freedom offered by the CSD3 model must, by definition, provide better fits than the CSD2 model. Therefore, we compared models using a log-likelihood ratio test that showed that the CSD3 model fit was significantly better than the CSD2 fit for each of the four subjects (P = 0.011 for S1, and P < 0.0001 for S2, S3, and S4). Not surprisingly, the difference in deviance between the CSD2 and CSD3 models changed the most for S4 — the subject who demonstrated the greatest degree of underconfidence. For this subject (S4), the log-likelihood ratio was significantly different between the CSD2 and CSD3 model fits. Similarly, for the other three subjects (S1, S2, and S3, all of whom were slightly underconfident), the log-likelihood ratios demonstrated a significant difference in the CSD2 and CSD3 model fits.
Table 3.
Goodness of fit parameters
| Subject | Deviance CSD2 | Deviance CSD3 | Log-Likelihood Ratio | P value |
|---|---|---|---|---|
| S1 | 977.2 (39.3) | 970.7 (38.7) | 6.44 (7.55) | P = 0.011 |
| S2 | 1019.3 (31.0) | 1000.2 (27.1) | 19.10 (18.59) | P < 0.0001 |
| S3 | 1086.8 (12.0) | 1063.3 (17.5) | 23.49 (20.60) | P < 0.0001 |
| S4 | 1048.9 (15.9) | 971.2 (18.5) | 77.66 (20.26) | P < 0.0001 |
Mean (and standard deviation) of deviance and log-likelihood ratio are shown for each subject for both CSD2 and CSD3 model fits.
Generally, the parameter fits were consistent with those reported in the previous article (Yi and Merfeld 2016). The parameter estimates obtained using conventional psychometric fits were more variable than the fits obtained using the CSD model. Subject 3 binary parameter estimates were especially noisy and never showed clear stabilization (Figs. 3C and 4C). (We will discuss the methodological explanation of subject 3’s irregularity in the discussion. In brief, the initial estimate of this subject’s threshold obtained using a standard adaptive procedure, underestimated the actual threshold, leading to suboptimal sampling in the main nonadaptive test series, which led, in turn, to noisier fits.) The estimates of the psychometric function width parameter () obtained via the confidence fit (Figs. 3 and 4) appeared much more stable (i.e., flatter, less variable) than those obtained via conventional fitting methods (Figs. 3 and 4), especially for subject 3. Furthermore, the precision of the psychometric width estimate using the confidence model was almost always better (i.e., had smaller error bars) than the conventional psychometric fit estimate.
Figure 5 shows confidence histograms that show the probability of different confidence reports at each of 5 stimulus levels. The gray bars show the empirical human data for each of the four subjects. Consistent with predictions shown in Fig. 2, we see that the confidence distributions skew toward confidence judgments of zero for large negative stimuli (leftmost columns) and skew toward confidence judgments of 1 for large positive stimuli (rightmost columns). For small (subthreshold) stimuli, the confidence judgments span the entire range for all four subjects — mimicking the predicted distributions shown in the center column of Fig. 2.
Using the CSD2 and CSD3 model fits, we can predict what these confidence distributions should look like if each model captured the true experimental variations. In other words, these theoretical predictions (calculated using the methods demonstrated for Fig. 2 but now using empirically determined fit parameters for each subject) demonstrate the predicted confidence distributions when the fitted confidence function and fitted psychometric function accurately represent the underlying processes. The CSD3 model fit the confidence histogram density very well. The model fits were excellent for stimuli near threshold (σ) and two times threshold (2σ) when confidence was higher (e.g., usually >70%). While still good, the model fits were not as superb for very small stimuli (e.g., 0.1σ), when confidence was lower (e.g., often <70%). The CSD3 fit was significantly better than the CSD2 fit for 310 of the 480 confidence bins (4 subjects × 12 stimulus levels × 10 confidence bins each); this difference with the CSD3 fit being better than the CSD2 fit was statistically significant (P < 0.05). The Spearman correlation between the observed confidence distributions and CSD2 fits was 0.616 with a 95% confidence interval of [0.559, 0.667], and the Spearman correlation between the observed confidence distributions and CSD3 fits was 0.825 with a 95% confidence interval of [0.795, 0.851]. This means that CSD3 fits correlated with the observed distributions significantly better than CSD2 fits (difference = 0.209, P < 0.001).
Consistent with these findings, Table 4 shows the unexplained variance for the CSD2 and CSD3 models fitted for each subject (i.e., the squared difference between the model predictions and empirical value for each bin). There is a statistically significant difference between the mean confidence density and both the CSD3 (P = 0.0077 with Bonferroni correction) and the CSD2 (P = 0.0270 with Bonferroni correction) fits. While both models captured the general characteristics observed in the confidence distributions (Fig. 5), the unexplained variance for the CSD3 model was lower than for the CSD2 model for all four subjects. Not surprisingly, the biggest relative difference (+35% for CSD2 relative to CSD3) was found for the most underconfident subject (S4) with the smallest relative difference (+3.0%) found for the most well-calibrated subject (S1). Since three of the subjects (S1, S2, and S3) were fairly well calibrated, meaning that the differences between CSD2 and CSD3 models for these three subjects are not large, the mean difference was not statistically significant (P = 0.8728 with Bonferroni correction).
Table 4 also shows the % variance explained by the CSD2 and CSD3 models relative to the mean of the confidence histogram density (y-axis amplitude shown in Fig. 2). Both explain the majority of the variance present in each subject’s confidence distributions, but the CSD3 model explains more of the variation than the CSD2 model. Specifically, across the four subjects, CSD2 and CSD3 fits, explained 53.9 and 69.4% of the variance, respectively. For the three subjects whose confidence was well calibrated (S1, S2, and S3), the difference between the variance explained by the CSD3 and CSD2 models was not large. For the most underconfident subject (S4), the CSD3 fit explained 77.9% of the variance while the CSD2 fit explained 42.6% of the variance — a 235% improvement by the CSD3 model fit. We also note that, if we remove the subject (S3) who received mostly subthreshold stimuli, the CSD3 model explained over 84% of the confidence histogram density variance across the three remaining subjects.
We also provide conventional analyses of confidence. Figure 6 shows confidence calibration curves that plot confidence versus binary performance — quantified as the percentage of binary decisions that are “positive.” To take advantage of large numbers, Fig. 6 combines the data across all 30 tests (120 trials each, yielding 3,600 total trials), so each of the 12 data points represents 300 trials. Figure 6, A–D plots the mean confidence data versus average performance. Between 50 and 100%, the mean confidence falls below performance, which demonstrates “underconfidence.” Between 0 and 50%, the mean confidence falls above performance for all subjects, which is also representative of underconfidence. Hence, the mean confidence data suggest that all subjects were underconfident.
Fig. 6.
Confidence calibration plots. Each column represents one of the four subjects in same order as Fig. 3. Top row shows conventional calibration plots where average confidence for each stimulus level is plotted versus average accuracy. Bottom row shows calibration plots but with median confidence replacing mean confidence. For comparison, Fig. 9 shows similar plots for simulated subjects. Errors bars represent standard deviation.
Before proceeding, note that taking the mean of the confidence data does not fully consider the skewed confidence distributions (e.g., Fig. 5). To take the skewed distributions into account, Fig. 6, E–H plots the exact same data in the same calibration plot format; the only difference is that this time the median confidence is plotted instead of the mean. (See simulated data in Fig. 9 for further support for plotting the median.)
Fig. 9.
Simulated confidence calibration plots. Errors bars represent standard deviation. First column shows confidence calibration plots for simulated subject S1. Second column shows same confidence data for simulated subject S1 plotted versus stimulus level in a psychometric function plot format where the actual psychometric function is plotted as a dashed curve. Third column shows confidence calibration plots for simulated subject S4. Fourth column shows same confidence data for simulated subject S4 plotted versus stimulus level in a psychometric function plot format where the actual psychometric function is plotted as a dashed curve. Top row shows conventional plots with average confidence plotted versus average accuracy. For bottom row, median confidence replaces mean confidence for y-axis of all four subplots.
As another conventional confidence analyses, we calculated the Brier score and three components of its decomposition: reliability, resolution, and uncertainty (Brier 1950; Murphy 1973). Table 5 shows these metrics for each of the four subjects. Uncertainty is not shown in the table because the value was exactly 0.25 (by definition), since we provided exactly same number of stimuli for both motion directions [p(1−p) = 0.5(1−0.5) = 0.25]. Note that the substantively underconfident subject (S4) had better metrics — a lower Brier score and higher resolution and reliability — than the well-calibrated subject (S1).
Table 5.
Brier scores and two subcomponents for each of the four human subjects and the two simulated subjects
| BS | REL | RES | |
|---|---|---|---|
| S1 | 0.179 (0.018) | 0.082 (0.018) | 0.154 (0.018) |
| S2 | 0.193 (0.026) | 0.110 (0.025) | 0.167 (0.011) |
| S3 | 0.230 (0.014) | 0.097 (0.020) | 0.117 (0.019) |
| S4 | 0.170 (0.018) | 0.089 (0.015) | 0.169 (0.015) |
| Simulated S1 | 0.193 (0.021) | 0.123 (0.019) | 0.180 (0.014) |
| Simulated S4 | 0.173 (0.013) | 0.095 (0.015) | 0.172 (0.014) |
BS, Brier scores; REL, reliability; RES, resolution.
Simulations
To test CSD model confidence predictions, we also simulated tens of thousands of test sessions. To allow direct comparison to the human findings, analyses for the simulation data mimic those used for human studies. The first set of simulations was performed using average fit parameters from the well-calibrated subject (S1) — specifically, χ(x) = ϕ(x; μ = 0.13, kσ = 0.99) and Ψ(x) = ϕ(x; μ = 0.13, σ = 0.90) [i.e., the noise model is ε = N(μ = 0.13, σ = 0.90)]. These simulations are referred to as “Simulated S1.” Similarly, the second set of simulations was performed using average fit parameters from the most underconfident subject (S4), χ(x) = ϕ(x; μ = −0.02, kσ = 1.44) and Ψ(x) = ϕ(x; μ = −0.02, σ = 0.72), which we refer to as “Simulated S4.”
The simulated data (Fig. 7) mimicked the actual human data (Fig. 4). For the well-calibrated simulated subject, the Simulated S1 data (1st and 2nd column of Fig. 7) show that both CSD models yielded fit parameters that roughly matched those of the well-calibrated subject. While estimates of the width parameter () using the CSD2 and CSD3 models were roughly similar, the average estimated width parameter () using the CSD2 model was slightly greater than that estimated by the conventional binary fit (Fig. 7A), which converged to the correct value. On the other hand, the CSD3 model yielded an average width parameter estimate that matched that estimated via conventional binary fits (Fig. 7B), and both converged to the actual simulated value. The error bars for both CSD models were substantially less than those found for the conventional binary fit process.
Fig. 7.
CSD2 and CSD3 parameter fits for simulated subjects S1 and S4. First column shows fitted CSD2 parameters for simulated subject S1. Second column shows fitted CSD3 parameters for simulated subject S1. Third column shows fitted CSD2 parameters for simulated subject S4. Fourth column shows fitted CSD3 parameters for simulated subject S4. Top row (A–D) shows fitted psychometric width parameter (). Middle row (E and F) shows fitted confidence-scaling factor (). Bottom row (G–J) shows fitted psychometric function bias (). Thick black curves show average psychometric parameter estimates calculated using conventional forced-choice analyses. Thick red curves show average parameter estimates determined by fitting confidence probability judgment data. Errors bars (thin gray curves and thin red curves, respectively) represent standard deviation of parameter estimates.
The third and fourth columns of Fig. 7 show the fitted parameter estimates when the subject was underconfident (k = 2), which we refer to as Simulated S4. The average width parameter estimated using the CSD2 model was highly overestimated. The average width parameter estimated using the CSD3 model matched both the binary fit parameters and the actual simulated value. For all simulations, the precision of the width parameter that utilized the CSD model was smaller than the precision of the width parameter that utilized the conventional estimation. The estimates of the noise bias parameter of the psychometric functions () showed a qualitatively similar pattern; the estimates that utilized confidence reached stable levels a little sooner and were more precise than the estimates provided by the conventional analysis.
Figure 8 shows simulated confidence distributions for Simulated S1 and Simulated S4 using the same format as the human data (Fig. 5). For quantitative analyses, including χ2 analyses, we used 10,000 bootstrap samples. Consistent with human data, we see that the confidence distributions skew toward confidence judgments of zero for large negative stimuli (leftmost columns) and skew toward confidence judgments of 1 for positive stimuli (rightmost columns). For small (subthreshold) stimuli, the confidence judgments span the entire range. As for the human data, we matched the simulated confidence distributions with model predictions made via a separate set of simulations using the methods outlined in Fig. 2. For the simulated well-calibrated subject (S1) shown as the top row of Fig. 8, both models fit the data fairly well, but even for this subject, the CSD3 model (VAR = 0.0000079) fit the confidence distribution data better than the CSD2 model (VAR = 0.00012). For the simulated underconfident subject (S4), the CSD2 model did not match the simulated confidence distributions very well (VAR = 0.0053). For example, it showed an especially poor fit for the small stimuli (Fig. 8, bottom row, middle panel). In comparison, the CSD3 model fit the confidence distributions very well (VAR = 0.0000085). To examine the goodness of fit, χ2 tests were performed for the CSD2 and CSD3 models fit (Table 6). For the CSD3 model fit, the difference between the simulated results and the empiric data was never statistically significant. On the other hand, for the CSD2 model fit, the difference between empiric data and the model showed a statistically significant difference in at all stimulus levels.
Table 6.
Statistical significance of the χ2 tests
| Stimulus Level |
|||||
|---|---|---|---|---|---|
| −2 | −1 | 0 | 1 | 2 | |
| Simulated S1 | |||||
| CSD2 | <0.001 | <0.001 | <0.001 | <0.001 | <0.001 |
| CSD3 | 0.389 | 0.522 | 0.588 | 0.896 | 0.582 |
| Simulated S4 | |||||
| CSD2 | <0.001 | <0.001 | <0.001 | <0.001 | <0.001 |
| CSD3 | 0.331 | 0.544 | 0.222 | 0.497 | 0.635 |
Values are P values, each representing significance probability at each stimulus level between the simulated data and the model fits.
To take advantage of large numbers for illustration purposes, Fig. 9 combines the data across all 10,000 tests (120 trials each, yielding 1,200,000 total trials), so each of the 12 data points represents 100,000 trials. The top row plots mean confidence data, and the bottom row plots median confidence data. Figure 9, A and E plots conventional confidence calibration curves for the simulated subjects by plotting average confidence versus average binary performance. Figure 9, C and G plots average calibration versus stimulus level in standard psychometric function format. Figure 9A suggests that this simulated subject (S1) was underconfident, but we know that this simulated subject was well calibrated. A similar effect is observed when this same simulated data set is plotted in standard psychometric function format (Fig. 9C). The appearance of underconfidence disappears when the median confidence is plotted (Fig. 9, B and D).
For the simulated underconfident subject (Simulated S4), the mean data again suggest underconfidence (Fig. 9, E and G), but, for this subject, plotting the median confidence (Fig. 9, F and H) still suggests underconfidence. This should not come as a surprise as we know that this subject was underconfident with a confidence-scaling factor of nearly 2.
For comparison to the human analysis presented earlier, we also calculate the Brier score and several of its subcomponents. See Table 5. Consistent with actual human data, our simulated underconfident subject (Simulated S4) showed a slightly lower (i.e., better) Brier score than our simulated well-calibrated subject (Simulated S1). Simulated values appear about the same as those previously reported for our four human subjects.
DISCUSSION
In this pair of articles, we introduced a new CSD model that combines a confidence model with a standard signal detection noise model. While the first article (Yi and Merfeld 2016) focused on psychometric function fits, this article focuses on confidence analysis. The confidence distributions (Fig. 5) match those that would be predicted to arise if the CSD3 model were to provide a reasonable reflection of how humans estimate confidence while performing a vestibular self-motion direction-recognition task. More specifically, over 80% of the variance found in the empiric human confidence data is explained by our three-parameter CSD (i.e., CSD3) model.
Confidence Distributions
The primary findings reported herein are that the measured confidence distributions are consistent with those predicted by the CSD3 model that provides a confidence scaling factor to model underconfidence and/or overconfidence. As shown in Table 4, over 80% of the variance inherent in the empirical confidence histograms could be explained by the CSD3 model. This compares to just over 40% of the variance explained by mean confidence and 70% of the variance explained by the CSD2 model. Not surprisingly, the impact of the confidence scaling factor was particularly evident for the subject (S4) whose data showed the greatest deviation from perfect confidence calibration; for this subject, the CSD3 model explained 87% of the variance, while the CSD2 model explained 67% of the variance.
The impact of the CSD model was less for subject S3 than for the other three subjects. For this subject, the mean confidence density calculation explained 40% of the variations. In comparison, the CSD2 and CSD3 models explained just 45 and 55% of the variations. This is not surprising because subject S3 was forced to guess the most. As discussed elsewhere, for this subject, the vast majority of the applied stimuli were subthreshold, and our CSD model has little impact when the subject is just guessing (i.e., reports 50% confidence).
Confidence Calibration
For our simulations, we assumed that a single noisy decision variable yielded both a binary decision and a confidence rating. This directly coupled the confidence rating to the decision variable. Even with this tight coupling, we found that average confidence did not match average performance for a simulated well-calibrated subject (Fig. 9, A and C) when confidence was set to exactly match average performance for each sampled decision variable. At first glance, this may seem counterintuitive but can be explained as the direct result of distortions that arise during the nonlinear transformation of a decision variable to confidence.
Specifically, our simulated results showed that for a direction-recognition task, the predicted confidence distributions (e.g., Fig. 2) are not Gaussian even when the underlying signal detection noise model is Gaussian. In fact, except for one specific condition (stimulus plus mean noise near zero, Fig. 2, C and H), the predicted confidence distributions were not even symmetric about the mean. Since these confidence distributions are asymmetric, the median represents the distribution better than the mean. Consistent with this, simulated results showed that conventional confidence calibration analyses, which typically plot average confidence versus average binary task performance (e.g., Fig. 9A), will yield systematic biases that suggest that humans are underconfident (e.g., Fig. 9, A and C). Specifically, we showed that a standard confidence calibration plot (Fig. 9, A) for a direction-recognition task demonstrated underconfidence when, in fact, the simulated subject was well calibrated and behaving as a rational perfect signal detector (e.g., Fig. 9, A and C). When the same plot was made using median instead of mean, the apparent underconfidence disappeared.
This could — at least partially — explain reports that subjects performing perceptual discrimination tasks are often underconfident (e.g., Björkman et al. 1993; Festinger 1943; Juslin and Olsson 1997; Juslin et al. 1998; Stankov 1998; Stankov et al. 2012), but we do not discount that subjects might truly be underconfident for some tasks. In fact, our analysis showed that all four of our subjects appeared underconfident with one of the four (S4) substantially underconfident. We emphasize that this must be evaluated empirically for each specific task/experiment. As elaborated below, we further emphasize that such empirical studies should fully consider the full confidence data set — as we did 1) by fitting psychometric and confidence functions to the data and 2) by plotting confidence calibration plots using median confidence at each stimulus level — and not simply calculate average confidence. While specific solutions will depend on task specific details, we show that fully considering the confidence distributions is important. For the Gaussian noise example detailed herein, Fig. 9, B and D showed that plotting the median instead of the mean would yield a good representation of confidence calibration for a well-calibrated subject.
Why does a calibration plot using median confidence work? Individual confidence probability judgments result from a nonlinear transformation of each sample from the decision variable distribution, which we represent as Gaussian. When symmetric Gaussian noise is transformed by a monotonic nonlinear function (i.e., any sigmoidal cumulative distribution function), the resultant distribution is distorted (i.e., is no longer Gaussian). This was shown via the confidence distributions shown in the various panels of Fig. 2, which each show the resultant distribution at different stimulus levels. Nonetheless, the transformed median value of the original distribution will remain the median of the confidence distribution because the transformation is monotonic. To help put this in context, we remind the reader of two additional facts: 1) Average performance is represented by the psychometric function. 2) By definition, confidence should equal average performance (i.e., the psychometric function) for a well-calibrated subject. When considered together, these facts explain why median confidence matches average performance for our simulated well-calibrated subject.
Why does average confidence not match average performance even for such a well-calibrated subject? A short answer is that the nonlinear conversion of a decision variable (Fig. 1B) to confidence (Fig. 1D) distorts the distribution such that the average transformed distribution does not equal the transformed average value. In comparison, for any monotonic transformation, the median of the transformed distribution equals the transformed median value.
Confidence Studies
In this section, we will briefly describe some earlier approaches to understand confidence. [Those interested in historical coverage of the empirical studies of confidence will likely find the review by Lichtenstein and colleagues (1982) useful.] A number of models have suggested that confidence can be linked to signal detection with confidence directly related to the distance between the sampled decision variable and the decision criteria (sometimes called the decision boundary). For example, Ferrell and McGoey (1980) assumed that the decision variable modeled by signal detection theory for decision making was also used to in the confidence determination process and that confidence should increase with distance from a decision boundary. The paper compares model predictions with data from the literature and suggests that such a model might help explain both subjective overconfidence and underconfidence. Later a model having similar characteristics was developed and used to help explain the fact that subjects are often underconfident (Björkman et al. 1993); that this model matched portions of the earlier Ferrell and McGoey model was explicitly noted in a paper that expanded the potential use of such models beyond sensory discrimination to cognitive judgments (Ferrell 1995). A later study used confidence ratings to show that human decision making was best modeled a distance-from-criterion model (like those noted above) (Balakrishnan and Ratcliff 1996).
A mechanistic model of confidence was later created (Pleskac and Busemeyer 2010) by combining the distance concept from the previous paragraph with a standard drift diffusion model (e.g., Laming 1968; Ratcliff 1978) of decision making. The authors judged the model to provide solid predictions of choice accuracy, response time, and confidence, which together form the three pillars of empirical decision making. Predictions from a family of drift diffusion models similar to Pleskac’s were later compared with published response-time data (Ratcliff et al. 1994) to predict human behavior that previously had not been modeled (Ratcliff and Starns 2013).
While human confidence is the focus of this paper, the CSD3 model (and variants derived from it) may prove applicable for animals too. Recently, techniques have been developed that allow confidence to be studied in animals. For example, one study showed that the willingness of rats to wait for a reward increased with confidence as quantified via neuronal activity in their orbitofrontal cortex (Kepecs et al. 2008). Another recent study showed that monkey parietal cortex neurons might represent the degree of uncertainty on an individual trial (Kiani and Shadlen 2009).
For those interested in more information, several recent papers review and discuss confidence models and implications of such models (Fetsch et al. 2014; Grimaldi et al. 2015), present new computational and theoretical confidence analyses (Drugowitsch et al. 2014; Kepecs and Mainen 2012, 2014), and develop and test new methods to study confidence under the assumption of a signal detection model (e.g., Maniscalco and Lau 2012; Massoni et al. 2014).
Our CSD Assumptions
It is important to list the assumptions underlying the general CSD model as well as the basis for the general model; we then discuss the assumptions/bases that underlie the specific models developed herein and in the previous article.
General assumptions.
Following standard signal detection theory approaches described in the previous section, we assume that a noisy decision variable provides the basis for a decision when compared with the location of a stable consistent decision boundary; the psychometric function provides an empirical construct that characterizes the statistical features of this noisy decision variable given a stable decision boundary. Furthermore, like the majority of threshold studies that use signal detection approaches, we specifically assume an additive signal and noise model and further assume that the variance of the noise is independent of the signal — at least for the small signals (i.e., those having a magnitude no larger than several times the threshold) typically utilized to investigate thresholds. We also specifically assume that the decision boundary is set rationally (i.e., motion sensed as rightward is reported as rightward, etc.).
We further generally assume that the same decision variable used to make the decision provides the basis for a confidence rating. If the assumption were false, it would imply a degree of independence between confidence and the pertinent decision that has never (to our knowledge) been reported. Such issues have been identified as important (e.g., Ratcliff 2006). Studies that address such issues of confidence and decision variable distribution are certainly crucial to the development of our understanding of confidence and our ability to analyze confidence-rating data. The model presented herein may help address such questions.
The other general assumptions are 1) that there is a neural process that maps each decision variable into confidence and 2) that we can model this process mathematically using a general confidence function, χ(x). Just as the psychometric function is an experimental/theoretical construct that does not assume that such a function is directly accessible to the subject, we do not assume that the subject has access to a confidence function per se — just that there is a neural process that maps each decision variable into confidence and that we can model the result of this neural process. For the general model, no assumptions about this function (e.g., symmetric, nonlinear, etc.) are made, other than it exists. Generally speaking, we do not even need to assume that the confidence function is deterministic (i.e., not random), but this model would not be very useful if the neural processes underlying the mapping of a decision variable to a confidence rating were not, at least to a large extent, deterministic. While we focused on a specific model appropriate to a direction-recognition task, these general assumptions seem applicable to a large variety of other tasks.
Specific assumptions.
For our specific implementation, we always assumed Gaussian noise; this noise model can be changed to any appropriate noise model when warranted. The assumption of Gaussian noise for symmetric (i.e., left/right symmetric) tasks like our direction-recognition task seems reasonable. First, statistics’ central limit theorem (Larsen and Marx 1986; Lyapunov 1901) states that the sum of a large number of independent random variables will be approximately normally distributed regardless of the distribution of each of the random variables. Certainly, for processes as complex as perceptual decision making, it seems reasonable to consider that numerous independent noise sources — from stimulus noise to transduction to neural processing to decision making, including attention, and every synapse in between — might contribute.
Gaussian noise leads directly to the psychometric function being represented by a Gaussian cumulative distribution, Ψ(x) = ϕ(x), but this could be replaced by any CDF appropriate for the pertinent noise model, Ψ(x) = Fx(x). For most applications investigated herein, we assumed that the confidence function had the same form (i.e., were both Gaussian CDFs) as the psychometric function, Ψ(x) = ϕ(x; μ, σ) and χ(x) = ϕ(x; μ, kσ). This seems like a reasonable assumption, since both depend directly on the same noise distribution. On the other hand, it is not essential that the confidence and psychometric functions have the same form. In fact, in the previous article, we specifically simulated cases where the confidence function was linear while the psychometric function was a cumulative Gaussian and showed that the fitting algorithm still converged to reasonable parameter estimates.
We focus exclusively on a direction-recognition task throughout. We do not consider this a fundamental limitation; application of our CSD model should generalize to other tasks. The more similar the task to that investigated herein, the more likely that our results will directly apply. Results should be generally applicable to two-alternative forced-choice (2AFC) tasks, but since such psychometric functions typically range from 0.5 to 1 instead of from 0 to 1, specific findings may vary from those reported herein. Results should also generalize to detection (yes/no) tasks. Finally, the model should even generalize to tasks that have subjects categorize the responses into four (e.g., “guessing,” “uncertain,” “confident,” “certain”) or more categories, but at least one additional free parameter would be required for each additional category. Each additional free parameter would almost certainly have some impact on the efficiency of confidence fit methods. See appendix b for a little more discussion of this issue.
Signal Detection Theory and Confidence Ratings
We explicitly note that we are not the first to consider the application of signal detection noise models to confidence ratings. As discussed earlier, a number of earlier studies have considered signal detection theory to help interpret confidence data. Our primary contributions are 1) to recognize (and show) that confidence calibration needs to be considered when analyzing confidence data, 2) to provide a way to calculate/estimate confidence calibration via a fitting procedure (Yi and Merfeld 2016), and 3) to show that confidence calibration plots need to consider confidence response distributions.
Some of the earliest studies described earlier (e.g., Björkman et al. 1993; Ferrell 1995; Ferrell and McGoey 1980) applied such models and analyses to either the signal detection task (i.e., Is a stimulus present or not?) or to 2AFC tasks (i.e., In which interval/location did the stimulus appear?). A number of early studies used confidence ratings to generate receiver operating characteristic (ROC) curves (e.g., Decker and Pollack 1958; Gescheider et al. 1971; Green and Swets 1966; Macmillan and Creelman 2005; Peli et al. 1991; Pollack and Decker 1958; Swets et al. 1961; Watson et al. 1964). Since humans have limits to their ability to classify information (e.g., Miller 1956), many of these studies reasonably provided subjects with a limited number of confidence classification choices. For example, integer scales between 1 and 6 (Peli et al. 1991; Sawides et al. 2013) have been used to indicate confidence or subjects are provided an option to indicate when they are “uncertain” (e.g., García-Pérez and Alcalá-Quintana 2011; Hockley and Murdock 1987; Okamoto 2012; Vickers 1979). While such methods provide valuable information, our model cannot be direct applied to such data. To apply the methods described herein to such data, one would need to assume (or measure) how these classification ratings convert to confidence, which offers additional free parameters. More specifically, our methods only apply directly to confidence probability judgments that are provided as — or directly map to — a fraction (i.e., 0 to 1 or 0.5 to 1) or a percentage (i.e., 0 to 100% or 50 to 100%). (See appendix b for more discussion of this issue.)
Nonetheless, these earlier studies are highly pertinent as they provide evidence that subjects can report reasonably well-calibrated confidence ratings. One vibration study (Gescheider et al. 1971) reported that d′ (“d-prime”) estimates obtained from standard yes/no detection procedures and confidence ratings were not statistically different from one another. Another study (Pollack and Decker 1958) reported that performing a confidence rating did not interfere with task accuracy and that the confidence rating was related to average accuracy over a range of signal-to-noise ratios (Pollack and Decker 1958). This latter finding would be consistent with findings expected for well-calibrated observers. Another study showed that trained listeners can perform both rating tasks and a binary decision procedure (Egan et al. 1959) with little discrepancy in the results obtained via the two different methods (Green and Swets 1966). And a final study showed that confidence ratings convey pertinent information as such ratings substantially improved message reception (Decker and Pollack 1959).
Direction Recognition
To help simplify our analysis of confidence, we chose a simple direction-recognition task and model (Merfeld 2011) where the subject is provided a single motion — e.g., either rightward or leftward — on each trial and is asked to classify the motion as either rightward or leftward. Direction-recognition tasks utilize stimuli ranging from minus infinity to infinity. Motions, either visual or whole-body, are commonly used stimuli for direction recognition tasks, which are also referred to as direction discrimination tasks. Recognition tasks are not generally applicable to stimuli having amplitudes ranging from 0 to infinity (i.e., stimuli characterized by magnitude), like sound intensity, brightness of light, etc. [See the Introduction and Appendix B of Chaudhuri et al. (2013) for more detailed descriptions and definitions.] Such direction-recognition tasks are not uncommon (e.g., Benson et al. 1986, 1989; Britten et al. 1992; Crane 2012; Fetsch et al. 2009; Grabherr et al. 2008; Karmali et al. 2014; Lagacé-Nadon et al. 2009; MacNeilage et al. 2010; Roditi and Crane 2012; Swensson 1972; Valko et al. 2012). Before proceeding, we note that we are not saying that direction recognition is better than other tasks. We are simply noting that this task is simpler in certain respects than others; such simplicity is often pertinent when developing a new model as the tenets of the model can be tested more readily before generalizing to more complex applications.
As a specific comparison, this task is simpler in the sense that the decision boundary is not arbitrary — as it is for a yes/no detection task — unless the stimulus amplitude is known to the subject (Green and Swets 1966). Objectively, in the absence of biases (e.g., that the objective and subjective a priori probabilities for each motion direction and the costs for all decisions are equal), the decision boundary for a direction-recognition task should be placed at subjective “zero.” In other words, the subject should respond “right” if that trial yields a decision variable to the side of the decision boundary representing rightward motion and “left” if that trial yields a decision variable to the side the decision boundary representing leftward motion. This simplification means that there is reason to consider the decision boundary fixed (i.e., constant) for direction-recognition tasks — as opposed to yes/no tasks where variations in the decision boundary are often pertinent (e.g., Erev 1998; Kac 1969; Mueller and Weidemann 2008)
As a second comparison, this task is also simpler than traditional 2AFC methods where the subject is presented stimuli in one of two intervals (or one of two locations) and the subject must identify the interval with (or location of) the stimuli. Specifically, there is no relative comparison as required for two-interval forced-choice methods. Fits are also subtly different for direction recognition as compared with 2AFC. For our direction recognition task, the variability in the binary response goes to zero as the proportion of positive responses approaches one. And, as the stimulus gets more and more negative, the response variability again goes to zero — in this case as the proportion of positive responses goes to zero. (For an example, see Fig. 7d in Merfeld 2011). In contrast, for a traditional 2AFC task, the variability remains finite as the stimulus gets smaller. (For an example, see Fig. 7b in Merfeld 2011).
Taken together, under the assumption of a perfect signal detector, the various simplifications that accrue for a direction-recognition task lead to average accuracy being determinate in the sense that accuracy does not depend on the arbitrary placement of a decision boundary. This contrasts with yes/no detection tasks where accuracy depends on arbitrary placement of the decision boundary (e.g., Clarke et al. 1959; Galvin et al. 2003; Pollack 1959). This leads to the conclusion that, theoretically, there is no distinction between Type 1 confidence ratings and Type 2 confidence ratings for a symmetric unbiased direction recognition task, where Type 1 tasks distinguish events independent of the observer and Type 2 tasks evaluate correctness of an observer’s Type 1 decision. [For a detailed description of Type 1 versus Type 2 tasks, which may be particularly pertinent for tasks other than our unbiased direction-recognition task, see Galvin et al. (2003).] We emphasize that this does not imply or suggest that Type 1 and Type 2 task instructions cannot lead to different results for a direction-recognition (or any other) task as such experimental findings depend on both training and interpretation of the instructions.
Nonadaptive Sampling
In this paper, we utilized nonadaptive sampling so that we could obtain multiple trials at the exact same stimulus level. This is one advantage of nonadaptive sampling. However, our results accidentally highlighted one disadvantage, which is that the preselected stimuli may end up being far from optimal levels. This impacted the data collected for S3. Specifically, the original adaptive data collection estimated that the psychometric width parameter () for S3 was 0.28°/s. This was determined using the conventional binary fitting process. This led to the nonadaptive stimulus levels listed in Table 2 with a maximum stimulus magnitude of 0.59°/s. However, S3’s width parameter estimated following extensive nonadaptive data collection (3,600 trials) was 0.83°/s. Therefore, the maximal stimuli were actually below threshold (~70% of threshold) with all other stimuli well below threshold (<25% of threshold).
Table 2.
Fitted linear regression slopes
| μ | σ | κ | |
|---|---|---|---|
| S1 | 0.0035 (0.144) | −0.0047 (0.473) | 0.0033 (0.380) |
| S2 | 0.0007 (0.649) | 0.0011 (0.733) | 0.0065 (0.030) |
| S3 | 0.0030 (0.416) | 0.0265 (0.037) | 0.0038 (0.800) |
| S4 | −0.0016 (0.689) | −0.0096 (0.209) | −0.0010 (0.878) |
Each fit parameter was analyzed through the sequence of 30 tests for each subject (S1 through S4). P values are given in parentheses.
By accident, this highlights an advantage of our confidence modeling and fitting algorithm. After 75 trials, the width parameter found via the CSD3 fit was 0.44°/s. If we had chosen to design the nonadaptive test protocol using the width parameter estimated by the CSD3 fit, we would have increased all stimuli by 60%. While this would still have yielded sampling at somewhat lower stimulus levels than desired, this sampling would have yielded stimuli much closer to the desired levels.
Summary
In summary, our data show confidence reports whose distribution is consistent with that predicted by our CSD3 model. In fact, on average across our four subjects, over 80% of the variations found in the confidence histogram density values could be explained by the CSD3 model. In addition, the CSD3 fits found that three of our four subjects (S2, S3, and S4) were, on average, underconfident with the remaining subject (S1) fairly well calibrated (Fig. 4); this finding was confirmed by confidence calibration plots of the same data when the median confidence data was plotted versus average performance (Figs. 7, E–H).
GRANTS
This research was supported by a MedEl contract with Massachusetts Eye and Ear Infirmary as well as by three different NIH/National Institute of Deafness and Other Communications Disorders grants (R01-DC04158, R01-DC014924, and R56-DC12038).
DISCLOSURES
No conflicts of interest, financial or otherwise, are declared by the authors.
AUTHOR CONTRIBUTIONS
Y.Y., W.W., and D.M.M. conceived and designed research; Y.Y. performed experiments; Y.Y., W.W., and D.M.M. analyzed data; Y.Y., W.W., and D.M.M. interpreted results of experiments; Y.Y. and D.M.M. prepared figures; Y.Y., W.W., and D.M.M. drafted manuscript; Y.Y., W.W., and D.M.M. edited and revised manuscript; Y.Y., W.W., and D.M.M. approved final version of manuscript.
ACKNOWLEDGMENTS
We appreciate the participation of our anonymous subjects. We thank Sho Chaudhuri, Torin Clark, Raquel Galvan-Garza, Faisal Karmali, and Koeun Lim for helpful comments on earlier manuscript drafts. We also thank Bob Grimes and Wangsong Gong for technical support.
APPENDIX A: GENERALIZING OUR DIRECTION-RECOGNITION CSD MODEL
To keep the presentation focused on the confidence aspects of this CSD model, we introduced our model without explicitly defining a psychophysical function that maps the physical space (i.e., stimuli) onto the perceptual space. We were able to do this because the psychophysical function for our direction recognition task can be represented as an identity function, ψ(x) = x, where x represents the stimulus. For a given stimulus (x =sj) — under the assumption of additive Gaussian noise, ε = N(μ,σ)—this yields the following probability distribution (e.g., Fig. 1B):
| (A1) |
Via the process outlined in Fig. 1, this defines the psychometric function shown earlier:
| (A2) |
[Before proceeding, to minimize potential confusion due to the somewhat similar standard names and standard functional symbols, we explicitly note that notation for the psychometric function Ψ(x) subtly differs from the psychophysical function that we introduced above, ψ(x).]
To generalize, a psychophysical function other than an identity function can be defined, ψ(x) = g(x), where g(x) could in principle be any deterministic function of the stimulus that maps perception of the physical stimulus. Under the same additive Gaussian noise model, this yields the following probability distribution:
| (A3) |
Since human psychophysical studies provide no way to distinguish g(0) from μ, we will redefine our psychometric function as ψ(x) = μ(x) = g(x)+μ and represent the noise as zero-mean, ε = N(0,σ). We represent the psychophysical function as μ(x) because this nomenclature is consistent with published literature (e.g., García-Pérez and Alcalá-Quintana 2013) and as a reminder that it includes any noise bias that may be present. This substitution yields the following probability distribution and psychometric function:
| (A4) |
| (A5) |
Before proceeding, it is pertinent to note that the value of the psychophysical function at zero [μ(0)] includes three overlapping impacts that are empirically difficult (if not impossible under most circumstances) to isolate from one another: 1) shifts of the subjective zero, known only to the subject, relative to objective zero, known only to the operator, 2) any bias in the physiologic noise, and 3) the value of the psychophysical function when no stimulus is provided. This issue is typically handled by assuming that two of the three forementioned impacts are zero and lumping the cumulative impact of the three effects as the remaining (third) effect. In the main body, we chose to lump all three effects into the noise bias term, while Eqs. A4 and A5 lump these effects into the value of the psychometric function, μ(x), at x = 0.
While the identity function is often assumed to map physical space onto psychological space, the units of physical space must differ from the units of perceptual space. This, at a minimum, requires a scaling factor (n′). To elaborate, we assume a psychophysical function of the form: ψ(x) = n′x+x0 with additive Gaussian noise, ε′ = N(μ′,σ′). Substituting, and solving for the probability distribution for a specific stimulus, sj, yields the following probability distribution:
| (A6) |
Letting y = y′/n′, σ = σ′/ n′, and μ = (x0+μ′)/n′, we obtain:
| (A7) |
which is identical to Eq. A1. There is no way to identify both the scaling factor, n′, and the noise standard deviation, σ′. This issue is widely known and is typically handled in one of two ways: 1) Simply estimate the noise standard deviation, σ = σ′/n′, which represents the ratio of noise to scaling factor in physical units (i.e., ˚/s in this paper). 2) Assume that the noise variance is unity (σ′ = 1) and estimate the scaling factor, n′ = σ′/σ = 1/σ.
Each approach is commonly used and each has advantages for certain application. In the previous paragraph, we showed these approaches to be mathematically equivalent. We choose the first approach (e.g., Fig. 1) here because of its relative simplicity for our direction recognition application.
APPENDIX B: CONFIDENCE RATING SCALES
Subjective indications of confidence have typically been called confidence ratings (e.g., Björkman et al. 1993; Decker and Pollack 1958; Galvin et al. 2003; Gescheider 1985; Green and Swets 1966; Macmillan and Creelman 2005; Peli et al. 1991; Pollack and Decker 1958; Stankov et al. 2012; Watson et al. 1964). Since our CSD model cannot directly utilize all types of confidence ratings, we call the subset of confidence ratings that allow direct application of our quantitative CSD model “direct quantitative confidence scales,” where we define a direct quantitative confidence scale as any confidence rating that can be directly mapped to the range 0 through 1 (or a subset of this range like 0.5 through 1, or a percentage 0 through 100%, etc.) without offering any free parameters (i.e., degrees of freedom) to the investigator when analyzing/fitting the data. For example, the probability judgments we focus on herein provide a standard direct quantitative confidence scale. On the other hand, ratings that offer the operator the freedom to change how the confidence rating is mapped to the confidence range for any reason (e.g., different mapping for different subjects, different mapping for different experimental conditions, etc.) would not provide a direct quantitative confidence scale. To help make this distinction clear, we illustrate using examples.
Example 1: Three buttons labeled 1, 2, and 3 without explicit quantitative instructions would not provide a direct confidence scale as both the subject and the operator would have to decide how to map these buttons to the range 0 through 100% (or 50 through 100%). As one mapping, the operator might assume that buttons 1, 2, and 3 represent 0 to 33%, 34 through 66%, and 67 through 100%, respectively. Depending on specific instructions provided to the subject, this may not be an unreasonable assumption but such post hoc fitting assumptions offer additional degrees of freedom when fitting the data. (To generalize this example, we note that N buttons without the subject knowing the defined mapping would provide at least N−1 additional free parameters when fitting the data with a CSD model.)
Example 2: The same three buttons labeled 1, 2, and 3 could provide a direct confidence scale if the operator explicitly trained every subject before data collection commenced to use buttons 1, 2, and 3 to indicate confidence between 0 through 33%, 34 through 66%, and 67 through 100%, respectively.
Example 3: The same three buttons could provide a direct confidence scale if the buttons were labeled “0 through 33%,” “34 through 66%,” and “67 through 100%” and subjects were trained to use the buttons to indicate confidence probability judgments.
Example 4: Three buttons, with the bottom button labeled “guessing” and the top button labeled “certain” would not provide a direct confidence scale as the operator would need to assume a quantitative mapping of these buttons to scale the acquired data.
Example 5: A mechanical (or virtual) slider would provide a direct confidence scale if the anchors provided were directly convertible to a confidence range of 50 through 100% (or 0 through 100%). For example, if anchors at the two extremes were clearly labeled 50 and 100%, with 75% midway between the two anchors, the mechanical (or virtual) measurement could be directly converted to a percentage between 50 and 100%.
Example 6: A mechanical (or virtual) slider would not provide a direct confidence scale if the anchors provided to the subject were numbers like 1, 2, and 3 without further instruction.
Example 7: A mechanical (or virtual) slider would not provide a direct confidence scale if the anchors provided to the subject were words like “guessing,” “confident,” “certain,” or “uncertain.”
We emphasize that while examples 2, 3, and 5 meet our definition of a direct quantitative confidence scale, it is not clear which of these provide good quantitative confidence scales; quality of the quantitative confidence scale is a separate matter. The quality of any confidence scale (like the quality of any confidence rating, more generally) would need to be determined empirically for the task/modality being investigated. Furthermore, we explicitly note that indirect confidence scales can legitimately be mapped onto our CSD model, but we highlight the difference between direct and indirect confidence scales because indirect confidence scales offer additional free parameters to be fit to data.
While resolution is not the characteristic that defines a direct quantitative confidence scale, we emphasize that resolution of confidence scaling is important for practical applications. For example, one would not generally want less measurement resolution (e.g., 40% resolution) than the subject’s underlying confidence information resolution (e.g., better than 20% resolution for at least some subjects). In this sense, resolution of a confidence scale is somewhat analogous to the resolution of an analog to digital converter. Having too little measurement resolution discards information. On the other hand, having too much measurement resolution can detract, if a high-resolution scale distracts, confuses, or overworks a subject. We illustrate using two examples. Asking an individual to press 1 of 25 (or more) buttons would be a challenging task, which is why such psychophysical tasks are seldom used. On the other hand, asking a subject to use a mechanical (or virtual) slider that provides 25 (or more) levels (e.g., Watson et al. 1964) has been successfully used.
To minimize occurrence(s), we briefly highlight a potential misunderstanding. Analogous to digitally sampled signals, there is a danger than an inexperienced operator (e.g., a graduate student) could confuse measurement resolution with the resolution of the measured subjective data, but this is a training issue that is a fundamental topic in most signal processing (e.g., Oppenheim and Schafer 1975) classes and is not a fundamental measurement limitation per se. The bottom line is that we do not assume that subject’s provide responses having the same 1% resolution as the scale we provide; we intentionally provide a scale having higher resolution than the subject’s responses to minimize impact of the chosen scale resolution.
We cannot know each individual subject’s resolution in advance, so we do not assume herein that humans necessarily provide confidence reports with any specific resolution. Probably most important is 1) that the scale be relatively easy for subjects to understand and master and 2) that it provide resolution high enough to capture any subjective categorization performed by the subject.
We emphasize that we do not think that subjects performing perceptual threshold tests are typically capable of providing confidence probability judgments having 1% resolution. When the measurement resolution is 1%, it may make sense to emphasize that we do not necessarily expect subjects to be able to report their confidence with 1% resolution, but this is likely a matter that depends on the task being performed as well as the subject’s background, expertise, and experience.
Footnotes
Note that the psychophysical function, which is the function that maps physical stimuli to a perceptual variable, is represented by an identity function in this psychometric function. This seems a reasonable assumption for our self-motion direction recognition task, since to first order vestibular responses are well represented as linear. See appendix a for an extensive discussion of this issue including equations that show how to include nonidentity psychophysical functions when desired.
REFERENCES
- Balakrishnan JD, Ratcliff R. Testing models of decision making using confidence ratings in classification. J Exp Psychol Hum Percept Perform 22: 615–633, 1996. doi: 10.1037/0096-1523.22.3.615. [DOI] [PubMed] [Google Scholar]
- Benson AJ, Hutt EC, Brown SF. Thresholds for the perception of whole body angular movement about a vertical axis. Aviat Space Environ Med 60: 205–213, 1989. [PubMed] [Google Scholar]
- Benson AJ, Spencer MB, Stott JR. Thresholds for the detection of the direction of whole-body, linear movement in the horizontal plane. Aviat Space Environ Med 57: 1088–1096, 1986. [PubMed] [Google Scholar]
- Björkman M, Juslin P, Winman A. Realism of confidence in sensory discrimination: the underconfidence phenomenon. Percept Psychophys 54: 75–81, 1993. doi: 10.3758/BF03206939. [DOI] [PubMed] [Google Scholar]
- Brier GW. Verification of forecasts expressed in terms of probability. Mon Weather Rev 78: 1–3, 1950. doi: 10.1175/1520-0493(1950)078<0001:VOFEIT>2.0.CO;2. [DOI] [Google Scholar]
- Britten KH, Shadlen MN, Newsome WT, Movshon JA. The analysis of visual motion: a comparison of neuronal and psychophysical performance. J Neurosci 12: 4745–4765, 1992. doi: 10.1523/JNEUROSCI.12-12-04745.1992. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chaudhuri SE, Karmali F, Merfeld DM. Whole body motion-detection tasks can yield much lower thresholds than direction-recognition tasks: implications for the role of vibration. J Neurophysiol 110: 2764–2772, 2013. doi: 10.1152/jn.00091.2013. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chaudhuri SE, Merfeld DM. Signal detection theory and vestibular perception: III. Estimating unbiased fit parameters for psychometric functions. Exp Brain Res 225: 133–146, 2013. doi: 10.1007/s00221-012-3354-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Clarke FR, Birdsall TG, Tanner WP Jr. Two types of ROC curves and definitions of parameters. J Acoust Soc Am 31: 629–630, 1959. doi: 10.1121/1.1907764. [DOI] [Google Scholar]
- Crane BT. Fore-aft translation aftereffects. Exp Brain Res 219: 477–487, 2012. doi: 10.1007/s00221-012-3105-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Decker L, Pollack I. Confidence ratings and message reception for filtered speech. J Acoust Soc Am 30: 432–434, 1958. doi: 10.1121/1.1909638. [DOI] [Google Scholar]
- Decker LR, Pollack I. Multiple observers, message reception, and rating scales. J Acoust Soc Am 31: 1327–1328, 1959. doi: 10.1121/1.1907628. [DOI] [Google Scholar]
- Drugowitsch J, Moreno-Bote R, Pouget A. Relation between belief and performance in perceptual decision making. PLoS One 9: e96511, 2014. doi: 10.1371/journal.pone.0096511. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Egan J, Schulman AI, Greenberg GZ. Operating characteristics determined by binary decisions and by ratings. J Acoust Soc Am 31: 768–773, 1959. doi: 10.1121/1.1907783. [DOI] [Google Scholar]
- Erev I. Signal detection by human observers: a cutoff reinforcement learning model of categorization decisions under uncertainty. Psychol Rev 105: 280–298, 1998. doi: 10.1037/0033-295X.105.2.280. [DOI] [PubMed] [Google Scholar]
- Ferrell WR. A model for realism of confidence judgments: implications for underconfidence in sensory discrimination. Percept Psychophys 57: 246–254, 1995. doi: 10.3758/BF03206511. [DOI] [PubMed] [Google Scholar]
- Ferrell WR, McGoey PJ. A model of calibration for subjective probabilities. Organ Behav Hum Perform 26: 32–53, 1980. doi: 10.1016/0030-5073(80)90045-8. [DOI] [Google Scholar]
- Festinger L. Studies in decision: I. Decision-time, relative frequency of judgment and subjective confidence as related to physical stimulus difference. J Exp Psychol 32: 291–306, 1943. doi: 10.1037/h0056685. [DOI] [Google Scholar]
- Fetsch CR, Kiani R, Shadlen MN. Predicting the accuracy of a decision: a neural mechanism of confidence. Cold Spring Harb Symp Quant Biol 79: 185–197, 2014. doi: 10.1101/sqb.2014.79.024893. [DOI] [PubMed] [Google Scholar]
- Fetsch CR, Turner AH, DeAngelis GC, Angelaki DE. Dynamic reweighting of visual and vestibular cues during self-motion perception. J Neurosci 29: 15601–15612, 2009. doi: 10.1523/JNEUROSCI.2574-09.2009. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Galvin SJ, Podd JV, Drga V, Whitmore J. Type 2 tasks in the theory of signal detectability: discrimination between correct and incorrect decisions. Psychon Bull Rev 10: 843–876, 2003. doi: 10.3758/BF03196546. [DOI] [PubMed] [Google Scholar]
- García-Pérez MA, Alcalá-Quintana R. Interval bias in 2AFC detection tasks: sorting out the artifacts. Atten Percept Psychophys 73: 2332–2352, 2011. doi: 10.3758/s13414-011-0167-x. [DOI] [PubMed] [Google Scholar]
- García-Pérez MA, Alcalá-Quintana R. Shifts of the psychometric function: distinguishing bias from perceptual effects. Q J Exp Psychol (Hove) 66: 319–337, 2013. doi: 10.1080/17470218.2012.708761. [DOI] [PubMed] [Google Scholar]
- Gescheider GA. Psychophysics: Method, Theory, and Application. Hillsdale, NJ: Erlbaum, 1985. [Google Scholar]
- Gescheider GA, Wright JH, Polak JW. Detection of vibrotactile signals differing in probability of occurrence. J Psychol 78: 253–260, 1971. doi: 10.1080/00223980.1971.9916910. [DOI] [PubMed] [Google Scholar]
- Grabherr L, Nicoucar K, Mast FW, Merfeld DM. Vestibular thresholds for yaw rotation about an earth-vertical axis as a function of frequency. Exp Brain Res 186: 677–681, 2008. doi: 10.1007/s00221-008-1350-8. [DOI] [PubMed] [Google Scholar]
- Green D, Swets J. Signal Detection Theory and Psychophysics. New York: Wiley, 1966. [Google Scholar]
- Grimaldi P, Lau H, Basso MA. There are things that we know that we know, and there are things that we do not know we do not know: confidence in decision-making. Neurosci Biobehav Rev 55: 88–97, 2015. doi: 10.1016/j.neubiorev.2015.04.006. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hockley WE, Murdock BB. A decision model for accuracy and response latency in recognition memory. Psychol Rev 94: 341–358, 1987. doi: 10.1037/0033-295X.94.3.341. [DOI] [Google Scholar]
- Juslin P, Olsson H. Thurstonian and Brunswikian origins of uncertainty in judgment: a sampling model of confidence in sensory discrimination. Psychol Rev 104: 344–366, 1997. doi: 10.1037/0033-295X.104.2.344. [DOI] [PubMed] [Google Scholar]
- Juslin P, Olsson H, Winman A. The calibration issue: theoretical comments on Suantak, Bolger, and Ferrell (1996). Organ Behav Hum Decis Process 73: 3–26, 1998. doi: 10.1006/obhd.1998.2749. [DOI] [PubMed] [Google Scholar]
- Kac M. Some mathematical models in science. Science 166: 695–699, 1969. doi: 10.1126/science.166.3906.695. [DOI] [PubMed] [Google Scholar]
- Karmali F, Lim K, Merfeld DM. Visual and vestibular perceptual thresholds each demonstrate better precision at specific frequencies and also exhibit optimal integration. J Neurophysiol 111: 2393–2403, 2014. doi: 10.1152/jn.00332.2013. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kepecs A, Mainen ZF. A computational framework for the study of confidence in humans and animals. Philos Trans R Soc Lond B Biol Sci 367: 1322–1337, 2012. doi: 10.1098/rstb.2012.0037. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kepecs A, Mainen ZF. A computational framework for the study of confidence across species. In: The Cognitive Neuroscience of Metacognition, edited by Fleming S and Frith C. Heidelberg: Springer, 2014, p. 115–145. [Google Scholar]
- Kepecs A, Uchida N, Zariwala HA, Mainen ZF. Neural correlates, computation and behavioural impact of decision confidence. Nature 455: 227–231, 2008. doi: 10.1038/nature07200. [DOI] [PubMed] [Google Scholar]
- Kiani R, Shadlen MN. Representation of confidence associated with a decision by neurons in the parietal cortex. Science 324: 759–764, 2009. doi: 10.1126/science.1169405. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lagacé-Nadon S, Allard R, Faubert J. Exploring the spatiotemporal properties of fractal rotation perception. J Vis 9: 3, 2009. doi: 10.1167/9.7.3. [DOI] [PubMed] [Google Scholar]
- Laming D. Information Theory of Choice Reaction Time. New York: Wiley, 1968. [Google Scholar]
- Larsen RJ, Marx ML. An Introduction to Mathematical Statistics and Its Applications. Englewood Cliffs, NJ: Prentice-Hall, 1986. [Google Scholar]
- Lichtenstein S, Fischhoff B, Phillips L. Calibration of probabilities: the state of the art to 1980. In: Judgement Under Uncertainty: Heuristics and Biases, edited by Kahneman D, Slovic P, Tverski A. New York: Cambridge University Press, 1982. [Google Scholar]
- Lyapunov A. Nouvelle forme du théorème sur la limite de probabilité. Memoires de l’Acad de St-Pétersbourg 12: 1–24, 1901. [Google Scholar]
- Macmillan NA, Creelman CD. Detection Theory: A User’s Guide. Mahwah, NJ: Erlbaum, 2005. [Google Scholar]
- MacNeilage PR, Banks MS, DeAngelis GC, Angelaki DE. Vestibular heading discrimination and sensitivity to linear acceleration in head and world coordinates. J Neurosci 30: 9084–9094, 2010. doi: 10.1523/JNEUROSCI.1304-10.2010. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Maniscalco B, Lau H. A signal detection theoretic approach for estimating metacognitive sensitivity from confidence ratings. Conscious Cogn 21: 422–430, 2012. doi: 10.1016/j.concog.2011.09.021. [DOI] [PubMed] [Google Scholar]
- Massoni S, Gajdos T, Vergnaud JC. Confidence measurement in the light of signal detection theory. Front Psychol 5: 1455, 2014. doi: 10.3389/fpsyg.2014.01455. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Merfeld DM. Signal detection theory and vestibular thresholds: I. Basic theory and practical considerations. Exp Brain Res 210: 389–405, 2011. doi: 10.1007/s00221-011-2557-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Miller GA. The magical number seven plus or minus two: some limits on our capacity for processing information. Psychol Rev 63: 81–97, 1956. doi: 10.1037/h0043158. [DOI] [PubMed] [Google Scholar]
- Mueller ST, Weidemann CT. Decision noise: an explanation for observed violations of signal detection theory. Psychon Bull Rev 15: 465–494, 2008. doi: 10.3758/PBR.15.3.465. [DOI] [PubMed] [Google Scholar]
- Murphy AH. A new vector partition of the probability score. J Appl Meteorol 12: 595–600, 1973. doi: 10.1175/1520-0450(1973)012<0595:ANVPOT>2.0.CO;2. [DOI] [Google Scholar]
- Okamoto Y. An experimental analysis of psychometric functions in a threshold discrimination task with four response categories. Jpn Psychol Res 54: 368–377, 2012. doi: 10.1111/j.1468-5884.2012.00513.x. [DOI] [Google Scholar]
- Oppenheim A, Schafer R. Digital Signal Processing. Englewood Cliffs, NJ: Prentice-Hall, 1975. [Google Scholar]
- Peli E, Goldstein RB, Young GM, Trempe CL, Buzney SM. Image enhancement for the visually impaired. Simulations and experimental results. Invest Ophthalmol Vis Sci 32: 2337–2350, 1991. [PubMed] [Google Scholar]
- Pleskac TJ, Busemeyer JR. Two-stage dynamic signal detection: a theory of choice, decision time, and confidence. Psychol Rev 117: 864–901, 2010. doi: 10.1037/a0019737. [DOI] [PubMed] [Google Scholar]
- Pollack I. On indices of signal and response discriminability. J Acoust Soc Am 31: 1031, 1959. doi: 10.1121/1.1907802. [DOI] [Google Scholar]
- Pollack I, Decker LR. Confidence ratings, message reception, and the receiver operating characteristic. J Acoust Soc Am 30: 286–292, 1958. doi: 10.1121/1.1909571. [DOI] [Google Scholar]
- Ratcliff R. A theory of memory retrieval. Psychol Rev 85: 59–108, 1978. doi: 10.1037/0033-295X.85.2.59. [DOI] [Google Scholar]
- Ratcliff R. Modeling response signal and response time data. Cognit Psychol 53: 195–237, 2006. doi: 10.1016/j.cogpsych.2005.10.002. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ratcliff R, McKoon G, Tindall M. Empirical generality of data from recognition memory receiver-operating characteristic functions and implications for the global memory models. J Exp Psychol Learn Mem Cogn 20: 763–785, 1994. doi: 10.1037/0278-7393.20.4.763. [DOI] [PubMed] [Google Scholar]
- Ratcliff R, Starns JJ. Modeling confidence judgments, response times, and multiple choices in decision making: recognition memory and motion discrimination. Psychol Rev 120: 697–719, 2013. doi: 10.1037/a0033152. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Roditi RE, Crane BT. Directional asymmetries and age effects in human self-motion perception. J Assoc Res Otolaryngol 13: 381–401, 2012. doi: 10.1007/s10162-012-0318-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rosner B, Wang W, Eliassen H, Hibert E. Comparison of dependent Pearson and Spearman correlation coefficients with and without correction for measurement error. J Biom Biostat 6: 226, 2015. doi: 10.4172/2155-6180.1000226. [DOI] [Google Scholar]
- Sawides L, Dorronsoro C, Haun AM, Peli E, Marcos S. Using pattern classification to measure adaptation to the orientation of high order aberrations. PLoS One 8: e70856, 2013. doi: 10.1371/journal.pone.0070856. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Stankov L. Calibration curves, scatterplots and the distinction between general knowledge and perceptual tasks. Learn Individ Differ 10: 29–50, 1998. doi: 10.1016/S1041-6080(99)80141-1. [DOI] [Google Scholar]
- Stankov L, Pallier G, Danthiir V, Morony S. Perceptual underconfidence: a conceptual illusion? Eur J Psychol Assess 28: 190–200, 2012. doi: 10.1027/1015-5759/a000126. [DOI] [Google Scholar]
- Swensson RG. The elusive tradeoff: Speed vs accuracy in visual discrimination tasks. Percept Psychophys 12: 16–32, 1972. doi: 10.3758/BF03212837. [DOI] [Google Scholar]
- Swets J, Tanner WP Jr, Birdsall TG. Decision processes in perception. Psychol Rev 68: 301–340, 1961. doi: 10.1037/h0040547. [DOI] [PubMed] [Google Scholar]
- Valko Y, Lewis RF, Priesol AJ, Merfeld DM. Vestibular labyrinth contributions to human whole-body motion discrimination. J Neurosci 32: 13537–13542, 2012. doi: 10.1523/JNEUROSCI.2157-12.2012. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Vickers D. Decision Processes in Visual Perception. New York: Academic, 1979. [Google Scholar]
- Vickers D. Where does the balance of evidence lie with respect to confidence? In: Fechner Day 2001: Proceedings of the Seventeenth Annual Meeting of the International Society of Psychophysics, edited by Sommerfeld E, Kompass R, Lachmann T. Lengerich, Germany: Pabst Science, 2001, p. 148–153. [Google Scholar]
- Watson CS, Rilling ME, Bourbon WT. Receiver‐operating characteristics determined by a mechanical analog to the rating scale. J Acoust Soc Am 36: 283–288, 1964. doi: 10.1121/1.1918947. [DOI] [Google Scholar]
- Yi Y, Merfeld DM. A quantitative confidence signal detection model: 1. Fitting psychometric functions. J Neurophysiol 115: 1932–1945, 2016. doi: 10.1152/jn.00318.2015. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yu S, Pleskac TJ, Zeigenfuse MD. Dynamics of postdecisional processing of confidence. J Exp Psychol Gen 144: 489–510, 2015. doi: 10.1037/xge0000062. [DOI] [PubMed] [Google Scholar]









