Abstract
This study represents the first network meta-analysis (NMA) comparing the efficacy and safety of different concentrations of carbamide peroxide (CP) for at-home bleaching in permanent dentition. A comprehensive search was conducted across PubMed, Cochrane Central, LILACS/BBO, SCOPUS, Web of Science, EMBASE, and gray literature. Randomized controlled trials (RCTs) comparing at least two CP concentrations were included. Color change (ΔE, ΔSGU), risk and intensity of tooth sensitivity TS was assessed Cochrane RoB 2.0 and GRADE were used to evaluate the risk of bias (RoB) and certainty of evidence. Thirteen RCTs were included, most with high RoB. CP5 % and 37 % appeared in a single trial each. For ΔSGU, no significant differences were detected between concentrations. For ΔE, all concentrations were superior to CP 5 %, with the largest difference observed for CP 37 %; intermediate concentrations showed minimal variations. Regarding TS, CP10 % was associated with a 67 % lower risk compared to CP 20–22 % (RR 1.67; 95 % CrI: 1.15–2.57), and CP 20–22 % also caused significantly higher TS intensity on the NRS scale. Conclusion: Lower concentrations of CP reduced the risk of TS while achieving bleaching efficacy comparable to higher concentrations, suggesting that less concentrated gels is safer and effective for at-home bleaching.
Keywords: Tooth bleaching, Tooth Bleaching Agents, Dentin sensitivity, Systematic review, Network meta-analysis
1. Introduction
In recent years, there has been a growing focus on achieving an attractive smile, with tooth color is a primary concern for many individuals [1]. Dissatisfaction with tooth color is a common reason individuals seek cosmetic dental treatments [2], [3], [4], [5], contributing to the increased popularity of dental bleaching treatments. Among the available protocols, dentist-supervised at-home protocol is one of the most recommended protocols, with carbamide peroxide (CP) being the most frequently used bleaching agent.
The efficacy of at-home bleaching with 10 % CP is well documented in the literature [6], [7], [8], [9]. However, nowadays manufacturers produce CP gels in varied concentrations ranging from as low as 5 % to as high as 22 % making product selection challenging for clinicians who search for the most effective option with minimal tooth sensitivity (TS) and gingival irritation (GI).
Clinicians might intuitively expect that higher concentrations yield faster and more noticeable results but at the cost of increased side effects. Conversely, lower concentrations are assumed to be safer but may require longer treatment periods to achieve the desired outcome. This trade-off between bleaching speed and side effects should not rely solely on intuition or common sense but rather on the best available evidence to guide clinical recommendations.
The literature on the efficacy and safety of different CP concentrations for at-home bleaching remains inconclusive. Some studies report that higher CP concentrations result in faster whitening [10], [11], [12], while others find no significant difference between 10 % CP and higher concentration products [6], [13]. Similarly, findings on TS vary, with some studies reporting comparable TS levels across different CP concentrations [6], [10], [13], [14], [15], while others observe increased TS with higher CP concentrations [9], [16], [17], [18].
These discrepancies in the outcomes may be due to factors related to study bias, imprecision, publication bias, and variations in follow-up periods after treatment. A comprehensive approach to resolving these inconsistencies is through a systematic review and network meta-analysis (NMA). A systematic review can compile and critically appraise existing studies, while a network meta-analysis allows for the comparison of multiple treatments even in the absence of direct head-to-head trials [19]. By synthesizing the available evidence, this approach would provide a clearer understanding of which concentrations offer the best balance of efficacy and safety, assisting clinicians in making informed decisions for at-home dental bleaching. Therefore, this current systematic review aims to answer the following focused research question: Is there any difference in bleaching efficacy and the risk of TS and GI across different CP concentrations for at-home bleaching in adults?
2. Material and methods
2.1. Protocol and registration
The study protocol was registered in the International Prospective Register of Systematic Reviews (PROSPERO) under identify number CRD42023455435, and followed the recommendations of the Preferred Reporting Items for Systematic Reviews (PRISMA-NMA) guidelines with extension for NMA [20].
2.2. Eligibility criteria
We included published and unpublished parallel, and split-mouth randomized controlled trials (RCTs) that compared any two at-home bleaching protocols with different CP concentrations. RCTs were excluded if they compared (1) bleaching products that did not employ any physical barrier to keep the product in intimate contact with the teeth, such as paint-on products, rinses, dentifrices, and chewing gums, (2) compared the same bleaching product with very close concentrations or placebo 3) bleaching product from different brands with the same concentration, (3) combined in-office/at-home bleaching treatment, (4) peroxide-free tooth bleaching agents, and (5) bleaching agents with an overall treatment time lower than two weeks. No restrictions were placed regarding language or publication date.
2.3. Definition of intervention
Based on our preliminary search and clinical judgment, we considered lumping products with very close concentrations as a single node in the network meta-analysis before study initiation. Thus, the bleaching products were categorized into (1) CP 5 %, (2) CP 10 %, (3) CP 15–16 %, (4) CP 20–22 %, (6) CP 37 %. While lumping interventions has been criticized in the literature [21], [22], [23], this approach makes it more possible to obtain a connected and denser network, facilitating statistical analysis. Given the minimal differences in concentrations among the grouped products, lumping would maximize similarity within the corresponding nodes and minimize it across the nodes [24].
Additionally, we merged (1) bleaching agents with the same active agent and concentration but different delivery methods (2) groups with the same active agent and concentration but with different daily application times and (3) products with and without desensitizing agents in their composition. Previous literature findings have reported that these variables do not affect bleaching efficacy [25], [26], [27], [28].
2.4. Information sources and search strategy
The search strategy combined controlled vocabulary (MeSH and Emtree terms) with free-text keywords, using the Boolean operators “OR” and “AND” to capture relevant studies. The strategy was structured around two core concepts: the target population (individuals with permanent dentition) and the intervention (at-home dental bleaching systems). The initial search was conducted on August 11, 2023, and subsequently updated on April 12, 2025.
A comprehensive search was carried out in the databases MEDLINE via PubMed, EMBASE, Cochrane Central Register of Controlled Trials (CENTRAL), Latin American and Caribbean Health Sciences Literature database (LILACS), Brazilian Library in Dentistry (BBO), Scopus. The citation databases Web of Science and Scopus were also consulted. Gray literature was inspected by searching the abstracts of the annual conferences of the International Association for Dental Research and its regional divisions (2001–2022), the System for Information on Grey Literature in Europe database, dissertations and theses using the ProQuest Dissertations and Theses full-text database, and the Periódicos Capes Theses database. The first ten pages of Google Scholar were also consulted. The complete search strategies for all databases are provided in the supplementary material (Supplementary Table S1).
To identify unpublished or ongoing studies relevant to the review question, additional searches were performed in major clinical trial registries, including: EU Clinical Trials Register (https://www.clinicaltrialsregister.eu), Current Controlled Trials (www.controlled-trials.com), International Clinical Trials Registry Platform (http://apps.who.int/trialsearch/), ClinicalTrials.gov (www.ClinicalTrials.gov), Brazilian Registry of Clinical Trials (www.rebec.gov.br), and Australian New Zealand Clinical Trials Registry (https://www.anzctr.org.au/TrialSearch).
2.5. Study selection and data collection process
The retrieved records were managed using reference management software (EndNote X9, Clarivate Analytics). Duplicate entries were eliminated through a combination of automatic detection by software and manual review by sorting article titles alphabetically. Relevance screening was conducted in three stages: title, abstract, and full-text review, using the online tool Rayyan (Qatar Computing Research Institute). Studies with unclear eligibility at any stage were retained for further evaluation in the subsequent phase. All screening steps were independently performed by three reviewers (D.C.F.C., R.N.V., L.C.C.).
The reviewers summarized and categorized data, such as study design, number of patients, interventions, and outcomes. In case of disagreements, the decision was reached via consulting a fourth reviewer (R.M.O.T.). When multiple reports referred to the same study, such as duplicate data within a single publication, results from different follow-up periods, or a combination of published and unpublished findings from the same population, all relevant information was consolidated into a single data extraction form to prevent duplicate entries. A second reviewer independently checked extracted data for accuracy and completeness. Any discrepancies were resolved through discussion until consensus was achieved.
2.6. Data items and outcomes
The instrumental color evaluation (ΔEab CIEL∗a∗b∗ system) [29] obtained from spectrophotometers, colorimeters, and photographs, as well as the visual color assessment in shade guide units (ΔSGU and final SGU measurements using the Vita Classical shade guide), were the primary study outcomes, while the intensity and risk of TS were the secondary outcomes.
To ensure comparability across studies, color change outcomes were extracted from time points closest to 30 days after the completion of bleaching. Due to varying follow-up durations among included studies, assessment periods ranged from 1 week to 3 months post-treatment. When multiple follow-up data points were reported, the closest to the one-month period was selected. If post-treatment evaluation timing was unstated, the last available assessment after the end of treatment was used.
Certain methodological adjustments were implemented to enable the meta-analysis: When studies omitted ΔSGU but provided final shade guide values, those were extracted instead, assuming mean differences in endpoint values approximate mean changes [30]. In cases where color was measured in more than one tooth, a predefined hierarchy was applied: central incisors were prioritized, followed by canines, and then other teeth. If outcomes covered both arches, maxillary data were preferred. For studies where lower numerical values indicated greater whitening, we reversed the effect direction by multiplying the outcome by −1. Quantitative extraction was not performed when outcomes were expressed as raw L*, a*, b* coordinates rather than ΔE, or when ΔSGU was calculated using shade guides not compatible with the Vita Classical or Vita Lumin-Vacuum systems.
To estimate absolute risk of tooth sensitivity (TS) and gingival irritation (GI), the number of events and total participants in each group were recorded. TS intensity was evaluated using either a visual analog scale (VAS, 0–10) or a numerical rating scale (NRS). For these outcomes, the extracted data included group means, standard deviations, and sample sizes. TS risk was treated as a binary outcome, indicating presence or absence of sensitivity.
When studies reported TS intensity at multiple time points, the highest recorded score, regardless of treatment week was used. In cases where the NRS scale differed from the 0–4 format, such as NRS 0–3 or NRS 1–5, values were rescaled through a linear transformation to align proportionally with the 0–4 range, ensuring comparability across studies. For studies presenting TS or GI data for both arches, the most severe value between maxilla and mandible was selected. Studies were excluded from quantitative synthesis if they: (1) assessed TS under stimulated conditions; (2) reported aggregate TS and GI data without separating treatment groups; (3) presented a combined incidence of TS and GI as a general adverse effect rate; or (4) employed pain measurement tools other than VAS or NRS.
For all outcomes, when conflicting data were identified across different sources from the same study, extraction followed a predefined hierarchy: (1) peer-reviewed journal articles, (2) theses, and (3) conference abstracts. This order reflects the assumption that published articles undergo rigorous peer review and are more likely to provide accurate and complete information.
2.6.1. Dealing with missing data
When numerical data were not explicitly reported but available graphically, values were estimated using a digital screen ruler. If group-specific sample sizes were unspecified, equal allocation across groups was assumed. In cases of attrition without specifying affected groups, the baseline sample size was retained for analysis.
When key numerical outcomes such as the mean or standard deviation (SD) were missing, appropriate estimation methods were applied as outlined by Furukawa et al. [31]. If only medians and interquartile ranges were reported, the median was treated as the mean, and the SD was estimated accordingly. When studies provided standard errors, confidence intervals, or p-values instead of SDs, conversions followed Cochrane Handbook guidelines [30].
If none of these strategies enabled SD calculation, imputation was performed by calculating the average coefficient of variation from the remaining studies and multiplying it by the corresponding study mean. The influence of including imputed data was examined through sensitivity analysis to assess its potential impact on the overall findings.
2.7. Geometry of the network
A network plot was generated to provide a qualitative overview of the network geometry. In this visualization, each node represented a specific concentration of CP, while connections between nodes indicated direct comparisons reported in the included studies. Node size was proportional to the total number of participants assigned to each intervention, and the thickness of connecting lines reflected the number of studies contributing to each comparison. The structure of the network was visually inspected to identify dominant comparisons, assess overall connectivity, and detect any interventions that were isolated from the network.
2.8. Risk of bias within individual studies
Risk of bias was independently assessed for each study outcome by three reviewers (D.C.F.C., R.N.V., L.C.C.) using the Cochrane Risk of Bias tool version 2.0 (RoB 2.0) for randomized controlled trials [32]. The tool evaluates five domains: bias arising from the randomization process (D1), deviations from intended interventions (D2), missing outcome data (D3), measurement of the outcome (D4), and selection of the reported result (D5), and overall risk of bias judgment (OVERALL). Each domain was rated as ‘low risk,’ ‘some concerns,’ or ‘high risk’ based on responses to signaling questions from the tool. Studies were judged to have low risk of bias only if all domains were rated as low; if one domain raised concerns, the study was classified as having ‘some concerns’; and if at least one domain was rated high risk or if two or more domains had ‘some concerns’ the study was considered at high risk of bias. Any discrepancies among reviewers were resolved through discussion or, when necessary, by consulting a fourth reviewer (A.R.).
2.9. Summary measures and planned methods of analysis
The mean difference (MD) with 95 % credible intervals (CrIs) were calculated for the continuous data from the eligible studies (ΔE*ab, ΔSGU/final SGU, and intensity of TS). Additionally, the risk ratio (RR) with 95 % CrIs was calculated for the dichotomous outcome risk of TS. The meta-analysis was conducted with the studies that reported outcomes in a way that could be extracted, calculated or converted in the appropriate format for data analysis. Data from eligible studies were analyzed as follows: (1) for the dichotomous outcome, the risk ratio with 95 % confidence intervals (CrIs) was calculated; (2) for continuous outcomes, the mean difference with 95 % CrIs was calculated. Disagreements in the data extraction were resolved by discussion and consensus.
Meta-analysis focused on short-term data. Despite being part of the study protocol, the meta-analysis of medium-term data (around 1-year post-bleaching) and the meta-analysis of ΔE using the formula CIEDE2000 [33] were not feasible due to a reduced number of studies available for analysis.
When encountering a study with multiple treatment arms and a limited sample size, it was not feasible to divide the study sample for multiple comparisons for the network meta-analysis. Then, two groups were randomly selected by coin toss to allow the inclusion of at least one comparison in the network meta-analysis. For studies with multiple treatment arms, and for those in which data from the control group were compared with more than one group, the number of participants (n) from the control group was divided by the number of comparisons.
2.10. Statistical analysis
Initially, a traditional meta-analysis was performed for each pairwise comparison where evidence was available for two or more studies, using RevMan web software (Cochrane Collaboration). Heterogeneity was assessed using the Cochran Q test and I2 statistics.
A mixed treatment comparison (MTC), commonly reported as network meta-analysis, was subsequently performed using a Bayesian framework to integrate both direct and indirect treatment comparisons [19]. Given the expected variability across studies, a random-effects model was employed to account for heterogeneity inherent in literature-based data. The analysis was conducted using MetaInsight (version 6.2.0) [34], a web-based platform that applies Bayesian estimation via Markov Chain Monte Carlo (MCMC) simulation. The model used non-informative priors for treatment effects and followed default settings: 5000 burn-in iterations followed by 20,000 sampling iterations (iterations 5001–25,000), generating 20,000 data points per chain.
Model convergence was evaluated according to MetaInsight guidance [34]. Gelman-Rubin diagnostic plots for each mixed treatment comparison were visually assessed, and convergence was confirmed by ensuring the potential scale reduction factor (Rc) remained below the threshold of 1.1.
The results of the MTC were displayed as point estimates, and 95 % CrIs. When significant differences were detected between comparisons, we calculated the relative ranking for each intervention using the Surface Under the Cumulative Ranking curve (SUCRA), estimated within the Bayesian framework. The larger the SUCRA value, the higher the rank of an intervention in the network. As SUCRA is not accompanied by the 95 % CrI, the values should be interpreted alongside with differences obtained in the network meta-analysis.
Furthermore, sensitivity analyses were conducted to explore the impact of potentially important effect modifiers on findings from MTC. These included separate analyses that involved exclusion of the following: (1) studies with high risk of bias and (2) studies with missing data and imputed values.
2.11. Assessment of inconsistency
Local inconsistency was evaluated using the node-splitting approach [35], which compares direct and indirect estimates for each treatment comparison individually. A non-significant difference between these sources of evidence for any split node (p ≥ 0.05) suggests insufficient evidence to reject the assumption of consistency; conversely, a statistically significant difference indicates inconsistency within that comparison.
A global assessment of inconsistency was also undertaken. As the MetaInsight platform does not currently support implementation of the Unrelated Mean Effects (UME) model or a saturated model for model fit comparisons using the Deviance Information Criterion (DIC), alternative approaches were used. Leverage plots were reviewed to detect influential data points, with a threshold of c = 3 [36]. In addition, residual deviance plots were examined, comparing each data point’s contribution under the consistency model (x-axis) and the UME inconsistency model (y-axis). Optimal consistency is indicated when most values align near 1 on the horizontal axis. Model complexity was also evaluated by verifying that the effective number of parameters (pD) was less than the total number of data points, ensuring model parsimony.
When local or global inconsistency was identified, a structured approach was used to investigate and address its potential sources. Initially, we re-examined the extracted data and input files to rule out errors or misclassification. We then explored heterogeneity across studies by reviewing differences in participant characteristics, intervention protocols, and outcome definitions. Subgroup and sensitivity analyses were also conducted to determine whether excluding outlier or poorly fitting studies improved model consistency. Additionally, we analyzed the network geometry to identify unusual or weakly connected comparisons. Finally, we re-estimated the models in R using alternative model specifications to assess the robustness of the findings. In cases where the source of inconsistency remained unclear or was attributable to a specific study, we proceeded with the consistency model while transparently reporting this limitation in the manuscript.
2.12. Assessment of the quality of evidence using grading of recommendations: assessment, development, and evaluation
We evaluated the quality of evidence for each comparison and outcome using the Grading of Recommendations Assessment, Development and Evaluation (GRADE) approach [37], [38]. Certainty for direct estimates was initially assessed using classic GRADE domains: risk of bias, inconsistency, indirectness, and publication bias. For indirect estimates, the certainty was determined by identifying the lowest-rated first-order loop and considering potential intransitivity. For network estimates, certainty was derived by comparing the direct and indirect estimates, selecting the higher rating, and addressing incoherence and imprecision at the network level. Imprecision was evaluated using a minimally contextualized framework, in which the minimally important difference (MID) was applied as a decision threshold to interpret certainty of effect estimates [39], [40]. The findings were presented in GRADE evidence profiles, accompanied by detailed explanatory footnotes to support transparency and interpretation [38], [41].
3. Results
3.1. Study selection
A total of 10,551 articles were retrieved from electronic databases. After the removal of duplicates, and the title and abstract screening, 28 articles remained (Fig. 1). Of these, 12 articles were excluded, which left 16 articles eligible for this systematic review. From these 16 articles, 3 articles were different follow-up periods of the same study [6], [42], [43], [44], therefore they received the same study ID, giving a total number of 13 studies (Fig. 1). We only imputed the standard deviation of TS of one study [13].
Fig. 1.
Flow diagram of study identification. The same study ID may include more than one article due to multiple follow-up publication or presentation in different scientific documents.
3.2. Study characteristics
The characteristics of the included RCTs are presented in Table 1. The most often study design (n = 9) was parallel [6], [9], [10], [15], [16], [42], [43], [44], [45], [46], [47], [48], [49], while four studies used the split-mouth design [13], [18], [47], [50].
Table 1.
Characteristics of the included articles.
| Study ID | Study Design (Setting) | Patients Age Mean ± SD (Range) | No. of men (women) | Color Assessment | Drop-outs |
|---|---|---|---|---|---|
| Alonso, 2014 | Parallel (University) | 25.9 ± 5.6 (18 – n.r.) | 28 (68) | Shade Guide Unit (Vita Classical Guide) and Spectrophotometer | 0 |
| Basting, 2012 | Parallel (University) | n.r. ± n.r. (18 – n.r.) | 18 (76) | Shade Guide Unit (Vita Pan Classical) | 0 |
| Bernardon, 2015 | Split-mouth (n.r) | n.r. ± n.r. (18 – 40) | 10 (20) | Shade Guide Unit (Vita Classical Guide) and Spectrophotometer | n.r. |
| Bernardon, 2016 | Split-mouth (n.r) | n.r. ± n.r. (18 – 40) | n.r. | Shade Guide Unit (Vita Classical and Zahnfabrik Shade Guide) and Spectrophotometer | n.r. |
| Gerlach, 2000 | Parallel (n.r) | 39.02 ± 8.43 (24 – 57) | 6 (30) | Spectrophotometer and digital images | 4 |
| Khin, 2000 | Parallel Double-blind (University) | n.r. ± n.r. (18 – 65) | n.r. | Shade Guide Unit (Vita Lumin) | n.r. |
| Matis, 2000 | Split-mouth, blind (n.r.) | 50.4 ± n.r. (26 – 73) | 8 (17) | Spectrophotometer | n.r. |
| Meireles, 2008a | Parallel Double-blind (University) | 25.3 ± 7.9 (18 – 55) | 61 (31) | Shade Guide Unit (Vita Classical Guide) and Spectrophotometer | 1 |
| Meireles, 2008b | 3 | ||||
| Meireles,2009 | 3 | ||||
| Meireles, 2010 | 11 | ||||
| Moreira, 2014 | Split-mouth (University) | n.r. ± n.r. (18 – 40) | n.r. | Shade Guide Unit (Vita Classical Guide) and Spectrophotometer | n.r. |
| Piknjac, 2021 | Parallel Double-blind (University) | n.r. ± n.r. (18 – 70) | 19 (41) | Spectrophotometer | n.r. |
| Santana, 2014 | Parallel (n.r) | n.r. ± n.r. (18 – n.r.) | 23 (77) | Shade Guide Unit (Vita Bleachguide), digital images, and spectrophotometer | n.r. |
| Silva, 2014 | Parallel Single-blind (University) | n.r. ± n.r. (18 – 28) | n.r. | Spectrophotometer | n.r. |
| Sutil, 2020 | Parallel Single-blind (University) | 24.4 ± 6.8 (n.r.) | 26 (54) | Shade Guide Unit (Vita Classical Guide and Bleachguide), and spectrophotometer | n.r. |
Abbreviations: ID – identification; SD – standard deviation; n.r. – not reported
All included studies evaluated color change (Table 2). Among these, 11 studies used instrumental evaluation [6], [13], [15], [16], [18], [42], [43], [44], [45], [46], [47], [48], [49], [50], while 10 studies used visual evaluation [6], [9], [10], [13], [15], [18], [42], [43], [44], [46], [47], [48], [50].
Table 2.
Experimental groups, protocol, and outcomes of the included studies.
| Study ID | Groups/Materials (No. of patients per group) |
Bleaching Protocol – Daily Application (Treatment Time in days) | Color Assessment (Outcome) | Tooth Sensitivity Scale (Outcome) | Gingival Irritation |
|---|---|---|---|---|---|
| Alonso, 2014 | CP 10 % (24), CP 15 % (24) | 1 h daily (14) | ΔSGU and ΔE | Absolute risk and Intensity (NRS 0–4) | n.r. |
| Basting, 2012 | CP 10 % (19), CP 20 % (21) | 2 h daily (21) | ΔSGU | Absolute risk and Intensity (NRS 0–3) | n.r. |
| Bernardon, 2015 | CP 10 % (30), CP 22 % (30) | 2 h daily (42) | ΔSGU and ΔE | n.r | n.r. |
| Bernardon, 2016 | CP 10 %a (25), CP 10 %b (25), CP 15 % (25), CP 16 % (25) | 2 h daily (45) | ΔSGU and ΔE | Absolute risk and Intensity (VAS 0–10) | Löe Index 0–3 |
| Gerlach, 2000 | CP 10 % (10), CP 15 % (11), CP 20 % (5) | 2 h daily (14) | ΔE | Absolute risk | Absolute risk |
| Khin, 2000 | CP 10 % (26), CP 15 % (26) | 4 h daily (14) | ΔSGU | Absolute risk and Intensity (VAS 0 – 20) | n.r. |
| Matis, 2000 | CP 10 % (25), CP 15 % (25) | Overnight (14) | ΔSGU and ΔE | Intensity (NRS 1–5) | Scale 1–5 |
| Meireles, 2008 (a, b), 2009 and 2010 | CP 10 % (46), CP 16 % (46) | 2 h daily (21) | ΔSGU and ΔE | Absolute risk and Intensity (NRS 0 – 4) | n.r. |
| Moreira, 2014 | CP 10 % (18), CP 5 % (18), CP 2 % (18), CP 1 % (18) | 2 h daily (45) | ΔSGU and ΔE | Absolute risk and Intensity (VAS 0 – 10) | n.r. |
| Piknjac, 2021 | CP 10 % (20), CP 16 % (20) | 6 h daily (14) | ΔE | Absolute risk and Intensity (NRS 0–10) | n.r. |
| Santana, 2014 | CP 10 %a (20), CP 20 %a (20), 10 %b (20), 20 %b (20) | 2 h daily (14) | ΔSGU and ΔE | Absolute risk and Intensity (VAS 0 – 10) | Löe Index 0–3 |
| Silva, 2014 | CP 10 % (12), CP 20 % (12), | CP 10 % Overnight, CP 20 % 1 h daily (14) | ΔE | Absolute risk and Intensity (NRS 0–5) | n.r |
| Sutil, 2020 | CP 10 % (40), CP 37 % (40), | 30 min daily (21) | ΔSGU and ΔE | Absolute risk and Intensity (NRS 0–4, and VAS 0–10) | VAS 0–10 |
Abbreviations: ID – identification; CP – carbamide peroxide; n.r – not reported
Of the 13 studies included, the risk of TS was evaluated in 11 studies [6], [9], [10], [15], [16], [18], [42], [43], [44], [45], [46], [47], [48], [49]. For intensity of self-reported TS, four studies used the VAS scale [10], [46], [48], [49], and five studies used the NRS scale [6], [9], [13], [15], [42], [43], [44], [45], [48] (Table 2).
The patients’ ages in the studies ranged from 18 to 73 years old, but four studies did not report this information [9], [15], [46], [48]. The mean age of all participants included in the RCTs that reported this information was approximately 30 years, showing a predominance of young adults. Women (81.2 %) were predominant in all studies that reported this characteristic (Table 1).
Regarding the bleaching protocol, CP concentrations ranged from 5 % to 37 %, and on average participants used the products for two hours daily. A total bleaching period of 14 days was the most reported, but periods of 21, 42 and 45 days were also observed in some studies (Table 2).
3.3. Risk of bias (RoB) within studies and across studies
Most studies exhibited a high RoB (Fig. 2). One study [6], [42], [43], [44] was classified as having some concerns, and another study [48] was classified as having a low RoB for the outcome color change and TS. Bias due to deviations from intended interventions (D2) and bias in outcome measurement (D4) were the ones that presented more concerns (Fig. 2).
Fig. 2.
Summary of risk of bias assessment for the color change in (a) ΔE, (b) ΔSGU, (c) Risk of TS, and (d) Intensity of TS according to RoB 2.0 (Cochrane Collaboration tool).
3.4. Network meta-analysis
3.4.1. Instrumental Color change in ΔE
The traditional pairwise meta-analysis results are presented in Supplementary Figure S1. The network geometry plot (Fig. 3a) illustrates the structure of the evidence, showing that the CP 5 % and CP 37 % nodes were each informed by only one direct comparison, reflecting their limited representation in the included studies. In contrast, CP 10 % was the most frequently evaluated concentration, supported by 11 studies [6], [13], [15], [16], [18], [42], [43], [44], [45], [46], [47], [48], [49], [50] and a total of 301 patients.
Fig. 3.
Network plot geometry for color change in (a) ΔE, (b) ΔSGU, (c) Risk of TS, (d) TS Intensity in NRS, (e) TS Intensity in VAS.
Fig. 4 displays the forest plot from the Bayesian random-effects consistency model for color change measured by ΔE, with estimates reported as mean differences (MD) and 95 % credible intervals (CrI). All concentrations outperformed CP 5 %, with a statistically significant difference observed between CP 37 % and CP 5 % (MD = –3.66; –6.24 to –0.93).
Fig. 4.
Bayesian random effect consistency model forest plot of the pooled effects estimates of the color change in ΔE expressed in mean difference and respective 95 % CrI. CP: carbamide peroxide.
The SUCRA rankings for color change in ΔE are shown in Fig. 5a. CP 20–22 % had the highest probability of being the most effective bleaching concentration (83.4 %), followed by CP 37 % (78.1 %). In contrast, CP 5 % had the lowest ranking (0.3 %), while CP 10 % (43.6 %) and CP 15–16 % (44.6 %) had intermediate and similar probabilities. However, these rankings should be interpreted cautiously, as CP 5 % was the only concentration shown to be significantly less effective than others in the NMA, and this conclusion was based on a single study evaluating CP 5 % [47].
Fig. 5.
The rank probability of different interventions based on the SUCRA for color change in ΔE (a), Risk of tooth sensitivity (b), the intensity of tooth sensitivity – NRS scale (c) and VAS scale (d).
3.4.2. Color change in ΔSGU
Of the ten studies that evaluated color change using visual methods, two were excluded for employing shade guide other than the Vita Shade Guide/Vita Lumin-Vacuum [13], [46]. The final network included eight studies [6], [9], [10], [15], [18], [42], [43], [44], [47], [48], [50] encompassing five interventions and eight possible pairwise comparisons, with a total of 511 patients.
The network plot is presented in Fig. 3b. As observed in the instrumental analysis, CP 5 % and CP 37 % treatment were supported by only one direct comparison. In contrast, CP 10 % was the condition with the highest number of direct comparisons and patients included.
Results from the traditional paired meta-analysis can be seen in the Supplementary Figure S2. Fig. 6 presents the forest plot from the Bayesian random-effect consistency model for color change measured in ΔSGU, in mean difference (MD). No significant differences were found between concentrations. As a result of this lack of between-groups, SUCRA analysis was not performed.
Fig. 6.
Bayesian random effect consistency model forest plot of the pooled effects estimates of the color change in ΔSGU expressed in mean difference and respective 95 % CrI. CP: carbamide peroxide.
3.4.3. Risk of Tooth Sensitivity (TS)
Of the 13 studies included in this systematic review, 11 assessed the risk of TS studies [6], [9], [10], [15], [16], [18], [42], [43], [44], [45], [46], [47], [48], [49]. As shown in Fig. 3c, the CP 5 % and CP 37 % nodes were informed by only one direct comparison each, reflecting limited evidence. In contrast, the CP 10 % node had the highest number of comparisons, derived from the 11 studies, with a total of 296 patients.
Results from the traditional pairwise meta-analyses are provided in Supplementary Figure S3. Fig. 7 presents the forest plot from the Bayesian random-effects consistency model for the risk of TS, expressed as risk ratio (RR, 95 % CrI). A statistically significant difference was observed between CP 10 % and CP 20 % (RR = 1.67; 1.15–2.57), indicating a higher risk of TS associated with CP 20 %. No significant differences were observed in the other concentration comparisons.
Fig. 7.
Bayesian random effect consistency model forest plot of the pooled effects estimates of the Risk of tooth sensitivity expressed in risk ratio and respective 95 % CrI. CP: carbamide peroxide.
Fig. 5b shows the SUCRA values for all bleaching agents. CP 10 % had the highest probability of presenting the lowest risk of TS (81.8 %), followed by CP 5 % (64.2 %), CP 15–16 % (50.6 %), CP 37 % (46.5 %), and CP 20–22 % (6.9 %).
3.4.4. Intensity of TS – NRS scale
Of the 13 studies included in this review, nine evaluated the intensity of TS, with six of them [6], [9], [13], [15], [42], [43], [44], [45], [48] using the NRS scale. The network plot in Fig. 3d presents 4 interventions, with 6 possible pairwise comparisons, with a total of 276 patients.
Traditional pairwise meta-analyses are presented in Supplementary Figure S4. Fig. 8 illustrates the forest plot from the Bayesian random-effects consistency model for pooled estimates of TS intensity using the NRS 0–4 scale as MD. Statistically significant differences were found in the comparisons of CP 10 % vs. CP 20–22 % (MD = 0.579; 0.176–0.980) and CP 15–16 % vs. CP 20–22 % (MD = −0.503; 0.0705–0.941). Unexpected results were noted in the comparison involving CP 37 %, which did not follow the overall trend of products with higher concentration yielding higher TS intensity. These findings are likely due to the limited evidence from a single study evaluating CP 37 %.
Fig. 8.
Bayesian random effect consistency model forest plot of the pooled effects estimates of the Intensity of tooth sensitivity - NRS scale expressed in mean difference and respective 95 % CrI. CP: carbamide peroxide.
Fig. 5c displays the SUCRA rankings, which revealed that CP 10 % (82.26 %) exhibited the lowest TS intensity, followed by CP 37 % (58.14 %), CP 15–16 % (56.73 %), and CP 20–22 % (2.87 %). Both analysis in VAS scale and NRS suggested that CP 10 % is the most likely concentration to be associated with minimal or no TS.
3.4.5. Intensity of TS – VAS scale
Of the 13 included studies, eleven valuated TS intensity, with four of them [18], [46], [47], [48] using the VAS scale. The network plot in Fig. 3e, shows 4 interventions through 6 possible pairwise comparisons, involving 212 patients.
Traditional pairwise meta-analyses were performed and can be found in Supplementary Figure S5. Fig. 9 presents the forest plots from the Bayesian random-effects consistency model for TS intensity on the VAS 0–10 scale. A significant MD was observed only between CP 10 % vs. CP 15–16 % (MD = 0.704; 0.242–1.17).
Fig. 9.
Bayesian random effect consistency model forest plot of the pooled effects estimates of the Intensity of tooth sensitivity - VAS scale expressed in mean difference and respective 95 % CrI. CP: carbamide peroxide.
SUCRA rankings (Fig. 5d) showed that CP 10 % (93.01 %) exhibited the lowest TS intensity, followed by CP 37 % (54.91 %), CP 15–16 % (39.31 %), and CP 20–22 % (12.77 %).
3.4.6. Gingival irritation (GI)
Only two studies assessed GI. Matis et al. [13] used a 0–5 scale and presented the daily mean change over the course of the protocol, with higher GI intensity (1.2–1.5 VAS units) between days 3 their data as means and standard deviations preventing data extraction.1.5 VAS units, no SD reported) in the first two weeks of treatment. Sutil et al. [48] used VAS 0–10 scale, showing no significant difference between the CP 10 % (0.9 ± 1.4) and CP 37 % (0.8 ± 1.2) groups.
3.5. Assessment of consistency
An inconsistency was identified in the network analysis of the risk of TS, attributed by the study of Bernardon et al. [18], which showed a high residual deviance in the leverage plot (Supplementary Figure S6). However, sensitivity analysis (Supplementary Figure S7) revealed that excluding this study had minimal impact on the overall results. While slight variations in estimates were observed, the statistical significance of the findings remained unchanged, supporting the robustness of the network.
For color change in ΔE, the node-splitting approach (Supplementary Figure S8) comparing CP 20–22 % and CP 15–16 % showed no evidence of inconsistency (p-valor 0.1026). Similarly, the leverage plots for ΔSGU (Supplementary Figure S9a) and intensity of TS measured on both the NRS (Supplementary Figure S9b) and VAS scales (Supplementary Figure S9c) demonstrated no evidence of inconsistency as the distribution of data points closely followed the model’s expected pattern.
3.6. Sensitivity analysis
Standard deviations for the intensity of TS were imputed in only one study [13]. Although a sensitivity analysis excluding studies at high RoB was initially planned, it was not feasible since most included studies fell into this category. The impact of this were assessed in the certainty of the evidence.
As previously noted, a sensitivity analysis excluding the Bernardon et al. [18] study, identified as a source of inconsistency in the leverage plot for TS risk, did not alter the overall results (Supplementary Figure S7).
3.7. Assessment of the quality of evidence using grading of recommendations: assessment, development, and evaluation
The Supplementary Material (Tables S2a-c; Tables S3a-c; Tables S4a-c; Tables S5a-c; and Tables S6a-c) presents quality ratings for the direct, indirect, and network evidence for each outcome. Most comparisons were classified as having low or very low certainty due to high RoB and significant imprecision.
Certainty of evidence was assessed following the GRADE approach [37], [38]. For direct comparisons, the evaluation considered the standard GRADE domains: RoB, inconsistency, indirectness and publication bias. Certainty in indirect comparisons was determined by identifying the lowest-rated first-order loop and evaluating potential intransitivity. Then the overall certainty of the network estimates was derived from the higher certainty between the direct and indirect estimates, while also accounting for incoherence and imprecision at the network estimate level.
Imprecision was assessed using a minimally contextualized approach, with minimally important difference (MID) serving as the threshold [39], [40]. MID judgments were based on published estimates and authors consensus. We adopted conservative MID threshold of 1 point for NRS scale and 2 points for VAS 0–10 for intensity of TS, while setting it at 100 per 1000 patients for the risk of TS [39]. For color change, we used the established acceptability threshold (ΔE*ab = 2.7) [51] for the instrumental color evaluation and a 2.0 shade guide unit for visual assessment [28]. We reported the results using GRADE evidence tables with detailed explanatory footnotes [38], [41] (Tables S2a-c; Tables S3a-c; Tables S4a-c; Tables S5a-c; and Tables S6a-c).
4. Discussion
4.1. Methodological limitations and risk of bias of the eligible studies
Several challenges were encountered during this study, primarily due to inconsistencies in data presentation across studies. These inconsistencies often made it difficult to extract data reliably or interpret the study conclusions with confidence.
A major issue was the lack of standardization in collecting and reporting TS outcomes. While some studies addressed only the risk of TS, others focused solely on its intensity. In many cases the threshold used to define the presence of TS was not reported. This prevents reviewers from calculating the risk of TS between different groups. Among studies measuring TS intensity there was often insufficient detail on how the mean values were calculated, specifically whether the mean represented the average score per participant over time or the highest TS score reported by each participant during the treatment period.
Additional limitations included the absence of group-specific means or measures of dispersions, such as standard deviations. Some studies failed to disclose the method used to assess TS or employed questionable methods, such as artificially stimulating TS. In several cases, data on adverse effects were reported in aggregate, combining TS and GI, or reporting overall TS outcomes across all treatment groups, which prevented extraction of group-level data.
These methodological limitations were reflected in the RoB assessments. Among the 13 eligible studies, only one was graded as having a low RoB across all domains for all outcomes (ΔE, ΔSGU, and risk and intensity of TS). The most common reasons of high RoB were inappropriate randomization procedures and failure to report allocation concealment, which raises concerns about potential baseline imbalances between groups that could compromise internal validity.
Additionally, many studies failed to report whether blinding was implemented for patient, operators and examiners. Lack of patient blinding can influence behavior or adherence to the treatment protocols, as participant’s expectations may unconsciously affect outcomes. Similarly, examiner blinding is equally crucial for subjective measures such as TS risk/intensity or color change measurements with shade guide units, where knowledge of group allocation may introduce measurement bias. Inadequate blinding and flawed randomization procedures can distort effect estimates, potentially leading to either over- or underestimation of treatment efficacy and safety.
Another recurrent issue was the lack of reporting on patient withdrawal or treatment discontinuations. Few studies explained the reasons for these losses which raises concerns about the completeness and reliability of the data. If discontinuations were related to adverse effects, such as tooth sensitivity, the omission of this information may result in underreporting of both the risk and intensity of TS across treatment groups, thereby compromising the validity of the findings.
Although robust statistical analyses were conducted using the available this systematic review highlights the urgent need to improve methodological quality of RCTs on dental bleaching. In particular, there is a need to develop a standardized reporting guidelines for outcomes related to color change and adverse effects (TS/GI) in bleaching studies. Adherence to such a guideline, yet to be established, would facilitate data synthesis, reduce heterogeneity, and improve the comparability of outcomes in future meta-analyses.
Some practical recommendations can be highlighted. For color outcomes, studies should report means, standard deviations and total number of participants evaluated in each time assessment for all color change tools. Commonly used tools include ΔEab [29], ΔE00 [33], whiteness index (WID) [52] and ΔSGU measured with color shade guides (Vita Classical and Vita Bleachedguide).
Furthermore, studies should report changes from baseline measures rather than the final outcome after treatment. Authors should clearly state the timing and number of participants in the study follow-ups, provide information about the number and the reasons for patient dropouts and report how missing data were handled. Ideally, authors should use the intention-to-treat protocol, including all randomized participants in the groups they were allocated. Per-protocol or as-available analysis can be performed only with clear justification for this approach.
Data should always be reported at the group level rather than at the aggregate study level, as the latter limits the ability the usability of data for secondary analyses. When presenting results in graphical form, standard deviations must be clearly displayed to allow accurate data extraction.
Although this remains a topic of ongoing debate, the authors of the present review recommend collecting data on spontaneous tooth sensitivity (TS), without the use of any stimulation. Stimulated assessments may artificially provoke responses that do not reflect clinically relevant outcomes.
4.2. Color change evaluation
When examining the color change in ΔE, we observed that all concentrations were like one another, except for CP 5 % which were inferior to all others. However, only one study [47] evaluated the bleaching efficacy of CP 5 %, which prevents us from drawing definitive conclusions about its potential inferiority. Further RCTs should explore the color change of lower CP concentration levels, particularly CP 5 %.
In the instrumental analysis (ΔE) SUCRA rankings indicated that CP 20 % (83.4 %) and CP 37 % (78.1 %) had the highest probabilities of whitening teeth. However, caution should be exercised when interpreting this result. SUCRA values reflect the relative ranking of interventions but do not account for the magnitude of effect differences between treatments [53]. As such, a top-ranked product may offer only a marginal advantage over others.
To avoid misinterpretations, SUCRA findings must be interpreted alongside with the MTC results. Except for CP 5 %, no statistically significant differences were observed across concentrations in terms of ΔE. This interpretation is supported by the findings from the visual evaluation (ΔSGU), where no concentration showed statistically or clinically meaningful superiority by the end of treatment.
These findings challenge the common assumption that higher concentrations of carbamide peroxide lead to greater whitening. In fact, the evidence from this review suggests that lower concentrations, such as CP 10 %, may offer similar whitening efficacy without the need for higher dosages.
It is important to note that this analysis included only ΔEab and ΔSGU values based on the Vita Classical shade guide, as these were the most frequently reported color parameters in the eligible studies. Future research should prioritize the inclusion of ΔE₀₀ and the Whiteness Index (WID), along with ΔSGU values obtained using the Vita Bleachedguide shade guide, to allow comparability of other color measurement tools.
4.3. Adverse effects
Both the risk and intensity of TS were found to be lower with CP10 % compared to higher concentrations. A statistically significant reduction in the risk of TS was observed between CP 10 % and CP 20–22 %, with CP 10 % demonstrating a 67 % lower average risk. Similarly, the analysis of TS intensity using both the NRS (0–4) and VAS (0–10) scales confirmed that CP 10 % had the highest probability of producing the least sensitivity, while CP 20–22 % had the highest likelihood of causing more intense TS. These results were further supported by SUCRA rankings, with CP 10 % showing the highest probability of being associated with the lowest risk and intensity of TS outcomes.
These findings are consistent with in vitro studies that reported reduced diffusion of HP into the pulp chamber with less concentrated products [54] and with biological evidence suggesting lower concentrations of carbamide peroxide induce less irritation, cellular damage and inflammatory responses in the pulp tissue [55], [56], causing the less discomfort observed by the patients with less concentrated products.
The gingival irritation was found to be a poorly reported event in the included clinical studies. Only two studies mentioned this outcome, and data could be extracted from just one of them. From these studies [13], [48], gingival irritation was minor, suggesting a favorable safety profile when properly indicated and applied according to professional guidance.
4.4. Certainty of the evidence
Unfortunately, most comparisons in this systematic review were rated as having low or very low certainty of evidence due to the high RoB of the studies and substantial imprecision as the primary studies are generally low powered.
4.5. Final considerations
There is an urgent need for well-designed RCTs with a low RoB to enhance the certainty and credibility of research findings. Improving methodological quality is essential to support more reliable conclusions in systematic reviews, particularly given that most comparisons in the present analysis were rated as having low or very low certainty due to high risk of bias and substantial imprecision.
5. Conclusion
The analyses of both color change and the risk and intensity of tooth sensitivity (TS) consistently indicated that 10 % carbamide peroxide (CP) can achieve comparable bleaching efficacy while causing less sensitivity, although the certainty of this evidence is low. Future research should further investigate the effects of more extreme concentrations, such as 5 % and 37 % CP, which were underrepresented in this systematic review and currently lack sufficient evidence to support reliable conclusions.
Funding
none
Declaration of Competing Interest
The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
Acknowledgements
Special thanks to the Bleaching&Bond group (Instagram @bleachingbond; Brazil) for their invaluable assistance throughout all stages of this systematic review. This study received partial support from the National Council for Scientific and Technological Development (CNPq) under grants 304817/2021–0 and 308286/2019–7, as well as the Coordenação de Aperfeiçoamento de Pessoal de Nível Superior, Brazil (CAPES), Finance Code 001.
Footnotes
Scientific field of dental Science: Esthetic Dentistry
Supplementary data associated with this article can be found in the online version at doi:10.1016/j.jdsr.2025.10.001.
Contributor Information
Deisy Cristina Ferreira Cordeiro, Email: 240308800007@uepg.br.
Romina Ñaupari-Villasante, Email: 4100122817005@uepg.br.
Letícia Caroline Condolo, Email: 240208700022@uepg.br.
Renata Maria Oleniki Terra, Email: renata.mot@hotmail.com.
Michael Willian Favoreto, Email: michael.favoreto@utp.br.
Juliana Larocca de Geus, Email: ju_degeus@hotmail.com.
Ana Claudia Chibinski, Email: anachibinski@uepg.br.
Alessandra Reis, Email: alereis@uepg.br.
Appendix A. Supplementary material
Supplementary material
References
- 1.Godinho J., Gonçalves R.P., Jardim L. Contribution of facial components to the attractiveness of the smiling face in male and female patients: a cross-sectional correlation study. Am J Orthod Dentofac Orthop. 2020;157:98–104. doi: 10.1016/j.ajodo.2019.01.022. [DOI] [PubMed] [Google Scholar]
- 2.Alkhatib M., Holt R., Bedi R. Prevalence of self-assessed tooth discolouration in the United Kingdom. J Dent. 2004;32:561–566. doi: 10.1016/j.jdent.2004.06.002. [DOI] [PubMed] [Google Scholar]
- 3.Akarslan Z., Sadik B., Erten H., Karabulut E. Dental esthetic satisfaction, received and desired dental treatments for improvement of esthetics. Indian J Dent Res. 2009;20:195–200. doi: 10.4103/0970-9290.52902. [DOI] [PubMed] [Google Scholar]
- 4.Goulart M.D., Condessa A.M., Hilgert J.B., Hugo F.N., Celeste R.K. Concerns about dental aesthetics are associated with oral health related quality of life in Southern Brazilian adults. Cienc Saude Coletiva. 2018;23:3957–3964. doi: 10.1590/1413-812320182311.24172016. [DOI] [PubMed] [Google Scholar]
- 5.Isiekwe G.I., Aikins E.A. Self-perception of dental appearance and aesthetics in a student population. Int Orthod. 2019;17:506–512. doi: 10.1016/j.ortho.2019.06.010. [DOI] [PubMed] [Google Scholar]
- 6.Meireles S.S., Heckmann S.S., Leida F.L., dos Santos Ida S., Della Bona A., Demarco F.F. Efficacy and safety of 10% and 16% carbamide peroxide tooth-whitening gels: a randomized clinical trial. Oper Dent. 2008;33:606–612. doi: 10.2341/07-150. [DOI] [PubMed] [Google Scholar]
- 7.Grobler S.R., Hayward R., Wiese S., Moola M.H., van W.K.T.J. Spectrophotometric assessment of the effectiveness of opalescence PF 10%: a 14-month clinical study. J Dent. 2010;38:113–117. doi: 10.1016/j.jdent.2009.09.009. [DOI] [PubMed] [Google Scholar]
- 8.Jadad E., Montoya J., Arana G., Gordillo L.A., Palo R.M., Loguercio A.D. Spectrophotometric evaluation of color alterations with a new dental bleaching product in patients wearing orthodontic appliances. Am J Orthod Dentofac Orthop. 2011;140:e43–e47. doi: 10.1016/j.ajodo.2010.11.021. [DOI] [PubMed] [Google Scholar]
- 9.Basting R.T., Amaral F.L., França F.M., Flório F.M. Clinical comparative study of the effectiveness of and tooth sensitivity to 10% and 20% carbamide peroxide home-use and 35% and 38% hydrogen peroxide in-office bleaching materials containing desensitizing agents. Oper Dent. 2012;37:464–473. doi: 10.2341/11-337-C. [DOI] [PubMed] [Google Scholar]
- 10.Kihn P.W., Barnes D.M., Romberg E., Peterson K. A clinical evaluation of 10 percent vs. 15 percent carbamide peroxide tooth-whitening agents. J Am Dent Assoc. 2000;131(1939):1478–1484. doi: 10.14219/jada.archive.2000.0061. [DOI] [PubMed] [Google Scholar]
- 11.Matis B.A., Wang Y., Eckert G.J., Cochran M.A., Jiang T. Extended bleaching of tetracycline-stained teeth: a 5-year study. Oper Dent. 2006;31:643–651. doi: 10.2341/06-6. [DOI] [PubMed] [Google Scholar]
- 12.Braun A., Jepsen S., Krause F. Spectrophotometric and visual evaluation of vital tooth bleaching employing different carbamide peroxide concentrations. Dent Mater. 2007;23:165–169. doi: 10.1016/j.dental.2006.01.017. [DOI] [PubMed] [Google Scholar]
- 13.Matis B.A., Mousa H.N., Cochran M.A., Eckert G.J. Clinical evaluation of bleaching agents of different concentrations. Quintessence Int. 2000;31:303–310. [PubMed] [Google Scholar]
- 14.Leonard R.H., Jr., Garland G.E., Eagle J.C., Caplan D.J. Safety issues when using a 16% carbamide peroxide whitening solution. J Esthet Restor Dent. 2002;14:358–367. doi: 10.1111/j.1708-8240.2002.tb00178.x. [DOI] [PubMed] [Google Scholar]
- 15.de la Peña Alonso, López Ratón V. M. Randomized clinical trial on the efficacy and safety of four professional at-home tooth whitening gels. Oper Dent. 2014;39:136–143. doi: 10.2341/12-402-C. [DOI] [PubMed] [Google Scholar]
- 16.Gerlach R.W., Gibb R.D., Sagel P.A. A randomized clinical trial comparing a novel 5.3% hydrogen peroxide whitening strip to 10%, 15%, and 20% carbamide peroxide tray-based bleaching systems. Compend Contin Educ Dent Suppl. 2000 S22-8; quiz S42-3. [PubMed] [Google Scholar]
- 17.Krause F., Jepsen S., Braun A. Subjective intensities of pain and contentment with treatment outcomes during tray bleaching of vital teeth employing different carbamide peroxide concentrations. Quintessence Int. 2008;39:203–209. [PubMed] [Google Scholar]
- 18.Bernardon J.K., Vieira Martins M., Branco Rauber G., Monteiro Junior S., Baratieri L.N. Clinical evaluation of different desensitizing agents in home-bleaching gels. J Prosthet Dent. 2016;115:692–696. doi: 10.1016/j.prosdent.2015.10.020. [DOI] [PubMed] [Google Scholar]
- 19.van Valkenhoef G., Lu G., de Brock B., Hillege H., Ades A.E., Welton N.J. Automating network meta-analysis. Res Synth Methods. 2012;3:285–299. doi: 10.1002/jrsm.1054. [DOI] [PubMed] [Google Scholar]
- 20.Hutton B., Salanti G., Caldwell D.M., Chaimani A., Schmid C.H., Cameron C., et al. The PRISMA extension statement for reporting of systematic reviews incorporating network meta-analyses of health care interventions: checklist and explanations. Ann Intern Med. 2015;162:777–784. doi: 10.7326/M14-2385. [DOI] [PubMed] [Google Scholar]
- 21.Gotzsche P.C. Why we need a broad perspective on meta-analysis. It may be crucially Important Patients. BMJ. 2000;321:585–586. doi: 10.1136/bmj.321.7261.585. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Grimshaw J., McAuley L.M., Bero L.A., Grilli R., Oxman A.D., Ramsay C., et al. Systematic reviews of the effectiveness of quality improvement strategies and programmes. Qual Saf Health Care. 2003;12:298–303. doi: 10.1136/qhc.12.4.298. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Salanti G. Indirect and mixed-treatment comparison, network, or multiple-treatments meta-analysis: many names, many benefits, many concerns for the next generation evidence synthesis tool. Res Synth Methods. 2012;3:80–97. doi: 10.1002/jrsm.1037. [DOI] [PubMed] [Google Scholar]
- 24.Chaimani A., Caldwell D.M., Li T., Higgins J.P., Salanti G. In: Cochrane Handbook for Systematic Reviews of Interventions version 6.5 ed. Higgins, TJ J.P.T., Chandler J., Cumpston M., Li T., Page M.J., Welch V.A., editors. Cochrane; 2024. Chapter 11: undertaking network meta-analyses [last updated October 2019] [Google Scholar]
- 25.Cardoso P.C., Reis A., Loguercio A., Vieira L.C.C., Baratieri L.M. Clinical effectiveness and tooth sensitivity associated with different bleaching times for a 10 percent carbamide peroxide gel. J Am Dent Assoc. 2010;141:1213–1220. doi: 10.14219/jada.archive.2010.0048. [DOI] [PubMed] [Google Scholar]
- 26.Cordeiro D., Toda C., Hanan S., Arnhold L.P., Reis A., Loguercio A.D., et al. Clinical evaluation of different delivery methods of At-Home bleaching gels composed of 10% hydrogen peroxide. Oper Dent. 2019;44:13–23. doi: 10.2341/17-174-C. [DOI] [PubMed] [Google Scholar]
- 27.Maran B.M., Vochikovski L., Hortkoff D.R.A., Stanislawczuk R., Loguercio A.D., Reis A. Bleaching sensitivity with a desensitizing in-office bleaching gel: a randomized double-blind clinical trial. Quintessence Int. 2020;51:788–797. doi: 10.3290/j.qi.a45173. [DOI] [PubMed] [Google Scholar]
- 28.Terra R.M.O., Sutil E., Ferreira Cordeiro D.C., Favoreto M.W., Faria E.S.A., Best A.M., et al. Different daily times for at-home bleaching with 10% carbamide peroxide: a randomized single-blind, noninferiority controlled trial. J Am Dent Assoc. 2025;156:57–67.e5. doi: 10.1016/j.adaj.2024.10.010. [DOI] [PubMed] [Google Scholar]
- 29.Robertson A.R. The CIE 1976 color-difference formulae. Color Res Appl. 1977;2:7–11. [Google Scholar]
- 30.Higgins J.P.T., Thomas, J., Chandler J., Cumpston M., Li T., Page M.J., Welch V.A.Cochrane handbook for systematic reviews of interventions version 6.5 (updated August 2024) (editors). Cochrane; 2024. [DOI] [PMC free article] [PubMed]
- 31.Furukawa T.A., Barbui C., Cipriani A., Brambilla P., Watanabe N. Imputing missing standard deviations in meta-analyses can provide accurate results. J Clin Epidemiol. 2006;59:7–10. doi: 10.1016/j.jclinepi.2005.06.006. [DOI] [PubMed] [Google Scholar]
- 32.Sterne J.A.C., Savović J., Page M.J., Elbers R.G., Blencowe N.S., Boutron I., et al. RoB 2: a revised tool for assessing risk of bias in randomised trials. BMJ. 2019;366:l4898. doi: 10.1136/bmj.l4898. [DOI] [PubMed] [Google Scholar]
- 33.Luo M.R., Cui G., Rigg B. The development of the CIE 2000 colour-difference formula: CIEDE2000. Color Research & Application: Endorsed by Inter-Society Color Council, The Colour Group (Great Britain), Canadian Society for Color, Color Science Association of Japan, Dutch Society for the Study of Color, The Swedish Colour Centre Foundation, Colour Society of Australia, Centre Français de la Couleur. 2001;26:340-350.
- 34.Owen R.K., Bradbury N., Xin Y., Cooper N., Sutton A. MetaInsight: an interactive web-based tool for analyzing, interrogating, and visualizing network meta-analyses using R-shiny and netmeta. Res Synth Methods. 2019;10:569–581. doi: 10.1002/jrsm.1373. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Dias S., Welton N.J., Caldwell D.M., Ades A.E. Checking consistency in mixed treatment comparison meta-analysis. Stat Med. 2010;29:932–944. doi: 10.1002/sim.3767. [DOI] [PubMed] [Google Scholar]
- 36.Dias S., Ades A.E., Welton N.J., Jansen J.P., Sutton A.J. John Wiley & Sons; 2018. Network meta-analysis for decision-making. [Google Scholar]
- 37.Puhan M.A., Schünemann H.J., Murad M.H., Li T., Brignardello-Petersen R., Singh J.A., et al. A GRADE working group approach for rating the quality of treatment effect estimates from network meta-analysis. Bmj. 2014;349:g5630. doi: 10.1136/bmj.g5630. [DOI] [PubMed] [Google Scholar]
- 38.Izcovich A., Chu D.K., Mustafa R.A., Guyatt G., Brignardello-Petersen R. A guide and pragmatic considerations for applying GRADE to network meta-analysis. Bmj. 2023;381 doi: 10.1136/bmj-2022-074495. [DOI] [PubMed] [Google Scholar]
- 39.Brignardello-Petersen R., Guyatt G.H., Mustafa R.A., Chu D.K., Hultcrantz M., Schünemann H.J., et al. GRADE guidelines 33: addressing imprecision in a network meta-analysis. J Clin Epidemiol. 2021;139:49–56. doi: 10.1016/j.jclinepi.2021.07.011. [DOI] [PubMed] [Google Scholar]
- 40.Zeng L., Brignardello-Petersen R., Hultcrantz M., Mustafa R.A., Murad M.H., Iorio A., et al. GRADE guidance 34: update on rating imprecision using a minimally contextualized approach. J Clin Epidemiol. 2022;150:216–224. doi: 10.1016/j.jclinepi.2022.07.014. [DOI] [PubMed] [Google Scholar]
- 41.Santesso N., Carrasco-Labra A., Langendam M., Brignardello-Petersen R., Mustafa R.A., Heus P., et al. Improving GRADE evidence tables part 3: detailed guidance for explanatory footnotes supports creating and understanding GRADE certainty in the evidence judgments. J Clin Epidemiol. 2016;74:28–39. doi: 10.1016/j.jclinepi.2015.12.006. [DOI] [PubMed] [Google Scholar]
- 42.Meireles S.S., Heckmann S.S., Santos I.S., Della Bona A., Demarco F.F. A double blind randomized clinical trial of at-home tooth bleaching using two carbamide peroxide concentrations: 6-month follow-up. J Dent. 2008;36:878–884. doi: 10.1016/j.jdent.2008.07.002. [DOI] [PubMed] [Google Scholar]
- 43.Meireles S.S., dos Santos Ida S., Della Bona A., Demarco F.F. A double-blind randomized controlled clinical trial of 10 percent versus 16 percent carbamide peroxide tooth-bleaching agents: one-year follow-up. J Am Dent Assoc. 2009;140:1109–1117. doi: 10.14219/jada.archive.2009.0337. [DOI] [PubMed] [Google Scholar]
- 44.Meireles S.S., Santos I.S., Bona A.D., Demarco F.F. A double-blind randomized clinical trial of two carbamide peroxide tooth bleaching agents: 2-year follow-up. J Dent. 2010;38:956–963. doi: 10.1016/j.jdent.2010.08.003. [DOI] [PubMed] [Google Scholar]
- 45.Silva M.B.D. Universidade Federal do Rio Grande do Sul; 2014. Avaliação do clareamento dental caseiro com diferentes protocolos de Utilização: ensaio Clínico randomizado [Trabalho de Conclusão de Curso] [Google Scholar]
- 46.Santana R.S.D. Universidade Estadual Paulista; 2014. Efetividade e sensibilidade com uso de géis clareadores experimental e comercial à base de peróxido de carbamida [Mestrado] [Google Scholar]
- 47.Moreira J.M. Universidade Federal de Santa Catarina; 2014. Avaliação clínica de agentes clareadores de baixa concentração [Trabalho de Conclusão de Curso] [Google Scholar]
- 48.Sutil E., da Silva K.L., Terra R.M.O., Burey A., Rezende M., Reis A., et al. Effectiveness and adverse effects of at-home dental bleaching with 37% versus 10% carbamide peroxide: a randomized, blind clinical trial. J Esthet Restor Dent. 2020;34:313–321. doi: 10.1111/jerd.12677. [DOI] [PubMed] [Google Scholar]
- 49.Piknjac A., Soldo M., Illeš D., Zlataric D.K. Patients' assessments of tooth sensitivity increase one day following different whitening treatments. Acta Stomatol Croat. 2021;55:280–290. doi: 10.15644/asc55/3/5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Bernardon J.K., Ferrari P., Baratieri L.N., Rauber G.B. Comparison of treatment time versus patient satisfaction in at-home and in-office tooth bleaching therapy. J Prosthet Dent. 2015;114:826–830. doi: 10.1016/j.prosdent.2015.05.014. [DOI] [PubMed] [Google Scholar]
- 51.Paravina R.D., Ghinea R., Herrera L.J., Bona A.D., Igiel C., Linninger M., et al. Color difference thresholds in dentistry. J Esthet Restor Dent. 2015;27(1):S1–S9. doi: 10.1111/jerd.12149. [DOI] [PubMed] [Google Scholar]
- 52.Pérez M.M., Herrera L.J., Carrillo F., Pecho O.E., Dudea D., Gasparik C., et al. Whiteness difference thresholds in dentistry. Dent Mater. 2019;35:292–297. doi: 10.1016/j.dental.2018.11.022. [DOI] [PubMed] [Google Scholar]
- 53.Mbuagbaw L., Rochwerg B., Jaeschke R., Heels-Andsell D., Alhazzani W., Thabane L., et al. Approaches to interpreting and choosing the best treatments in network meta-analyses. Syst Rev. 2017;6:79. doi: 10.1186/s13643-017-0473-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Eachempati P., Kumbargere Nagraj S., Kiran Kumar Krishanappa S., Gupta P., Yaylali I.E. Home-based chemically-induced whitening (bleaching) of teeth in adults. Cochrane Database Syst Rev. 2018 doi: 10.1002/14651858.CD006202.pub2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Soares D.G., Basso F.G., Hebling J., de Souza Costa C.A. Concentrations of and application protocols for hydrogen peroxide bleaching gels: effects on pulp cell viability and whitening efficacy. J Dent. 2014;42:185–198. doi: 10.1016/j.jdent.2013.10.021. [DOI] [PubMed] [Google Scholar]
- 56.Markowitz K. Pretty painful: why does tooth bleaching hurt? Med Hypotheses. 2010;74:835–840. doi: 10.1016/j.mehy.2009.11.044. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Supplementary material









