Skip to main content
Diagnostics logoLink to Diagnostics
. 2026 Sep 6;16(17):2865. doi: 10.3390/diagnostics16172865

Artificial Intelligence for Sleep Bruxism Detection: A Technical Pipeline Synthesis

Özge Çekirge 1
Editor: Kaan Orhan1
PMCID: PMC13564490  PMID: 42739295

Abstract

Sleep bruxism is the rhythmic or nonrhythmic masticatory muscle activity that occurs during sleep. It is conventionally diagnosed by a rule-based algorithm. Candidate muscle bursts are flagged against a fixed electromyographic (EMG) amplitude threshold, classified by duration into phasic, tonic, or mixed episodes, and counted into an hourly rate compared against a second threshold. These thresholds generalize poorly across individuals and nights. This has motivated a growing body of artificial intelligence (AI) research aimed at replacing one or more stages of this classical algorithm with a learned decision function. This paper synthesizes that literature as a technical pipeline across 28 reviewed AI-based sleep bruxism studies, identified through Web of Science and Google Scholar searches for “sleep bruxism” combined with artificial intelligence terms, conducted in 2026, and screened for a learned, trained classification approach reported in English. These accuracy figures should be read alongside several complicating factors. Ten studies draw on the same subset of two patients from a public polysomnography database. Several studies framed as sleep bruxism detection were validated only on awake, simulated grinding. In addition, the ground-truth labels these classifiers are trained against remain contested within the field’s own consensus literature. The paper’s contribution is therefore twofold. It assembles the pipeline synthesis itself, spanning ground truth, recording modality, preprocessing, feature extraction, and architecture, as a technical reference not previously brought together in this form. It also uses that synthesis to set a concrete research agenda, including validation benchmarks across nights and across subjects, a shift from binary toward continuous severity targets, and an extension of detection into real-time treatment devices and clinical prognosis.

Keywords: sleep bruxism, artificial intelligence, machine learning, electromyography, signal processing, diagnosis

1. Introduction

Sleep bruxism is defined as a recurring pattern of jaw muscle activity that manifests as grinding or clenching of the teeth during sleep [1]. It affects a considerable share of the adult population. Reported prevalence varies widely. According to a systematic review and meta-analysis of research published between 2003 and 2023, global sleep bruxism prevalence is approximately 21%, while the occurrence of sleep bruxism based on polysomnography is approximately 43% [2]. Its authors trace this wide-ranging estimate to differences in the populations studied, the diagnostic thresholds applied, and how each underlying study was designed. Its clinical footprint extends well beyond a cosmetic nuisance. Documented consequences include tooth wear, fractured teeth or implants, temporomandibular joint disorders, and morning headache, together with masticatory muscle pain, masseter muscle hypertrophy, and measurable reductions in quality of life related to oral health, along with poorer sleep quality [3]. Sleep bruxism’s underlying cause remains incompletely understood. There has been a progressive shift in understanding its origin, from peripheral explanations toward central mechanisms involving a central pattern generator and the autonomic nervous system [4]. Recent biopsychosocial models echo this shift, framing bruxism episodes as a discharge of accumulated neural load shaped by the interaction of biological, psychological, and social factors [5]. This is consistent with a demonstrated genetic component of still unclear mechanism and stress-related psychosocial contributors to bruxism [6]. Sleep bruxism is also not a clinically isolated phenomenon. Nearly half of a large obstructive sleep apnea cohort met criteria for sleep bruxism [7], and it is further associated with several other sleep-related conditions, including restless legs syndrome and REM behavior disorder [8], as well as conditions outside the sleep disorder spectrum such as chronic migraine and attention deficit/hyperactivity disorder [9].

Polysomnography (PSG) with concurrent audiovisual recording is widely treated as the reference standard for sleep bruxism diagnosis [10], as it is for most sleep disorders more broadly [11]. It is, however, expensive and of limited availability [1]. It also requires a dedicated sleep laboratory and specialized expertise, and burdens patients with a single, unfamiliar night of monitoring that may not represent their habitual sleep [12,13]. Waiting times for such testing have been reported to extend to up to a year and a half in some healthcare systems [12]. Beyond PSG, researchers have explored a range of sensing modalities to work around these limitations, including electroencephalographic, accelerometer based, acoustic, and other multimodal sensing approaches. Electromyography, however, remains the dominant and most extensively studied signal source.

Portable electromyography offers a more accessible alternative to PSG. A scoping review identified 78 studies using ambulatory EMG devices for sleep bruxism assessment published between 1977 and 2020, with the number of such studies growing markedly over the past decade [14]. Its data quality, however, is contested, because ambulatory EMG loggers structurally lack simultaneous audio or video confirmation. Even full polysomnography is not immune to this same problem once its audio-video support is withheld. Removing audio-video support from an otherwise rich portable polysomnographic recording overestimated bruxism activity by nearly a quarter and reversed the diagnosis in two of ten participants [15]. This was independently confirmed elsewhere, with the overestimation traced to coughing, swallowing, and yawning, producing masseter electromyographic bursts mistaken for genuine bruxism once video confirmation was unavailable [16]. The problem compounds further when the recording is reduced to electromyography alone, as in genuine ambulatory devices: electromyography only setups produced six to twenty times more false positives than any polysomnography-based configuration when evaluated against the same full audiovisual standard [17]. The mechanism each of these studies converges on is orofacial and other nonbruxism motor activity that looks like bruxism bursts, not a flaw specific to any one device.

Beyond these orofacial and movement-related confounds, the electromyographic signal can also be shaped by comorbid physiology, as in obstructive sleep apnea. Sleep bruxism and obstructive sleep apnea appear to share an arousal related mechanism, so apnea-related arousals can themselves shape the timing and intensity of recorded masticatory bursts [18]. This could plausibly complicate the separation of primary bruxism activity from apnea-driven muscle activation within the same recording, although this specific inference has not been directly tested. Comorbid physiology is not the only population-specific limitation. Portable instrumentation validated in the general adult population may not generalize to every patient group, such as edentulous patients, whose severity cutoffs have never been validated in a larger sample [19]. At a more basic technical level, surface electromyography captures the superimposed electrical activity of many motor units simultaneously, and the resulting signal is inherently noisy. Crosstalk from adjacent muscles, movement of the electrode relative to the skin, and power line interference all contaminate the recording before any classification step begins [20]. Given the confound landscape just described, separating true bruxism from this envelope of overlapping signals is a genuine signal discrimination problem rather than a matter of setting a single amplitude cutoff correctly. This problem is not unique to electromyography. Comparable confounds and noise sources are reported across the other modalities. This makes careful preprocessing, including filtering, artifact rejection, and confound removal, a necessary step throughout the detection pipeline regardless of the sensing approach.

Recording modality is not the only axis of difficulty. Scoring itself has long posed a separate challenge. Bruxism patterns were traditionally extracted manually according to a fixed EMG amplitude threshold, most often set at 20% of the patient’s maximum voluntary contraction (MVC), although a multiple of background muscle activity has also been used as an alternative convention [10,14]. Even under the MVC criterion, however, trained scorers could not fully agree on how to classify episodes into the field’s own phasic, tonic, and mixed categories. One early reliability study reached only moderate agreement, with scorers concurring on roughly 62–63% of episodes [21]. Manual annotation was, in addition, highly labor-intensive, creating a clear need for automation. Early automated approaches followed the same logic, simply adopting one of these same fixed amplitude thresholds. However, a fixed threshold of this kind still cannot fully account for genuine night-to-night variability. Across two consecutive nights of laboratory polysomnography, one study reported a coefficient of variation of approximately 30% in the bruxism episode index, with over a third of patients shifting severity category entirely between the two nights [22]. A separate three-night home polysomnography study found even greater fluctuation, with a mean individual coefficient of variation of 50.7% and the number of patients meeting diagnostic criteria rising from 2 to 7 out of 16 across the three nights [23]. Together, this phenotype disagreement and night-to-night variability point to the need for an updated approach to thresholding. This is the specific technical difficulty that motivates a learned, multivariate approach in place of a fixed threshold.

These combined sources of variability point to a clear need for a more generalizable approach than a fixed amplitude threshold of either kind can offer. Artificial intelligence has been proposed as just such a remedy, on the reasoning that a learned, adaptive decision boundary should discriminate genuine bruxism from these confounds more robustly than a fixed threshold. This proposal is not unique to bruxism. Automated, learned approaches to sleep disorder detection have proliferated across sleep medicine generally, from wearable sensor-based sleep stage classification [13] to broader efforts spanning multiple sleep disorder categories [24]. Bruxism detection is therefore best understood as one instance of this broader movement toward learned, data driven sleep diagnostics, rather than an isolated technical proposal. A growing number of bruxism-specific studies published over the past decade test this proposition using electromyographic, electroencephalographic, and increasingly multimodal signals. One brief review, by Gan and Yao [25], has already attempted to synthesize this bruxism-specific literature, observing that most existing studies involve modest cohort sizes and lack large-scale, multicenter clinical validation, and that the field has not yet converged on a unified evaluation standard. Unlike the present analysis, however, it does not examine specifically why these limitations hold or whether they affect the interpretation of individual findings it cites.

This paper undertakes that examination, tracing the technical pipeline shared, explicitly or implicitly, across 28 artificial intelligence-based sleep bruxism detection studies, from recording modality through preprocessing, feature extraction, and model architecture to the reported detection decision. It offers the pipeline synthesis itself as a technical reference that the existing literature has not yet provided in depth or brought together in one sustained treatment, organized so that future work can be positioned against a shared framework. It also offers a comparison of each reviewed study across three dimensions: the ground truth criterion the classifier was trained against, the validation context under which it was tested, and the specific classification target it was asked to predict.

2. Technical Pipeline Synthesis

This review was designed as a structured narrative technical synthesis of artificial intelligence approaches to sleep bruxism detection, comparing how existing studies implement successive stages of the detection pipeline. The predefined dimensions of comparison were ground truth, recording modality, preprocessing, feature extraction, model architecture, classification target, sample size, reported performance, and validation context.

Primary studies were identified through a Web of Science search in early 2026 using the query “sleep bruxism” AND (“machine learning” OR “artificial intelligence” OR “deep learning” OR “classifier” OR “EMG” OR “EEG” OR “wearable”), which returned 453 records; 21 of the 28 studies discussed here were identified this way. The same query was used as a supplementary search on Google Scholar, and reference lists of reviews and included studies were examined through citation chaining to identify the remaining seven primary studies. The search prioritized the literature from the last decade, although earlier foundational studies found in this way were also retained.

Studies were retained if they addressed sleep bruxism detection or a technical system explicitly presented in that context, applied a trained classifier, reported original research in English and provided sufficient methodological information for technical extraction. Studies were excluded if they were the following: (i) review articles, (ii) without a trained classifier, (iii) applying machine learning only to treatment calibration, or (iv) not based on device or sensor recordings for detection, although such studies could still be cited for contextual discussion. The final primary set comprised 28 studies. Studies using awake or simulated tasks were retained but identified separately by validation context; this does not treat these conditions as equivalent to naturally occurring sleep bruxism.

Because these studies differed substantially in ground truth definition, sensing modality, classification target, and validation context, their reported performance values were not treated as estimates of a common diagnostic task and were not quantitatively pooled. Instead, each study was mapped onto the same technical pipeline and interpreted within its stated validation context. Figure 1 summarizes this framework.

Figure 1.

Figure 1

AI-based sleep bruxism detection pipeline, synthesized from the reviewed literature. AI-based sleep bruxism detection pipeline, synthesized from the reviewed literature. The ground truth definition chain shown follows Lavigne et al. [10], Lobbezoo et al. [1], Lobbezoo et al. [26], Manfredini et al. [27], and Verhoeff et al. [28].

2.1. Ground Truth: The Clinical Scoring Framework AI Systems Are Trained Against

Every classifier discussed in this review is trained to reproduce an operational definition of a bruxism event, and that definition has a documented history of instability. Lobbezoo et al. [1] unified three competing verbal definitions of bruxism and adopted Lavigne et al.’s [10] criteria, EMG amplitude-based burst detection classified by duration into phasic, tonic, or mixed episodes and then expressed as a rate per hour of sleep, as the operational basis for a “definite” diagnosis, while noting that reliable diagnostic tools remained scarce. Five years later, the same author group reversed course. Lobbezoo et al. [26] explicitly recommended against using standard cutoff points, such as a fixed percentage of maximum voluntary contraction, to determine whether bruxism is present or absent at all, arguing that the widely cited thresholds had been derived from a super-selected study sample and contained an element of circularity, and proposing instead that masticatory muscle activity be assessed along a scale, rather than treated as simply present or absent. Manfredini et al. [27] reaffirmed this position five years later, describing the definitional project as still the work in progress and citing Thymi et al. [14] as empirical confirmation of the underlying instability. Verhoeff et al. [28], reporting the same 2024 IADR consensus meeting from which Manfredini et al.’s [27] note draws, describe two further revisions agreed there. They removed the qualifier “in otherwise healthy individuals” from both the sleep and awake bruxism definitions, and replaced the earlier hierarchical possible/probable/definite grading scheme with three nonhierarchical assessment modes, subject-based, clinically based, and device-based, explicitly described as potentially assessing “different aspects of bruxism” rather than forming a validity ladder. As of the 2024 consensus, there is no single accepted numeric threshold for a bruxism event; the field’s current standard is this nonhierarchical, multi-mode assessment framework itself. The fact that the field’s core definitional group revised its grading terminology again as recently as 2024 [28] indicates the instability this section documents is an ongoing revision process, not a closed historical episode.

This definitional instability is compounded by a second, more granular layer of inconsistency in the numeric amplitude threshold used to flag a candidate bruxism burst, before any episode counting or rate calculation begins. A segment of masseter, and sometimes temporalis, electromyographic activity is flagged as a candidate burst when its amplitude exceeds a chosen reference value, commonly a percentage of maximum voluntary contraction or a multiple of background muscle activity [29]; consecutive bursts within roughly three seconds of each other count as one episode, with burst duration then determining the type (three or more bursts of 0.25–2.0 s is phasic, a single burst over two seconds is tonic, a mix of both is mixed [10]), counted across a night and expressed as a rate per hour of sleep. It is this first, burst-level threshold, not the episode or rate stages that follow it, which varies so inconsistently across the literature: Thymi et al. [14] catalogued 25 distinct thresholds used to define a scorable bruxism event across 78 studies, ranging from 3% to 50% of maximum voluntary contraction, alongside incompatible multiples-of-background or absolute microvolt alternatives. These differences are so incompatible that Raja et al. [30] conclude results across devices are not generalizable: polysomnography conventionally requires an amplitude above 20% of maximum voluntary contraction, the BiteStrip device uses a threshold of 30%, standard electromyography requires an amplitude more than twice the baseline, and the Bruxoff device uses a threshold of 10%.

This instability is not confined to the clinical and device validation literature. It propagates directly into the ground-truth labels of the artificial intelligence classifiers this review covers. Ten of the twenty-eight studies reviewed here [31,32,33,34,35,36,37,38,39,40] all draw their bruxism-positive training examples from the same two patients in the public Cyclic Alternating Pattern Sleep Database (CAP), without independently stating, or apparently having access to, the original diagnostic criterion behind that patient level label. One of these ten [38] supplements its negative class with a second public database, whose bruxism-specific arousal annotations were assigned by certified sleep technologists, but whose exact scoring criterion for distinguishing a bruxism arousal from the ten other arousal types it also annotates is likewise undocumented in accessible sources. A further study [41] sources its bruxism cases from a different database entirely but is equally silent on the criterion used to diagnose them. Among the studies that instead generate their own recordings, task instruction substitutes for any physiological confirmation in roughly a third of the remaining set [42,43,44,45,46,47,48,49]. Each labels a segment as bruxism because the subject was instructed to clench or grind on cue, and several explicitly note that the intensity, duration, or timing of this simulated activity was not controlled across participants. Two studies [50,51] instead specify a target intensity for this cued grinding activity, 10% of maximum voluntary contraction, matching the figure the same two studies cite from the literature as characteristic of natural bruxism bursts. A third study by the same group [52] moves to real sleep recordings rather than a cued task, but documents only a minimum event duration of eight seconds, calibrated against each subject’s own waking maximum voluntary contraction reference. The criterion that subsequently assigns a qualifying event to one of its output classes is not stated. Where bruxism is not cued at all, ground truth is defined differently again. One study [53] flags a candidate event directly from continuous, unprompted recordings whenever amplitude exceeds 10% of maximum voluntary contraction together with a concurrent heart rate rise. Only two studies anchor their ground truth to an established clinical standard. One [54] scores rhythmic masticatory muscle activity against criteria consistent with AASM and ICSD-3, confirmed by audio-video review, and the other [55] uses polysomnography scored by a specialist physician to AASM criteria with the same audio-video confirmation. One further study [56] relies on self-report and clinical examination alone, a lower diagnostic tier than any instrumentally confirmed label in this review. Two further studies [57,58] repurpose the teeth-grinding category of an unrelated public nonverbal sound corpus as their bruxism proxy, with no bruxism patients, clinicians, or even instructed simulation involved at all.

Every architecture this review covers is, in effect, trained against one of these disparate ground truth definitions rather than a single, stable target. These definitions range from inherited database labels and task-instructed proxies to idiosyncratic thresholds or genuine clinical scoring, and that diversity is the baseline every accuracy figure in this review must be read against.

2.2. Recording Modalities

Polysomnography has long functioned as the benchmark method for diagnosing sleep bruxism, but its cost, time, and effort have motivated a search for alternatives, spanning non instrumental options such as patient self-report and instrumental, device-based methods [59]; the recording modalities surveyed here belong to that second category. Electromyography was the natural starting point for simplifying polysomnography into something portable, since masseter and temporalis EMG bursts are themselves the core diagnostic signal full polysomnography already scores, so recording just that one channel was the most direct way to strip the gold standard down to something usable at home. Within this device literature, portable, non-wearable devices appeared before 2000, with wearable devices emerging afterward and eventually dominating [60]; a meta-analysis found the Bruxoff EMG/ECG holter, also used by Castroflorio et al. [53], to have the best diagnostic validity among portable options [61], and a more recent review of the commercial landscape, spanning holters such as Bruxoff, single-channel devices such as BiteStrip, biofeedback systems such as GrindCare, and multichannel systems such as NOX T3 and Sleep Profiler, reached a similar conclusion: These devices show promise as a screening tool but still require further validation before substituting for polysomnography [62]. This progression from portable to wearable to increasingly multimodal ambulatory systems mirrors a broader shift in sleep disorder diagnosis generally [13,63]. The newer modalities this review catalogues beyond electrophysiological signals, inertial, acoustic, imaging, and textile-based follow the same trajectory within bruxism specifically: Figure 2 plots this against publication year, with EMG-based detection dating back to 2013, EEG-based work dating back to 2019, and every newer modality study published in 2021 or later.

Figure 2.

Figure 2

Historical emergence of recording modalities across the 28 reviewed AI-based sleep bruxism detection studies. The dashed line marks 2020, separating the earlier electrophysiological literature from the newer modalities that emerged in 2021 and later. Dot color and background shading distinguish electrophysiological modalities from other modalities.

Electrophysiological signals dominated the reviewed literature, but neither EMG nor EEG showed a standardized recording configuration. EMG ranged from single- or bilateral-masseter recordings to multimuscle and EMG/ECG combinations [39,42,43,49,50,51,53,54], whereas EEG ranged from selected one- or two-channel derivations to multichannel montages of up to 27 channels [31,32,33,34,35,36,37,38,40,41]. The distinction is nevertheless clinically important. EMG configurations overlap with sensing principles already used in portable clinical devices, whereas EEG-based bruxism detection remains confined to study-specific research configurations.

Beyond electrophysiological sensing, the literature has expanded toward motion, physiological, acoustic, imaging, pressure, and textile-based approaches. These include accelerometry and inertial sensing [47,52,54], intraoral pressure sensing [44], PPG/HRV and fNIRS [45,55], ultrasonography [56], contactless radar [46], audio [57,58], and textile strain sensing [48]. Audio-based bruxism recognition is further evidenced as a real, technically distinct signal class by a broader, nonbruxism-specific sleep sound system that treats bruxism as one input category among others [64]. Unlike EMG, however, each of these modalities is represented by only one or a small number of studies and remains at the proof-of-concept stage rather than clinical deployment.

The modality landscape therefore shows expansion rather than convergence. Newer sensors broaden the range of measurable correlates of bruxism, but none has yet displaced EMG as the clinically established portable signal source.

Several additional studies sit adjacent to this review’s 28 primary studies. Four kinds of evaluated acquisition hardware for audio, combined inertial and acoustic sensing, or piezoelectric bite force sensing, without a trained bruxism classifier [65,66,67,68], and two applied machine learning to treatment calibration rather than bruxism detection [69,70]. They nevertheless indicate that, for these emerging modalities, sensor development and deployment remain active technical problems upstream of classification.

2.3. Signal Preprocessing

Preprocessing strategies differed substantially across recording modalities, reflecting the distinct noise sources and signal characteristics of EEG, EMG, and alternative sensing approaches. Accordingly, the studies are considered separately by modality below. Across modalities, preprocessing primarily served three functions: frequency restriction, artifact suppression, and signal normalization or segmentation.

2.3.1. Electroencephalography

EEG preprocessing followed three broad strategies: fixed cutoff filtering, frequency band decomposition, and minimal or no additional frequency-domain filtering.

Fixed cutoff filtering was relatively consistent. Three studies used a 25 Hz lowpass cutoff, implemented using a Hamming windowed, linear phase FIR filter, or closely related high order filters [32,36,37]. One further study applied a 1–35 Hz FIR bandpass filter with a Kaiser window, followed by zero mean/unit variance normalization and 30 s segmentation [40].

Other studies incorporated frequency separation directly into preprocessing rather than applying a single passband. These approaches included DB5 wavelet decomposition over 1–30 Hz [31], decomposition into nine sub-bands spanning delta to gamma [34], and independent component analysis combined with maximum overlap discrete wavelet transform for selective artifact suppression [35]. In contrast, three studies applied no comparable frequency-domain filtering: one relied on downsampling, segmentation, and z score normalization [33], another used externally preprocessed EEG data [41], and a third primarily addressed dataset harmonization and class imbalance rather than signal denoising [38].

Thus, EEG preprocessing shows some convergence around restricting or decomposing the analyzed frequency range, but no common pipeline. Filtering, sub-band decomposition, normalization, artifact handling, and even the extent of preprocessing differ across studies, making preprocessing another source of methodological variation when model performance is compared.

2.3.2. Electromyography

Bandpass filtering was the dominant EMG preprocessing strategy, although cutoff frequencies and implementation varied substantially. Five studies used either an explicit bandpass filter or paired highpass and lowpass filtering. Reported passbands ranged from 10–400 Hz during acquisition [53], through 10–500 Hz [50,51] and 30–500 Hz [43] to a narrower 150–300 Hz digital stage following 30–1000 Hz acquisition filtering in one study [42]. Notch filtering at 50 Hz was additionally used in two of these pipelines [42,43].

The remaining studies used single-sided frequency restriction. A 25 Hz lowpass filter was applied to a combined EMG/ECG recording [39], whereas an EMG reference channel accompanying an intraoral pressure sensor was highpass-filtered at 20 Hz [44]. One further study applied only a narrowband notch filter to remove power line interference from its EMG and sound channels, with no lowpass, highpass, or bandpass stage [49]. Amplitude normalization against maximum voluntary contraction was also used in one EMG study [51].

Overall, EMG studies agree on the need for frequency-based noise suppression but not on a standardized passband or processing sequence. The substantial variation in cutoff frequencies, filter order, acquisition stage versus digital filtering, and normalization means that downstream feature sets are derived from signals that have undergone materially different transformations. Reported classifier performance therefore reflects not only differences in algorithms but also differences in the signal presented to them.

2.3.3. Other Modalities

Preprocessing outside EEG and EMG was necessarily modality-specific, with each modality’s approach developed independently across audio, motion, physiological, pressure, and imaging signals.

Audio studies used either amplitude-based silence removal and overlapping segmentation [57] or simple 1 s windowing without frequency-domain filtering [58]. Motion-based systems showed greater processing diversity: accelerometry used 5 Hz highpass filtering and participant-specific normalization [52]; a nine-axis IMU combined smoothing, rectification, 10 Hz lowpass filtering, and z score normalization [47]; textile strain sensing used 10 s segmentation without an explicit denoising stage [48]; and millimeter wave radar used range FFT, chirp accumulation, phase extraction, and phase differencing to isolate facial micromotion from larger respiratory and head movements [46].

Modality specific physiological processing was similarly distinct. fNIRS signals were bandpass-filtered at 0.01–0.1 Hz, screened for motion and optode artifacts, and converted to oxy/deoxyhemoglobin concentrations [45], whereas PPG preprocessing combined DC offset removal, 20 Hz lowpass filtering, smoothing, and peak detection for subsequent heart rate variability estimation [55]. For direct jaw sensing, intraoral pressure data underwent resampling and baseline wander removal [44], while mandibular movement data were segmented through a proprietary pipeline with artifact exclusion [54]. Ultrasonography required a different workflow altogether, involving manual masseter delineation and image standardization before radiomic analysis [56].

The heterogeneity in this group therefore reflects differences in sensing principle more than competing preprocessing conventions. Unlike EEG and EMG, where several studies can be compared within broadly similar signal processing frameworks, most alternative modalities have been investigated only through modality-specific pipelines. Consequently, preprocessing choices in these studies are difficult to separate from the sensing technology itself when interpreting downstream classification performance.

2.4. Feature Extraction

Feature extraction was more heterogeneous than preprocessing, but the variation was strongly modality-dependent. Across the reviewed studies, representations ranged from conventional spectral and statistical descriptors to nonlinear measures, multimodal feature banks, event-derived summaries, and modality-specific representations. The sections below therefore retain the signal specific organization used for preprocessing while emphasizing recurring feature design strategies rather than individual study pipelines.

2.4.1. Electroencephalography

EEG studies followed three broad feature extraction strategies: standalone spectral measures, nonlinear or recurrence-based representations, and mixed feature vectors combining several feature families.

A normalized Welch power spectral density estimate was the most repeated single spectral representation, appearing in three studies [32,36,37]. Nonlinear approaches were more diverse. Katz fractal dimension and the Lyapunov exponent, computed across nine sub-bands and five sleep stages, formed one such set, described as the first application of this combination to bruxism classification, with the resulting ninety candidate features narrowed to twenty [34]. A later study extended this combination with spectral entropy [40]. Other nonlinear representations included phase−amplitude and amplitude−amplitude coupling across frequency band pairs [35], and a recurrence-based approach reducing two EEG channels to recurrence rate and a single coupling index [41].

Two studies instead combined multiple feature families. One used a 28-feature vector containing spectral, time-domain, and nonlinear descriptors, including Sample Entropy [31]. Another combined spectral features, thirteen statistical measures, Hjorth parameters, and two entropy measures across wavelet derived bands, followed by group-wise feature ranking [33]. One further study computed features across its full multimodal montage rather than EEG alone, combining time-domain EMG statistics, EEG-derived spectral band power, and ECG-derived heart rate variability as per channel features [38].

In sum, EEG feature engineering shows substantial methodological diversity but no clearly dominant representation beyond the repeated use of spectral information. Nonlinear and mixed feature sets broaden the information extracted from EEG, yet the reviewed studies do not provide controlled, independent comparisons sufficient to determine whether these representations consistently outperform simpler spectral features. This limitation is particularly important because several EEG studies were evaluated on overlapping source data.

2.4.2. Electromyography

EMG feature extraction was comparatively more concentrated around conventional signal descriptors, although several studies incorporated autoregressive, wavelet-based, or event-derived representations.

A normalized Welch power spectral density estimate was used in one combined EMG/ECG study [39]. Two studies computed autoregressive or wavelet-derived descriptors: twelve autoregressive coefficients combined with eight wavelet Shannon entropy values, most accurate when combined rather than used singly [51], and thirteen-dimensional Mel frequency cepstral coefficients computed via a 100 ms Hamming window with a 50 ms shift, borrowed from automatic speech recognition and described by its authors as a novel application to this signal type [49]. Three further studies built comparably sized candidate feature banks but treated reduction very differently, from retaining 95% of variance through principal component analysis, to skipping reduction entirely, to narrowing a thirteen-feature bank to five via regression-based ranking [42,43,50].

A final study replaced raw signal statistics entirely with event-derived metafeatures computed from a rule-based bruxism detector’s own output, such as contraction and episode counts, contraction type, mean EMG amplitude, and sleep duration [53].

Overall, EMG studies show greater convergence than EEG around compact handcrafted representations, particularly amplitude- and distribution-based descriptors. However, differences in preprocessing, feature selection procedures, and validation context make it difficult to attribute performance differences specifically to feature choice. The evidence therefore supports the feasibility of several EMG feature families, but not the superiority of any one representation for naturally occurring sleep bruxism.

2.4.3. Other Modalities

Feature extraction in the remaining modalities was necessarily sensor-specific and therefore showed little cross-study standardization.

Audio-based approaches diverged: One computed a small set of explicit acoustic features, including Mel frequency cepstral coefficients, zero crossing rate, RMS energy, and spectral centroid [57], while the other dispensed with handcrafted features altogether, feeding a Mel-scaled spectrogram directly into a deep learning model as a two-dimensional image input [58]. Motion-based systems derived kinematic or statistical measures from accelerometers, inertial measurement units, or radar signals, with features reflecting movement amplitude, temporal variation, orientation, or frequency content [46,47,52]. A textile strain sensor study instead extracted no handcrafted features at all, with its deep learning model performing feature extraction internally via a ResNet component [48]. Physiological sensing approaches produced modality-specific variables: PPG-derived heart rate variability descriptors, reduced from ninety-six candidate features to ten via principal component analysis and Fisher score ranking [55], and a twelve feature statistical and spectral summary, including mean, peak, and dominant frequency, computed from the converted oxy/deoxyhemoglobin fNIRS signal, with five competing reduction techniques tested against one another rather than a single committed pipeline [45].

Other studies derived features directly from the physical phenomenon being measured. Jaw movement and pressure sensing systems diverged in specificity: One derived sixty-four candidate features per epoch from an intraoral pressure sensor, retaining five task-specific features each for its clench and grind detectors [44], while a separate study’s accessible text disclosed no specific feature set at all, describing only an undisclosed “feature generating module” [54]. The ultrasonography study extracted 2818 radiomic features from standardized masseter images, narrowing to just 6 through a variance threshold/SelectKBest/LASSO pipeline, the most extreme reduction ratio in the review [56].

The absence of a recurring feature family across these modalities should therefore not be interpreted as methodological inconsistency in the same sense as within EEG or EMG. Rather, feature design is largely determined by the sensing principle itself. The main evidential limitation is that most of these representations have been evaluated in only one or a small number of studies, so their apparent performance currently reflects proof of concept feasibility more than replicated superiority over established electrophysiological features.

A related limitation cuts across all three modality groups. Dimensionality reduction tracks the size of the original candidate feature set far more than any settled methodological consensus: The largest raw feature banks see the most aggressive reduction, while small, purpose-built sets are often kept intact. Which features survive, and by what criterion, remains a largely unreported degree of freedom behind any accuracy figure, and a systematic head-to-head comparison across feature families on a shared, independently collected dataset would meaningfully advance the field.

2.5. Architecture, Detection Decision, and Reported Performance: A Critical Synthesis

Classical machine learning methods, rather than deep learning, dominate the studies reviewed here. Eight studies commit to a single reported architecture without an internal comparison. These are decision trees in [31,32,39], random forest in [35,46], linear discriminant analysis in [44], a shallow neural network (MLP) in [51], and gradient boosted XGBoost in [54].

The remaining sixteen studies all involve more than one classifier, although they differ in what role that plurality plays. Thirteen simply compare several classifier types and report whichever wins, with a kernel-based or tree-based method prevailing in most cases. Support vector machines emerge as the best of four to ten classical alternatives in [33,34,56], random forest in [41,42], an artificial neural network in [50,52], a decision tree ensemble in [55], gradient boosted CatBoost in [47], and k-nearest neighbors in [45]. A further study compares just two classifiers, with support vector machines again winning over random forest [57]. Another compares three, with a radial basis function support vector machine winning over k-nearest neighbors and a shallow artificial neural network [40]. Three further studies draw on the same ten-classifier pool, but combine it differently. Ref. [38] keeps only the single best-performing classifier from the pool, linear discriminant analysis, while Refs. [36,37] instead take a majority vote across all ten. One further study combines a rule-based threshold detector with a small multilayer perceptron applied only to a secondary subtype classification task [53].

Beyond these classical and hybrid approaches, deep learning appears in some form in four of these studies [33,43,48,58]. Two of these, Refs. [33,43], directly compare deep learning architectures against classical alternatives on the same data. This means a two-dimensional convolutional neural network in the first case, and a convolutional network, long short-term memory model, and a recurrent neural network in the second. Of the four, Ref. [48] is the most architecturally complex, pairing a textile strain sensor garment with a hybrid architecture, SleepNet, which combines a BiLSTM encoder, Transformer-derived attention, and residual convolutional feature extraction, to classify six sleep states, including bruxism and both apnea types. It is the only architecture here to use attention.

Table 1, Table 2 and Table 3 collect the recording modality, feature family, architecture, classification target, sample size, best reported accuracy, and validation context for every artificial intelligence bruxism detection study discussed in this review, grouped by recording modality, so that the pattern-linking modality-to-architecture choice is visible directly rather than only once combined across studies. Two points of interpretation apply to these columns. First, the Best Accuracy column reproduces whichever headline metric each study reported, and these are not all the same metric. These quantities are not interchangeable, so the column should be read as a per study summary of each study’s own stated result, not as a directly comparable ranking across rows or across tables. Second, the Architecture column reports each study’s best-performing configuration where multiple architectures were compared internally, not necessarily the only architecture attempted.

Table 1.

EEG-based studies (n = 10).

Study Modality Feature Family Architecture Classification Target N Best Accuracy Validation Context
Piritu & Taralunga [38] EEG (CAP DB) Frequency band LDA Bruxism vs. nonbruxism 2 bruxism vs. 10 nonbruxism 80–82% Real sleep
Wang et al. [31] EEG (CAP DB) Time-domain statistical, spectral, nonlinear Decision Tree Bruxism vs. healthy 2 bruxism + 4 healthy 97.84% Real sleep
Bin Heyat et al. [32] EEG (CAP DB) Spectral Decision Tree Bruxism vs. healthy 2 bruxism + 6 healthy 81.25% Real sleep
Kidwai & Khan [41] EEG Nonlinear Random Forest Bruxism vs. normal vs. epilepsy vs. Alzheimer’s 10 bruxism + 30 others 98.82% Not specified
Dimitriadis et al. [35] EEG (CAP DB) Nonlinear Random Forest 7 sleep disorders + healthy 2 bruxism + 105 others 74% Real sleep
Khan et al. [33] EEG (CAP DB) Spectral, statistical, nonlinear Cubic SVM Bruxism vs. nonbruxism 2 bruxism + 4 normal 98.39% Real sleep
Tiwari et al. [34] EEG (CAP DB) Nonlinear SVM (RBF) Bruxism vs. healthy 2 bruxism + unspecified healthy 93.5% Real sleep
Tiwari et al. [40] EEG (CAP DB) Nonlinear SVM (RBF) Bruxism vs. healthy 2 bruxism + unspecified healthy 93% Real sleep
Bin Heyat et al. [36] EEG (CAP DB) Spectral 10 classifier ensemble Bruxism vs. healthy 2 bruxism + 6–7 healthy 94% Real sleep
Tripathi et al. [37] EEG (CAP DB) Spectral 10 classifier ensemble Bruxism vs. healthy 2 bruxism + 14 healthy 94% Real sleep

Table 2.

EMG-based studies (n = 8).

Study Modality Feature Family Architecture Classification Target N Best Accuracy Validation Context
Castroflorio et al. [53] EMG + ECG Event derived Threshold + MLP Bruxism vs. nonbruxism 21 bruxism + 21 healthy 99% Real sleep, home environment
Lai et al. [39] EMG + ECG (CAP DB) Spectral Decision Tree Bruxism vs. normal 2 bruxism + 7 healthy 97.21% Real sleep
Gul et al. [42] EMG, 3 postures Time-domain + PCA Random Forest Bruxism vs. nonbruxism 10 healthy 93.33% Awake/simulated
Martinot et al. [54] EMG Not specified XGBoost RMMA episodes vs. micro arousals vs. breathing rate jaw movement 67 suspected OSA patients Kappa = 0.799 Real sleep
Sönmezocak & Kurt [50] EMG Time-domain, spectral → top 5 ANN Relaxed vs. clenching vs. grinding vs. fatigue/pain 20 98.8% Awake/simulated
Sönmezocak & Kurt [51] EMG Statistical, nonlinear MLP Relaxed vs. clenching vs. grinding vs. fatigue/pain 10 100% Awake/simulated
Minakuchi et al. [49] EMG + sound Spectral HMM Bruxism with vs. without tooth contact vs. nonbruxism 12 healthy F = 0.83, F = 0.43, F = 0.94, Awake/simulated
Ishtiaq et al. [43] EMG Time-domain RNN Bruxism vs. nonbruxism 30 99% Awake, supine simulated

Table 3.

Other recording modalities (n = 10, plus one study cross-listed from Table 2).

Study Modality Feature Family Architecture Classification Target N Best Accuracy Validation Context
O’Hare et al. [44] Intraoral MEMS pressure Time-domain statistical, spectral LDA Bruxism vs. nonbruxism events 8 82.2% Awake/simulated
Fatima et al. [45] fNIRS Time-domain, spectral → LDA kNN Bruxism vs. rest vs. talking vs. chewing 10 92.07% Awake/simulated
Shen et al. [46] mmWave radar Time-domain, spectral, statistical Random Forest Grinding vs. no grinding 3 96.1% Awake/simulated
Eris et al. [55] PPG + HRV Time-domain, statistical → Fisher Decision Tree Ensemble Bruxism vs. control 17 (12 sleep apnea + 5 healthy) 92.7% Real sleep
Martinot et al. [54] Jaw movement sensor (MJM) Not specified XGBoost RMMA episodes vs. micro arousals vs. breathing rate jaw movement 67 suspected OSA patients 86.6% Real sleep
Schlaepfer et al. [47] 9-axis IMU None—preprocessed raw signal CatBoost Grinding vs. opening/closing vs. no movement 3 96% Awake/simulated
Takamura et al. [57] Audio Time-domain, spectral SVM Bruxism vs. snoring vs. breathing Not specified F = 0.894 Not specified
Orhan et al. [56] Ultrasonography + clinical variables Imaging-derived, questionnaire-derived SVM Bruxism vs. controls 78 bruxism + 24 controls AUC 0.986 (train), 0.848 (test) Clinically diagnosed bruxism vs. controls; not a temporal sleep/wake recording
Sönmezocak & Kurt [52] Single-axis accelerometer Time-domain, spectral ANN Relaxed vs. clenching vs. rhythmic grinding 5 99% Real sleep
Peruzzi et al. [58] Audio Spectral CNN Bruxism vs. other audio Not specified 79.13% (complete)/80% (quantised) Not specified
Tang et al. [48] Textile strain sensor array None—preprocessed raw signal SleepNet (BiLSTM + Transformer attention + ResNet) nasal breath, mouth breath, snoring, bruxism, CSA, OSA 7 healthy (+2 apnea patients) 98.6% (6 class, incl. bruxism); 95% new user few shot Awake/simulated (breathing/snoring on real sleep)

Eight of these ten EEG-based studies draw on the CAP database’s subset of two bruxism patients without qualification. A ninth study, ref. [38], combines this same subset with an additional, independently collected database, so its evidence is only partly not independent. The sole full exception is ref. [41], who source EEG from elsewhere entirely. Architecture choice here is also uniformly classical, with no deep architectures, making this simultaneously the review’s most database-dependent and architecturally conservative cluster.

As seen in Table 2, three of these eight EMG-based studies were validated on real overnight sleep. These are ref. [39], drawn from the same CAP database subset as Table 1, and refs. [53,54] recorded in participants’ home environment. Ref. [54] is cross-listed here because EMG was used only to validate its automated output against manual scoring. Its primary classifier input is actually the mandibular jaw movement (MJM) sensor shown in Table 3, so it counts once toward the review’s total of 28 rather than twice. The remaining five rely on awake, voluntary, or simulated grinding and clenching.

As seen in Table 3, this heterogeneous group of emerging sensing modalities is where architecture diversifies most, and where both of the review’s non-EEG/EMG deep learning studies sit. These are ref. [58]’s convolutional network and ref. [48]’s SleepNet. Sample sizes here also skew small, with two notable exceptions. Refs. [54,56], at 67 and 102 participants respectively, are large enough to support adequately powered validation, while the rest of this cluster remains at the scale of early proof-of-concept demonstrations.

Across all three tables, the review’s three central validity concerns map onto three separate modality clusters, each with its own defining weakness. These are database non-independence for EEG studies, scope conflation for EMG studies, and small, proof-of-concept sample sizes for the remaining modalities. All three concerns appear to some degree in every cluster, but each cluster is best characterized by just one.

A fourth, cross-cutting concern is what these classifiers actually predict, which the Classification Target column added to Table 1, Table 2 and Table 3 makes explicit in one place for the first time. Only 15 of the 28 studies frame their task as a straightforward binary discrimination between bruxism and a healthy or nonbruxism comparison state. Nine replace bruxism versus healthy with a multi-way split among bruxism’s own sub-behaviors, severity subgroups, or physiological correlates instead, e.g., jaw relaxed, clenching, and grinding states, tooth contact status, or nonbruxer/low-frequency/high-frequency severity subgroups, rather than bruxism-affected people versus healthy ones. Four more embed bruxism as a single class within a broader, nonbruxism-specific classification problem. Ref. [41] separates bruxism from epilepsy, Alzheimer’s disease, and normal controls in a single four-way task, Ref. [35] separates it from six other sleep disorders and a healthy control group in an eight-way task, Ref. [48] places it alongside unrelated sleep monitoring targets, and Ref. [57] separates it from snoring and breathing sounds in a three-way task. None of this means the remaining studies’ reported numbers are wrong, but it is why the Classification Target column matters for comparing them properly. Read in isolation, a bare accuracy figure makes more of these studies look directly comparable than they really are, adding to the database non independence and scope conflation concerns named above.

Table 1, Table 2 and Table 3’s recording modality column also shows that EEG-based and EMG-based studies together account for 18 of the 28 (64%), the field’s two most established recording modalities. The remaining 10 are spread thinly across several newer categories, most of which are single-study demonstrations with no independent replication yet.

Figure 3 extends this pattern to architecture choice itself, showing how many studies tried each classifier family as one candidate in an internal comparison, against how many report it as their final system, ordered from the simplest classical methods to deep learning. Classical machine learning accounts for 25 of the 28 studies and deep learning for 3.

Figure 3.

Figure 3

Architectures tried vs. architectures kept.

Every tried and kept classifier was confirmed directly against each study’s own text, and studies testing multiple named variants of the same family, such as Decision Tree and CART, are counted once per family. k-nearest neighbors is tried by 11 of the 28 studies but reported as the final architecture in only 1. Decision Tree variants are tried by 14 but kept in just 3. Support vector machine, tried by 14, is kept in 4. Naive Bayes and Logistic Regression sit at the extreme end of this pattern. Each is tried by roughly a quarter of the reviewed studies, most often as one candidate among several in a head-to-head comparison, yet neither is ever reported as the final architecture anywhere in this literature. This suggests they function more as a baseline reference point in these comparisons than as genuine contenders. Random Forest converts most efficiently among the heavily tried families and is also the single most common final architecture overall, kept in 5 of the 14 studies that test it, ahead of Shallow ANN/MLP/BPN and SVM, which tie at 4 each. The two studies that adopt a ten-classifier majority vote ensemble instead [36,37] sit outside this comparison logic entirely. Rather than testing candidates and keeping a single winner, the ensemble itself is the chosen strategy, so their perfect tried-to-kept ratio reflects a different research design rather than unusual predictive strength. Read together, the range of methods actually kept as final architectures is considerably narrower than the range tested, concentrated mostly in tree based and kernel-based methods.

Set against that concentration, the reported accuracies themselves offer little reassurance. They are uniformly high across the reviewed studies, ranging from 74% on an 8-way, nonbruxism-specific task to a perfect 100% on a four-way jaw state task (relaxed, clenching, grinding, and fatigue/pain), with the latter reported by Sönmezocak and Kurt [51] on a sample of ten participants performing awake, simulated bruxism, a sample size and validation context that itself warrants caution before treating the figure as representative. With most studies clustering between 80% and 99%, taken at face value, this range would suggest the detection problem is close to solved. Most studies report only a single held-out or cross-validated figure, but where a train test split is reported explicitly, the gap can be substantial. Orhan et al. [56] report a training data area under the curve of 0.986 that falls to 0.848 once evaluated on a held-out test set for their best-performing classifier, a rare direct window onto the overfitting that a single reported accuracy figure elsewhere in this literature could just as easily be concealing. A related caveat applies to Piritu and Taralunga [38]. The EEG-only figure of 80–82% shown in Table 1 is not the study’s highest reported number. A three-signal fusion of EMG, EEG, and ECG reaches 94% in one specific cross-dataset condition, yet the study’s own stated conclusion favors a two-signal EMG-and-EEG combination as the more consistently effective approach across all four train test conditions it tests, a reminder that even within a single study, which combination of signals counts as best depends on which condition is highlighted.

3. Discussion and Future Directions

3.1. Database Non-Independence, Scope Conflation, and Ground Truth Instability

A meaningful share of the evidence is not statistically independent. Ten studies draw at least partly on the same two patients from the CAP database, meaning that agreement in reported performance cannot be interpreted as ten independent replications. This is more consequential than small sample size alone because nominally separate studies repeatedly evaluate models on essentially the same biological source data.

Validation context is also not consistent across studies. Several investigations described as sleep bruxism detection were evaluated only on voluntary or simulated awake activity. Such experiments can demonstrate that a sensing modality and classifier distinguish instructed jaw states, but they do not establish detection of naturally occurring sleep bruxism, which can be shaped by sleep-specific physiology absent from an awake recording, including its documented associations with obstructive sleep apnea, restless legs syndrome, periodic limb movement, and REM behavior disorder. Reported performance should therefore be interpreted together with validation context rather than treated as directly comparable across awake simulation and natural sleep. Figure 4 summarizes this distinction alongside database dependence across the reviewed studies.

Figure 4.

Figure 4

Database dependence(a) and validation context (b) across the 28 reviewed studies.

Nor is the target variable itself standardized. As detailed in Section 2.1, the reviewed classifiers are trained against heterogeneous ground truths, including inherited database labels, instructed proxy tasks, study-specific amplitude criteria, and clinically confirmed sleep events. Only two studies anchor their labels to an established clinical standard with audiovisual confirmation, while the diagnostic criterion underlying the repeatedly used CAP cases is not reported. A model that achieves 97% accuracy against one study’s ten percent maximum voluntary contraction threshold and a model that achieves 93% accuracy against another study’s twofold baseline multiplication threshold cannot be straightforwardly ranked against each other, because they are not, strictly, predicting the same thing. This instability is not unique to this review. Bibliometric analysis identifies diagnostic criteria as one of the field’s most actively contested open questions, still drawing sustained citation attention as of 2021 [71].

These three issues, data reuse, validation outside natural sleep, and heterogeneous ground truth, limit the extent to which currently reported accuracy can be interpreted as evidence of generalizable sleep bruxism detection. The central methodological requirement for future work is therefore not simply higher accuracy, but independent validation on new participants under natural sleep conditions using an explicitly defined and clinically defensible reference standard.

3.2. What Artificial Intelligence Is Actually for and What Remains to Be Shown

The most clinically meaningful role for artificial intelligence in sleep bruxism detection is not simply to outperform a fixed threshold on the same dataset, but to distinguish genuine bruxism from physiologically or technically similar nonbruxism events. This capability has not yet been demonstrated directly by the reviewed AI studies. Existing evidence nevertheless shows why it matters. A non-AI algorithm designed to discriminate rhythmic masticatory activity from sixteen confounding oral tasks found coughing was the single most frequent source of false-positive detections even under controlled conditions, misclassified in three of eleven subjects [72]. This discrimination logic is not unique to bruxism, either: A narrative review of temporomandibular dysfunction diagnosis reports a directly comparable case, where an artificial neural network outperformed general dentists at diagnosing orofacial pain of nondental origin, including cases clinicians themselves had misdiagnosed as bruxism [73]. A more informative benchmark would therefore evaluate AI models on real sleep recordings containing clinically plausible confounds, explicitly benchmarked against these documented false positives [16,17], rather than primarily on clearly separated voluntary or simulated tasks.

Generalization should also be tested directly rather than inferred from within-sample accuracy. The reviewed literature provides little evidence of robustness across nights, participants, datasets, or recording environments, despite substantial night-to-night variability already documented in sleep bruxism assessment [22,23]. Future benchmarks should therefore prioritize independent participants, repeated-night recordings, and external or cross-dataset validation under natural sleep conditions. Real-world deployment should be evaluated separately, since performance obtained under controlled acquisition conditions may not persist in home monitoring environments, as shown by one earable classifier whose accuracy dropped substantially once it left the controlled setting [74].

A second opportunity concerns the definition of the prediction target itself. None of the 28 reviewed studies estimates sleep bruxism as a continuous variable; all use binary or discrete categories, with the closest precedent being a three-level frequency classification [53]. This contrasts with the field’s broader move away from rigid diagnostic cutoffs toward graded assessment [26]. Continuous severity estimation would not resolve the underlying ground truth problem, but it could reduce dependence on arbitrary binary thresholds and provide a more clinically informative target for future AI models.

Finally, translation beyond retrospective classification will require models that can operate within real-time clinical or wearable systems. Improved confound discrimination could make closed-loop biofeedback more reliable by reducing inappropriate interventions triggered by nonbruxism activity. However, deployment feasibility is rarely evaluated in the reviewed literature: Only one study explicitly targets embedded implementation [58], leaving computational requirements such as power, memory, and latency largely unaddressed. Lightweight and tinyML approaches therefore represent a practical next step, but not yet an established capability of current sleep bruxism AI systems.

What unites these directions is a shift from demonstrating high accuracy on an existing sample toward testing a specific, falsifiable claim about generalization, confound discrimination, deployment, ground truth definition, or clinical utility. The pipeline synthesis presented earlier gives each of these claims a shape specific enough to test, leaving what comes next to future work.

4. Conclusions

This paper asked whether artificial intelligence has resolved the well-documented generalization problem of fixed-threshold sleep bruxism detection. The answer is qualified: The case made for AI’s broader success in the literature rests on much thinner ground than its accuracy figures suggest. Ten of the twenty-eight studies reviewed here share the same two patients. Several were validated only on awake, voluntary clenching rather than sleep. In addition, the ground truth every classifier is trained against is itself unstable, a problem the field’s own consensus body has never fully resolved.

What this paper adds is not a new algorithm but a shared reference point. It offers a pipeline synthesis spanning ground truth, recording modality, preprocessing, feature extraction, and architecture, assembled here for the first time, and a framework that lets any future study be positioned against it rather than judged in isolation. A researcher can now be asked directly whether their data overlap with the same subset of two patients already used ten times over, whether their validation happened during sleep or simulated wakefulness, and which of the field’s competing thresholds defined their ground truth. The same synthesis points past these questions toward a fuller agenda. This includes benchmarks that test generalization across nights, subjects, and presumed etiology, a shift from binary labels toward continuous severity, and an extension of detection into real-time treatment and clinical prognosis. Taken together, this agenda is what it would actually mean for sleep bruxism detection to move from a fixed threshold to a learned boundary. Whether that boundary ends up drawn on evidence solid enough to trust is the question this paper leaves for the field to answer.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

This paper reports a literature synthesis and did not generate new primary data. All source studies discussed are cited in the reference list and are publicly available through their respective publishers or repositories.

Conflicts of Interest

The author declares no conflicts of interest.

Funding Statement

This work received no external funding.

Footnotes

Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

References

  • 1.Lobbezoo F., Ahlberg J., Glaros A.G., Kato T., Koyano K., Lavigne G.J., De Leeuw R., Manfredini D., Svensson P., Winocur E. Bruxism Defined and Graded: An International Consensus. J. Oral Rehabil. 2013;40:2–4. doi: 10.1111/joor.12011. [DOI] [PubMed] [Google Scholar]
  • 2.Zieliński G., Pająk A., Wójcicki M. Global Prevalence of Sleep Bruxism and Awake Bruxism in Pediatric and Adult Populations: A Systematic Review and Meta-Analysis. J. Clin. Med. 2024;13:4259. doi: 10.3390/jcm13144259. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Câmara-Souza M.B., De Figueredo O.M.C., Rodrigues Garcia R.C.M. Association of Sleep Bruxism with Oral Health-Related Quality of Life and Sleep Quality. Clin. Oral Investig. 2019;23:245–251. doi: 10.1007/s00784-018-2431-0. [DOI] [PubMed] [Google Scholar]
  • 4.Thomas D.C., Manfredini D., Patel J., George A., Chanamolu B., Pitchumani P.K., Sangalli L. Sleep bruxism: The past, the present, and the future—Evolution of a concept. J. Am. Dent. Assoc. 2024;155:329–343. doi: 10.1016/j.adaj.2023.12.004. [DOI] [PubMed] [Google Scholar]
  • 5.Manfredini D., Lobbezoo F. The Biopsychosocial Model of Bruxism. CRANIO®. 2026:1–9. doi: 10.1080/08869634.2026.2655187. [DOI] [PubMed] [Google Scholar]
  • 6.Caivano T., Felipe-Spada N., Roldán-Cubero J., Tomàs-Aliberas J. Influence of Genetics and Biopsychosocial Aspects as Etiologic Factors of Bruxism. CRANIO®. 2021;39:183–185. doi: 10.1080/08869634.2021.1904181. [DOI] [PubMed] [Google Scholar]
  • 7.Li D., Kuang B., Lobbezoo F., de Vries N., Hilgevoord A., Aarab G. Sleep Bruxism Is Highly Prevalent in Adults with Obstructive Sleep Apnea: A Large-Scale Polysomnographic Study. J. Clin. Sleep Med. 2023;19:443–451. doi: 10.5664/jcsm.10348. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Kuang B., Li D., Lobbezoo F., de Vries R., Hilgevoord A., de Vries N., Huynh N., Lavigne G., Aarab G. Associations between Sleep Bruxism and Other Sleep-Related Disorders in Adults: A Systematic Review. Sleep Med. 2022;89:31–47. doi: 10.1016/j.sleep.2021.11.008. [DOI] [PubMed] [Google Scholar]
  • 9.Balasubramaniam R., Manfredini D., Guan G. Sleep Bruxism: A Narrative Review of Current Concepts, Mechanisms, and Clinical Implications. J. R. Soc. N. Z. 2026;56:e70046. doi: 10.1002/snz2.70046. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Lavigne G.J., Rompre P.H., Montplaisir J.Y. Sleep Bruxism: Validity of Clinical Research Diagnostic Criteria in a Controlled Polysomnographic Study. J. Dent. Res. 1996;75:546–552. doi: 10.1177/00220345960750010601. [DOI] [PubMed] [Google Scholar]
  • 11.Palombini L.O., Assis M., Drager L.F., Mello L.I.L.D., Pires G.N., Zancanella E., Santos-Silva R. 2024 Position Statement on the Use of Different Diagnostic Methods for Sleep Disorders in Adults—Brazilian Sleep Association. Sleep Sci. 2024;17:e476–e492. doi: 10.1055/s-0044-1800887. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Vitazkova D., Kosnacova H., Turonova D., Foltan E., Jagelka M., Berki M., Micjan M., Kokavec O., Gerhat F., Vavrinsky E. Transforming Sleep Monitoring: Review of Wearable and Remote Devices Advancing Home Polysomnography and Their Role in Predicting Neurological Disorders. Biosensors. 2025;15:117. doi: 10.3390/bios15020117. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Mogavero M.P., Lanza G., Bruni O., Ferini-Strambi L., Silvani A., Faraguna U., Ferri R. Beyond the Sleep Lab: A Narrative Review of Wearable Sleep Monitoring. Bioengineering. 2025;12:1191. doi: 10.3390/bioengineering12111191. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Thymi M., Lobbezoo F., Aarab G., Ahlberg J., Baba K., Carra M.C., Gallo L.M., de Laat A., Manfredini D., Lavigne G., et al. Signal Acquisition and Analysis of Ambulatory Electromyographic Recordings for the Assessment of Sleep Bruxism: A Scoping Review. J. Oral Rehabil. 2021;48:846–871. doi: 10.1111/joor.13170. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Carra M.C., Huynh N., Lavigne G.J. Diagnostic Accuracy of Sleep Bruxism Scoring in Absence of Audio-Video Recording: A Pilot Study. Sleep Breath. 2015;19:183–190. doi: 10.1007/s11325-014-0986-9. [DOI] [PubMed] [Google Scholar]
  • 16.Smardz J., Wieckiewicz M., Michalek-Zrabkowska M., Gac P., Poreba R., Wojakowska A., Blaszczyk B., Mazur G., Martynowicz H. Is Camera Recording Crucial for the Correct Diagnosis of Sleep Bruxism in Polysomnography? J. Sleep Res. 2023;32:e13858. doi: 10.1111/jsr.13858. [DOI] [PubMed] [Google Scholar]
  • 17.Miettinen T., Myllymaa K., Muraja-Murro A., Westeren-Punnonen S., Hukkanen T., Töyräs J., Lappalainen R., Mervaala E., Sipilä K., Myllymaa S. Polysomnographic Scoring of Sleep Bruxism Events Is Accurate Even in the Absence of Video Recording but Unreliable with EMG-Only Setups. Sleep Breath. 2020;24:893–904. doi: 10.1007/s11325-019-01915-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Martynowicz H., Marschollek K., Nowacki D., Kanclerska J., Lachowicz G., Jodkowska A., Macek P., Frosztęga W., Madziarska K., Waliszewska-Prosół M. The Effect of the Arousal Threshold on Sleep Bruxism Intensity in Coexisting Sleep Apnea and Sleep Bruxism: A Polysomnographic Study. Sci. Rep. 2025;15:17044. doi: 10.1038/s41598-025-01687-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Orlandi D.B., Feldmann A., Polmann H., de Oliveira J.M.D., Pauletto P., Stefani C.M., Gonçalves T.M.S.V., de Luca Canto G. Bruxism in Completely Edentulous Patients: A Scoping Review. J. Oral Rehabil. 2025;52:1912–1921. doi: 10.1111/joor.14046. [DOI] [PubMed] [Google Scholar]
  • 20.Al-Ayyad M., Owida H.A., de Fazio R., Al-Naami B., Visconti P. Electromyography Monitoring Systems in Rehabilitation: A Review of Clinical Applications, Wearable Devices and Signal Acquisition Methodologies. Electronics. 2023;12:1520. doi: 10.3390/electronics12071520. [DOI] [Google Scholar]
  • 21.Gallo L.M., Lavigne G., Rompré P., Palla S. Reliability of Scoring EMG Orofacial Events: Polysomnography Compared with Ambulatory Recordings. J. Sleep Res. 1997;6:259–263. doi: 10.1111/j.1365-2869.1997.00259.x. [DOI] [PubMed] [Google Scholar]
  • 22.Hasegawa Y., Lavigne G., Rompré P., Kato T., Urade M., Huynh N. Is There a First Night Effect on Sleep Bruxism? A Sleep Laboratory Study. J. Clin. Sleep Med. 2013;9:1139–1145. doi: 10.5664/jcsm.3152. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Miettinen T., Myllymaa K., Hukkanen T., Töyräs J., Sipilä K., Myllymaa S. Home Polysomnography Reveals a First-Night Effect in Patients With Low Sleep Bruxism Activity. J. Clin. Sleep Med. 2018;14:1377–1386. doi: 10.5664/jcsm.7278. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Xu S., Faust O., Seoni S., Chakraborty S., Barua P.D., Loh H.W., Elphick H., Molinari F., Acharya U.R. A Review of Automated Sleep Disorder Detection. Comput. Biol. Med. 2022;150:106100. doi: 10.1016/j.compbiomed.2022.106100. [DOI] [PubMed] [Google Scholar]
  • 25.Gan C., Yao D. Proceedings of the 2025 5th International Conference on Internet of Things and Machine Learning, Nanchang, China, 16 May 2025. ACM; New York, NY, USA: 2025. Artificial Intelligence in Sleep Bruxism Diagnosis and Treatment; pp. 171–175. [Google Scholar]
  • 26.Lobbezoo F., Ahlberg J., Raphael K.G., Wetselaar P., Glaros A.G., Kato T., Santiago V., Winocur E., de Laat A., de Leeuw R., et al. International Consensus on the Assessment of Bruxism: Report of a Work in Progress. J. Oral Rehabil. 2018;45:837–844. doi: 10.1111/joor.12663. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Manfredini D., Ahlberg J., Lavigne G.J., Svensson P., Lobbezoo F. Five Years after the 2018 Consensus Definitions of Sleep and Awake Bruxism: An Explanatory Note. J. Oral Rehabil. 2024;51:623–624. doi: 10.1111/joor.13626. [DOI] [PubMed] [Google Scholar]
  • 28.Verhoeff M.C., Lobbezoo F., Ahlberg J., Bender S., Bracci A., Colonna A., Fabbro C.D., Durham J., Glaros A.G., Häggman-Henrikson B., et al. Updating the Bruxism Definitions: Report of an International Consensus Meeting. J. Oral Rehabil. 2025;52:1335–1342. doi: 10.1111/joor.13985. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Iber C., Ancoli-Israel S., Chesson A., Quan S.F. The AASM Manual for the Scoring of Sleep and Associated Events: Rules, Terminology and Technical Specifications. American Academy of Sleep Medicine; Westchester, IL, USA: 2007. [Google Scholar]
  • 30.Raja H.Z., Saleem M.N., Mumtaz M., Tahir F., Iqbal M.U., Naeem A. Diagnosis of bruxism in adults: A systematic review. J. Coll. Physicians Surg. Pak. 2024;34:1221–1228. doi: 10.29271/jcpsp.2024.10.1221. [DOI] [PubMed] [Google Scholar]
  • 31.Wang C., Verma A.K., Guragain B., Xiong X., Liu C. Classification of Bruxism Based on Time-Frequency and Nonlinear Features of Single Channel EEG. BMC Oral Health. 2024;24:81. doi: 10.1186/s12903-024-03865-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Heyat M.B.B., Lai D., Khan F.I., Zhang Y. Sleep Bruxism Detection Using Decision Tree Method by the Combination of C4-P4 and C4-A1 Channels of Scalp EEG. IEEE Access. 2019;7:102542–102553. doi: 10.1109/ACCESS.2019.2928020. [DOI] [Google Scholar]
  • 33.Khan A.A.S., Fattah S.A., Quamruzzaman M., Saquib M. Detection of bruxism using inverse discrete wavelet transformed reconstructed band limited EEG signals by group wise feature ranking. IEEE Access. 2024;12:88086–88110. doi: 10.1109/ACCESS.2024.3409441. [DOI] [Google Scholar]
  • 34.Tiwari S., Arora D., Nagar V. Supervised Approach Based Sleep Disorder Detection Using Non—Linear Dynamic Features (NLDF) of EEG. Meas. Sens. 2022;24:100469. doi: 10.1016/j.measen.2022.100469. [DOI] [Google Scholar]
  • 35.Dimitriadis S.I., Salis C.I., Liparas D. An Automatic Sleep Disorder Detection Based on EEG Cross-Frequency Coupling and Random Forest Model. J. Neural Eng. 2021;18:046064. doi: 10.1088/1741-2552/abf773. [DOI] [PubMed] [Google Scholar]
  • 36.Bin Heyat M.B., Akhtar F., Khan A., Noor A., Benjdira B., Qamar Y., Abbas S.J., Lai D. A Novel Hybrid Machine Learning Classification for the Detection of Bruxism Patients Using Physiological Signals. Appl. Sci. 2020;10:7410. doi: 10.3390/app10217410. [DOI] [Google Scholar]
  • 37.Tripathi P., Ansari M.A., Gandhi T.K., Albalwy F., Mehrotra R., Mishra D. Computational Ensemble Expert System Classification for the Recognition of Bruxism Using Physiological Signals. Heliyon. 2024;10:e25958. doi: 10.1016/j.heliyon.2024.e25958. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Piritu D.A.I., Taralunga D.D. Proceedings of the 2025 IEEE 31st International Symposium for Design and Technology in Electronic Packaging (SIITME), Brasov, Romania, 22 October 2025. IEEE; Piscataway, NJ, USA: 2025. Robust Sleep Bruxism Detection Using Multimodal Physiological Signal Classification and Crossdataset Machine Learning; pp. 380–386. [Google Scholar]
  • 39.Lai D., Heyat M.B.B., Khan F.I., Zhang Y. Prognosis of Sleep Bruxism Using Power Spectral Density Approach Applied on EEG Signal of Both EMG1-EMG2 and ECG1-ECG2 Channels. IEEE Access. 2019;7:82553–82562. doi: 10.1109/ACCESS.2019.2924181. [DOI] [Google Scholar]
  • 40.Tiwari S., Arora D., Bhardwaj B. High-fidelity EEG feature-engineered taxonomy for bruxism and PLMS prognostication through pioneering and avant-garde ML frameworks. Meas. Sens. 2025;39:101868. doi: 10.1016/j.measen.2025.101868. [DOI] [Google Scholar]
  • 41.Kidwai M.S., Khan M.Z. Proceedings of the 2021 International Conference on Advance Computing and Innovative Technologies in Engineering (ICACITE) IEEE; New York, NY, USA: 2021. A new perspective of detecting and classifying neurological disorders through recurrence and machine learning classifiers; pp. 200–206. [Google Scholar]
  • 42.Gul J.Z., Fatima N., Din Z.M.U., Khan M., Kim W.Y., Rehman M.M. Advanced Sensing System for Sleep Bruxism across Multiple Postures via EMG and Machine Learning. Sensors. 2024;24:5426. doi: 10.3390/s24165426. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Ishtiaq A., Gul J., Din Z.M.U., Imran A., El Hindi K. Wearable Regionally Trained AI-Enabled Bruxism-Detection System. IEEE Access. 2025;13:15503–15528. doi: 10.1109/ACCESS.2025.3532360. [DOI] [Google Scholar]
  • 44.O’Hare E., Cogan J.A., Dillon F., Lowery M., O’Cearbhaill E.D. An Intraoral Non-Occlusal MEMS Sensor for Bruxism Detection. IEEE Sens. J. 2022;22:153–161. doi: 10.1109/JSEN.2021.3128246. [DOI] [Google Scholar]
  • 45.Fatima N., Din Z.M.U., Aishan A.A., Gul J.Z. Artificial Intelligence-Driven Hemodynamic Monitoring of Simulated Bruxism Using Functional Near-Infrared Spectroscopy: A Preliminary Study. CNS Neurosci. Ther. 2025;31:e70619. doi: 10.1111/cns.70619. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Shen Q., Cui Y., Yang J., Jing X., Feng Z., Jin S. ICC 2026—IEEE International Conference on Communications. IEEE; Glasgow, UK: 2026. Bruxism Recognition via Wireless Signal; pp. 1–6. [Google Scholar]
  • 47.Schlaepfer B., Langer J., Erni S., Colombo V. A New Approach for the Field Detection of Sleep Bruxism Based on Inertial Sensor Data and Machine Learning Classification. Sci. Rep. 2026;16:443. doi: 10.1038/s41598-025-29679-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Tang C., Yi W., Xu M., Jin Y., Zhang Z., Chen X., Liao C., Kang M., Gao S., Smielewski P., et al. A Deep Learning–Enabled Smart Garment for Accurate and Versatile Monitoring of Sleep Conditions in Daily Life. Proc. Natl. Acad. Sci. USA. 2025;122:e2420498122. doi: 10.1073/pnas.2420498122. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Minakuchi H., Nagasaki M., Đình L.H., Miki H., Omori K., Nishimura T., Kuboki T., Minematsu N. Experimental analysis of automatic discrimination performance between simulated bruxism and non-bruxism under conscious conditions using electromyography and machine learning. Int. J. Dent. 2026;2026:7874254. doi: 10.1155/ijod/7874254. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Sonmezocak T., Kurt S. Machine Learning and Regression Analysis for Diagnosis of Bruxism by Using EMG Signals of Jaw Muscles. Biomed. Signal Process. Control. 2021;69:102905. doi: 10.1016/j.bspc.2021.102905. [DOI] [Google Scholar]
  • 51.Sonmezocak T., Kurt S. Detection of EMG Signals by Neural Networks Using Autoregression and Wavelet Entropy for Bruxism Diagnosis. Elektron. IR Elektrotechnika. 2021;27:11–21. doi: 10.5755/j02.eie.28838. [DOI] [Google Scholar]
  • 52.Sönmezocak T., Kurt S. Detection of Lower Jaw Activities from Micro Vibration Signals of Masseter Muscles Using MEMS Accelerometer. Comput. Methods Biomech. Biomed. Eng. Imaging Vis. 2023;11:476–484. doi: 10.1080/21681163.2022.2079003. [DOI] [Google Scholar]
  • 53.Castroflorio T., Mesin L., Tartaglia G.M., Sforza C., Farina D. Use of Electromyographic and Electrocardiographic Signals to Detect Sleep Bruxism Episodes in a Natural Environment. IEEE J. Biomed. Health Inform. 2013;17:994–1001. doi: 10.1109/JBHI.2013.2274532. [DOI] [PubMed] [Google Scholar]
  • 54.Martinot J.-B., Le-Dong N.-N., Cuthbert V., Denison S., Gozal D., Lavigne G., Pépin J.-L. Artificial Intelligence Analysis of Mandibular Movements Enables Accurate Detection of Phasic Sleep Bruxism in OSA Patients: A Pilot Study. Nat. Sci. Sleep. 2021;13:1449–1459. doi: 10.2147/NSS.S320664. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Eris O., Recep Bozkurt M., Bulut Eris S., Bilgin C. An Artificial Intelligence-Based Approach With Photoplethysmogram and Heart Rate Variability for Sleep Bruxism Diagnosis. IEEE Access. 2025;13:40413–40428. doi: 10.1109/ACCESS.2025.3546720. [DOI] [Google Scholar]
  • 56.Orhan K., Yazici G., Önder M., Evli C., Volkan-Yazici M., Kolsuz M.E., Bağış N., Kafa N., Gönüldaş F. Development and Validation of an Ultrasonography-Based Machine Learning Model for Predicting Outcomes of Bruxism Treatments. Diagnostics. 2024;14:1158. doi: 10.3390/diagnostics14111158. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Takamura R., Li Y., Hamasaki T., Nakatoh Y. Proceedings of the 2025 IEEE International Symposium on Consumer Technology (ISCT) IEEE; New York, NY, USA: 2025. A study on the detection of sleep bruxism using acoustic features; pp. 97–101. [DOI] [Google Scholar]
  • 58.Peruzzi G., Galli A., Pozzebon A. Proceedings of the 2022 IEEE International Symposium on Measurements & Networking (M&N), Padua, Italy, 18 July 2022. IEEE; Piscataway, NJ, USA: 2022. A Novel Methodology to Remotely and Early Diagnose Sleep Bruxism by Leveraging on Audio Signals and Embedded Machine Learning; pp. 1–6. [Google Scholar]
  • 59.Cid-Verdejo R., Chávez Farías C., Martínez-Pozas O., Meléndez Oliva E., Cuenca-Zaldívar J.N., Ardizone García I., Martínez Orozco F.J., Sánchez Romero E.A. Instrumental Assessment of Sleep Bruxism: A Systematic Review and Meta-Analysis. Sleep Med. Rev. 2024;74:101906. doi: 10.1016/j.smrv.2024.101906. [DOI] [PubMed] [Google Scholar]
  • 60.Yamaguchi T., Mikami S., Maeda M., Saito T., Nakajima T., Yachida W., Gotouda A. Portable and Wearable Electromyographic Devices for the Assessment of Sleep Bruxism and Awake Bruxism: A Literature Review. CRANIO®. 2023;41:69–77. doi: 10.1080/08869634.2020.1815392. [DOI] [PubMed] [Google Scholar]
  • 61.Casett E., Réus J.C., Stuginski-Barbosa J., Porporatti A.L., Carra M.C., Peres M.A., de Luca Canto G., Manfredini D. Validity of Different Tools to Assess Sleep Bruxism: A Meta-analysis. J. Oral Rehabil. 2017;44:722–734. doi: 10.1111/joor.12520. [DOI] [PubMed] [Google Scholar]
  • 62.Li C., Yap S., Loh A., Yap Y., Kujan O., Balasubramaniam R. Ambulatory Devices to Detect Sleep Bruxism: A Narrative Review. Aust. Dent. J. 2024;69:S53–S62. doi: 10.1111/adj.13057. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Aziz S., Ali A.A.M., Aslam H., Abd-alrazaq A.A., AlSaad R., Alajlani M., Ahmad R., Khalil L., Ahmed A., Sheikh J. Wearable Artificial Intelligence for Sleep Disorders: Scoping Review. J. Med. Internet Res. 2025;27:e65272. doi: 10.2196/65272. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Fukui K., Ishimaru S., Kato T., Numao M. Sound-Based Sleep Assessment with Controllable Subject-Dependent Embedding Using Variational Domain Adversarial Neural Network. Int. J. Data Sci. Anal. 2025;20:369–379. doi: 10.1007/s41060-023-00407-7. [DOI] [Google Scholar]
  • 65.Nahhas M.K., Türp J.C., Cattin P., Gerig N., Wilhelm E., Rauter G. Toward wearables for bruxism detection: Voluntary oral behaviors sound recorded across the head depend on transducer placement. Clin. Exp. Dent. Res. 2024;10:e70001. doi: 10.1002/cre2.70001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.Alfieri D., Awasthi A., Belcastro M., Barton J., O’Flynn B., Tedesco S. Proceedings of the 2021 32nd Irish Signals and Systems Conference (ISSC), Athlone, Ireland, 10 June 2021. IEEE; Piscataway, NJ, USA: 2021. Design of a Wearable Bruxism Detection Device; pp. 1–5. [Google Scholar]
  • 67.Flores-Ramírez B., Suaste-Gómez E., García-Limón V., Angeles-Medina F. Flexible PVDF Sensors for Bruxism Bite Force Measurement: A Redefined Instrumental Approach. PLoS ONE. 2025;20:e0330422. doi: 10.1371/journal.pone.0330422. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Tajitsu Y., Shimda S., Nonomura T., Yanagimoto H., Nakamura S., Ueshima R., Kawanobe M., Nakiri T., Takarada J., Takeuchi O., et al. Application of Braided Piezoelectric Poly-l-Lactic Acid Cord Sensor to Sleep Bruxism Detection System with Less Physical or Mental Stress. Micromachines. 2023;15:86. doi: 10.3390/mi15010086. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Gao J., Liu L., Gao P., Zheng Y., Hou W., Wang J. Intelligent Occlusion Stabilization Splint with Stress-Sensor System for Bruxism Diagnosis and Treatment. Sensors. 2019;20:89. doi: 10.3390/s20010089. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.Al-Hamad K.A., Asiri A., Al-Qahtani A.M., Alotaibi S., Almalki A. Development and In-Vitro Validation of an Intraoral Wearable Biofeedback System for Bruxism Management. Front. Bioeng. Biotechnol. 2025;13:1572970. doi: 10.3389/fbioe.2025.1572970. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Zhong J., Gao X., Hu S., Yue Y., Liu Y., Xiong X. A Worldwide Bibliometric Analysis of the Research Trends and Hotspots of Bruxism in Adults during 1991–2021. J. Oral Rehabil. 2024;51:5–14. doi: 10.1111/joor.13577. [DOI] [PubMed] [Google Scholar]
  • 72.Farella M., Palla S., Gallo L.M. Time–Frequency Analysis of Rhythmic Masticatory Muscle Activity. Muscle Nerve. 2009;39:828–836. doi: 10.1002/mus.21262. [DOI] [PubMed] [Google Scholar]
  • 73.Moxley B., Stevens W., Sneed J., Pearl C. Novel Diagnostic and Therapeutic Approaches to Temporomandibular Dysfunction: A Narrative Review. Life. 2023;13:1808. doi: 10.3390/life13091808. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.Bondareva E., Hauksdóttir E.R., Mascolo C. Proceedings of the Adjunct Proceedings of the 2021 ACM International Joint Conference on Pervasive and Ubiquitous Computing and Proceedings of the 2021 ACM International Symposium on Wearable Computers, Virtual, 21 September 2021. ACM; New York, NY, USA: 2021. Earables for Detection of Bruxism: A Feasibility Study; pp. 146–151. [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

This paper reports a literature synthesis and did not generate new primary data. All source studies discussed are cited in the reference list and are publicly available through their respective publishers or repositories.


Articles from Diagnostics are provided here courtesy of Multidisciplinary Digital Publishing Institute (MDPI)

RESOURCES