Abstract
Objective
Monitoring the condition of stool is useful for daily health management and early detection of diseases; however, a self-assessment of stool form is often inaccurate. To address this issue, TOTO, developed a prototype toilet seat equipped with a sensor that classifies stool forms according to the Bristol Stool Form Scale (BSFS). This study aimed to verify the accuracy of this method for distinguishing stool types.
Methods
We recruited 38 healthy Japanese participants and obtained BSFS data using the device from November 2021 to March 2022. Meanwhile, we taught the participants how to assess their own stool using the BSFS and used their assessments as reference standards. We evaluated the strength of agreement between the participants' assessments and the device's classification results using the kappa coefficient, as well as the strength of correlation using Spearman's rank correlation coefficient.
Materials
We analyzed 305 samples from 38 participants after excluding 62 samples for which the participant's judgment was unreliable or the device data were of a low quality.
Results
The kappa coefficient between participants' assessments and the device's classification results based on BSFS was 0.55 [95% confidence interval (CI) 0.45-0.65, p<0.001], while the Spearman's rank correlation coefficient was 0.56 (95% CI 0.48-0.63, p<0.001).
Conclusion
The device appeared to possess limited but moderate validity, with fair to good agreement and moderate correlation with the participants' assessment results. This toilet seat with an automated classification of stool form by sensing falling stool may support daily healthcare and early disease detection.
Keywords: toilet seat, automatic stool classification, machine learning, Bristol Stool Form Scale (BSFS)
Introduction
Abnormalities in stool form can adversely affect the quality of life and labor productivity. For example, a large-scale study in Japan showed that patients with chronic constipation had a lower health-related quality of life and greater impairment of their ability to work than healthy controls (1). Furthermore, a web-based survey indicated that bloating and unpredictable timing of defecation were linked to reduced work productivity and less daily activity, with improvements observed after treatment (2). Similarly, it was reported that patients with inflammatory bowel disease (IBD) are highly impaired in their ability to work and bear the burden of indirect costs associated with their diseases (3).
Diseases such as irritable bowel syndrome (IBS), which is often associated with mental health (4), IBD, and colorectal cancer, may cause changes in the stool form. Stool consistency, which is associated with colonic transit time and gut microbial diversity, reflects the underlying changes in the gut ecosystem (5) and is frequently observed in gastrointestinal diseases (6). Several other studies have also reported alterations in stool form associated with IBS, IBD, and colorectal cancer (7-10).
It is thus important to be aware of stool characteristics to maintain both mental and physical health, but self-assessment of stool may lack accuracy. The Bristol Stool Form Scale (BSFS) is widely used to objectively classify stool forms into seven types (4,11) (Fig. 1). Previous studies have evaluated the BSFS as having substantial reliability and validity (12,13). Against this background, the development of a device that can automatically classify stool forms based on the BSFS would contribute to health management and early disease detection - in other words, to preventive medicine.
Figure 1.

The Bristol Stool Form Scale (BSFS). Modified from Reference 4, with permission from Elsevier Science & Technology Journals.
In pursuit of this goal, TOTO, (Research Institute, Chigasaki, Japan), developed a prototype toilet seat equipped with a sensor to automatically and objectively classify stool forms based on BSFS. The present study therefore evaluated whether or not this device can accurately determine the BSFS score.
Materials and Methods
Participants
We recruited healthy participants working at TOTO, and taught them how to classify their stool using the BSFS. Thus, we defined their assessments as reference standards against which to compare automatically determined BSFS scores. The participants were asked to use the prototype toilet seat voluntarily, and the device automatically recorded their BSFS scores from November 2021 to March 2022.
This study was approved by the Ethics Committee of the University of Tsukuba Hospital (reference #R03-102). Before accessing the data, full anonymization was ensured, and informed consent was obtained and approved by the Ethics Committee.
Device description
The evaluated device was a prototype toilet seat integrated with a sensor designed to detect stool (Fig. 2, 3). This sensor is linear in form and features light-receiving elements arranged in a single horizontal line positioned at the rear of the toilet seat. This configuration ensures that the field of view of the sensor is orthogonal to the trajectory of the stool, as it is excreted and falls into the toilet bowl. To facilitate detection, LEDs illuminate the falling stool. By continuously capturing the temporal sequence of light reflected from the falling stool, a comprehensive representation of the excreted stool was obtained. Stationary objects, including the person on the toilet seat, were not recognized as subjects for imaging because of the motion-based detection mechanism.
Figure 2.

Conceptual diagram of the prototype (left) and the prototype installed in a toilet for testing (right).
Figure 3.

Mechanism of prototype actuation illustrating line data capture and image construction from sequential line data. The top section shows the linear sensor and LED capturing the line data from a falling stool sequentially at different time points. The captured line data were then used to construct a complete image, as indicated at the bottom of the figure.
Regarding data capture control, the prototype was configured to capture data for 10 s immediately after the initial detection of stool under the condition that the subject was seated on the toilet seat. To account for the possibility of multiple defecation events during a single toilet use, this capture process was repeated while the subject remained seated, with a maximum of five 10-s datasets recorded.
Development of the classification algorithm
The classification algorithm for the prototype was developed through the following steps. First, a training dataset was acquired. Healthy volunteers used the prototype toilet seat to defecate, and the sensor recorded stool data throughout each instance of toilet use, from initiation to cessation of excretion. Immediately after defecation, the individual who excreted the stool photographed the remaining stool within the bowl from a position directly above it. Three board-certified gastroenterologists independently reviewed the photographs and classified them according to the BSFS. Ideally, stool morphology should be assessed under the same conditions that the device observes, i.e. while the stool is falling before it reaches the water. However, there is currently no validated method for assessing stool morphology during falling, and obtaining such images is technically difficult and raises ethical concerns. Therefore, in this study, we used photographs of stool that had settled in the toilet bowl as the reference standard for BSFS classification. In cases where the stool was stacked or its morphology was unclear, preventing unambiguous classification, three gastroenterologists convened to discuss the case and either excluded the sample from the dataset or reached a consensus on the BSFS score. This process resulted in a labeled dataset, where each set of sensor data, corresponding to a complete defecation event, was associated with a BSFS category determined by expert visual assessment of the post-defecation stool.
To address the limited availability of data for certain stool consistencies, simulated stool samples designed to mimic the physical characteristics of each BSFS type were dropped through the prototype toilet seat to generate additional training data. Subsequently, a classification algorithm was developed using the acquired dataset of 291 labeled sets (Type 1: 20, Type 2: 23, Type 3: 21, Type 4: 92, Type 5: 91, Type 6: 26, Type 7: 18). The classification algorithm was implemented based on a machine-learning decision tree model, with developers performing the final parameter tuning and optimization. This model was developed using Python version 3.7 (Python Software Foundation) with the scikit-learn machine learning library version 1.0.2 (14).
Statistical analyses
Kappa coefficient and Spearman's rank correlation coefficients were used to evaluate the validity of the device. First, we assessed the kappa coefficient to evaluate the strength of agreement between the participants' assessments and the device classification results. Because we considered adjacent agreement rather than perfect agreement to be clinically acceptable, we used the weighted kappa coefficient with quadratic weights for our analysis. Second, Spearman's rank correlation coefficient was used to assess the strength of correlation between the participants' assessments (reference standards) and the device's classification results. Statistical analyses were performed using the Bell Curve for Excel version 3.20 (Social Survey Research Information, Tokyo, Japan) and R version 4.5.0 (R Core Team, Vienna, Austria) software programs. Statistical significance was set at p<0.05.
Results
A total of 38 healthy Japanese participants were included in this study (Table 1). Out of the 367 samples, we excluded those with ambiguous evaluations by participants (n=5), brightness saturation (n=17), truncated in the longitudinal direction (n=32) (with 4 data entries overlapping between “brightness saturation” and “truncated in the longitudinal direction”), truncated in the width direction (n=5), and those too small to be assessed by the participants (n=1) and the devices (n=6) (Fig. 4). In total, 305 stool samples were included in the analysis. The median number of stool samples per participant was 4.0 (interquartile range, 2.25-9.0) (data not shown). First, we constructed a cross-tabulation table to assess the degree of agreement between the participants' assessments and the device's classification results, based on the BSFS (Table 2). The overall accuracy of the BSFS classification is 40.7%. However, the exact/adjacent agreement, defined as correct when the classification either matched or differed by only one BSFS type, was 87.9%. The kappa coefficient between these two sets of results was 0.55 [95% confidence interval (CI) 0.45-0.65, p<0.001]. Next, we created a scatter plot to assess the correlation between the two sets of results (Fig. 5). Spearman's rank correlation coefficient was 0.56 (95% CI 0.48-0.63) (p<0.001).
Table 1.
Participant Characteristics.
| Age | Number | Sex | |
|---|---|---|---|
| Male | Female | ||
| 20 | 10 | 7 | 3 |
| 30 | 10 | 9 | 1 |
| 40 | 8 | 4 | 4 |
| 50 | 6 | 6 | 0 |
| Unknown | 4 | 4 | 0 |
| 38 | 30 | 8 | |
Figure 4.

Flow chart of the study.
Table 2.
Cross-tabulation Table of the Device’s Classification Results and Participants’ Assessments.
| Device’s classification results | |||||||||
|---|---|---|---|---|---|---|---|---|---|
| Participants’ assessments |
Type | 1 | 2 | 3 | 4 | 5 | 6 | 7 | Total |
| 1 | 2 | 3 | 1 | 0 | 0 | 0 | 0 | 6 | |
| 2 | 1 | 5 | 3 | 7 | 6 | 0 | 0 | 22 | |
| 3 | 0 | 3 | 11 | 32 | 11 | 3 | 0 | 60 | |
| 4 | 0 | 1 | 3 | 48 | 30 | 5 | 0 | 87 | |
| 5 | 0 | 0 | 1 | 27 | 49 | 34 | 0 | 111 | |
| 6 | 1 | 0 | 0 | 1 | 5 | 9 | 1 | 17 | |
| 7 | 0 | 0 | 0 | 0 | 0 | 2 | 0 | 2 | |
| Total | 4 | 12 | 19 | 115 | 101 | 53 | 1 | 305 | |
Kappa coefficient 0.55 (95% CI 0.45-0.65, p<0.001).
Figure 5.

Scatter plot shows the device’s classification results based on BSFS on the x-axis and the participants’ assessments on the y-axis, where the BSFS types were converted to ordinal ranks. The size of the bubbles indicates the number of samples. The distribution of the bubbles shows an upward trend along the diagonal, indicating a positive correlation between the device’s classification results and participants’ assessments.
Discussion
This was an observational study aimed at evaluating the validity of a toilet seat intended to automatically classify the stool form by sensing falling stool. Among the 367 captured datasets, 50 were excluded because of inherent limitations in the data capture mechanism of the prototype. Specifically, 32 samples exhibited a truncation in the longitudinal direction. This issue stemmed from the prototype's configuration, which was set to capture data for a fixed 10-s period immediately following the initial detection of defecation. Truncation occurred when the overall defecation activity, which can be intermittent, continued beyond the 10-s capture window.
This challenge can be resolved by modifying the system to capture all the data throughout a defecation episode, eliminating the need for a predetermined time limit. Furthermore, the exclusion of samples caused by width truncation can be mitigated by adjusting the position and orientation of the sensor. Exclusions resulting from brightness saturation can also be addressed by appropriately adjusting the light emission intensity of LEDs.
We used the kappa coefficient and Spearman's rank correlation coefficient between the participants' assessments and device classification results based on the BSFS. The BSFS was originally developed to classify the stool form as observed when the stool is at rest after defecation. In contrast, our device captured stool morphology in air while the stool was falling. Although the overall shape is generally preserved, the appearance during falling may vary depending on the consistency of the stool, and stool may have a different appearance in the air than when they are at rest. As no validated scale currently exists for stool observed while falling, we used the BSFS as the most practical reference for this study. The kappa coefficient was 0.55 and Spearman's rank correlation coefficient was 0.56, both of which were statistically significant (p<0.001). This kappa coefficient value can be interpreted as reflecting fair to good agreement (15), whereas the Spearman's rank correlation coefficient value reflects a moderate correlation. These findings suggest that the device has a limited but moderate validity. The fair to good kappa agreement likely reflects a combination of statistical factors, study design, and device-related factors. From a statistical perspective, the stool types in this study were unevenly distributed, with many normal stool specimens and a few hard or watery specimens. Such an imbalance increases the likelihood of agreement by chance and can lower the kappa coefficient. In addition, because the participants' own BSFS assessments were used as reference standards, some subjectivity in these ratings may have affected the results. Furthermore, the limited image quality of the monochrome sensor and the limited and relatively uniform training dataset may also have reduced the ability of the device to accurately classify stool form. These factors suggest that the present results reflect not only the characteristics of the dataset but also the limits of the reference standard and the technical limits of the prototype device. Future studies that include more cases with rare stool types, larger and more varied training data, and continued improvements to the device, will be important for achieving better classification performance.
Reports on the classification of stool forms based on the BSFS using machine learning have been published. For example, Choy et al. used deep convolutional neural networks (CNNs) to classify stool types (16). Their results showed that the ResNeXt-50 classifier achieved a classification accuracy of 94.35%, demonstrating decent performance. Pimentel et al. developed a smartphone application using a model obtained by applying transfer learning to a canonical convolutional image classification model. There were good correlations between the assessment of stool form by two experts and by artificial intelligence [intraclass correlation coefficient (ICC) 0.782-0.852] (17). Furthermore, Park et al. developed a “smart toilet” system (18) that operates autonomously using pressure and motion sensors and classifies stool forms based on BSFS using CNNs, with assessments by board-certified general surgeons used as a gold standard for comparison. The CNN achieved classification performance comparable to that of medical students with general medical training, with all area under the curve (AUC) values exceeding 0.91. In addition, the CNN predictions and the surgeons' assessments showed good agreement, which matched at a rate of approximately 75%. Regarding classification performance, a direct comparison with the present study was difficult because these previous studies used different evaluation metrics, such as ICC, AUC, and match rate. Choy et al. reported an accuracy of 94.35% (16). However, this was calculated using a modified BSFS with three categories (Types 1-2: constipated, Types 3-5: normal, Types 6-7: loose) instead of the standard seven-type BSFS used in our study. Therefore, a direct comparison with our results (40.7%) is not appropriate. For reference, we also calculated accuracy using a modified BSFS with these three categories to match the approach used in previous studies. The accuracy was 77.0% (data not shown), which indicates the current performance of the device. Nevertheless, because stool form changes gradually along a continuum, we also calculated the exact/adjacent agreement across the seven BSFS categories in this study, which was 87.9%. Thus, further studies using standardized evaluation criteria are required.
In addition to classification performance, other characteristics, including intrusiveness, are important for real-world applications. One major difference between the previous studies and the present study is the method used for classifying stool forms. Two previous studies used static color images of stool inside the toilet bowl (16,17), whereas another study used video recordings to capture the entire defecation process within the toilet bowl (18). In contrast, our study used monochrome images of stool sensed in air. Based on this difference, we describe three strengths of the present study.
First, the device captured the stool from start to finish. The stool type can vary even within a single defecation (13). By analyzing the entire stool in mid-air, our device identified the dominant stool type based on its overall composition. For example, if 70% of the stool was Type 3 and 30% Type 5, the stool was classified as Type 3. In contrast, two previously reported methods that use static images of the toilet bowl (16,17) may struggle to accurately classify physically overlapping stool, given the difficulty of observing their overall characteristics.
An additional strength of our device is that, because it analyzes stool in mid-air, it avoids the effects of toilet water, such as stool dissolution or interference caused by reflection on the water. In a previous report (16), it was noted that Type 7 stool often mix with toilet water, making classification difficult. Although the previous study did not address the accuracy of classification for Type 7, our study classified two samples of Type 7 stool as Type 6, which is also considered representative of diarrhea (19) (Table 2). Given the small sample size of these stool types in this study, there is a need for further validation of the device in future studies.
Third, our device uses monochromatic imaging without the use of cameras. In previous research, 30% of the survey respondents felt uncomfortable with the smart toilet system (18). Among the various features of smart toilets, camera-based modules, such as those that record the toilet bowl from the beginning to the end of defecation, were less preferred. Meanwhile, our use of monochrome images may help address privacy concerns. Furthermore, monochrome data are easier to handle than color images in terms of data processing, require less time for training and inference, and are more cost-effective.
Toilet-based systems for monitoring stool offer various advantages over portable smartphone-based systems. For example, they enable automated monitoring, reducing the user burden by eliminating the need to capture photos on a daily basis. Consequently, these systems may improve compliance and support long-term monitoring, which may contribute to a healthier lifestyle. Furthermore, by noticing changes in stool form more easily, people may be prompted to visit a doctor sooner, which could in turn lead to disease detection at an earlier stage. Another benefit of the system is that it supports individuals with constipation who treat themselves by adjusting their laxative use. In addition to personal home-based use, these systems can be used in healthcare facilities. For example, in hospitals, these systems would enable accurate monitoring of defecation, including its timing and stool form. Moreover, by sending data to a smartphone or electronic medical record, nurses can monitor patients more efficiently and save time during hospitalization. These systems would also be useful for preparation prior to colonoscopy when accurate tracking of stool form is required. In addition, in nursing care facilities, such systems could relieve residents of the need to report their bowel movements, which is particularly beneficial for people with dementia. The systems would also help preserve the dignity of residents by removing the need for them to talk about their bowel movements. However, toilet-based systems must be installed in toilets, which may limit their widespread adoption.
Several limitations associated with the present study warrant mention. First, only healthy participants were included, resulting in a small number of abnormal stool samples, including only six cases of Type 1 and two cases of Type 7 in the validation dataset. Therefore, we were unable to adequately assess the device classification performance across the full range of BSFS categories. The insufficient number of abnormal stool affected not only the validation dataset but also the training of the machine-learning algorithm. Reflecting this limitation, 16 out of 22 (72.7%) Type 2 samples were misclassified as Types 3-5, which are considered to represent normal stool (4,19,20). Because the device used in this study was a prototype, TOTO, has been continuously working on improving the accuracy of the machine learning algorithm by adding data with the goal of enhancing the performance of the final product. We are planning future clinical studies to collect additional samples of rare stool types, such as Type 7, by including more patients, such as those with IBS and IBD, to validate the device's performance. Second, the reference standard used was based on the participant's own stool assessment, but an evaluation by gastroenterologists would be more appropriate. However, owing to some participants' feelings of embarrassment and privacy concerns about having their stool observed by others, this approach was not feasible.
In conclusion, the device analyzed here demonstrated limited but moderate validity, showing fair to good agreement and a moderate correlation with the participants' assessment results using the BSFS. This work suggests that a toilet seat with automated stool form classification by sensing falling stool has the potential to support daily health monitoring and may aid in the early detection of diseases.
Author’s disclosure of potential Conflicts of Interest (COI).
Satoko Kizuka and Ryuji Kawazoe are employed by TOTO.
References
- 1.Tomita T, Kazumori K, Baba K, Zhao X, Chen Y, Miwa H. Impact of chronic constipation on health-related quality of life and work productivity in Japan. J Gastroenterol Hepatol 36: 1529-1537, 2021. [DOI] [PubMed] [Google Scholar]
- 2.Ota T, Kuratani S, Masaki H, Ishizaki S, Seki H, Takebe T. Impact of chronic constipation symptoms on work productivity and daily activity: a large-scale internet survey. JGH Open 8: e70042, 2024. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Youssef M, Hossein-Javaheri N, Hoxha T, Mallouk C, Tandon P. Work productivity impairment in persons with inflammatory bowel diseases: a systematic review and meta-analysis. J Crohns Colitis 18: 1486-1504, 2024. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Lacy BE, Mearin F, Chang L, et al. Bowel disorders. Gastroenterology 1393-1407.E5, 2016. [DOI] [PubMed] [Google Scholar]
- 5.Vandeputte D, Falony G, Vieira-Silva S, Tito RY, Joossens M, Raes J. Stool consistency is strongly associated with gut microbiota richness and composition, enterotypes and bacterial growth rates. Gut 65: 57-62, 2016. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Quaglio AEV, Grillo TG, De Oliveira ECS, Di Stasi LC, Sassaki LY. Gut microbiota, inflammatory bowel disease and colorectal cancer. World J Gastroenterol 28: 4053-4060, 2022. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Garrigues V, Mearin F, Badía X, et al.; RITMO Group . Change over time of bowel habit in irritable bowel syndrome: a prospective, observational, 1-year follow-up study (RITMO study). Aliment Pharmacol Ther 25: 323-332, 2007. [DOI] [PubMed] [Google Scholar]
- 8.Urayama S, Chang EB. Mechanisms and treatment of diarrhea in inflammatory bowel diseases. Inflamm Bowel Dis 3: 114-131, 1997. [PubMed] [Google Scholar]
- 9.Tashiro N, Budhathoki S, Ohnaka K, et al. Constipation and colorectal cancer risk: the Fukuoka Colorectal Cancer Study. Asian Pac J Cancer Prev 12: 2025-2030, 2011. [PubMed] [Google Scholar]
- 10.Park JY, Mitrou PN, Luben R, Khaw KT, Bingham SA. Is bowel habit linked to colorectal cancer? - results from the EPIC-Norfolk study. Eur J Cancer 45: 139-145, 2009. [DOI] [PubMed] [Google Scholar]
- 11.Arasaradnam RP, Brown S, Forbes A, et al. Guidelines for the investigation of chronic diarrhoea in adults: British Society of Gastroenterology, 3rd edition. Gut 67: 1380-1139, 2018. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Chumpitazi BP, Self MM, Czyzewski DI, Cejka S, Swank PR, Schulman RJ. Bristol Stool Form Scale reliability agreement decreases when determining Rome III stool form designations. Neurogastroenterol Motil 28: 443-448, 2016. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Blake MR, Raker JM, Whelan K. Validity and reliability of the Bristol Stool Form Scale in healthy adults and patients with diarrhoea-predominant irritable bowel syndrome. Aliment Pharmacol Ther 44: 693-703, 2016. [DOI] [PubMed] [Google Scholar]
- 14.Pedregosa F, Varoquaux G, Gramfort A, et al. Scikit-learn: machine learning in Python. arXiv [cs.LG], 2012 [Internet]. http://arxiv.org/abs/1201.0490.
- 15.Feinstein AR. Clinimetrics. Yale University Press, New Haven, 1987. [Google Scholar]
- 16.Choy YP, Hu G, Chen J. Detection and classification of human stool using deep convolutional neural networks. IEEE Access 9: 160485-160496, 2021. [Google Scholar]
- 17.Pimentel M, Mathur R, Wang J, et al. A smartphone application using artificial intelligence is superior to subject self-reporting when assessing stool form. Am J Gastroenterol 117: 1118-1124, 2022. [DOI] [PubMed] [Google Scholar]
- 18.Park SM, Won DD, Lee BJ, et al. A mountable toilet system for personalized health monitoring via the analysis of excreta. Nat Biomed Eng 4: 624-635, 2020. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.The Japanese Gastroenterological Association . Evidence-Based Clinical Practice Guidelines for Chronic Diarrhea 2023 (in Japanese). Nankodo, Tokyo. [Google Scholar]
- 20.The Japanese Gastroenterological Association . Evidence-Based Clinical Practice Guidelines for Chronic Constipation 2023 1st ed. (in Japanese). Nankodo, Tokyo, 2023. [Google Scholar]
