Abstract
Social presence has been known to impact eating behavior among people with obesity; however, the dual study of eating behavior and social presence in real-world settings is challenging due to the inability to reliably confirm the co-occurrence of these important factors. High-resolution video cameras can detect timing while providing visual confirmation of behavior; however, their potential to capture all-day behavior is limited by short battery lifetime and lack of autonomy in detection. Low-resolution infrared (IR) sensors have shown promise in automating human behavior detection; however, it is unknown if IR sensors contribute to behavior detection when combined with RGB cameras. To address these challenges, we designed and deployed a low-power, and low-resolution RGB video camera, in conjunction with a low-resolution IR sensor, to test a learned model’s ability to detect eating and social presence. We evaluated our system in the wild with 10 participants with obesity; our models displayed slight improvement when detecting eating (5%) and significant improvement when detecting social presence (44%) compared with using a video-only approach. We analyzed device failure scenarios and their implications for future wearable camera design and machine learning pipelines. Lastly, we provide guidance for future studies using low-cost RGB and IR sensors to validate human behavior with context.
Keywords: human activity recognition, wearable camera, deep learning
1. INTRODUCTION
A recent study found that strategies for preventing dietary lapses may differ based on social presence (i.e., the presence of another individual during the eating episode), which is known to influence many health habits [31]. For example, eating with a companion was shown to increase food intake by 28% [26]. On the other hand, one possible criterion for defining a binge eating episode (overeating while sensing loss of control) is eating alone because of feeling embarrassed by how much one is eating. Human behavior is complex and multifactorial, and it is necessary to understand the interplay between eating and social presence to model their relationship more accurately within conceptual models of human behavior.
Determinants of eating behaviors and the importance of the social context in which the behaviors occur have been a long-standing interest for researchers [16, 44, 43, 55, 67]. These contexts have been studied together in laboratory conditions [44, 43, 25, 40, 67]; however, there is an unmet need for studies in real-world settings as in-laboratory studies cannot be generalized to real-life owing to the myriad social and environmental factors that affect human behaviors [38, 42, 50, 66, 53, 72, 19].
With the advancement of wearable computing, several devices have emerged that can automatically capture and model human behaviors, such as wrist and neck sensors. However, wearer adherence to multiple devices in the wild is a known problem, preventing the practical utility of these devices in longitudinal studies. Expecting participants to wear more than one device reliably reduces the feasibility of longitudinal studies while simultaneously increasing chances for multiple points of device failure [10] and challenges with synchronization [79]. One promising device that can monitor multiple human behaviors is a wearable camera. Wearable cameras can provide visual confirmation of human activities, and advances in image processing and machine learning models can automatically detect whether an individual is performing an activity such as eating or is in the presence of others (Figure 1).
Figure 1:

WildCam system comprises an IR sensor array and an RBG camera. WildCam enables monitoring of eating behaviors and the environmental context related to eating, such as social presence.
The camera’s benefit over other sensing modalities is the ability to visually confirm and validate activities such a eating [11, 18] or screen time [1]. Researchers have used these cameras to visually confirm several behaviors, including an individual’s physical activity [35, 36, 49], sleep pattern [22, 46], and more recently eating behaviors [21, 45, 77]. SenseCam, one of the earliest cameras used to confirm human behavior in the wild visually, captured one image every 30 seconds and built a memory log for the wearer but had a storage limitation of 1 GB [35]. This intermittent capture of behavior helps ensure an all-day battery lifetime; however, the inability to continuously record video prevents its utility in capturing fine-grained behavior related to eating (e.g., feeding gestures) and social presence (e.g., someone coming into and out of the field of view). This fine-grained capture is needed not only for timely visual confirmation but also for automated detection of activity [33].
Moreover, while SenseCam was capable of egocentric video capture, its lack of mechanical customization made it challenging to capture certain wearer activity, particularly for varying body shapes and sizes. For some activities, like detecting feeding gestures for eating, it increases reliability when the camera is oriented towards the wearer’s mouth. Activity-oriented cameras are particularly well suited to capture a specific activity of the wearer as they are designed and oriented to capture the wearer [5]. Participants can conveniently wear activity-oriented cameras to collect fine-grained activity data (e.g., hand-to-mouth gestures, eating, food types consumed) and environmental context associated with the specific activity. However, most wearable cameras tested in detecting feeding gestures are high-resolution cameras that consume significant amounts of energy and prevent the capture of all-day activity. The use of low-resolution RGB video is necessary to address power constraints without significantly increasing battery weight and size; however, low-resolution RGB video may compromise the machine learning model’s ability to reliably detect the behavior of interest [39]. Recent literature has shown the potential for thermal imagers to augment RGB data, thereby enhancing a camera’s ability to detect salient objects [71]. Simultaneously, to mitigate power constraints, researchers have shown the ability of low-resolution thermal imagers to detect human activities such as lying, sitting, standing, and walking [75].
A thermal sensor provides unique information about objects with thermal signatures (e.g., hot and cold objects), including the human body. For instance, thermal cameras can successfully capture a person’s silhouette as the average body temperature is usually higher than the surrounding temperature. The use of thermal images for human activity recognition via wearables has become increasingly common in recent years, as low-resolution, power-efficient thermal cameras are now available [3, 8, 4]. A further application of thermal images is to facilitate the segmentation of human silhouettes in RGB images to enable privacy-preserving applications such as face obfuscation [7, 62]. More recently, researchers used thermal sensors to reduce the power consumption of wearable devices by developing a triggering mechanism that uses thermal images to execute a high-power-consuming task only when needed [63].
In this paper, we focus on the ability of a low-resolution RGB and thermal sensor to detect two critical activities of interest – eating gestures and social presence.
The key contributions of this work are:
Activity detection using a low-resolution, wearable activity-oriented camera comprising RGB and thermal sensing: We present a framework combining a low-resolution, wearable activity-oriented camera with a low-resolution IR sensor array as a tool and corresponding data analytic pipeline to detect eating gestures and social presence. Based on a study with 10 participants and 80 hours of video across 3 days of device wear, we detect eating episodes and social presence with an F1-score of 70% (5% increase using IR) and 74% (44% increase using IR), respectively. In addition, we report results with and without the IR sensor array.
Implications for device design and real-world studies: We study the false positives and false negatives for each activity and provide guidance on the design of future multimodal activity-oriented devices for detecting eating and social presence and for deploying such devices in real-world studies.
2. RELATED WORK
This section discusses recently developed methods and systems to automate the detection of eating episodes and faces/social presence. We primarily focus on image-based approaches. Additionally, since we use an IR sensor array, we provide a brief background of the promise of this sensor in capturing our behaviors of interest.
2.1. Automated Eating-Detection Approaches
The ability to automatically detect eating in free-living settings has been a long-standing research problem in the mobile and wearable computing community. Over the years, researchers have explored various proxies to detect eating episodes, such as the use of inertial sensor data from a smartwatch to detect feeding gestures [12, 28, 59, 69], audio sensor data to detect chewing or swallowing sounds [21, 51, 74], cameras to capture images of food items [17, 61, 68], and a variety of other sensing techniques [77, 24, 57]. Although inertial-based sensors can detect feeding gestures, automated models are often confounded by similar signals from other hand-to-mouth gestures [76], and they are unable to provide data regarding the presence of other individuals. Similar to inertial-based sensors, audio-based methods have shown promise in detecting eating activity, but they provide little utility in detecting social presence.
More recently, with the availability of wearable cameras, visual confirmation of behavior became possible to monitor fine-grained details of the eating activity. This is advantageous for the rich sensor information in video signals about an eating episode’s characteristics, allowing for the detection and confirmation of multiple behaviors. However, annotating and coding the eating activity based on camera visuals is a timely, costly, and labor-intensive process. Methods that automatically detect these details will significantly improve our ability to design studies that understand factors contributing to overeating in real-world environments.
High-resolution image- and video-based methods have been explored by several researchers to detect feeding- and eating-related gestures. Reddy et al. [52] and Thomaz et al. [68] suspended a smartphone across the neck using a lanyard to capture images throughout the day, including moments of eating. However, suspending a smartphone across the neck can be uncomfortable for long-term wear. To overcome this, Bedri et al. [17] designed FitByte, a smartglass with an embedded camera, and Echterhoff et al. [29] designed the personal activity radius (PAR) device that could be attached to any glass frame. In addition, both devices used a camera to monitor and detect eating moments. More recently, Jia et al. [37] used the eButton device to capture egocentric images of food being consumed. To enable data collection from individuals in a privacy-preserving approach, Schiboni et al. [58] designed a head-mounted camera system that points downward from a cap and captures eating moments without affecting user privacy. Recently, Rouast and Adam [54] used a 360° video camera placed at the center of a dinner table to record participant meals and built feeding gesture–detection models using deep neural networks. However, this approach is limited to a controlled laboratory setting. In contrast, our approach adopts a low-resolution and low-power, wearable activity-oriented device [5] with a fish-eye lens (oriented toward the mouth), which enables visualization of food and drink from table to mouth.
2.2. Automated Social-Detection Approaches
To detect social presence, one approach is that of face detection, a specific type of object-class detection that has been an important part of image processing and machine learning research in the past three decades [9]. Shetty et al. [64] have shown the utility of detecting faces using a thermal sensing array, the “Grid-EYE,” combined with a Kalman filter to detect human motion. Other modes of detection have also shown promise, including RFID tags [23, 14], Bluetooth Low Energy [15], wearable sensor (GNSS/IMU)-based human tracking and motion capture [34], and even WiFi signals, which enable detection of people from behind walls [2]. These modalities are useful but are either too costly or not wearable and hence are limited to a specific region, require access to devices on multiple individuals, or are unable to provide visual confirmation of the behavior, thereby preventing validation of the activity of interest.
Although sensing faces has been a long-standing research topic, given the ubiquitous nature of video cameras, understanding social presence and interaction are of more recent interest. In one of the earliest egocentric camera-based social interaction studies, Fathi et al. [30] used a Go-Pro camera to detect and characterize social interactions by focusing on assigning roles to individuals engaged in a conversation and localizing them in a 3D scene. Alternatively, Alletto et al. [9] clustered people into different socially related groups using a novel head pose estimation framework and then localized them in a 3D scene. In both of these studies, the researchers assumed that people would likely look at each other, whereas in the real world people can interact without constantly orienting themselves toward each other. These studies all focused largely on classifying the type of interactions rather than identifying interactions in a free-living context. In addition, Zhao et al. [80] used high-resolution images from multiple chest-worn cameras to analyze human interaction. However, this study did not focus on capturing the wearer’s social activities. All of these studies used a Go-Pro or similar high-resolution camera. Go-Pro (and similar) cameras provide high-definition images with several pixels on target where substantial detail can be captured, simplifying the task of face detection. A recent study by Okuno et al. [47] counted the number of faces to build an index that quantifies the amount of social activity; however, they still relied on high-quality RGB images and used out-of-the-box face detection models to take the initial step. Continuously capturing images at a high resolution is battery intensive—the Go-Pro lasts only a few hours [32]. Since we are interested in capturing interactions throughout the day, we investigated using lower-resolution images [41].
3. SYSTEM DESIGN
3.1. Activity Definitions
We defined two constructs, eating episode and social presence, and we describe each construct and the method used to annotate each of these constructs below. Two trained annotators independently annotated each video frame and then met to resolve any conflicts. Remaining conflicts were resolved in a group meeting with expert annotators. Figure 2 presents sample images captured by the WildCam system.
Figure 2:

Different activities captured by the camera. Blue boxes indicate eating, and red boxes represent social presence.
Eating episode:
We defined feeding as an activity involving a hand-to-mouth gesture where the hand brings food to the mouth. The gesture starts when the hand (with food) begins traveling toward the mouth (seen in the images) and ends when either (1) the hand returns to a position of rest, (2) the hand is no longer visible in the captured images after putting the food in the mouth, or (3) the participant immediately starts the next feeding gesture by picking up the next morsel of food. Eating was determined by the annotator viewing the image (food or hand visible in the low-resolution image and act of eating present). An eating episode was a combination of one or more feeding gestures, each separated by no more than 15 minutes (as defined by recent studies [21, 77]) from adjacent feeding gestures.
Social presence:
We defined the wearer to be “alone” when no other visible face was present in the scene, as captured by the wearable camera. Social presence was defined as when one or more faces other than the participant’s were visible in the scene captured by the wearable camera. We considered a person’s face to be visible if at least one eye could be visually confirmed by the annotator. For cases where the back of the head of a bystander was seen, we assumed the bystander was not engaged with the wearer, and we did not consider that face to be visible, thus lacking social presence. This could become one of the limitations of our approach, which we address in Section 7.
3.2. Hardware Apparatus
We developed WildCam (Figure 1), an activity-oriented camera, which is the modified version of an activity-oriented camera system that has been previously envisioned by Alharbi et al. [7]. WildCam is a dual-stream sensing video camera with a fish-eye lens that records the wearer’s activities and part of the surrounding environment. WildCam is built around an STM32L4 microcontroller (ARM Cortex-M4), with two imagers: an RGB camera without IR filter (OmniVision OV26401 with an 180° fish-eye lens), and a 8 × 8 IR sensing array (GridEYE AMG8833). We also add night vision capability via an IR illumination LED. A user of WildCam wears the device in such a way as to ensure that their head is partially captured in the IR sensor’s field of view. The Panasonic GridEYE AMG8833 IR array has a narrower field of view (i.e., 60°) focused in the upper portion (facing the user) of the OV2640 view, allowing us to capture hand-to-head-related gestures (i.e., eating). WildCam gathers the raw video and IR data streams at 5 frames per second, encrypts these streams on the fly using a stream cipher (salsa20 [20]), and stores these images onto a microSD card. We extracted and processed the data streams offline when the participants returned the device.
3.3. Participants and Study Dataset
We recruited 10 participants (5 male, 5 female; mean [SD] age, 45.2 [14.8] years). Participants wore WildCam for 3 days, during which we collected data from the device. This study was approved by the Northwestern Institutional Review Board, and all participants provided written informed consent before enrolling in this study. As we were interested in understanding the eating behavior of individuals with obesity, we recruited participants with a body mass index >30 kgs/m2. Although recent studies have collected data with larger sample sizes, there are limited data for individuals with obesity. Recent work has shown that models developed using data from adults without obesity perform poorly for adults with obesity [77], thus making a case for building models with data from participants with obesity.
Participants were instructed to wear the WildCam device throughout the day while doing their everyday activities. We instructed participants to ensure the IR sensor could capture their faces while wearing the device by pointing the lens upward toward their chin. To ensure uniformity of data across participants, we randomly selected 8 unique hours of data from each participant across the 3 days, which resulted in labeling and analyzing 80 hours of data in the wild. Overall, our dataset had 1,440,000 frames of RGB and IR sensor images, and each frame in the video was labeled to indicate whether the participant was (1) eating or not eating and (2) alone or not alone in that frame. Figure 3 shows the distribution of the images across the two activities.
Figure 3:

Activities captured by WildCam and their overlap. Green indicates the number of frames labeled as eating, and red indicates the number of frames labeled as containing social presence.
4. SYSTEM PROCESSING FRAMEWORK
4.1. Overview
Figure 4 presents a functional overview of WildCam. Our proposed framework comprises the following stages:
Figure 4:

Functional overview of WildCam. Raw images are passed into the activity recognition models and used to manually annotat the ground truth. DBSCAN-based clustering is independently performed on the output of the activity recognition module and the manually annotated ground truth. A comparison of the clustering of the machine-learned (ML) model output is compared against the labeled ground truth’s clustering output to determine the performance of each model.
Labeling: Annotators label whether there is a face and/or eating gesture at the frame level (discussed in Section 3.1).
Activity recognition: Our system augments the RGB images with the IR sensor data and trains a model to predict (either a face or eating gesture) at the frame level using RGB data. Section 4.2 provides details of the activity recognition stage.
Post processing and episode generation: Our system uses the IR sensor array to remove false positives for each activity. At this stage, WildCam applies clustering techniques to generate start and end times of both ground truth and predicted events. Details of the episode generation stage and IR utility are provided in Section 4.2.3.
Evaluation: We evaluate the performance of our system by comparing the results of the machine learning approach and the ground truth at the episode level. Details of the evaluation are provided in Section 5.
4.2. Activity Recognition
We performed the activity recognition task independently for each of the two activities. For each activity, the overall activity recognition task comprised a preprocessing stage, a machine learning model creation stage, and a postprocessing and classification stage to generate episode level predictions.
4.2.1. Preprocessing.
The first step in detecting activities and context such as eating episode or social presence is to preprocess the raw images. The following steps were taken on the images for each of the activities and context recognition.
Image size and rotation (used in social presence recognition):
The WildCam hardware captures RGB images of 320 × 240 pixels. We tested each image at two scales, the original and double scale (i.e., 640 × 480 pixels). We realized that resizing the image had a substantial positive effect during inference, mainly due to the enlarged faces. Moreover, due to the distortion properties of a fish-eye lens, objects captured by the lens may not be oriented properly in the image (Figure 2). Therefore, this adds a layer of difficulty in detecting faces and objects at different orientations, which most face detection models fail to address. To address this, we rotated each image by 45°, 90°, 135°, and 180° and used the original image and the four rotated versions of it in creating the face detection model that is described later.
Dataset augmentation (used in social presence recognition):
Training data for the social presence recognition model consisted of data from 10 participants and around 37,000 images labeled with a bounding box around the faces. Although this sample size appeared to be enough to train a face detection model, there was low diversity in identity and facial characteristics in the dataset. To this end, the training set was augmented with WiderFace [73], a large face detection dataset with a high degree of variability in scale, pose, and occlusion. Images included in this dataset are mostly high-resolution images with more details compared with WildCam images. To account for this, we first modified images in the WiderFace dataset by using an averaging filter of size of 4 × 4 to simulate blurriness caused by the motion. Next, we performed the same JPEG compression used on-device with WildCam to diminish more details from faces and create similar faces in our dataset (Figure 5a, 5b and 5c).
Figure 5:

Data augmentation for face: Using image processing techniques, normal images in a public dataset (a) were degraded. As a result, more training data (b) similar to the data from our dataset (c) were generated.
Optical flow calculation (used in eating activity recognition):
Eating gesture is not an instantaneous activity but includes temporal information that one has to consider for reliable detection performance. One approach to quantifying this temporal information is by using optical flow. We relied on the deep-learning optical flow estimator FlowNet22, which can process up to 140 frames per second when estimating the optical flow.
4.2.2. Machine Learning Model Creation.
The preprocessed data are then fed into a machine learning model that can recognize specific activities.
Eating activity:
As mentioned before, there are two levels at which we detected the eating activity: (1) at a per-frame level and (2) at an episode level (described in Section 4.2.3). For the per-frame level, our model aimed to determine whether a participant was performing an eating gesture in frame Fi. We used a two-stream CNN architecture to detect the eating gesture in each frame, where the first stream represents the spatial channel and the second stream the temporal channel [65]. The spatial and temporal channels were trained independently with ResNet101 as the backbone.3
To train the network, we used a binary cross-entropy loss function. To address the class imbalance problem, we leveraged the approach used by Rouast et al. [54]. That is, in each mini-batch, we scaled the loss for class i by wi according to:
where m is the number of labels (y = {y1, …, ym}), n is the number of classes, and C(i) is the number of elements of y that equal yi.
Overall, the model classified the ith frame as eating or not eating based on a context of 12 other frames surrounding the target frame. The spatial stream accepted a single 320 × 240 × 3 frame Fi as input and classified it as eating or not eating. The temporal stream accepted stacked optical flows from the sequence of 13 input frames [Fi−6, Fi+6] to classify the Fi frame as eating or not eating. The final classification was found by late fusion in averaging the outputs of the spatial and temporal channels as suggested by Simonyan et al. [65].
Social presence:
Due to the camera’s low-resolution output and distance to bystanders, WildCam mainly captures tiny faces. Moreover, motion artifacts and the camera’s on-device JPEG compression further reduce important details from the images, posing a challenge to the face detection task. We thus used RetinaFace [27], a state-of-the-art tiny face–detection approach. RetinaFace uses extra supervision and multitask learning to overcome the challenge of detecting faces at various scales, including tiny ones. RetinaFace achieves a mean average precision of 91.4% on WiderFace, outperforming the state-of-the-art face detection model. However, this model was trained on the WiderFace [73] dataset, which comprises higher-resolution images and differs significantly from our dataset. Figure 5a and 5c shows sample images from WiderFace and our dataset. To account for the differences, we retrained RetinaFace with degraded versions of images from the WiderFace dataset and additional labeled RGB images from our camera’s RGB images. More details regarding dataset augmentation and preprocessing methods can be found in Section 4.2.1. All hyperparameters in the architecture and weights of the backbone network (i.e., Resnet-152), were kept the same as in the original work to keep the training process short. Using our model, we classified an image as having a face (i.e., not being alone) when there was at least one detected face above a confidence threshold across the five possible combinations of orientations; otherwise, the image was classified as alone.
4.2.3. Episode Generation: Postprocessing and Classification.
The final step in our system pipeline is the step that enables clustering data from the activity recognition stage to create activity-related episodes. However, before creating the episodes, the data are post-processed and cleaned as described below.
Postprocessing:
The IR sensor is capable of detecting the heat emitted by humans and objects in close vicinity to the camera. We utilized this capability of the IR sensor array to remove false positives (i.e., falsely detected humans and human-related activities) generated in the machine learning step described in Section 4.2.2, thus increasing the precision of the detected activity or event.
Episode generation:
Several researchers have generated episodes of activities and events to understand the specific temporal order of such activities/events. For eating activities, a number of researchers focused on evaluating eating episodes [77, 24, 60]. We focused on creating similar episodes by using a density-based approach (DB-SCAN) on frame-level ground truths and predictions. This method required two parameters, namely eps and minpts. Parameter eps was defined as the minimum distance needed between two clusters for the clusters to be merged into one, and minpts was defined as the minimum number of points required for a cluster to be created.
Eating episode and postprocessing:
We used a density-based (DBSCAN) approach, similar to one proposed by Zhang et al. [78], to create eating episodes. Similar to prior works, we used a 15-minute gap between gestures as the boundary between eating episodes [21, 78]. At the prediction stage, frame-level outputs were clustered with the same parameters used for ground truth to prepare predicted episodes for evaluation. We then used the IR data from each predicted episode to decide if it was a false positive. An eating episode consists of multiple hand-to-mouth gestures that can be captured and identified by the IR sensor array. Using this fact, false positive eating episodes initially detected by the RGB images were filtered consequently because of the absence of IR activity.
Social presence episode and postprocessing:
Episode-level predictions were achieved by clustering frame-level predictions via DBSCAN. Since there is no clear definition of an episode with or without social presence, we took a data-driven approach to determine DBSCAN parameters. First, we calculated the cumulative distribution function (CDF) of the number of continuous frames with faces (a proxy for minpts) and the number of frames between each set of successive frames with faces (a proxy for eps). We then identified the saturation point using a technique introduced by Ville et al. [56] for each CDF and chose minpts and eps accordingly. Due to the camera’s orientation and fish-eye lens, the wearer’s face was visible in the frames and was confounded with bystanders’ faces. As a result, we benefited from the IR sensor as a postprocessing step to estimate the position of the wearer’s head and filter false positive faces that overlapped with the wearer’s head. Our IR sensor captured the participant’s head in its small field of view; using the image, a mask was created on the participant’s head location by thresholding the body’s temperature. This was done by Otsu’s method [48], which automatically bins pixels into two classes based on pixel intensity. In our case, the classes were human and background. Then, overlap was calculated between the human mask and the bounding box area around the predicted face. The face was discarded if the overlap was above a threshold (empirically set to 200 pixels).
5. EVALUATION
5.1. Dataset Split
When splitting the data for training and testing purposes, we considered the variability in how people spend their time on different activities in a real-world daily-living situation. Figure 3 shows the imbalance in distribution of activities in our dataset, with some participants (e.g., P2, P3) having almost no social presence and some (e.g., P1) having approximately 50% of their total time being social presence. Due to this imbalance, common dataset split approaches such as person-independent training, validation, and testing might not have been appropriate. To this end, we split the dataset into three folds by manually picking participants such that there was sufficient data for each activity in each fold. Table 1 shows participant distribution in each fold to create three balanced splits. We ensured that we did not bias our testing process by performing cross-validation.
Table 1:
Study summary. Percentages show each participant’s share from total amount of frames labelled as that task.
| Participant | Eating | Socializing | |
|---|---|---|---|
| First Fold | P1 | 14% | 50% |
| P2 | 7% | 0% | |
| P3 | 2.3% | 0% | |
| Second Fold | P4 | 9.8% | 8.5% |
| P5 | 19.7% | 5.6% | |
| P6 | 5.1% | 3.1% | |
| P7 | 11.4% | 7.2% | |
| Third Fold | P8 | 7.3% | 11.4% |
| P9 | 8.2% | 11% | |
| P10 | 14.7% | 2.8% |
5.2. Evaluation Metric
For each activity and event, we were interested in understanding the prediction capabilities of the RGB images and observing improvement in results with the introduction of the IR sensor array. Thus, our evaluation metric revolved around the performance of detecting activities using RGB-only data and RGB+IR data at an episode level. Specifically, for eating episodes and social presence, we computed the precision, the recall, and the F1-score of detecting the episodes. This computation was done separately for RGB only and RGB+IR.
5.3. Episode-Level Evaluation
We compared our prediction approach against the ground truth to determine WildCam’s performance. False positivess occur when WildCam incorrectly predicts an episode when there is no episode, whereas false negatives occur when an episode present in the ground truth is not predicted. By comparing the occurrence of events in the ground truth episode and the prediction, we determine if a false positive or a false negative has occurred. If an episode is present and overlaps the ground truth and prediction, then it is a true positive. A true positive can span multiple ground truth events, or multiple true positives can occur for a single ground truth entry. We sum all true positives, false positives, and false negatives and use these values to compute precision, recall, and F1-score.
Eating detection:
Overall, there were 72 eating episodes in our dataset (ground truth). When we used the RGB-only approach and performed three-fold cross-validation, we attained a recall, precision, and F1-score of 67.16%, 63.38%, and 65.21%, respectively. Table 2 shows the performance for each participant independently. For one participant, we could precisely detect all recalled eating episodes. We next computed the performance of WildCam when we included the IR sensor as a postprocessing step. Table 2 also presents the performance of detecting eating episodes when we used the RGB+IR approach. There was a substantial increase (>10%) in the precision, indicating that the IR sensor successfully removed several false positives. Using the RGB+IR approach, the updated recall, precision, and F1-score were 67.16%, 73.77%, and 70.31%, respectively. Using the RGB+IR approach, for three participants, at a per-person level, we could achieve 100% precision, whereas for all participants, the precision and F1-score remained the same or improved compared with the RGB-only approach. Ultimately, the IR sensor helped remove 11 false positive episodes from the prediction.
Table 2:
Episode-level prediction result for eating episodes. RGB indicates that only RGB images were used for evaluation; RGB+IR indicates that both RGB and IR images were used for evaluation.
| RGB | RGB+IR | # Episodes | |||||
|---|---|---|---|---|---|---|---|
| Recall | Precision | F1-score | Recall | Precision | F1-score | ||
| P1 | 0.5 | 0.4 | 0.44 | 0.5 | 0.5 | 0.5 | 5 |
| P2 | 0.75 | 1 | 0.86 | 0.75 | 1 | 0.86 | 9 |
| P3 | 0.67 | 0.5 | 0.57 | 0.67 | 1 | 0.8 | 3 |
| P4 | 0.67 | 0.75 | 0.71 | 0.67 | 1 | 0.8 | 10 |
| P5 | 0.86 | 0.75 | 0.8 | 0.86 | 0.86 | 0.86 | 8 |
| P6 | 0.71 | 0.63 | 0.67 | 0.71 | 0.83 | 0.77 | 7 |
| P7 | 0.56 | 0.63 | 0.59 | 0.56 | 0.63 | 0.59 | 9 |
| P8 | 0.67 | 0.44 | 0.53 | 0.67 | 0.5 | 0.57 | 6 |
| P9 | 0.8 | 0.5 | 0.62 | 0.8 | 0.57 | 0.67 | 5 |
| P10 | 0.56 | 0.71 | 0.63 | 0.56 | 0.71 | 0.63 | 10 |
| Total | 0.67 | 0.63 | 0.65 | 0.67 | 0.74 | 0.7 | 72 |
Episodes indicates the number of eating episodes that were present in the ground truth.
Social presence:
In our dataset, 33 episodes indicated social presence. However, for two participants, there were no social events at all. Overall, of these 33 episodes, we could detect 27 episodes using the RGB approach. Table 3 presents the per-participant performance. Overall, we attained a recall, precision, and F1-score of 81.81%, 18.12%, and 29.67%, respectively. We had excessive false positives, which adversely affected our precision and F1-score and introduced substantial noise in our prediction.
Table 3:
Episode-level evaluation result for social presence. RGB indicates that only RGB images were used for evaluation; RGB+IR indicates that both RGB and IR images were used for evaluation.
| RGB | RGB+IR | # Episodes | |||||
|---|---|---|---|---|---|---|---|
| Recall | Precision | F1-score | Recall | Precision | F1-score | ||
| P1 | 0.67 | 0.3 | 0.41 | 0.67 | 0.86 | 0.75 | 9 |
| P2 | - | - | - | - | - | - | 0 |
| P3 | - | - | - | - | - | - | 0 |
| P4 | 1 | 0.17 | 0.29 | 1 | 0.75 | 0.86 | 3 |
| P5 | 0.5 | 0.17 | 0.25 | 0.5 | 0.5 | 0.5 | 2 |
| P6 | 1 | 0.06 | 0.11 | 1 | 1 | 1 | 1 |
| P7 | 1 | 0.4 | 0.57 | 1 | 1 | 1 | 2 |
| P8 | 0.75 | 0.33 | 0.46 | 0.5 | 0.8 | 0.62 | 8 |
| P9 | 1 | 0.33 | 0.5 | 1 | 0.89 | 0.94 | 7 |
| P10 | 1 | 0.07 | 0.13 | 1 | 0.29 | 0.44 | 1 |
| Total | 0.82 | 0.18 | 0.3 | 0.77 | 0.71 | 0.74 | 33 |
Episodes indicates the number of social episodes that were present in the ground truth.
However, with the introduction of the IR sensor in conjunction with RGB, we substantially reduced these false positives. As a result, the precision improved from 18.12% to 71.05%, an approximately 53% improvement. Due to this improvement in precision, the overall F1-score also improved to 73.97%. Thus, the IR sensor successfully captured the wearer’s head and face and distinguished the wearer’s face from the bystander’s face.
Summary:
Overall, the RGB-only approach of detecting eating episodes and social presence was able to attain reasonable performance. More importantly, eating episode and social presence detection can substantially benefit from using an IR sensor array in its postprocessing step. Indeed, the precision for detecting social presence increased by almost 53% when the IR sensor data were used, thus showing the benefit of the RGB+IR approach.
6. IMPLICATIONS FOR DESIGN: AN ASSESSMENT OF FALSE POSITIVES AND FALSE NEGATIVES
In this section we further elaborate on false positive and false negative cases in which WildCam failed. By identifying the problems, we identify potential solutions as a guide for the design of a future all-day camera that can validate eating and social presence. Table 4 categorizes the reasons for false positives and false negatives for each class label along with their prevalence.
Table 4:
Analysis of false positives and false negatives for activities and events monitored by WildCam as observed in our dataset.
| False Positive Cases | False Negative Cases | |||
|---|---|---|---|---|
| Reason | Share | Reason | Share | |
| Eating Episode | Confounding hand gestures | 0.57 | Drinking gestures | 0.9 |
| Device motion artifacts | 0.43 | Partial occlusion | 0.1 | |
| Social Presence | Video calls | 0.19 | ||
| I see ghosts | 0.32 | Tiny faces | 0.69 | |
| I see myself | 0.31 | Sensor design | 0.31 | |
| Mirror mirror on the wall | 0.19 | |||
6.1. Eating Episodes
We observed two primary causes of false positives in eating episode detection: non-eating-related hand-to-mouth gestures and significant camera motion. Drinking episodes were the main cause of false negatives.
6.1.1. Non-eating-related hand-to-mouth gestures—57% of all false positives.
One common problem of false positives is detecting hand-to-mouth gestures from gestures not involving food intake. Common techniques such as optical flow often confuse actions like playing with hair, touching the mouth, scratching the face, and putting on headphones or glasses with eating gestures. Prior work has explored the role of food and food detection models in determining eating moments [70, 13]. Given the current perspective of our camera and its resolution, food detection becomes an even harder challenge. In the future, we can explore removing such false positives by including a chewing detection module that can filter out hand-to-mouth events that do not lead to a chewing action. Alternatively, we could combine a food detection module with a hand-to-mouth detection module separately to prevent such false positives. Figure 6a–d shows examples of hand-to-head gestures that are confused with eating.
Figure 6:

Sample cases causing false positives and/or false negatives in the system. a–d: Hand-to-mouth gestures that are confused with eating. e–g: Sensor position can affect the field of view captured by the IR sensor. Red boxes indicate field of view captured by the IR sensor.
6.1.2. Device movement—43% of all false positives.
Detecting an eating episode relies on computing the optical flow. To compute the optical flow, we considered a set of continuous images captured by WildCam and computed the distance by which the points in the image had moved. We observed that, at times, unintended device movement (due to the unstable position on the participant’s shirt) caused motion in the optical flow image that appeared to be similar to hand motion during a feeding gesture. Additionally, in some cases we observed motion artifacts in the camera causing confusion to the model.
6.1.3. Drinking gestures—90% of all false negatives.
As drinking a beverage contributes toward caloric consumption, we included drinking as an eating activity and annotated the dataset accordingly. However, several prior efforts exclude drinking gestures [78, 24]. Drinking gestures, however, have unique temporal patterns and in our data occurred less frequently than other eating gestures. They also often extended beyond the hand and occluded a large portion of the hand, making automated detection of the hand a challenge. Future work should look into designing a separate beverage detection model to ensure reliable detection of drinking gestures.
6.1.4. Partial occlusion—10% of all false negatives.
At times the wearer’s clothing or the table would partially occlude the image. Annotators were still able to visually confirm the food in the image; however, the model was not able to recognize the hand in the images.
6.2. Social Presence
6.2.1. Video calls—18.75% of all false positives.
Our current definition of social presence excluded the presence of video calls. However, we observed that several participants spent time in video calls on the smartphone. This highlights the need to either modify the definition to take into account the fact that social presence may occur through video calls or ensure our model is able to distinguish faces detected on screens from faces in real life. Alternatively, we could use data from the person’s smartphone to filter out these false positives or combine screen detection with social presence to filter faces bound within screens.
6.2.2. I see ghosts—31.75% of all false positives.
Due to the low-resolution nature of the RGB camera used, we trained the model using tiny faces with very few pixels on the face. This resulted, at times, to the model “seeing” faces that were not really there. We could potentially maintain a minimum number of pixels on target for detection, or we could slightly increase the resolution of the RGB camera to prevent these false positives.
6.2.3. I see myself—30.75% of all false positives.
As shown in Figure 6f and g, sometimes the participant wore the device incorrectly, with the IR sensor fully covered by the participant’s body or with the participant leaning over the IR sensor with his/her face appearing on the other side, resulting in a false positive. To mitigate this problem, a wider IR field of view can allow us to fully detect the participant’s body. Integrating a model that can track the participant’s face can also help address this problem.
6.2.4. Mirror mirror on the wall—18.75% of all false positives.
Some participants spent a lot of time in front of a mirror doing tasks such as putting on makeup or brushing their teeth. These activities became a part of the individual’s daily routines and cannot be ignored. Our face detection model was not able to distinguish between a wearer’s face and their face in a mirror, which appeared as a bystander’s face. Therefore, we ended up capturing a significant amount of false positives related to this scenario. Needed is a method to detect mirrors to filter out these false positives.
6.2.5. Tiny faces—69.25% of all false negatives.
Given the tiny faces in the image, there were scenarios where the faces were at the edge of the RGB camera. These faces were classified as hard cases according to the WiderFace dataset [73]. The only way to solve this problem would be to increase the RGB camera’s field of view or to increase the camera’s resolution.
7. DISCUSSION AND FUTURE WORK
Our analysis of RGB and IR data over 3 days from 10 participants with obesity shows that our system can accurately classify eating episodes and social presence in people with obesity in natural settings. Given a clip of RGB and IR sensor data from a camera oriented toward the wearer’s mouth, generating images at a low resolution, we address two key technical challenges of variable orientation and partial multi-sensor overlap. Variable orientation and resolution refer to the fact that the number of pixels on an object for an activity of interest varies as the participant moves in the scene. For example, given the varying positioning of the device, fish-eye lens, and motion of other people in the scene, the bystander’s face can appear at several different resolutions and orientations. This is in contrast to prior work that assumes higher-resolution images of people, with faces appearing in the same orientation throughout the video footage. We address this by testing our classifiers with images of the object at various resolutions and testing the faces at different rotations to ensure the classifiers can detect tiny faces at different orientations in the scene. The second challenge of partial sensor overlap has to do with the varying resolution of the IR sensor in relation to the RGB camera, which results in objects appearing entirely in one sensor and partially in another. As shown in Figure 1, the IR sensor overlaps with the center of the RGB frame. In contrast, prior work on object detection using multiple sensing modalities assume similar resolutions [71]. To address this, we used IR sensor data during the postprocessing step of the data pipeline to confirm or refute a detected object in the RGB image only if the object overlaped with the field of view of the IR sensor. Therefore, we could infer the IR sensor’s utility in confirming or refuting the object in the scene. Moroever, we show how future behavioral researchers could use the WildCam system for defining and validating various behavioral constructs.
WildCam currently uses images captured from a wearable camera primarily intended to capture the wearer and activities around the wearer. However, the privacy of both the wearer and bystanders is a concern with camera images. In the future, we plan to only record and train models using obfuscated images. To this end, we will run image obfuscation approaches to remove human faces from the images to further protect bystander privacy [6].
8. CONCLUSION
Studying human behavior has traditionally been theory based, deriving a set of constructs or complex concepts (e.g., self-efficacy or attitude toward a behavior) and relationships between variables to explain or predict behaviors. These approaches are often not based on objective, timely measures of behavior but rather on a limited set of tools, implemented as questionnaires administered in structured settings like a classroom or laboratory or over the internet.
Wearable mobile sensors can enable quantitative representation of coarse-grained (e.g., eating alone or with friends) and fine-grained (e.g., number of feeding gestures) behaviors, enabling the development of new behavioral models that more precisely explain and predict problematic behaviors. This will enable behavioral scientists to construct models that more closely represent reality and enable the improved design of interventions that change behavior. In this paper, we presented how researchers can use a low-power device that consists of an RGB camera and IR sensor to detect eating episodes and social presence via a free-living user study with 80 hours of data. Overall, we showed that the WildCam framework could detect these activities and events reliably when the RGB and IR sensor were used in conjunction. There are several scenarios where the WildCam framework is affected by false positives and false negatives, and we discuss possible modifications to the framework that can help minimize these types of errors. Overall, the WildCam framework will be a powerful tool for behavioral researchers interested in studying human behaviors related to eating and beyond.
Footnotes
REFERENCES
- [1].Adate Amit, Shahi Soroush, Alharbi Rawan, Sen Sougata, Gao Yang, Katsaggelos Aggelos K, and Alshurafa Nabil. 2022. Detecting screen presence with activity-oriented rgb camera in egocentric videos. In 2022 IEEE International Conference on Pervasive Computing and Communications Workshops and other Affiliated Events (PerCom Workshops), 403–408. doi: 10.1109/PerComWorkshops53856.2022.9767433. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [2].Adib Fadel and Katabi Dina. 2013. See through walls with wifi! In Proceedings of the ACM SIGCOMM 2013 Conference on SIGCOMM, 75–86. [Google Scholar]
- [3].Alharbi Rawan, Feng Chunlin, Sen Sougata, Jain Jayalakshmi, Hester Josiah, and Alshurafa Nabil. 2021. Heatsight: wearable low-power omni thermal sensing. In 2021 International Symposium on Wearable Computers (ISWC ‘21). Association for Computing Machinery, Virtual, USA, 108–112. isbn: 9781450384629. doi: 10.1145/3460421.3478811. [DOI] [Google Scholar]
- [4].Alharbi Rawan, Sen Sougata, Ng Ada, Alshurafa Nabil, and Hester Josiah. 2022. Actisight: wearer foreground extraction using a practical rgb-thermal wearable. In 2022 IEEE International Conference on Pervasive Computing and Communications (PerCom), 237–246. doi: 10.1109/PerCom53586.2022.9762385. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [5].Alharbi Rawan, Stump Tammy, Vafaie Nilofar, Pfammatter Angela, Spring Bonnie, and Alshurafa Nabil. 2018. I Can’t Be Myself: Effects of Wearable Cameras on the Capture of Authentic Behavior in the Wild. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, 2, 3, (Sept. 2018), 1–40. doi: 10.1145/3264900. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [6].Alharbi Rawan, Tolba Mariam, Petito Lucia C, Hester Josiah, and Alshurafa Nabil. 2019. To mask or not to mask? balancing privacy with visual confirmation utility in activity-oriented wearable cameras. Proceedings of the ACM on interactive, mobile, wearable and ubiquitous technologies, 3, 3, 1–29. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [7].Alharbi Rawan, Tolba Mariam, Petito Lucia C., Hester Josiah, and Alshurafa Nabil. 2019. To Mask or Not to Mask? Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, 3, 3, 1–29. doi: 10.1145/3351230. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [8].Alharbi Rawan et al. 2023. Smokemon: unobtrusive extraction of smoking topography using wearable energy-efficient thermal. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol, 6, 4, Article 155, (Jan. 2023), 25 pages. doi: 10.1145/3569460. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [9].Alletto Stefano, Serra Giuseppe, Calderara Simone, Solera Francesco, and Cucchiara Rita. 2014. From ego to nos-vision: detecting social relationships in first-person views. In Conference on Computer Vision and Pattern Recognition Workshops. IEEE, 580–585. [Google Scholar]
- [10].Alshurafa Nabil, Wen Lin Annie, Zhu Fengqing, Ghaffari Roozbeh, Hester Josiah, Delp Edward, Rogers John, and Spring Bonnie. 2019. Counting bites with bits: expert workshop addressing calorie and macronutrient intake monitoring. Journal of Medical Internet Research, 21, 12, (Dec. 2019), e14904. doi: 10.2196/14904. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [11].Alshurafa Nabil, Annie Wen Lin Fengqing Zhu, Ghaffari Roozbeh, Hester Josiah, Delp Edward, Rogers John, and Spring Bonnie. 2019. Counting bites with bits: Expert workshop addressing calorie and macronutrient intake monitoring. (Dec. 2019). doi: 10.2196/14904. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [12].Amft Oliver and Troster Gerhard. 2009. On-body sensing solutions for automatic dietary monitoring. IEEE pervasive computing, 8, 2, 62–70. [Google Scholar]
- [13].Arab L, Estrin D, Kim DH, Burke J, and Goldman J. 2011. Feasibility testing of an automated image-capture method to aid dietary recall. European Journal of Clinical Nutrition, 65, 10, (May 2011), 1156–1162. doi: 10.1038/ejcn.2011.75. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [14].Atzmueller Martin, Thiele Lisa, Stumme Gerd, and Kauffeld Simone. 2018. Analyzing group interaction on networks of face-to-face proximity using wearable sensors. In International Conference on Future IoT Technologies (Future IoT). IEEE, 1–10. [Google Scholar]
- [15].Barsocchi Paolo, Crivello Antonino, Girolami Michele, and Mavilia Fabio. 2020. Detecting social interactions in indoor environments with the red-HuP algorithm. In International Conference on Pervasive Computing and Communications Workshops (PerCom Workshops). IEEE, (Mar. 2020). doi: 10.1109/percomworkshops48775.2020.9156095. [DOI] [Google Scholar]
- [16].Beaglehole Robert et al. 2011. Priority actions for the non-communicable disease crisis. The Lancet, 377, 9775, 1438–1447. doi: 10.1016/S0140-6736(11)60393-0. [DOI] [PubMed] [Google Scholar]
- [17].Bedri Abdelkareem, Li Diana, Khurana Rushil, Bhuwalka Kunal, and Goel Mayank. 2020. FitByte: Automatic Diet Monitoring in Unconstrained Situations Using Multimodal Sensing on Eyeglasses. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems. Vol. 20. Association for Computing Machinery (ACM), New York, NY, USA, (Apr. 2020), 1–12. doi: 10.1145/3313831.3376869. [DOI] [Google Scholar]
- [18].Bell Brooke M., Alam Ridwan, Alshurafa Nabil, Thomaz Edison, Mondol Abu S., de la Haye Kayla, Stankovic John A., Lach John, and Spruijt-Metz Donna. 2020. Automatic, wearable-based, in-field eating detection approaches for public health research: a scoping review. npj Digital Medicine, 3, 1, (Dec. 2020), 1–14. doi: 10.1038/s41746-020-0246-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [19].F Bellisle, Dalix AM, and Slama G. 2004. Non food-related environmental stimuli induce increased meal intake in healthy women: comparison of television viewing versus listening to a recorded story in laboratory settings. Appetite, 43, 2, 175–180. [DOI] [PubMed] [Google Scholar]
- [20].Bernstein Daniel J. 2008. The salsa20 family of stream ciphers. In New stream cipher designs. Springer, 84–97. [Google Scholar]
- [21].Bi Shengjie et al. 2018. Auracle: Detecting Eating Episodes with an Ear-mounted Sensor. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, 2, 3, (Sept. 2018), 1–27. doi: 10.1145/3264902. [DOI] [Google Scholar]
- [22].Borazio Marko and Van Laerhoven Kristof. 2012. Combining wearable and environmental sensing into an unobtrusive tool for long-term sleep studies. In ACM SIGHIT International Health Informatics Symposium, 71–80. doi: 10.1145/2110363.2110375. [DOI] [Google Scholar]
- [23].Cattuto Ciro, Van den Broeck W outer, Barrat Alain, Colizza Vittoria, Pinton Jean-François, and Vespignani Alessandro. 2010. Dynamics of person-to-person interactions from distributed RFID sensor networks. PLoS ONE, 5, 7, (July 2010), e11596. Neylon Cameron, (Ed.) doi: 10.1371/journal.pone.0011596. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [24].Chun Keum San, Bhattacharya Sarnab, and Thomaz Edison. 2018. Detecting eating episodes by tracking jawbone movements with a non-contact wearable sensor. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol, 2, 1, Article 4, (Mar. 2018), 21 pages. doi: 10.1145/3191736. [DOI] [Google Scholar]
- [25].De Castro John M. 2000. Eating behavior: lessons from the real world of humans. Nutrition, 16, 10, 800–813. [DOI] [PubMed] [Google Scholar]
- [26].de Castro John M. and Brewer E.Marie. 1992. The amount eaten in meals by humans is a power function of the number of people present. Physiology & Behavior, 51, 1, 121–125. doi: 10.1016/0031-9384(92)90212-K. [DOI] [PubMed] [Google Scholar]
- [27].Deng Jiankang, Guo Jia, Zhou Yuxiang, Yu Jinke, Kotsia Irene, and Zafeiriou Stefanos. 2019. Retinaface: single-stage dense face localisation in the wild. (2019). arXiv: 1905.00641 [cs.CV]. [Google Scholar]
- [28].Dong Bo and Biswas Subir. 2013. Wearable diet monitoring through breathing signal analysis. In Proceedings of the Annual International Conference of the IEEE Engineering in Medicine and Biology Society, EMBS, 1186–1187. isbn: 9781457702167. doi: 10.1109/EMBC.2013.6609718. [DOI] [PubMed] [Google Scholar]
- [29].Echterhoff Jessica Maria and Wang Edward J. 2020. Par: personal activity radius camera view for contextual sensing. arXiv preprint arXiv:2008.07204, -, -, 15 pages. [Google Scholar]
- [30].Fathi Alircza, Hodgins Jessica K, and Rehg James M. 2012. Social interactions: a first-person perspective. In 2012 IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 1226–1233. [Google Scholar]
- [31].Forman Evan M, Schumacher Leah M, Crosby Ross, Manasse Stephanie M, Goldstein Stephanie P, Butryn Meghan L, Wyckoff Emily P, and Thomas J Graham. 2017. Ecological momentary assessment of dietary lapses across behavioral weight loss treatment: characteristics, predictors, and relationships with weight change. Annals of Behavioral Medicine, 51, 5, 741–753. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [32].Haines Dena. 2020. How long does a gopro battery last. https://clicklikethis.com/how-long-does-a-gopro-battery-last/. [Online; accessed 8-April-2021]. (Sept. 2020).
- [33].Harjanto Fredro, Wang Zhiyong, Lu Shiyang, Tsoi Ah Chung, and Feng David Dagan. 2016. Investigating the impact of frame rate towards robust human action recognition. Signal Processing, 124, 220–232. Big Data Meets Multimedia Analytics. doi: 10.1016/j.sigpro.2015.08.006. [DOI] [Google Scholar]
- [34].Heravi Behzad M, Gibson Jenny L, Hailes Stephen, and Skuse David. 2018. Playground social interaction analysis using bespoke wearable sensors for tracking and motion capture. In Proceedings of the 5th International Conference on Movement and Computing, 1–8. [Google Scholar]
- [35].Williams Lyndsay Hodges Steve, Berry Emma, Izadi Shahram, Srinivasan James, Butler Alex, Smyth Gavin, Kapur Narinder, and Wood Ken. 2006. SenseCam: A retrospective memory aid. In International conference on ubiquitous computing. [Google Scholar]
- [36].Intille Stephen S, Larson Kent, Beaudin JS, Nawyn Jason, Tapia E Munguia, and Kaushik Pallavi. 2005. A living laboratory for the design and evaluation of ubiquitous computing technologies. In Extended abstracts on Human factors in computing systems (CHI’05), 1941–1944. [Google Scholar]
- [37].Jia Wenyan et al. 2019. Automatic food detection in egocentric images using artificial intelligence technology. Public health nutrition, 22, 7, 1168–1179. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [38].Kissileff Harry R. 1992. Where should human eating be studied and what should be measured? Appetite, 19, 1, 61–68. [DOI] [PubMed] [Google Scholar]
- [39].Koziarski Michal and Cyganek Boguslaw. 2018. Impact of low resolution on image recognition with deep neural networks: an experimental study. International Journal of Applied Mathematics and Computer Science, 28, 4. [Google Scholar]
- [40].Lenard Natalie R and Berthoud Hans-Rudolf. 2008. Central and peripheral regulation of food intake and physical activity: pathways and genes. Obesity, 16, S3, S11–S22. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [41].Likamwa Robert, Priyantha Bodhi, Philipose Matthai, Zhong Lin, and Bahl Paramvir. 2013. Energy characterization and optimization of image sensing toward continuous mobile vision. In MobiSys 2013 - Proceedings of the 11th Annual International Conference on Mobile Systems, Applications, and Services, 69–81. isbn: 9781450316729. doi: 10.1145/2462456.2464448. [DOI] [Google Scholar]
- [42].Meiselman Herbert L. 1992. Methodology and theory in human eating research. Appetite, 19, 1, 49–55. [DOI] [PubMed] [Google Scholar]
- [43].Meiselman Herbert L. 1992. Obstacles to studying real people eating real meals in real situations.
- [44].Meiselman Herbert L et al. 2006. The role of context in food choice, food acceptance and food consumption. Frontiers in nutritional science, 3, 179. [Google Scholar]
- [45].Mirtchouk Mark, Lustig Drew, Smith Alexandra, Ching Ivan, Zheng Min, and Kleinberg Samantha. 2017. Recognizing Eating from Body-Worn Sensors. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, 1, 3, (Sept. 2017), 1–20. doi: 10.1145/3131894. [DOI] [Google Scholar]
- [46].Nam Yunyoung, Kim Yeesock, and Lee Jinseok. 2016. Sleep Monitoring Based on a Tri-Axial Accelerometer and a Pressure Sensor. Sensors, 16, 5, (May 2016), 750. doi: 10.3390/s16050750. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [47].Okuno Akane and Sumi Yasuyuki. 2019. Social activity measurement by counting faces captured in first-person view lifelogging video. In Proceedings of the 10th Augmented Human International Conference 2019, 1–9. [Google Scholar]
- [48].Otsu Nobuyuki. 1979. A threshold selection method from gray-level histograms. IEEE Transactions on Systems, Man, and Cybernetics, 9, 1, 62–66. doi: 10.1109/TSMC.1979.4310076. [DOI] [Google Scholar]
- [49].Pirsiavash Hamed and Ramanan Deva. 2012. Detecting activities of daily living in first-person camera views. In IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2847–2854. doi: 10.1109/CVPR.2012.6248010. [DOI] [Google Scholar]
- [50].Pliner Patricia. 1992. Let’s not throw out the barley with the dishwater: comments on meiselman’s” methodology and theory in human eating research.”.
- [51].Rahman Tauhidur, Adams Alexander T., Zhang Mi, Cherry Erin, Zhou Bobby, Peng Huaishu, and Choudhury Tanzeem. [n. d.] BodyBeat: A mobile system for sensing non-speech body sounds. In Proceedings of the Annual International Conference on Mobile Systems, Applications, and Services (Mobisys’14). Association for Computing Machinery. isbn: 9781450327930. doi: 10.1145/2594368.2594386. [DOI] [Google Scholar]
- [52].Reddy Sasank, Parker Andrew, Hyman Josh, Burke Jeff, Estrin Deborah, and Hansen Mark. 2007. Image browsing, processing, and clustering for participatory sensing: Lessons from a DietSense prototype. In Proceedings of the 4th Workshop on Embedded Networked Sensors, EmNets 2007, 13–17. isbn: 9781595936943. doi: 10.1145/1278972.1278975. [DOI] [Google Scholar]
- [53].Rolls Barbara J and Shide David J. 1992. Both naturalistic and laboratory-based studies contribute to the understanding of human eating behavior. [DOI] [PubMed]
- [54].Rouast Philipp V. and Adam Marc T. P.. 2020. Learning deep representations for video-based intake gesture detection. IEEE Journal of Biomedical and Health Informatics, 24, 6, (June 2020), 1727–1737. doi: 10.1109/jbhi.2019.2942845. [DOI] [PubMed] [Google Scholar]
- [55].Sallis James F, Owen Neville, and Fotheringham Michael J. 2000. Behavioral epidemiology: a systematic framework to classify phases of research on health promotion and disease prevention. Annals of behavioral medicine, 22, 4, 294–298. [DOI] [PubMed] [Google Scholar]
- [56].Satopaa Ville, Albrecht Jeannie, Irwin David, and Raghavan Barath. 2011. Finding a “kneedle” in a haystack: detecting knee points in system behavior. In 2011 31st International Conference on Distributed Computing Systems Workshops, 166–171. doi: 10.1109/ICDCSW.2011.20. [DOI] [Google Scholar]
- [57].Sazonov Edward, Schuckers Stephanie, Lopez-Meyer Paulo, Makeyev Oleksandr, Sazonova Nadezhda, Melanson Edward L, and Neuman Michael. 2008. Non-invasive monitoring of chewing and swallowing for objective quantification of ingestive behavior. Physiological measurement, 29, 5, 525. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [58].Schiboni Giovanni, Wasner Fabio, and Amft Oliver. 2018. A privacy-preserving wearable camera setup for dietary event spotting in free-living. In 2018 IEEE International Conference on Pervasive Computing and Communications Workshops (PerCom Workshops). IEEE, 872–877. [Google Scholar]
- [59].Sen Sougata, Subbaraju Vigneshwaran, Misra Archan, Balan Rajesh, and Lee Youngki. 2020. Annapurna: an automated smartwatch-based eating detection and food journaling system. Pervasive and Mobile Computing, 68, 101259. [Google Scholar]
- [60].Sen Sougata, Subbaraju Vigneshwaran, Misra Archan, Balan Rajesh, and Lee Youngki. 2020. Annapurna: an automated smartwatch-based eating detection and food journaling system. Pervasive and Mobile Computing, 68, 101259. doi: 10.1016/j.pmcj.2020.101259. [DOI] [Google Scholar]
- [61].Sen Sougata, Subbaraju Vigneshwaran, Misra Archan, Balan Rajesh Krishna, and Lee Youngki. 2015. The case for smartwatch-based diet monitoring. In IEEE International Conference on Pervasive Computing and Communication Workshops, PerCom Workshops 2015. Institute of Electrical and Electronics Engineers Inc., (June 2015), 585–590. isbn: 9781479984251. doi: 10.1109/PERCOMW.2015.7134103. [DOI] [Google Scholar]
- [62].Shahi Soroush, Alharbi Rawan, Gao Yang, Sen Sougata, Katsaggelos Aggelos K, Hester Josiah, and Alshurafa Nabil. 2022. Impacts of image obfuscation on fine-grained activity recognition in egocentric video. In 2022 IEEE International Conference on Pervasive Computing and Communications Workshops and other Affiliated Events (PerCom Workshops), 341–346. doi: 10.1109/PerComWorkshops53856.2022.9767447. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [63].Shahi Soroush, Pedram Mahdi, Fernandes Glenn, and Alshurafa Nabil. 2022. Smartact: energy efficient and real-time hand-to-mouth gesture detection using wearable rgb-t. In 2022 IEEE-EMBS International Conference on Wearable and Implantable Body Sensor Networks (BSN), 1–4. doi: 10.1109/BSN56160.2022.9928492. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [64].Shetty Akshaya D, Shubha B, Suryanarayana K, et al. 2017. Detection and tracking of a human using the infrared thermopile array sensor—“grid-eye”. In 2017 International Conference on Intelligent Computing, Instrumentation and Control Technologies (ICICICT). IEEE, 1490–1495. [Google Scholar]
- [65].Simonyan Karen and Zisserman Andrew. 2014. Two-stream convolutional networks for action recognition in videos. arXiv preprint arXiv:1406.2199. [Google Scholar]
- [66].Stellar Eliot. 1992. Real eating and the measurement of real physiological and behavioral variables. [DOI] [PubMed]
- [67].Stroebele Nanette and De Castro John M. 2004. Effect of ambience on food intake and food choice. Nutrition, 20, 9, 821–838. [DOI] [PubMed] [Google Scholar]
- [68].Thomaz Edison. 2013. Practical food journaling. In UbiComp 2013 Adjunct - Adjunct Publication of the 2013 ACM Conference on Ubiquitous Computing, 355–360. isbn: 9781450322157. doi: 10.1145/2494091.2501089. [DOI] [Google Scholar]
- [69].Thomaz Edison, Essa Irfan, and Abowd Gregory D.. 2015. A practical approach for recognizing eating moments with wrist-mounted inertial sensing. In Ubi- Comp 2015 - Proceedings of the 2015 ACM International Joint Conference on Pervasive and Ubiquitous Computing. Association for Computing Machinery, Inc, (Sept. 2015), 1029–1040. isbn: 9781450335744. doi: 10.1145/2750858.2807545. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [70].Thomaz Edison, Parnami Aman, Essa Irfan, and Abowd Gregory D.. 2013. Feasibility of identifying eating moments from first-person images leveraging human computation. In Proceedings of the 4th International SenseCam & Pervasive Imaging Conference on - SenseCam ‘13. ACM Press. doi: 10.1145/2526667.2526672. [DOI] [Google Scholar]
- [71].Tu Zhengzheng, Ma Yan, Li Zhun, Li Chenglong, Xu Jieming, and Liu Yongtao. 2020. Rgbt salient object detection: a large-scale dataset and benchmark. (2020). arXiv: 2007.03262 [cs.CV]. [Google Scholar]
- [72].Tuorila Hely and Lähteenmäki L. 1992. When is eating “real”? response to meiselman. Appetite, 19, 1, 80–83. [DOI] [PubMed] [Google Scholar]
- [73].Yang Shuo, Luo Ping, Loy Chen Change, and Tang Xiaoou. 2016. Wider face: a face detection benchmark. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR). [Google Scholar]
- [74].Yatani Koji and Truong Khai N. 2012. Bodyscope: a wearable acoustic sensor for activity recognition. In Conference on Ubiquitous Computing (UbiComp). ACM, New York, NY, 341–350. [Google Scholar]
- [75].Yin Cunyi, Chen Jing, Miao Xiren, Jiang Hao, and Chen Deying. 2021. Device-free human activity recognition with low-resolution infrared array sensor using long short-term memory neural network. Sensors, 21, 10. doi: 10.3390/s21103551. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [76].Zhang Shibo, Alharbi Rawan, Nicholson Matthew, and Alshurafa Nabil. 2017. When generalized eating detection machine learning models fail in the field. In Proceedings of the 2017 ACM International Joint Conference on Pervasive and Ubiquitous Computing and Proceedings of the 2017 ACM International Symposium on Wearable Computers (UbiComp ‘17). Association for Computing Machinery, Maui, Hawaii, 613–622. isbn: 9781450351904. doi: 10.1145/3123024.3124409. [DOI] [Google Scholar]
- [77].Zhang Shibo, Zhao Yuqi, Nguyen Dzung Tri, Xu Runsheng, Sen Sougata, Hester Josiah, and Alshurafa Nabil. 2020. NeckSense: A Multi-Sensor Necklace for Detecting Eating Activities in Free-Living Conditions. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, 4, 2, (June 2020), 1–26. doi: 10.1145/3397313. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [78].Zhang Shibo, Zhao Yuqi, Nguyen Dzung Tri, Xu Runsheng, Sen Sougata, Hester Josiah, and Alshurafa Nabil. 2020. Necksense: a multi-sensor necklace for detecting eating activities in free-living conditions. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, 4, 2, 1–26. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [79].Zhang Yun C., Zhang Shibo, Liu Miao, Daly Elyse, Battalio Samuel, Kumar Santosh, Spring Bonnie, Rehg James M., and Alshurafa Nabil. 2020. Syncwise: window induced shift estimation for synchronization of video and accelerometry from wearable sensors. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol, 4, 3, Article 107, (Sept. 2020), 26 pages. doi: 10.1145/3411824. [DOI] [Google Scholar]
- [80].Zhao Jiewen, Han Ruize, Gan Yiyang, Wan Liang, Feng Wei, and Wang Song. 2020. Human identification and interaction detection in cross-view multi-person videos with wearable cameras. In Proceedings of the 28th ACM International Conference on Multimedia, 2608–2616. [Google Scholar]
