Skip to main content
Scientific Reports logoLink to Scientific Reports
. 2026 May 7;16:17684. doi: 10.1038/s41598-026-48086-1

Anomaly-based data reduction for energy-efficient edge computing in IoT with LoRa

Fadime Karadas 1,✉,#, Bilal Usanmaz 1,✉,#
PMCID: PMC13246739  PMID: 42098196

Abstract

Edge computing, a key component of Internet of Things (IoT) systems, enables data processing close to the data source. In low-power, resource-constrained IoT environments, it can reduce dependency on centralized cloud systems, lower communication load, and minimize overall energy consumption. However, transmitting all sensor data from edge devices still incurs significant communication and energy costs. The key research gap is that existing approaches rarely address this inefficiency through selective transmission; specifically, few studies have explored how filtering and sending only anomalous data can improve energy efficiency. To bridge this gap, we propose an anomaly detection-based data reduction method operating at the edge using LoRa technology. We apply two unsupervised learning algorithms, DBSCAN and Isolation Forest, to identify and transmit only anomalous sensor instances. The proposed methods are evaluated against a full-data transmission (FDT) baseline. The key contributions are demonstrated through experimental results: DBSCAN reduces data volume by 98.19% and energy consumption by 98.10%, while Isolation Forest achieves reductions of 97.32% and 97.32%, respectively. These findings confirm that anomaly-driven selective transmission significantly reduces the communication load while ensuring high energy efficiency in edge computing systems.

Keywords: Edge Computing, LoRa, DBSCAN, Isolation Forest, Anomaly Detection, Energy Efficiency, IoT

Subject terms: Energy science and technology, Engineering, Mathematics and computing

Introduction

Edge Computing has become a key component in modern distributed systems, especially for Internet of Things (IoT) applications. By processing data close to where it is generated, this paradigm enables low-latency responses and improves overall system efficiency. In domains such as industrial automation, autonomous vehicles, and health monitoring, Edge Computing offers clear benefits: reduced latency, better data privacy, and lower network congestion1, 2.

One major challenge in edge-based IoT systems is the high volume of data generated by sensors. This increases energy consumption on resource-constrained devices and places additional load on network resources. For applications that require continuous operation, this situation threatens both efficiency and sustainability. To address this, researchers have focused on data reduction strategies that transmit only meaningful information. Anomaly detection algorithms play a central role here. By identifying and filtering out abnormal data, these algorithms prevent unnecessary transmissions and significantly reduce energy consumption.

The literature presents various approaches to improve energy efficiency through selective data transmission. Some studies focus on hardware-level optimizations, such as implementing anomaly detection on low-power FPGAs3 or microcontrollers4. Others examine communication protocols like LoRa and investigate how parameters such as spreading factor and channel selection affect energy use5, 6. A third group of studies explores algorithm-level innovations, including compression techniques7, federated learning8, and adaptive transmission strategies9. While these studies provide valuable insights, most of them evaluate either detection performance (accuracy, F1-score) or energy consumption separately. The energy cost of transmitting anomaly-detected data via low-power wireless protocols remains underexplored.

This study addresses this gap by offering a comparative evaluation of energy consumption during data transmission. We compare two scenarios: full data transmission using LoRa, and selective transmission based on two anomaly detection algorithms, Isolation Forest and DBSCAN. LoRa was chosen because it combines low power requirements with reliable long-range communication, making it well suited for edge-based IoT systems. Unlike most previous work, we focus on transmission energy rather than computation energy. The main reason is that communication typically dominates energy use in IoT edge devices. We also measure energy consumption experimentally, using real hardware, rather than relying on simulations or estimations.

The main contribution of this study is twofold. First, we analyze how anomaly detection–supported selective transmission affects communication energy consumption in a real edge computing setup. Second, we compare two algorithms not only in terms of detection performance but also in terms of the energy saved by transmitting fewer data. This combined view provides a more complete picture of decision-supported transmission methods and their practical impact on energy efficiency.

Related work

Recent studies indicate that anomaly detection–based selective data transmission at the edge can improve communication efficiency and reduce energy consumption in IoT systems. By processing data locally and transmitting only relevant or anomalous information, these approaches minimize unnecessary data transfers. This is particularly beneficial for resource-constrained devices.

A significant body of research has explored the role of anomaly detection in reducing communication load. Ni et al.10 proposed an energy-aware optimization framework that dynamically balances local processing and computation offloading decisions. Their approach achieved up to 23.8% reduction in energy consumption and extended battery life by up to 165% in energy-constrained IoT scenarios. Similarly, Taurone et al.11 showed that anomaly detection–based selective transmission can achieve comparable detection performance while transmitting less than 1% of the total data stream. This results in substantial communication energy savings. Guan et al.12 investigated power-optimizing scheduling strategies for IoT communications. They demonstrated that efficient radio bandwidth allocation, often the dominant contributor to sensor energy consumption, can significantly prolong device lifetime.

Several studies have specifically examined lightweight algorithms suitable for edge deployment. Sani13 demonstrated that lightweight machine learning models, such as decision trees and one-class SVMs, can provide real-time anomaly detection at the edge with minimal computational overhead. Solano et al.14 and Putrada et al.15 discussed the use of the DBSCAN algorithm in IoT systems, noting that its low processing cost makes it suitable for resource-limited devices. However, in both studies, the energy benefits were mentioned only hypothetically and were not supported by experimental measurements. Hoseinpur et al.16 also stated that DBSCAN is suitable for edge architectures in terms of both accurate detection and low energy usage, but again, this claim was based solely on classification performance.

Other researchers have focused on congestion mitigation and latency reduction in LoRa-based edge networks. Charif and Aknin17 proposed adaptive transmission strategies that anticipate network congestion by processing data at the edge, thereby reducing collisions and latency. Aliagas et al.18 introduced dynamic spreading factor and channel selection mechanisms that improve communication efficiency by minimizing contention and optimizing transmission parameters such as power level and message size.

Some studies have proposed system-level architectures rather than experimental evaluations. Kralevec et al.19 presented a lightweight, layered security architecture for IoT devices with limited processing power and energy capacity. While the framework aims to perform essential security functions without straining hardware resources, it has not been tested in practice. The claim of energy efficiency remains at the architectural level, and experimental results are needed to evaluate actual system performance. Dimara et al.20 suggested that the SEDGE system, which features semantic interoperability and self-healing capabilities, could improve energy efficiency. However, the study did not directly examine the impact of transmission decisions, such as data transmission frequency or protocol selection, on energy consumption.

A few studies have attempted to model energy consumption mathematically. Ni et al.21 presented an optimization-based computation and communication model aimed at improving energy efficiency in IoT networks. Their experimental tests on real edge devices reduced the communication load. However, energy consumption was estimated through mathematical modeling based on system parameters rather than directly measured through hardware. In this respect, the study focuses on system-level theoretical optimization and does not provide a physical measurement-based energy evaluation.

Despite these advances, existing studies often address either communication optimization or anomaly detection in isolation. Many rely on specific network configurations and application scenarios, and most lack experimental energy measurements on real hardware. Motivated by these limitations, our work focuses on anomaly detection–supported selective data transmission at the edge, with an explicit emphasis on analyzing and reducing communication energy consumption through hardware-based experiments.

Data set

The vibration data set was obtained by placing an MPU6050 vibration sensor in a 3D printer. The MPU6050 is a sensor that combines a three-axis gyroscope and a three-axis accelerometer and is used for motion and position tracking. It connects to microcontrollers via an Inline graphic interface and provides angular velocity and acceleration data22. As vibration data would be used, only accelerometer data was evaluated. As no significant changes were observed on the X and Y axes, only Z-axis values were recorded.

Figure 1 shows the 3D printer from which the dataset was obtained. During the printer’s operation, the initial phase involved collecting only normal, or healthy, data. Subsequently, abnormal conditions were introduced by deliberately loosening the screws at three specific locations, as illustrated in the figure. In this procedure, the central screws were loosened by three full turns and the edge screws by five. The final dataset consists of 5029 data samples in total, comprising 4971 normal and 58 abnormal instances(Table 1).

Fig. 1.

Fig. 1

The 3D printer used for dataset acquisition.

Table 1.

Dataset contents.

Dataset Normal Abnormal Total
Vibration Data 4971 58 5029

During the data preprocessing phase, a filtering technique was employed to reduce sensor-induced noise. In this context, a high-pass filter was applied to suppress the low-frequency components of the accelerometer data. This filter aims to eliminate constant or slowly varying influences (e.g., gravity-induced acceleration), thereby retaining only the high-frequency acceleration variations caused by sudden movements within the dataset23.

After filtering, a statistical method called Modified Z-Score analysis, shown in Fig. 2, was used. Modified Z-Score is a statistical method based on the median and median absolute deviation (MAD) developed to reduce the sensitivity of the classic Z-score method to extreme values. The classical Z-score, which is based on parametric measures such as the arithmetic mean and standard deviation, can produce misleading results, especially in skewed distributions or datasets containing outliers. The modified Z-score, developed to address this weakness, evaluates the central tendency and dispersion of the data using less susceptible measures.24.

Fig. 2.

Fig. 2

Anomaly data labeling based on the modified Z-score method.

Modified Z-score (Inline graphic) is calculated as follows:

graphic file with name d33e355.gif 1

In Eq.(1), Inline graphic represents the data point, Inline graphic denotes the median of the dataset, Inline graphic is the modified Z-score, and Inline graphic refers to the median absolute deviation. The coefficient 0.6745 is used to relate the Inline graphic coverage rate in a normal distribution to the standard deviation, thereby calibrating the Z-score based on the median instead of the mean. This technique is widely applied in univariate time series analysis and is effective in identifying local anomalies by reducing the impact of outliers in noisy and irregular datasets25.

Kanan et al.26, emphasized that the modified Z-score is a suitable method for anomaly detection in univariate datasets. Similarly, şahinler et al.25, in a comparable study, stated that the modified Z-score is an effective approach for detecting anomalies in time series data.

Using the Modified Z-Score method, anomalous values were identified and labeled as abnormal. As can be seen in Fig. 2, starting from about the middle of the dataset, there is a noticeable increase in the number of abnormal data points. This rise in abnormal values is intentionally caused by the anomalies generated on the 3D printer.

Materials and method

In order to test the applicability of the proposed approach, a data transmission structure based on anomaly detection with a focus on energy efficiency was designed. Within this scope, acceleration data were collected using an MPU6050 vibration sensor mounted on a 3D printer. The MPU6050, which integrates a three-axis gyroscope and a three-axis accelerometer, is widely used for motion and position tracking. The collected data were analyzed using two unsupervised anomaly detection algorithms: DBSCAN and Isolation Forest. Data points identified as anomalous were transmitted from one edge device to another via a LoRa module, which is characterized by its long-range and low-power communication capabilities. Throughout the transmission process, both the scenario in which only anomalous data were sent and the scenario involving full data transmission were comparatively evaluated, and energy consumption was directly measured through hardware-based monitoring.

Figure 3 shows the system architecture in which two different data transmission scenarios are applied. The data set consisting of vibration data was evaluated within the scope of these two scenarios. In the first scenario, Full Data Transmission (FDT), the entire dataset (both normal and abnormal data) was transmitted directly without any filtering. In the second scenario, only abnormal data was identified and transmitted using anomaly detection algorithms; two methods were used in this scenario. The methods are, respectively, the Isolation Forest (IF) and DBSCAN algorithms. In both scenarios (three methods), the data was wirelessly transmitted from one edge device to another via a LoRa module. During transmission, the system’s energy consumption was monitored using INA226 current sensors, and a comparative analysis between the scenarios was conducted.

Fig. 3.

Fig. 3

Experimental architecture.

Experiment materials

In the experiment, hardware components were used to measure power consumption during the transmission of vibration data. These components included the INA226 power monitoring sensor module, embedded circuit (Arduino Uno), edge devices (Raspberry Pi), and LoRa wireless communication module.

The INA226 is a power and current monitoring module used to track the power consumption during the transmission of data from the transmitting edge device to the receiving edge device. The sensor measures the voltage (Inline graphic) drop across a shunt resistor (Inline graphic), from which the current (I) is calculated. These measurements are converted into digital data via a 16-bit ADC integrated within the sensor. By interfacing with an Arduino or other microcontroller via IInline graphicC communication, the INA226 facilitates the analysis of electrical power consumption in embedded systems27.

graphic file with name d33e441.gif 2

The INA226 directly measures this shunt resistance. Inline graphic(Eq.(2)) is the voltage drop across the shunt resistance.

graphic file with name d33e454.gif 3

The sensor calculates the current using the measured voltage conversion according to (3).

In the experiment, a Raspberry Pi was utilized as the edge device, while an Arduino Uno served as the embedded system. Arduino is an open-source electronics platform consisting of a development board with a programmable microcontroller. It is capable of reading data from sensors and generating appropriate outputs based on these inputs, enabling real-time interaction with the physical environment28. Unlike Arduino, Raspberry Pi is a low-cost, single-board computer. It can run operating systems such as Android, Linux, and Windows 10 IoT. It can be used as a full-fledged computer by connecting a screen, keyboard, and mouse. It supports Wi-Fi, Ethernet, and Bluetooth29.

The LoRa wireless communication module was used to transmit vibration data from one edge device to another. LoRa is a wireless telecommunication system with long range, low power consumption, and low data rates30.

Table 2 summarizes the key technical specifications of the hardware components used in the system. The MPU6050 sensor, employed for vibration data acquisition, combines a three-axis accelerometer and a three-axis gyroscope to enable accurate motion detection. The INA226 module is responsible for precisely measuring current and voltage within the circuit, allowing for detailed analysis of power consumption. An Arduino Uno microcontroller is used to process data and manage sensor operations. Acting as the edge device, the Raspberry Pi 4B handles system-level control and data transmission, supported by its high processing power, various connectivity options, and operating system compatibility. Sensor data is wirelessly transmitted to a remote receiver via the SX1268 LoRa module, which operates with fixed parameters for Spreading Factor (SF), Coding Rate (CR), and Bandwidth (BW). Furthermore, the table specifies the communication protocol used by each component, thereby clarifying how data flows are organized within the system.

Table 2.

Specifications of hardware components.

Component Main Function Key Features Communication Protocol
MPU6050 Motion and Vibration Detection 3-axis accelerometer + 3-axis gyroscope IInline graphicC
INA226 Power and Current Monitoring 16-bit ADC, Inline graphic mV shunt measurement, direct power calculation, calibration supported IInline graphicC
Arduino Uno Microcontroller (Control Unit) ATmega328P-based, open source, sensor data reading and processing UART, IInline graphicC, SPI
Raspberry Pi 4B Single-Board Computer (Edge Device) 1.5 GHz quad-core ARM CPU, 2–8 GB RAM, Linux/Android support, Wi-Fi, BT, Ethernet UART, IInline graphicC, SPI, Ethernet
LoRa Module (SX1268) Wireless Data Transmission Long range, low power, low data rate; SF: 7, CR: Inline graphic, BW: 125 kHz SPI / UART

Experimental procedure

In order to determine the power consumption during the transmission of all points in the data set and only abnormal points, the power cable coming to the edge device transmitter was cut and an INA226 current and power monitoring module was placed in between. Thanks to this module, the current consumed by the edge device was recorded via the serial port of the embedded circuit. The experimental setup is shown in Fig. 4.

Fig. 4.

Fig. 4

Experimental setup.

Figure 4 shows the general layout of the hardware used in the experimental setup and the end-to-end data transmission flow. On the sender side, a Raspberry Pi and an Arduino module were used for data generation and transmission. The Raspberry Pi controls the system and collects real-time current data via the Power Sensor, which measures energy consumption. This data is processed via the Arduino and transmitted to the receiver at regular intervals via the LoRa module. Both sides are powered by fixed power source.

The data was sent at 1.5-second intervals during end-to-end transmission. This waiting period was found to be optimal in terms of transmission security compared to trials conducted at shorter intervals. In particular, in intervals shorter than 1.5 seconds (in the first scenario), data collisions occurred with the FDT method, leading to data integrity issues on the receiving end. Therefore, the 1.5-second transmission interval was preferred in terms of preventing collisions, ensuring transmission stability, and ensuring that data reaches the other side in a meaningful and complete manner.

Unsupervised machine learning algorithms

Unsupervised algorithms are machine learning algorithms that work on unlabeled data and attempt to discover hidden patterns, structures, or groupings in the data on their own31.

Isolation Forest (IF), IF randomly selects features and uses a threshold value to create isolation trees on data points. For a given data point x , the average path length across the trees, denoted as E[h(x)], is calculated. The anomaly score s(x, n), is then defined as follows:

graphic file with name d33e624.gif 4

In (4), c(n) denotes the average path length for a dataset of size n.

graphic file with name d33e642.gif 5

In Eq.(5), m denotes the number of data points entering the tree. h(n) is the n-th harmonic number, which can be approximated as Inline graphic, where Inline graphic is the Euler–Mascheroni constant.

Lower values of E[h(x)] indicate that the data point can be isolated with fewer splits in the isolation trees, corresponding to a higher anomaly score (i.e., Inline graphic). Conversely, higher values of E[h(x)] are associated with longer isolation paths, suggesting that the data point exhibits behavior closer to normal (i.e., Inline graphic)32.

DBSCAN (Density-Based Spatial Clustering of Applications with Noise), Clusters the points in the data set using a density-based approach. For each data pointInline graphic , the Inline graphic-radius neighborhood region Inline graphic is defined as follows:

graphic file with name d33e724.gif 6

In Eq.(6):

  • Inline graphic, the entire dataset,

  • Inline graphic, distance between data points (typically Euclidean distance),

  • Inline graphic, denotes the neighborhood radius..

A data point Inline graphic is considered a core point if the number of points within its neighborhood is greater than or equal to the value Inline graphic. This condition is expressed as in Eq.(7).

graphic file with name d33e764.gif 7

When this condition is satisfied, Inline graphic represents the minimum number of neighbors required for a data point to be defined as a core point.

The DBSCAN algorithm is based on the concept of direct density-reachability. A point Inline graphic is said to be directly density-reachable from a core point Inline graphic if:

graphic file with name d33e783.gif 8

If the conditions specified in Eqs.(7)-(8) are met, The algorithm forms clusters through direct or indirect connections between core points. Points that do not satisfy the density conditions and cannot be assigned to any cluster are classified as noise.

Anomaly (noise) detection is performed when a data point does not have a sufficient number of points within its neighborhood. This is calculated as shown in Eq.(9):

graphic file with name d33e793.gif 9

If this condition is satisfied, the point Inline graphic is considered an anomaly.

Through this method, DBSCAN is able to form clusters in dense data regions while effectively identifying data points in low-density areas as anomalies33.

In this study, DBSCAN and Isolation Forest algorithms were employed for anomaly detection, with careful selection of hyperparameters to optimize model performance. For DBSCAN, the neighborhood radius parameter eps was set to 0.12, while the minimum number of points required to form a core point, min_samples, was chosen as 10. These values were selected to effectively capture deviations in the data based on local density variations within the time series. In the case of the Isolation Forest algorithm, the contamination parameter, which indicates the expected proportion of anomalies in the dataset, was set to 0.03 (i.e., 3%). The number of decision trees used in the ensemble, n_estimators, was set to 100, and the max_samples parameter – which controls the number of data points sampled for each tree – was left at ”auto”, allowing the algorithm to adapt based on dataset size. To ensure reproducibility, the random_state was fixed at 42. These parameter choices helped both models to produce stable, consistent results and to identify anomalous patterns in the data with a high degree of accuracy.

Power consumption analysis

In this experimental study, the power consumption of an edge computing–based system was evaluated through measurements taken at 0,1-second intervals. Current readings were obtained using an INA226 sensor, acquired via an Arduino microcontroller, and directly recorded into Microsoft Excel. All calculations were carried out assuming a constant 5 V DC supply voltage for the device.

In total, the measurements of three different methods were evaluated. The first scenario covers the Full Data Transmission (FDT) method, while the second scenario involves data transmission using the DBSCAN and IF algorithms. In both scenarios, the data sending process was completed in the same duration. The scenarios were considered as if they were a real system, where the time to complete the transmission of all data is equal to the transmission time of the points identified by the anomaly detection algorithms. This is because, while transmitting the points identified by the algorithms, all points are still checked and only the abnormal points are sent. Each scenario was repeated three times; the second scenario was applied three times for each algorithm, making a total of nine experiments. By following the steps below, in each experiment, the peak points were identified, the peak current value was measured, instantaneous power was calculated, and the total energy consumed during the experiments was determined.

During the data analysis process, the findpeaks() function available in MATLAB software was used to detect sudden rises within the signal. This function determines local maximum positions in the signal based on certain threshold values. These peak points correspond to the data transmission points and are particularly critical in power and energy analyses, as they represent short-term high current draws in the system. The obtained values describe the most dominant components of the current, which are directly used in energy calculations.

The current data collected from the system are generally in milliampere (mA). However, energy and power calculations are performed in ampere (A) according to the International System of Units (SI). Therefore, each peak current value is divided by 1000 to convert from mA to A.

Instantaneous power calculation refers to the amount of energy consumed by a system per unit time and, in systems operating under a constant voltage, is calculated as the product of current and voltage. The calculation of energy was carried out by multiplying the instantaneous power measurement with the data transmission period of about 1.5 seconds. For the overall energy estimation, the consumptions corresponding to the identified peak points were then aggregated.

Based on the experimental results, the total energy consumption for each method was obtained by summing the energy values corresponding to the detected peak points. The mathematical formulas employed at each step are presented in Table 3 below.

Table 3.

Formulas and descriptions used in energy calculations.

Title Formula Description
Peak Current Calculation Inline graphic Since the current data were collected in milliamperes (mA), they were converted into amperes (A) by dividing by 1000, in order to perform the calculations in accordance with the SI unit system.
Instantaneous Power Calculation Inline graphic It was calculated by multiplying the voltage (V) by the current (I), representing the instantaneous power consumed by the system.
Energy Calculation Inline graphic (Here, Inline graphic) The instantaneous energy at each peak is calculated by multiplying the corresponding power value by a constant duration. Energy is expressed in joules.
Total Energy Inline graphic The sum of all energy values obtained during the experiments. n: number of peak points.

Statistical analysis

In the study, the effect of three different methods used for data transmission on energy consumption was statistically analyzed. One-way analysis of variance (ANOVA) was applied to examine the differences between groups, and in the cases with significant results, the Tukey HSD (Honestly Significant Difference) test was used to identify which groups had differences. The findings obtained were supported with graphs, contributing to the determination of the most suitable method in terms of energy efficiency.

Analysis of variance (ANOVA)

Analysis of Variance (ANOVA) is a parametric method used to determine whether there is a significant difference among the means of three or more groups. Its main purpose is to test the statistical significance of the differences by comparing the variances between groups and within groups34. ANOVA is structured within the framework of the General Linear Model (GLM) and is widely used, especially in experimental studies, to evaluate the effect of one or more independent variables on a dependent variable.

For ANOVA to provide statistically reliable results, certain assumptions must be satisfied. First, the observations comprising the sample should be entirely independent from one another. To ensure this, we conducted three independent experiments for each method. Second, the distribution spread, that is, the variances, should be equal across the groups.35.

One-way ANOVA

One-way analysis of variance (ANOVA) was employed to compare the total energy consumptions under two different data transmission scenarios. This analysis was conducted to examine the effect of a single factor, namely data transmission. Three different approaches, as defined below, were evaluated:

DBSCAN algorithm: is a method designed to identify abnormal data points and transmit only these points,

IF algorithm: is a method designed to identify abnormal data points and transmit only these points,

FDT: the method in which all data are transmitted without applying any filtering.

In the second scenario, the algorithms used were first compared with each other and then with the FDT method from the first scenario, in an effort to identify the most energy-efficient data transmission method.

Experimental results

This section presents the experimental results obtained to evaluate the impact of data transmission methods on energy consumption. A comparison between the two anomaly detection algorithms, as well as with the FDT method, is provided. The results obtained have also been assessed from a statistical perspective.

Anomaly detection results

The classification errors and performances of the anomaly detection algorithms were evaluated using a confusion matrix. In the confusion matrix, the horizontal axis represents the predicted values, while the vertical axis represents the actual values.

In Figure 5, the confusion matrix of the DBSCAN algorithm is presented, while Fig. 6 shows the confusion matrix of the IF algorithm. These confusion matrices demonstrate both numerically and visually the anomaly detection performance of the algorithms. The DBSCAN algorithm correctly classified 4,938 out of the 4,971 normal data points in the dataset, but misclassified 33 of them as anomalies, thus producing false positives. In the normal class, it achieved an F1-Score of 1.00. In anomaly classification, it successfully detected all 58 anomaly points in the dataset, achieving an F1-Score of 0.78. The low false positive rate prevents unnecessary data transmission, thus supporting energy efficiency on the edge device.

Fig. 5.

Fig. 5

Confusion matrix for DBSCAN.

Fig. 6.

Fig. 6

Confusion matrix for DBSCAN IF.

When examining the confusion matrix of the Isolation Forest (IF) algorithm, it can be seen that out of the 4,971 normal data points in the dataset, 4,894 were classified correctly, while 77 points were incorrectly marked as anomalies, resulting in false positives. With these prediction values, an F1-Score of 0.99 was achieved in the normal class. In anomaly detection classification, it identified all 58 anomaly data points correctly, reaching an F1-Score of 0.60. While it shows a high success in detecting anomalies without any omission, its false positive rate in the normal class is higher compared to DBSCAN. This, in turn, increases the energy consumption of the data transmitted to edge devices.

As a result of the anomaly detection, both the DBSCAN and IF algorithms have correctly identified all 58 anomaly data points. However, in normal classification, DBSCAN produced 33 false positive classifications, while IF made 77 false positives. Since the DBSCAN algorithm generates fewer false positives in normal classification, it will send less data compared to the IF algorithm, and therefore consume less energy.

Power consumption results

The current values recorded in each scenario (FDT, IF, and DBSCAN) reveal over time the effects of the data transmission methods used on energy consumption. The graphs visually support the gains in energy efficiency achieved by the anomaly detection approaches (IF and DBSCAN).

Figure 7 shows the time-dependent variation of the current (mA) values recorded during the operation of the system within two different scenarios. The first graph belongs to the scenario (FDT) where all data is transmitted without any filtering, and it can be clearly seen that in this case, the current consumption remains continuously at a high level; this indicates that intensive data transmission directly increases energy consumption.

Fig. 7.

Fig. 7

The current values recorded in the two scenarios performed.

The second and third graphs belong to the second scenario, in which the data is filtered using anomaly detection algorithms before being transmitted. When only the abnormal points are sent using the Isolation Forest (IF) and DBSCAN algorithms, there is a noticeable decrease in current consumption. This difference is demonstrated by the reduction in the intensity of current fluctuations in the graphs. The system became active only at specific time intervals, which shows that the transmission load decreased and energy efficiency increased. In addition, since the DBSCAN algorithm identified fewer data points as anomalies, the observed current values remained lower and more stable compared to the IF algorithm.

The difference between the two scenarios clearly demonstrates the impact of anomaly detection algorithms on system performance, not only in terms of anomaly detection success but also in the context of energy consumption related to data transmission.

Statistical analysis results

A one-way ANOVA test was applied to statistically evaluate the effect of data transmission scenarios on energy consumption values. In cases where significant differences were found between the groups, the Tukey HSD multiple comparison method was used to determine between which methods these differences occurred.

One-way ANOVA results

Within the first scenario, three independent experiments were conducted. In the second scenario, six independent experiments were carried out, since the data transmission method was used together with the DBSCAN and Isolation Forest (IF) algorithms. Thus, a total of nine experiments were performed. The average energy consumption values were recorded in joules, and the experimental observations are presented in the table below.

In the data presented in Table 4, for both scenarios, V = 5 and t = 1.5 were taken. The table also includes the average energy consumption of the methods and their values in watt-hours (Wh). In light of these data, the following hypothesis structure was adopted to test whether there is a significant difference between the energy consumption levels:

  • HInline graphic (Null Hypothesis): The mean energy consumption of the methods is equal.
    graphic file with name d33e1079.gif
  • HInline graphic (Alternative Hypothesis): At least one method differs significantly from the others in terms of mean energy consumption.
    graphic file with name d33e1093.gif

The results of the one-way ANOVA analysis are as follows:

graphic file with name d33e1097.gif

In light of the findings, the degrees of freedom for the test were determined as 2 (between groups) and 6 (within groups). The results clearly show that the FDT method leads to a much higher energy consumption compared to the other methods. The main reason for this is that all data is transmitted directly without any filtering. This approach is highly costly in terms of energy.

Table 4.

Average energy consumption values for different data processing methods.

Data Processing Method Voltage (V) Duration (s) Average Total Energy (J) Average Energy (Wh)
FDT 5 V 1.5 30,580 8.57
IF 5 V 1.5 820.57 0.229
DBSCAN 5 V 1.5 580.34 0.161

The fact that the F-value is very high and the p-value is well below the 0.05 significance level indicates that the Inline graphic hypothesis should be rejected. This confirms that there are statistically significant differences in terms of energy consumption among the evaluated methods.

The anomaly-based algorithms DBSCAN and IF have provided a significant reduction in total energy consumption since they transmit only unusual data points. This shows that these algorithms, used in the second scenario, offer an effective method in terms of energy efficiency.

ANOVA analysis only determines whether there are overall significant differences between groups, but it does not identify between which groups the difference occurs. Therefore, the analysis process was continued with the Tukey HSD multiple comparison test to determine the direction of the difference.

Tukey HSD multiple comparison test

After significant differences between groups were observed in the one-way ANOVA analysis, Tukey’s HSD (Honestly Significant Difference) test was applied to determine between which groups these differences occurred. This test allows for reliable statistical inferences by keeping the type I error rate under control in multiple comparisons. As a result of the analysis, statistically significant differences in energy consumption were found in all pairwise comparisons among the three methods (Inline graphic).

Table 5 below presents the mean differences between method pairs, their significance levels (p-value), and the 95% confidence intervals:

Table 5.

Differences in energy consumption between compared groups and their significance levels.

Compared Groups Mean Difference (J) 95% CI Lower Bound 95% CI Upper Bound p-value Significance
FDT – DBSCAN 29,999.66 28,879 30,119 Inline graphic Yes
FDT – IF 29,759.43 29,638 29,879 Inline graphic Yes
IF – DBSCAN 240.23 119.95 360.59 Inline graphic Yes

The multiple comparison results presented in Table 4 indicate that the differences in energy consumption between the data transmission methods are not only statistically significant but also practically important. The FDT method consumes significantly more energy compared to both the DBSCAN and IF algorithms. The mean difference between FDT and DBSCAN is 29,999.66 J, with a 95% confidence interval of [28,879, 30,119]. Similarly, the difference between FDT and IF is 29,759.43 J, with a confidence interval of [29,638, 29,879]. In both comparisons, the fact that the confidence intervals do not include zero shows that the differences are not merely due to sampling error, but also represent real and generalizable differences.

The difference between the DBSCAN and IF methods is smaller (240.23 J), but since the confidence interval [119.95, 360.59] does not include zero, this difference is also statistically significant. However, considering the size of the difference, the practical impact of the energy consumption gap between these two methods may be limited. Overall, the DBSCAN algorithm emerges as the method with the lowest energy consumption, while the FDT method, due to its high energy requirement, should be preferred only in scenarios where the complete transmission of data is mandatory. These findings suggest that in applications where saving energy is important, algorithms that transmit only the necessary data, particularly DBSCAN, should be used.

Discussion

This study addresses a gap that is frequently emphasized in the literature but rarely explored in practice: the experimental evaluation of anomaly detection–assisted data transmission in terms of energy efficiency on edge devices. In this context, we applied the DBSCAN and Isolation Forest algorithms offline on pre-collected data. Only the samples identified as anomalous were transmitted via a LoRa module. This design choice was intentional. By isolating the impact of selective transmission on communication energy, we excluded computational energy costs from the evaluation. This allowed for an objective assessment focused solely on the communication layer.

The energy efficiency of the proposed method was demonstrated both qualitatively and quantitatively. Statistical analyses, including ANOVA and Tukey HSD tests, confirmed the significance of the results. Furthermore, unlike many studies in the literature that rely on simulation-based analysis, this work performed end-to-end experimental measurements using a real edge device setup consisting of Raspberry Pi, Arduino, and LoRa modules. In this regard, the study provides a practical and reliable contribution to the literature by delivering reproducible energy measurements.

Table 6 presents the experimental results for each method. The FDT (Full Data Transmission) method consumed the most energy, with an average of 30,580 J. This high value results from transmitting all data without any filtering. In contrast, the DBSCAN algorithm transmitted only 91 data points, achieving a 98.19% data reduction rate and an average energy consumption of 580.34 J. This corresponds to approximately 98.10% energy savings compared to FDT. The Isolation Forest algorithm transmitted 135 data points, resulting in a 97.32% data reduction and an energy consumption of 820.57 J. All energy and data reduction rates were calculated using FDT as the reference point.

Table 6.

Percentages of data and energy reduction, and average energy consumption compared to FDT.

Method Number of Transmitted Data Points Data Reduction (%) Average Energy Consumption (J) Energy Reduction (%)
DBSCAN 91 98.19% 580.34 98.10%
IF 135 97.32% 820.57 97.32%
FDT 5029 0.00% 30,580 0.00%

The experimental findings indicate that the high energy consumption of the FDT method results from transmitting all data. DBSCAN achieved the lowest energy consumption by minimizing unnecessary transmissions. While Isolation Forest has the advantage of detecting anomalies without omission, it consumed more energy than DBSCAN due to its higher false positive rate for normal data. Statistical analyses (ANOVA and Tukey HSD) showed that the differences between the methods are not only statistically significant but also practically important.

To reduce the number of experimental variables, the system was implemented with a fixed LoRa configuration and a single application scenario: 3D printer vibration monitoring. During the experiments, energy consumption during data transmission was selected as the primary evaluation metric. Other factors such as computational cost, system scalability, and overall performance were excluded from this controlled setting. These aspects are planned for future studies.

The anomalies were generated synthetically, for example by loosening screws, and all experiments were conducted in a controlled environment. In dynamic real-world settings, maintaining the performance of the anomaly detection system may require adaptive parameter tuning or periodic retraining. The applicability of such approaches should be explored in greater depth in future research.

The experiments were conducted on a medium-sized dataset. However, the selected anomaly detection algorithms can handle increasing data volumes. The proposed approach does not depend on dataset-specific features. It can therefore be applied to larger IoT datasets without changes to the core workflow. Offline anomaly detection also reduces processing load at the edge, which supports scalability when data volume increases.

A fixed power supply was used in the experiments to ensure stable and consistent energy measurements. However, this choice does not fully reflect real-world edge device conditions, where battery-powered operation is common. In such cases, factors like voltage fluctuations, energy interruptions, and battery depletion can significantly affect energy consumption and system behavior.

In conclusion, this study presents both practical and statistically supported findings on anomaly detection–assisted selective data transmission for improving energy efficiency. The results provide valuable insights for developing decision-support mechanisms aimed at reducing communication costs in edge-based IoT systems.

Conclusion

In this study, we examined the effects of different data transmission scenarios on energy consumption through experiments conducted on an edge computing–based IoT architecture. The anomaly detection–based scenarios significantly reduced total energy consumption by minimizing unnecessary data transmission. The findings show that, especially in applications with high data volume and energy constraints, the choice of an appropriate algorithm directly influences system performance.

The performance evaluation of the algorithms was based on criteria such as the amount of data transmitted, data reduction rate, average energy consumption, and percentage of energy savings. DBSCAN achieved the lowest energy consumption among the tested methods. It transmitted only 91 data points and reduced energy use by approximately 98.10% compared to full data transmission. Isolation Forest also performed well, with a 97.32% data reduction and significant energy savings, though its higher false positive rate led to slightly higher energy consumption than DBSCAN. In contrast, transmitting all data without any filtering resulted in substantially higher energy use.

These findings have practical implications for edge-based IoT system design. Integrating an anomaly-based data transmission approach not only reduces communication load but also improves overall energy efficiency. This enables data to be processed closer to the network edge rather than being transferred to a central cloud, thereby reducing latency and optimizing resource usage. In this way, edge computing plays an important role in achieving efficiency and sustainability goals within IoT ecosystems.

Future work should address the limitations of this study. First, experiments with battery-powered operation would provide a more realistic picture of energy consumption in real-world deployments. Second, testing the approach on larger and more diverse datasets would help validate scalability. Third, exploring adaptive parameter tuning for anomaly detection algorithms could improve their performance in dynamic environments. Finally, investigating the trade-off between computational energy and communication energy more explicitly would provide a more complete understanding of overall system efficiency.

Acknowledgements

During the preparation of this work the author(s) used ChatGPT (OpenAI) and QuillBot in order to improve language clarity and readability. After using this tool/service, the author(s) reviewed and edited the content as needed and take(s) full responsibility for the content of the publication.

Author contributions

FK: Contributed to the conceptualization of the study, literature review, data collection, and preparation of visual materials. Also participated in the interpretation of the findings and contributed to the writing process of the manuscript. BU: Responsible for methodological design, data analysis, interpretation of the findings, and revision of the final manuscript. Both authors jointly contributed to the interpretation of the findings, reviewed all stages of the study, and approved the final version submitted for publication.

Funding

This research received no external funding.

Data and code availability

The dataset and source code used in this study are publicly available at: https://github.com/fadimekaradas/anomaly-data-reduction-lora-iot.

Declarations

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Fadime Karavaş and Bilal Usanmaz contributed equally to this work.

Contributor Information

Fadime Karadas, Email: fadimekaradas@atauni.edu.tr.

Bilal Usanmaz, Email: bilal@atauni.edu.tr.

References

  • 1.Tay, M. & Şentürk, A. Kenar, sis ve bulut bilişimin iot açısından ncelenmesi. Avrupa Bilim ve Teknoloji Dergisi 68–75 (2021).
  • 2.Vo, T., Dave, P., Bajpai, G. & Kashef, R. Edge, fog, and cloud computing: An overview on challenges and applications. arXiv preprint arXiv:2211.01863 (2022).
  • 3.Caba, J. et al. Low-power hyperspectral anomaly detector implementation in cost-optimized fpga devices. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens.15, 2379–2393 (2022). [Google Scholar]
  • 4.Moallemi, A., Burrello, A., Brunelli, D. & Benini, L. Exploring scalable, distributed real-time anomaly detection for bridge health monitoring. IEEE Internet Things J.9, 17660–17674 (2022). [Google Scholar]
  • 5.Keshmiri, H., Rahman, G. M. & Wahid, K. A. Lora resource allocation algorithm for higher data rates. Sensors25, 518 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Wu, D. & Liebeherr, J. A low-cost low-power lora mesh network for large-scale environmental sensing. IEEE Internet Things J.10, 16700–16714 (2023). [Google Scholar]
  • 7.Väänänen, O. & Hämäläinen, T. Efficiency of temporal sensor data compression methods to reduce lora-based sensor node energy consumption. Sens. Rev.42, 503–516 (2022). [Google Scholar]
  • 8.Sanchez, O. T. et al. Federated learning framework for lorawan-enabled iiot communication: A case study. IEEE Internet Things J.12, 24944–24957 (2025). [Google Scholar]
  • 9.Al-Sammak, K. A. et al. Optimizing iot energy efficiency: Real-time adaptive algorithms for smart meters with lorawan and nb-iot. Energies18, 987 (2025). [Google Scholar]
  • 10.Nia, A. M., Mozaffari-Kermani, M., Sur-Kolay, S., Raghunathan, A. & Jha, N. K. Energy-efficient long-term continuous personal health monitoring. IEEE Trans. Multi-Scale Comput. Syst.1, 85–98 (2015). [Google Scholar]
  • 11.Taurone, F., Dorsch, J., Lucani, D. & Zhang, Q. triaged: using compression for anomaly detection. In 2024 Data Compression Conference (DCC), 588–588 (IEEE, 2024).
  • 12.Guan, P., Dangwal, A., Taherkordi, A., Wolski, R. & Krintz, C. Energy-aware iot deployment planning. In Proceedings of the 21st ACM International Conference on Computing Frontiers, 61–70 (2024).
  • 13.Sani, G. Lightweight machine learning techniques for real-time anomaly detection in iot sensor networks. Electronics (2024),OSF Preprints. 10.31219/osf.io/u4yr6_v1.
  • 14.Solano, F., Krause, S. & Wöllgens, C. An internet-of-things enabled smart system for wastewater monitoring. IEEE Access10, 4666–4685 (2022). [Google Scholar]
  • 15.Putrada, A. G., Abdurohman, M., Perdana, D. & Nuha, H. H. Machine learning methods in smart lighting toward achieving user comfort: A survey. IEEE Access10, 45137–45178 (2022). [Google Scholar]
  • 16.Hoseinpur, F. Towards security and resource efficiency in fog computing networks. Ph.D. thesis, Lappeenranta-Lahti University of Technology LUT (2022).
  • 17.Charif, O. & Aknin, N. Enhancing lora network efficiency: Using edge computing for congestion mitigation and latency reduction. In 2025 International Conference on Computer Systems and Technologies (CompSysTech), 1–6 (IEEE, 2025).
  • 18.Aliagas, C., Pueyo Centelles, R., Meseguer, R., Millán, P. & Molina, C. Dynamic selection and detection of spreading factors and channels for end-node devices of lora networks. Electronics14, 3341 (2025). [Google Scholar]
  • 19.Kralovec, C. & Schagerl, M. Review of structural health monitoring methods regarding a multi-sensor approach for damage assessment of metal and composite structures. Sensors20, 826 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Dimara, A. et al. Self-healing of semantically interoperable smart and prescriptive edge devices in iot. Appl. Sci.12, 11650 (2022). [Google Scholar]
  • 21.Ni, C., Wu, J. & Wang, H. Energy-aware edge computing optimization for real-time anomaly detection in iot networks. Appl. Comput. Eng.139, 42–53 (2025). [Google Scholar]
  • 22.Fedorov, D., Ivoilov, A. Y., Zhmud, V. & Trubin, V. Using of measuring system mpu6050 for the determination of the angular velocities and linear accelerations. Automatics & Software Enginery11, 75–80 (2015). [Google Scholar]
  • 23.Gjoreski, H. & Gams, M. Accelerometer data preparation for activity recognition. In Proceedings of the International Multiconference Information Society, Ljubljana, Slovenia1014, 1014 (2011). [Google Scholar]
  • 24.Domański, P. D. Study on statistical outlier detection and labelling. Int. J. Autom. Comput.17, 788–811 (2020). [Google Scholar]
  • 25.Şahinler, R., Öztürk, C., Karazeybek, B. & Kahraman, B. Anomaly detection in time series data using unsupervised machine learning and statistical methods. In 2024 15th National Conference on Electrical and Electronics Engineering (ELECO), 1–5 (IEEE, 2024).
  • 26.Kannan, K. S., Manoj, K. & Arumugam, S. Labeling methods for identifying outliers. Int. J. Stat. Syst.10, 231–238 (2015). [Google Scholar]
  • 27.Hoss, M., Westmeier, F. & Akelbein, J.-P. Low cost high resolution ampere meter for automated power tests for constrained devices. In CERC, 81–88 (2019).
  • 28.Kim, S.-M., Choi, Y. & Suh, J. Applications of the open-source hardware arduino platform in the mining industry: A review. Appl. Sci.10, 5018 (2020). [Google Scholar]
  • 29.Jolles, J. W. Broad-scale applications of the raspberry pi: A review and guide for biologists. Methods Ecol. Evol.12, 1562–1579 (2021). [Google Scholar]
  • 30.Augustin, A., Yi, J., Clausen, T. & Townsley, W. M. A study of lora: Long range & low power networks for the internet of things. Sensors16, 1466 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Naeem, S., Ali, A., Anam, S. & Ahmed, M. M. An unsupervised machine learning algorithms: Comprehensive review. Int. J. Comput. Digit. Syst.13, 911–921 (2023). [Google Scholar]
  • 32.Lesouple, J., Baudoin, C., Spigai, M. & Tourneret, J.-Y. Generalized isolation forest for anomaly detection. Pattern Recognit. Lett.149, 109–119 (2021). [Google Scholar]
  • 33.Deng, D. Research on anomaly detection method based on dbscan clustering algorithm. In 2020 5th International Conference on Information Science, Computer Technology and Transportation (ISCTT), 439–442 (IEEE, 2020).
  • 34.Henson, R. Analysis of variance (anova). Brain Mapp.1, 477–481 (2015). [Google Scholar]
  • 35.McHugh, M. L. Multiple comparison analysis testing in anova. Biochem. Med.21, 203–209 (2011). [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

The dataset and source code used in this study are publicly available at: https://github.com/fadimekaradas/anomaly-data-reduction-lora-iot.


Articles from Scientific Reports are provided here courtesy of Nature Publishing Group

RESOURCES