Skip to main content
Scientific Reports logoLink to Scientific Reports
. 2026 Mar 14;16:13545. doi: 10.1038/s41598-026-43294-1

Human pose recognition and automated scoring detection for sports rehabilitation

Xiaoqian Peng 1, Xujiang Mao 2, Xiaomin Fang 3,✉
PMCID: PMC13121721  PMID: 41832253

Abstract

With the development of sports rehabilitation, accurate assessment of the patient’s rehabilitation process has become the key to enhance the rehabilitation effect. To solve the problems of inaccurate recognition and poor real-time performance of the rehabilitation human pose recognition model for traditional sports in complex environments, this study proposes an integrated framework for efficient and accurate human pose recognition and automated scoring in sports rehabilitation. The study constructs a human pose recognition model using a human pose tracking algorithm and achieves pose classification by extracting key points of the human skeleton and combining them with a random forest algorithm. Meanwhile, a siamese neural network and similarity metric algorithm are introduced to optimize the automated score detection model, accurately assessing the quality of rehabilitation movements. The outcomes indicated that the automatic scoring detection system achieved 98% accuracy in human body pose recognition. In terms of joint angle error, the error rate of the detection model designed in the study was below 6%, which was significantly better than the comparison method. In rehabilitation score correlation test, the correlation of the model was maintained at 92–98%, demonstrating higher scoring accuracy. The outcomes reveal that the model designed in the study has high recognition accuracy and evaluation stability. This makes it an efficient and accurate assessment tool for rehabilitation therapy. It can also effectively improve the effectiveness of rehabilitation training and the quality of life of patients.

Keywords: Sports rehabilitation, Human pose recognition, Automated scoring, BlazePose, Random forest algorithm, SNN

Subject terms: Computational biology and bioinformatics, Engineering, Health care, Mathematics and computing

Introduction

Under the background of the deep integration of modern sports and medicine, the field of sports rehabilitation is ushering in a wave of intelligent and precise development. Human body pose recognition (HBPR) technology, as the core means of rehabilitation training (RT) effect evaluation, has made great progress in recent years, but still faces many bottlenecks. Traditional methods of recognizing human body (HB) poses are difficult to use for capturing subtle changes in pose. Furthermore, their real-time performance is insufficient for providing immediate feedback in dynamic RT. RT assessment methods rely heavily on manual observation and empirical judgment. These methods are inefficient and subjective.1,2. On-device real-time body pose tracking (BlazePose) has strong real-time, high accuracy, light weight and good robustness to accurately predict the key points (KPs) of the HB in real-time3. A random forest classifier has strong overfitting resistance and can handle high-dimensional data. It is also interpretable, generalizable, and adaptable to data, and has a fast training speed4. Siamese neural network (SNN) is good at feature extraction and similarity learning with strong generalization ability5. Similarity metric algorithms can accurately quantify action differences and adapt to different scenarios6. This work is an applied systems integration effort aimed at engineering a practical clinical tool. The proposed integrated automated scoring and detection system (IASDS) is designed to meet the multifaceted demands of real-world RT by strategically combining established components: BlazePose for real-time keypoint extraction, random forest for robust pose categorization, and a SNN for similarity-based scoring. These components are integrated into a synergistic pipeline.

The innovation of the study is as follows: (1) This is a multimodal rehabilitation assessment framework that integrates BlazePose, random forest, and SNNs. It overcomes the limitations of single models in complex rehabilitation scenarios. (2) The research develops a collaborative optimization framework that balances real-time efficiency and accuracy. This model is specifically optimized to achieve an excellent balance between high recognition accuracy and low latency, making it ideal for providing instant feedback during RT. (3) The study introduces an automated scoring mechanism that simulates expert judgment. Using a similarity scoring method based on co-located neural networks, the mechanism provides an innovative, data-driven evaluation approach. This approach enables the creation of objective, quantifiable assessment tools in clinical practice that go beyond simple posture classification.

From the perspective of clinical application, this study provides a practical solution for rehabilitation, and its key contributions include: (1) This study provides a set of ready-to-use assessment tools that can automate the assessment process, thereby reducing the workload of medical staff. (2) This study provides objective and reliable quantitative assessment and a solid data basis for formulating and adjusting personalized rehabilitation plans. (3) This study enables real-time patient participation by providing immediate corrective guidance during the training process. This can effectively improve patient motivation, compliance and overall recovery. (4) This study verifies the effectiveness of clinical scenarios and confirms the efficiency and robustness of the system in simulating real medical environments. This provides support for future clinical translation and application.

Related works

HBPR technology accurately captures human pose information with the help of computer vision and artificial intelligence algorithms, which can accurately assess the rehabilitation status of patients in sports rehabilitation. It can help formulate personalized rehabilitation programs and improve the effectiveness and efficiency of rehabilitation7. Guan. S et al. designed a HBPR model for the three-dimensional (3D) human pose estimation (HPE) problem using a human pose generator and an unbiased learning strategy. Experiments indicated that the method achieved excellent performance on multiple datasets and improved estimation accuracy compared to other methods, especially in dealing with complex pose and occlusion situations8. He, S. et al. conducted a randomized controlled experiment on older sarcopenia patients and suggested a novel tele-rehabilitation technique based on 3D HPE. The outcomes revealed that the method accurately estimated the 3D pose of patients, providing effective technical support for tele-rehabilitation and improving the rehabilitation outcomes and quality of life of patients9. Martini. E et al. worked on enabling gait analysis in telemedicine practice through a portable and accurate 3D HPE technique. The experimental results indicated that the technique could accurately capture human pose changes during walking. It provided reliable data support for gait analysis in telemedicine and helped to improve diagnosis and rehabilitation10. In experiments on simulated depth photographs, Wang et al. demonstrated the effectiveness of a 3D human pose and form reconstruction method based on a hybrid deep learning and optimization method. The results indicated that the method could accurately reconstruct the 3D human pose and shape from depth images, which provided new ideas and methods for research and application in related fields11. Chen et al. designed a real-time pose detection model for mobile devices based on dynamic sparse graph convolutional networks in an attempt to address the challenge of real-time pose detection for mobile devices. The outcomes revealed that the model achieved 72 FPS on the dataset with an average correctness metric of 78.6% and a model size of only 4.2 MB12. In a narrative review, F. Roggio et al. provided a thorough examination of machine learning (ML) posture estimation models for human motion and pose analysis. Among these, BlazePose was employed as a convolutional neural network (CNN) architecture that was lightweight and able to execute HPE on mobile devices in real time. According to studies, BlazePose performed well across a range of application situations and had certain advantages in terms of resource usage and computing efficiency13.

In the field of sports rehabilitation, it can improve the accuracy and effect of RT, enhance the science and safety of RT, and guarantee the smooth progress of the rehabilitation process. V. Tsakanikas et al. proposed a data-driven scoring model for automated assessment of balance RT movements. The model utilized ML algorithms, classification and scoring of these movements. Experimental results indicated that the model classified sitting movements with an accuracy of 0.86%-0.90%, standing movements with 0.85%0.92%, and walking movements with 0.81%-0.90%, with a high degree of consistency in scoring among experts. It provided an effective automated assessment tool for balance RT14. Das et al. compared markerless and labeled motion capture techniques for balance and gait assessment. The experiments showed that the markerless motion capture technique could accurately capture human pose and gait characteristics with better convenience and generalizability in practical applications. This provided new perspectives and methods for the application of automated scoring detection technology in sports rehabilitation15. Raihan et al. proposed a rehabilitation exercise score prediction method that combines spatial features and graph structure learning. The method effectively captured spatial and temporal dependencies in rehabilitation exercise data by integrating the strengths of 2D CNNs and graph neural networks. Experiments demonstrated that the method achieved excellent performance on both publicly available datasets, providing a new idea for automated scoring of rehabilitation exercises16. SNNs excel in extracting features and learning similarities, and possess strong general applicability. The similarity metric technique is able to accurately quantify the difference between actions and is applicable to many different situations. L. Chen et al. proposed a neural network model that incorporated a lightweight multi-head attention module and a similarity metric mechanism. This model was designed to address the problem of inaccurate target localization in complex scenarios, such as those involving occlusion and fast motion17.

In summary, existing research has made some progress in the field of HBPR and automated score detection in sports rehabilitation. However, there are problems such as insufficient accuracy of motion capture, difficulty correlating multidimensional clinical rehabilitation indicators, and limited validity of scoring. Therefore, the study incorporates the BlazePose algorithm, random forest classifier, SNN, and similarity measure algorithm to propose an efficient automatic score detection model. It is anticipated to offer a more precise and trustworthy evaluation instrument for sports rehabilitation. It will meet the demand for accurate assessments in rehabilitation treatment and effectively improve rehabilitation outcomes and patients’ quality of life.

Methods and materials

Model construction of HBPR based on Blazepose algorithm

The HBPR model provides accurate data support for medical professionals through in-depth analysis of key indicators such as joint mobility, muscle strength performance, and body balance18. The data output from the model can help doctors and rehabilitators more accurately assess patients’ performance in the rehabilitation process. Then, they can adjust the treatment strategy based on the results to ensure the rehabilitation program is scientific and effective19. However, the traditional HBPR model used in sports rehabilitation has difficulty accurately recognizing and tracking poses in complex environments, such as those involving occlusion, multiple people, or cluttered backgrounds. Additionally, it fails to meet the real-time requirements. This leads to inaccurate or lost detection of critical points, which affects rehabilitation assessment and training effectiveness20. BlazePose is a lightweight CNN architecture designed for HPE. It has the advantages of high accuracy, high real-time performance, lightweight, good robustness, and easy integration and extension21. Therefore, this study constructs a HBPR model based on the BlazePose algorithm applied to sports rehabilitation. The model uses the BlazePose algorithm to extract KPs from the human skeleton and gait parameters. It also uses the random forest algorithm as a classifier to ultimately recognize, detect, and classify human poses for sports and sports rehabilitation22. The BlazePose architecture and its reasoning process are shown in Fig. 1.

Fig. 1.

Fig. 1

BlazePose architecture and its reasoning process.

In Fig. 1, the inference process for BlazePose begins with an input image that first passes through the face detection and pose alignment module to determine the pose orientation of the HB in the image. Next, the image is fed to the pose KP recognition module, which extracts features through a multilayer convolutional operation. The feature extraction is also enhanced by utilizing jump connections and stop-gradient connections. Finally, the model outputs 33 KPs and their visibility information for describing the human pose. In BlazePose, the hip centroid offset vector is used to dynamically update the human bounding box (BOB) to solve the tracking drift problem during motion. Its computational formula is shown in Eq. (1).

graphic file with name d33e338.gif 1

In Eq. (1), Inline graphic denotes the last frame hip center point coordinates. Inline graphic denotes the hip center point coordinates of the current frame. Inline graphic and Inline graphic denote the width and height of the head BOB. The BOB update is calculated as shown in Eq. (2).

graphic file with name d33e367.gif 2

In Eq. (2), Inline graphic denotes the center coordinates of the HB BOB. The whole body BOB calculation is generated based on the head detection results, as shown in Eq. (3).

graphic file with name d33e383.gif 3

In Eq. (3), Inline graphic and Inline graphic denote the horizontal and vertical coordinates of the center point of the head BOB. Inline graphic denotes the width of the HB BOB. Inline graphic denotes the height of the HB BOB. The result of HBPR based on BlazePose architecture is shown in Fig. 2. Figure 2 illustrates the topology and effect of the BlazePose architecture. Prior to inclusion in this study, informed consent was obtained from all subjects and/or their legal guardian(s), which included specific permission for the publication of their identifying information/images in an online open-access publication. The volunteer depicted in Fig. 2, Weichen Feng, has signed the Informed Consent form to this effect. Furthermore, confirmation was received that the guidelines outlined in the Declaration of Helsinki were followed, and Ethics Committee approval was obtained from the Institutional Ethics Committee of the 'Department of Information Engineering, Quzhou College of Technology’ prior to the commencement of the study. The figure shows the topology of a human skeleton with 33 key points (KPs) and their connections, along with an actual effect image where the KPs are marked in red and connected with white lines to clearly depict the human pose.

Fig. 2.

Fig. 2

BlazePose architecture based HBPR results.

Figure 2 shows the topology and effect of BlazePose. On the left is a topology of a human skeleton showing 33 KPs and how they are connected. The KPs are labeled with numbers and the connection lines indicate different parts of the HB. On the right side is a picture of an actual effect with BlazePose applied, in which the KPs are marked in red on the figure. It is also connected with white connecting lines, which clearly depict the human pose. BlazePose can effectively capture the human pose in the actual scene. However, BlazePose suffers from the problems of insufficient data adaptability and insufficient depth of feature extraction and utilization. Therefore, the study introduces random forests for optimization. As an integrated learning technique, random forests can handle high-dimensional data through the collaborative work of multiple decision trees (DTs). It can also provide feature importance rankings, which improve the efficacy and stability of classification decisions23. A DT consists of root nodes, internal nodes, and leaf nodes. Random forest technique is performed by randomly sampling the initial dataset in order to form several different subsets of data. Subsequently, each subset is trained independently to construct an independent DT24,25. Eventually, the classification results of all the DTs are summarized as the combined output of the random forest. The output of the regression task in the random forest is shown in Eq. (4).

graphic file with name d33e444.gif 4

In Eq. (4), Inline graphic denotes the predicted output of the Inline graphic th DT. The random forest classification task prediction calculation is shown in Eq. (5).

graphic file with name d33e464.gif 5

In Eq. (5), Inline graphic denotes the indicator function. Inline graphic denotes the candidate category label. The random forest out-of-bag error is calculated as shown in Eq. (6).

graphic file with name d33e484.gif 6

In Eq. (6), Inline graphic denotes the total number of out-of-bag samples. Inline graphic denotes the predicted value of sample Inline graphic. The study standardizes characteristics of extracted samples, which makes comparisons of different characteristics more accurate and fair. This enhances the efficacy of ML algorithms and the model’s wide applicability. The standardization is calculated as shown in Eq. (7).

graphic file with name d33e509.gif 7

In Eq. (7), Inline graphic is the original feature. Inline graphic is the minimum feature in the dataset. Inline graphic is the maximum feature in the dataset. The study refers to the HBPR model based on BlazePose architecture with random forest classifier as BR for short. Its framework structure is shown in Fig. 3.

Fig. 3.

Fig. 3

HBPR model framework.

In Fig. 3, after the detection video is input into the model, 33 KPs of the human skeleton are first extracted by BlazePose for human skeleton KP extraction. Then, the extracted KP data are processed to calculate the human pose feature parameters and complete the human pose feature parameter extraction. Finally, these feature parameters are input into the random forest classifier for human pose classification. This allows this study to determine if the pose is normal or abnormal. The framework designed in the study combines the powerful KP extraction capability of BlazePose and the efficient classification performance of the random forest classifier. It is able to effectively perform human pose recognition and classification.

Design of automatic scoring detection model for HBPR based on sports rehabilitation

The BR model designed in the study can meet the demand of an automatic score detection model that needs to collect and label a large amount of data containing human poses for training and optimization. It can improve scoring accuracy by accurately categorizing and identifying human pose. Rehabilitation movement quality assessment is the process of quantitatively or qualitatively analyzing the effectiveness, accuracy, strengths and weaknesses of rehabilitation movements when they are performed26,27. However, the traditional sports rehabilitation pose recognition auto-scoring model suffers from the shortcomings of limited recognition accuracy, poor real-time performance, strong dependence on the environment, weak generalization ability, and high cost of data annotation. SNN has powerful feature extraction capabilities and a similarity learning advantage. It can effectively extract gesture features under different conditions, accurately measure action similarity, and generalize strongly. The similarity metric algorithm can accurately quantify the action differences and comprehensively capture the detail information, and it is robust enough to stably calculate similarity28,29. Therefore, the study designs an automatic score detection model based on SNN and the similarity metric algorithm. The model provides an efficient solution to automatic score detection by combining the two. The target tracking recognition method of SNN is shown in Fig. 4.

Fig. 4.

Fig. 4

Target tracking and recognition method of SNN.

In Fig. 4, first, with the help of the established starting screen target region, its feature representation is extracted using a CNN. Subsequently, the current frame image is fed into the same CNN to extract its feature representation. Next, a similarity measure is applied to the features of the initial image and the current image. Then, the orientation of the target in the current image is determined based on the results of the similarity evaluation. The similarity metric algorithm primarily uses the Euclidean distance and dynamic time warping metrics, as well as other metrics, to calculate the similarity between actions and assess their quality30. The Euclidean distance is calculated as shown in Eq. (8).

graphic file with name d33e587.gif 8

In Eq. (8), Inline graphic denotes the Euclidean distance. Inline graphic denotes the number of similar actions. Inline graphic denotes the joint 2D coordinate information. The angular variation of the action is shown in Eq. (9).

graphic file with name d33e612.gif 9

In Eq. (9), Inline graphic denotes the joint angle difference of the mth action segment. The study utilizes the fully connected layer and Sigmoid function to map the feature vectors onto the similar and dissimilar binary classifications to obtain the quality scores of the actions, as shown in Eq. (10).

graphic file with name d33e628.gif 10

In Eq. (10), Inline graphic denotes weight. Inline graphic denotes action category. Inline graphic denotes action sequence number. Inline graphic denotes the total number of actions. Inline graphic denotes training action sequence vector. Late Fusion is calculated as shown in Eq. (11).

graphic file with name d33e661.gif 11

In Eq. (11), Inline graphic denotes the true score. Inline graphic denotes the mapping of the final score. The study is conducted under conventional indoor lighting conditions, and video data are collected from adult males using the lumbar disc herniation postoperative rehabilitation exercise as an example. The acquisition covers the standard template of postoperative rehabilitation exercises for lumbar disc herniation, as shown in Fig. 5.

Fig. 5.

Fig. 5

Standardized rehabilitation exercise templates for lumbar disc herniation.

Figure 5 presents various RT movements captured by the study, which are applicable to fields such as sports training, rehabilitation therapy, and motion recognition. The figure illustrates eight essential rehabilitation exercises: extension, abduction, flexion, horizontal adduction, external rotation, internal rotation, the bow step, and the squat. These movements are performed by an adult male undergoing postoperative RT for lumbar disc herniation. Skeletal models of each posture illustrate the expected positions of the joints and limbs. This provides a baseline for automated systems to evaluate and compare patient movements. The flow of behavioral action similarity assessment is shown in Fig. 6.

Fig. 6.

Fig. 6

Workflow for quantifying movement similarity.

Figure 6 systematically illustrates the complete computational workflow for human motion similarity assessment. The process begins with the acquisition of skeletal keypoint sequences. Segmentation via a sliding window is involved. Motion information is encoded. Feature vectors are normalized. The final results are aggregated through averaging to yield an overall motion similarity score. This workflow forms the core of the automated scoring mechanism, transforming raw pose data into a quantitative similarity assessment relative to the standard templates. RT movements are mainly intended to exercise the mobility of the human limbs. The main skeletal points of the HB are divided into the left and right upper limb and left and right lower limb regions. By accounting for the average mobility of the joints in these regions, it is possible to determine whether they are involved in RT. The calculation formula is shown in Eq. (12).

graphic file with name d33e711.gif 12

In Eq. (12), Inline graphic denotes the Inline graphic th region of the four limb regions. Inline graphic denotes the average movement amplitude of the joints in the region. Inline graphic denotes the Inline graphic th joint in the Inline graphic th limb region. Inline graphic denotes the joint movement angle of the Inline graphic th joint in the Inline graphic th limb region. Inline graphic is the number of all joints in the Inline graphic th limb region. The RT assessment method is the most direct evaluation of RT movements. The assessment calculation is shown in Eq. (13).

graphic file with name d33e769.gif 13

In Eq. (13), Inline graphic denotes the scoring result of a set of RT maneuvers. Inline graphic refers to the Inline graphic th movement in the set. Inline graphic denotes a total of Inline graphic RT movements. Inline graphic is the distance calculated between the Inline graphic th movement and the standard template movement. Inline graphic is an artificially set distance limit. The smaller the distance, the higher the similarity between the actions. When Inline graphic, the completion of the RT maneuver is in accordance with the standard. When Inline graphic, the completion of the RT action is substandard. For the RT model, its model performance is directly affected by the accuracy of the action matching and recognition algorithm. The average check accuracy rate is calculated as shown in Eq. (14).

graphic file with name d33e823.gif 14

In Eq. (14), Inline graphic is the number of correct patient RT action categories in the matching result. Inline graphic is the total number of action matches. The average checking accuracy mean is the accumulation of the checking accuracy of multiple standard template actions, and finally divided by the number of standard template actions. The average check accuracy mean value is calculated, as shown in Eq. (15).

graphic file with name d33e844.gif 15

In Eq. (15), Inline graphic is the sum of the checking accuracy of multiple standard template actions. Inline graphic is the number of standard template actions. The automatic scoring detection model for HBPR incorporating SNN and similarity metric algorithm is shown in Fig. 7.

Fig. 7.

Fig. 7

Automatic scoring and detection system for HBPR.

In Fig. 7, the automatic scoring model includes a data acquisition module, a HBPR module, an action recognition classification module, and an automatic recognition evaluation module. First, human pose data is collected with a camera and other devices, and HB KP information is extracted in real time using the BlazePose architecture. The pose estimation component of BlazePose is based on a heat map and regression approach, which can accurately predict KP locations. After obtaining the KP data of the human pose, an SNN is used to extract deep features and learn inter-sample similarities. This allows for an accurate measurement of the similarity between a patient’s rehabilitation action and a standard action. Meanwhile, the generalization ability of the model is enhanced so that it can adapt to different rehabilitation scenarios and individual differences. Next, the similarity algorithm is used to quantify the differences between the movements and fully capture the details of the movements to provide an accurate basis for subsequent scoring. Finally, the scoring results are output and fed back to the user, realizing the automatic scoring of rehabilitation movements.

Results

Performance testing of automatic score detection models for HBPR

To test the performance of the automatic score detection model designed by the research, the research method is abbreviated as improved the automatic scoring detection system (IASDS). This study develops two baseline methods based on the BlazePose keypoint extraction framework for performance comparison. The OP-SVM baseline uses linear support vector machines (SVMs) to classify keypoint features. However, it has limitations when it comes to capturing high-dimensional nonlinear relationships during complex motion recognition31. The OP-LSTM baseline uses a two-layer LSTM network to model temporal dependencies. However, it requires extensive training data and is prone to overfitting32. In contrast, the proposed IASDS method combines random forests with twin neural networks. This hybrid architecture employs ensemble learning to process high-dimensional data and detect subtle motion variations. It outperforms traditional baselines in terms of recognition accuracy, robustness, and consistency of scores. The software and hardware environment used in the experiment is shown in Table 1.

Table 1.

Experimental basic hardware and software environment setting.

Category Component Specification/version
Hardware Central processing unit (CPU) Intel Core i7-11800H @ 2.30 GHz
Graphics processing unit (GPU) NVIDIA GeForce RTX 3060 (6 GB VRAM)
Random access memory (RAM) 16 GB DDR4 3200 MHz
Software Operating system Ubuntu 20.04 Long Term Support (LTS)
Programming language Python 3.8
Deep learning framework TensorFlow 2.6
Machine learning library scikit-learn 1.0
Pose estimation toolkit MediaPipe 0.8.9
Performance Video input resolution 1920 × 1080 (Full High Definition)
Processing speed 30 Frames Per Second (FPS) real-time

Table 1 shows the configuration parameters of the basic hardware and software used for the experiments and the experimental environment setup. The datasets used in the study are AR dataset containing a large number of human poses and Yoga dataset containing only yoga and fitness poses for evaluating the performance of the model for specific sports rehabilitation action recognition. The study compares the missed recognition rate of different methods. Figure 8 displays the findings.

Fig. 8.

Fig. 8

Missed recognition rate of different methods.

In Fig. 8, the missed recognition rate of different methods gradually decreases as the number of model training times increases. In Fig. 8a, in the AR dataset, the missed recognition rate of IASDS is only 4.62% at the initial stage of training. When the number of training reaches 100 times, the missed recognition rate is the lowest, which is only 2.83%. OP-SVM has a missed recognition rate of 4.91% at the beginning of training. When the number of training reaches 100 times, the missed recognition rate is 4.97%. OP-LSTM has a missed recognition rate of 9.82% at the beginning of training. When the number of training reaches 100 times, the missed recognition rate is 7.14%. In Fig. 8b, in the Yoga dataset, IASDS has only 4.62% missed recognition rate at the initial training. The missed recognition rate is lowest when the number of training reaches 100, at only 4.54%. At the beginning of training, OP-SVM has a missed recognition rate of 11.08%. When the number of training reaches 100 times, the missed recognition rate is 7.14%. OP-LSTM has a missed recognition rate of 13.87% at the beginning of training. When the number of training reaches 100 times, the missed recognition rate is 9.11%. The result shows that the study method has lower missed recognition rate. The study tests the accuracy of HBPR extraction for different methods as shown in Fig. 9.

Fig. 9.

Fig. 9

Accuracy of HBPR and extraction.

Figure 9 shows the variation of HBPR extraction accuracy with the number of iterations for different methods on the AR dataset and Yoga dataset. In Fig. 9a, on the AR dataset, the accuracy of the IASDS method increases from about 93–98%, OP-SVM from about 88–96%, and OP-LSTM from about 89–95% as the number of iterations increases. In Fig. 9b, on the Yoga dataset, the accuracy of the IASDS method increases from about 90–97%, OP-SVM from about 87–94%, and OP-LSTM from about 88–93%. Overall, the accuracy of the IASDS method is higher than the other two methods on both datasets. Moreover, the advantage is gradually obvious as the iteration increases. It shows that it has better HBPR. The study tests the recognition accuracy of rehabilitation action features for different methods at different quantities as shown in Fig. 10.

Fig. 10.

Fig. 10

Recognition accuracy of rehabilitation action features under different numbers.

Figure 10 shows the recognition accuracy of different methods on the AR dataset and the Yoga dataset with different numbers of rehabilitation action features. In Fig. 10a, on the AR dataset, the recognition accuracy of the IASDS method is around 99%, and the overall curve is relatively smooth. The accuracy of OP-SVM increases slightly and then fluctuates and decreases, with an overall decreasing trend. The accuracy of OP-LSTM firstly increases and then fluctuates, but the overall recognition accuracy is lower than that of IASDS and higher than that of OP-SVM. In Fig. 10b, the recognition accuracy of the IASDS method on the Yoga dataset ranges from 98 to 99% with a relatively smooth curve. The accuracy of OP-SVM first rises and then fluctuates down. The accuracy of OP-LSTM also shows a trend of increasing and then fluctuating down, and the overall accuracy is lower than the former two. The findings display that the feature recognition accuracy of the IASDS method is higher than that of OP-SVM and OP-LSTM under both datasets. The curve is smoother and shows better stability and superiority.

To validate the importance of each module within the IASDS framework, comprehensive ablation experiments are conducted on the AR dataset. The results are summarized in Table 2. Specific model configurations are as follows: BlazePose Only: using only the keypoints extracted by BlazePose, cosine similarity is employed as the scoring metric to establish performance benchmarks. BlazePose + RF: a random forest classifier is added to BlazePose to classify actions, thereby verifying its role in improving recognition accuracy. BlazePose + SNN: SNN-based similarity scoring is integrated into BlazePose to enhance score correlation. IASDS (full model): This is the full model, which combines all modules and verifies their synergistic effects.

Table 2.

Ablation testing.

Model variant Recognition accuracy (%) Joint angle error (%) Score correlation (%)
BlazePose only 91.5 8.7 85.2
BlazePose + RF 96.8 5.9 90.1
BlazePose + SNN 95.2 7.1 93.
IASDS (full model) 98.0 5.2 96.5

The results clearly demonstrate that each component significantly contributes to the overall performance. When comparing rows 1 and 2, the random forest classifier drastically improves pose recognition accuracy. The SNN scoring module (comparing rows 1 and 3) greatly enhances the scoring correlation. The integration of all components (row 4) achieves the best results, validating the design choices and showing synergistic effects between the modules. The model characteristics of different methods are compared and analyzed, as shown in Table 3.

Table 3.

Characteristics of different pose estimation models used for comparison.

Model Recognition accuracy (%) Missed recognition rate (%) Joint angle error (%) Score correlation (%) Inference time (ms) References
OP-SVM

AR: 88 → 96

Yoga: 87 → 94

AR: 4.91 → 4.97

Yoga: 11.08 → 7.14

6–10 84–88 → 82 8–12 31
OP-LSTM

AR: 89 → 95

Yoga: 88 → 93

AR: 9.82 → 7.14

Yoga: 13.87 → 9.11

14–16 → 12 70–72 → 76–82 4–6 32
PoseGU 97.5 (est.) 3–5 (est.) 5.5 (est.) 85–90 (est.)  > 50 8
LightGraph 78.6 15–20 (est.) 8–12 (est.) 75–80 (est.) 13.9 12
ViTPose  > 99 (est.) 1–2 (est.) 4.0 (est.) 88–92 (est.) 30 23
IASDS

AR: 93 → 98

Yoga: 90 → 97

AR: 4.62 → 2.83

Yoga: 4.62 → 4.54

 < 6 92–98 4–22 This study

Table 3 presents a comprehensive performance comparison of various pose estimation models across five critical metrics for sports rehabilitation applications. The proposed IASDS model demonstrates excellent performance, particularly with regard to score correlation, which is crucial for automated rehabilitation assessments. The model excels in this area, achieving a correlation score of 92–98%. While ViTPose shows superior raw recognition accuracy (> 99%) and lower joint angle error (4.0%), it comes with higher computational cost (30 ms). LightGraph offers mobile efficiency (13.9 ms) but sacrifices accuracy across all metrics. Among the baseline models, OP-SVM and OP-LSTM demonstrate improved accuracy through training. However, they exhibit higher joint angle errors and lower score correlations than the IASDS framework. The IASDS framework strikes a balance between accuracy (less than 6% joint angle error and 93–98% recognition accuracy) and practical inference times (4–22 ms). This makes it suitable for real-time rehabilitation monitoring applications.

Application analysis of automatic score detection model for HBPR

To analyze the effectiveness of the automatic score detection model designed by the study when applied in practice, the study analyzes the loss function in recognizing human pose and score detection of IASDS in actual operation, as shown in Fig. 11.

Fig. 11.

Fig. 11

Loss function and accuracy of each training set.

Figure 11 shows the variation of the loss function when recognizing different human poses and scoring detection in a real run. In Fig. 11a, when recognizing different human poses, as the number of iterations increases from 0 to 20, the loss function decreases rapidly from about 1.9 at the beginning to about 0.5 after 2 iterations. After that, it tends to stabilize gradually, and finally stabilizes at about 0.1. Meanwhile, the accuracy rate increases rapidly from about 20% at the beginning, reaches about 70% in 2 iterations, and finally stabilizes at about 95%. In Fig. 11b, in terms of score detection, the loss function decreases rapidly from about 2.3 initially to about 0.8 in 2 iterations, and then stabilizes at about 0.1. The accuracy increases rapidly from about 10% initially, reaches about 50% in 2 iterations, and finally stabilizes at about 90%. Overall, the loss function decreases rapidly and stabilizes as the number of iterations increases, regardless of whether the task is recognizing different human poses or detecting scores. The accuracy rate also rises rapidly and tends to stabilize accordingly. It shows that the model gradually learns effective features during the training process, and the performance keeps improving and eventually stabilizes. The study compares the runtime consumption time and the credibility of the evaluation results of different methods with increasing samples, as shown in Fig. 12.

Fig. 12.

Fig. 12

Running operation time and assessing the reliability of results.

Figure 12 demonstrates the comparison of the run consumption time and the confidence of the assessment results of different methods with increasing samples. In Fig. 12a, the running time of the IASDS method shows a clear fluctuating upward trend with the increase of sample categories, and the overall running time is long. In particular, the running time peaks at sample categories A and E, which are about 22 ms and 20 ms, respectively. The OP-SVM method, on the other hand, has a relatively stable running time, which basically fluctuates between 8 and 12 ms. The OP-LSTM method has the most stable running time and the shortest overall running time, which is basically between 4 and 6 ms. In Fig. 12b, the IASDS method always has the highest confidence level and the overall trend is relatively smooth, gradually decreasing from about 98.5% to about 96%. The OP-SVM method has the next highest confidence, gradually decreasing from about 92.5% to about 88.5%. The OP-LSTM method exhibits the lowest confidence, with a gradual decrease from approximately 92% to 88%. The IASDS method performs best overall in terms of operational efficiency and clearly has an advantage in assessing the credibility of the results. This trade-off between higher confidence and longer runtime aligns well with the divergent demands of real-world rehabilitation settings. In clinical assessments, where accuracy is essential for making therapeutic decisions, the system can prioritize the high-confidence mode, even though processing takes longer. Conversely, maintaining patient engagement through immediate feedback is crucial for home-based rehabilitation. Therefore, switching to a faster, lower-latency mode is the best compromise for ensuring adherence without sacrificing too much reliability. The study analyzes the percentage of detected misidentified content and the score matching rate in the recognition results of different methods in practical application, as shown in Fig. 13.

Fig. 13.

Fig. 13

Proportion of misidentified content detected and matching rate of ratings.

Figure 13 demonstrates the percentage of detection misidentified content detected and score matching rate of different methods in real applications. In Fig. 13a, as the number of samples increases from 0 to 140, the misidentification rate of the IASDS method is always below 8%, which is the most stable performance. The misidentification rate of the OP-SVM method is between 6 and 11%, with relatively small fluctuations. The OP-LSTM method has the highest misidentification rate of 10–15% with large fluctuations. In Fig. 13b, the score matching rate of the IASDS method always stays above 90% and stabilizes after the sample number exceeds 60, maintaining at about 95%. The OP-SVM method has a score matching rate between 80 and 90%, which decreases slightly as the sample size increases. The OP-LSTM method has the lowest score matching rate, which is only 70–80% and fluctuates significantly. The results show that the IASDS method outperforms the other two methods in detecting the percentage of misidentified content detected and the score matching rate, and has higher accuracy and stability. Among them, the low error recognition rate and high matching score are the basis of system reliability. This reliability is the foundation for integrating the system into the user interface and establishing an automatic feedback mechanism. The study correlates the joint angle error and rehabilitation score of different methods in practical application, as shown in Fig. 14.

Fig. 14.

Fig. 14

Correlation between joint angle error and rehabilitation score.

In Fig. 14a, the performance of different methods in terms of joint angle error versus sample size variation is presented. Among them, the error rate of the IASDS method stays at a low level below 6% with the increase in the sample number. Furthermore, the fluctuation amplitude is small, showing high stability. The OP-SVM method has a higher error rate when the number of samples is small, which is about 6–8%. When the number of samples increases to 300 and above, the error rate decreases but still remains in the interval of 8–10%. The OP-LSTM method has the highest error rate and fluctuates significantly. The error rate is between 14 and 16% when the sample size is small. Figure 14b represents the correlation of rehabilitation scores of different methods with sample size. The correlation of the IASDS method is consistently high, ranging from about 96–98% when the sample size is small. As the sample size increases to 300 or more, the correlation decreases slightly, ranging from 92 to 94%. The correlation for the OP-SVM method is about 84–88% for smaller sample sizes and decreases to about 82% as the sample size increases to 300. The OP-LSTM method has the lowest correlation, which is about 70–72% for smaller sample sizes. As the sample size increases to 300 and 400, the correlation improves to about 76–82%, but then decreases slightly as the sample size reaches 500. The results show that the IASDS method is better than the other two methods in terms of joint angle error and rehabilitation score correlation, and has higher accuracy and stability. The results show that the IASDS method can measure the joint angle more accurately in practical applications. Moreover, the rehabilitation score has a strong correlation with other assessment indexes, which provides a more reliable basis for rehabilitation treatment. Among them, the low joint angle error and high score correlation validate the model’s effectiveness as an assessment tool. Subsequent clinical validation can be completed via concurrent validity tests with clinical scales.

Conclusion

This study successfully constructed an integrated framework (IASDS) for sports rehabilitation that leveraged BlazePose for real-time keypoint extraction, a random forest classifier for robust pose categorization, and an SNN with similarity metrics for automated scoring. Moreover, by combining with the random forest classifier, it realized the accurate recognition and classification of human pose in the complex environment. Meanwhile, an automatic scoring detection model based on SNN and similarity metric algorithm was designed to further improve the accuracy and efficiency of rehabilitation assessment. The experimental results indicated that the accuracy of IASDS for HBPR improved from 93 to 98% on the AR dataset and from 90 to 97% on the Yoga dataset. The joint angle error rate was below 6%. Rehabilitation score correlation was maintained at 92%-94% with larger sample size. The missed recognition rate was only 2.83% after 100 training sessions on the AR dataset and 4.54% after 100 training sessions on the Yoga dataset, both of which were better than the comparison methods. The missed recognition rate was consistently below 8%. The score matching rate stabilized at around 95% after the sample size exceeded 60. Practical application analysis revealed that the loss function of IASDS in HBPR and score detection decreased rapidly and leveled off. The accuracy rate increased rapidly and stabilized at a high level. Despite the long running time, the confidence of the evaluation results was always the highest, gradually decreasing from about 98.5–about 96%. The results indicated that the research method improved the performance and reliability of HBPR and automated scoring detection. This advancement is important for the development of intelligent and accurate sports rehabilitation. While the proposed model demonstrates high accuracy, several limitations must be addressed before practical deployment: (1) Limited environmental adaptability—Performance remains unvalidated under challenging conditions such as extreme lighting, dynamic backgrounds, or significant occlusion. (2) Constrained generalization capability—Using training data from specific demographic groups may limit the model’s applicability to diverse populations. (3) Suboptimal real-time performance—Further optimization is required for mobile deployment, which necessitates striking a careful balance between accuracy and latency.

Future work will focus on addressing the current limitations: (1) The model’s robustness can be improved by collecting and incorporating data captured under challenging conditions, such as varying lighting, complex backgrounds, and occlusions. Domain adaptation techniques can also be employed to enhance the model’s robustness. (2) To improve generalizability, the dataset should be expanded to include more diverse demographic groups and movement patterns. (3) The model’s efficacy and usability should be systematically validated through clinical trials in real-world rehabilitation scenarios.

Author contributions

X.Q.P. processed the numerical attribute linear programming of communication big data, and the mutual information feature quantity of communication big data numerical attribute was extracted by the cloud extended distributed feature fitting method. X.J.M. and X.M.F. Combined with fuzzy C-means clustering and linear regression analysis, the statistical analysis of big data numerical attribute feature information was carried out, and the associated attribute sample set of communication big data numerical attribute cloud grid distribution was constructed. X.Q.P. and X.M.F. did the experiments, recorded data, and created manuscripts. All authors read and approved the final manuscript.

Funding

The research is supported by the second batch of teaching reform projects for vocational education in Zhejiang Province during the 14th Five Year Plan period: Reform and Practice of Physical Education Curriculum in Higher Vocational Education Based on the Integration of Physical Education and Vocational Education—Taking the Creation of “Electrician” Outdoor Classroom as an Example (No:jg20240263); 2023 Guiding Science and Technology Research Project in Quzhou City: Research on Optimization Algorithm and Platform Design for Student Physical Health Evaluation (No. 2023ZD131).

Data availability

The datasets used and/or analysed during the current study available from the corresponding author on reasonable request.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Cai, H., Guo, S., Yang, Z. & Gao, J. A motor recovery training and evaluation method for the upper limb rehabilitation robotic system. IEEE Sens. J.23(9), 9871–9879. 10.1109/JSEN.2023.3258980 (2023). [Google Scholar]
  • 2.Zheng, C. et al. Deep learning-based human pose estimation: A survey. ACM Comput. Surv.56(1), 1–37. 10.1145/3603618 (2023). [Google Scholar]
  • 3.Chidambaram, V., Gopalsamy, M. M. & Kanchan, B. K. Ergonomic investigations on novel dynamic postural estimator using blaze pose and transfer learning. Ergonomics67(2), 240–256. 10.1080/00140139.2023.2221411 (2024). [DOI] [PubMed] [Google Scholar]
  • 4.Tariq, A., Yan, J., Gagnon, A. S., Khan, M. R. & Mumtaz, F. Mapping of cropland, cropping patterns and crop types by combining optical remote sensing images with decision tree classifier and random forest. Geo-spat. Inf. Sci.26(3), 302–320. 10.1080/10095020.2022.2100287 (2023). [Google Scholar]
  • 5.Gilakjani, P. and Al Osman, H. IoT firmware version identification using transfer learning with twin neural networks. Proc. 2nd ACM Conf. Internet Things Secur. Priv., pp. 1–10, Jan. 2024, 10.48550/arXiv.2501.06033
  • 6.Yang, F., Liu, J., Zhang, Q., Yang, Z. & Zhang, X. CNN-based two-branch multi-scale feature extraction network for retrosynthesis prediction. BMC Bioinformatics23(1), 476. 10.1186/s12859-022-04904-7 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Qi, J., Ma, L., Cui, Z. & Yu, Y. Computer vision-based hand gesture recognition for human-robot interaction: A review. Complex Intell. Syst.10(1), 1581–1606. 10.1007/s40747-023-01156-5 (2024). [Google Scholar]
  • 8.Guan, S., Lu, H., Zhu, L. & Fang, G. PoseGU: 3D human pose estimation with novel human pose generator and unbiased learning. Comput. Vis. Image Underst.233(7), 1–19. 10.1016/j.cviu.2023.103715 (2023). [Google Scholar]
  • 9.He, S. et al. Proposal and validation of a new approach in tele-rehabilitation with 3D human posture estimation: A randomized controlled trial in older individuals with sarcopenia. BMC Geriatr.24(1), 1–15. 10.1186/s12877-024-05188-7 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Martini, E. et al. Enabling gait analysis in the telemedicine practice through portable and accurate 3D human pose estimation. Comput. Meth. Program. Biomed.225(10), 107016. 10.1016/j.cmpb.2022.107016 (2022). [DOI] [PubMed] [Google Scholar]
  • 11.Wang, X. et al. Evaluation of hybrid deep learning and optimization method for 3D human pose and shape reconstruction in simulated depth images. Comput. Graph.115(6), 158–166. 10.1016/j.cag.2023.07.005 (2023). [Google Scholar]
  • 12.Chen, L., Xu, M. & Yang, J. LightGraph: Dynamic graph convolution for real-time 2D pose estimation. Int. J. Comput. Vis.131(9), 2408–2424. 10.1007/s11263-023-01818-6 (2023). [Google Scholar]
  • 13.Roggio, F. et al. A comprehensive analysis of the machine learning pose estimation models used in human movement and posture analyses: A narrative review. Heliyon10(7), e15660. 10.1016/j.heliyon.2024.e15660 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Tsakanikas, V. et al. Automated assessment of balance rehabilitation exercises with a data-driven scoring model: Algorithm development and validation study. JMIR Rehabil. Assist. Technol.9(3), e37229. 10.2196/37229 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Das, K., de Paula Oliveira, T. & Newell, J. Comparison of markerless and marker-based motion capture for balance and gait. Sci. Rep.13(1), 20441. 10.1038/s41598-023-49360-2 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Raihan, M. J., Mustafa, M. G. & Alammary, A. S. Rehabilitation exercise score prediction through integration of spatial feature with graph structural learning. Eng. Appl. Artif. Intell.110(1), 1–15. 10.1016/j.engappai.2022.104682 (2022). [Google Scholar]
  • 17.Chen, L., Liu, Y. & Wang, Y. Effective convolution mixed transformer Siamese network for robust visual tracking. Control Theory Technol.23(2), 345–360. 10.1007/s11768-025-00251-z (2025). [Google Scholar]
  • 18.Dai, Y. et al. MSEva: A musculoskeletal rehabilitation evaluation system based on EMG signals. ACM Trans. Sens. Netw.19(1), 1–23. 10.1145/352273 (2022). [Google Scholar]
  • 19.Xiang, Y., Zhang, Z., Chang, D. & Tu, L. The impact of gamified auditory-verbal training for hearing-challenged children at intermediate and advanced rehabilitation stages. Games Health J.13(5), 365–378. 10.1089/g4h.2023.021 (2024). [DOI] [PubMed] [Google Scholar]
  • 20.Li, W. et al. Development and evaluation of a wearable lower limb rehabilitation robot. J. Bionic Eng.19(3), 688–699. 10.1007/s42235-022-00172-6 (2022). [Google Scholar]
  • 21.Alsawadi, M. S., El-kenawy, E. S. M. & Rio, M. Advanced guided whale optimization algorithm for feature selection in BlazePose action recognition. Intell. Autom. Soft Comput.37(3), 2767–2782. 10.32604/iasc.2023.039440 (2023). [Google Scholar]
  • 22.Gündoğdu, S. Efficient prediction of early-stage diabetes using XGBoost classifier with random forest feature selection technique. Multimed. Tools Appl.82(22), 34163–34181. 10.1007/s11042-023-15165-8 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Miao, J. & Zhu, W. Precision-recall curve (PRC) classification trees. Evol. Intell.15(3), 1545–1569. 10.1007/s12065-021-00565-2 (2022). [Google Scholar]
  • 24.Yin, L., Li, B., Li, P. & Zhang, R. Research on stock trend prediction method based on optimized random forest. CAAI Trans. Intell. Technol.8(1), 274–284. 10.1049/cit2.12067 (2023). [Google Scholar]
  • 25.Subbiah, S., Anbananthen, K. S. M., Thangaraj, S., Kannan, S. & Chelliah, D. Intrusion detection technique in wireless sensor network using grid search random forest with Boruta feature selection algorithm. J. Commun. Netw.24(2), 264–273. 10.23919/JCN.2022.000002 (2022). [Google Scholar]
  • 26.Ma, Z., Qi, J., Xun, W. & Li, Y. Sports injury treatment and sports rehabilitation employing the nanoparticles containing zinc oxide. Adv. Nano Res.15(1), 67–74. 10.5755/j01.anr.15.1.48118 (2023). [Google Scholar]
  • 27.Prill, R., Królikowska, A., de Girolamo, L., Becker, R. & Karlsson, J. Checklists, risk of bias tools, and reporting guidelines for research in orthopedics, sports medicine, and rehabilitation. Knee Surg. Sports Traumatol. Arthrosc.31(8), 3029–3033. 10.1007/s00167-023-07525-1 (2023). [DOI] [PubMed] [Google Scholar]
  • 28.Vuyyuru, L. R. et al. Advancing automated street crime detection: A drone-based system integrating CNN models and enhanced feature selection techniques. Int. J. Mach. Learn. Cybern.16(2), 959–981. 10.1007/s13042-024-02315-z (2025). [Google Scholar]
  • 29.Pal, S., Roy, A., Shivakumara, P. & Pal, U. Adapting a Swin transformer for license plate number and text detection in drone images. Artif. Intell. Appl.1(3), 145–154. 10.47852/bonviewAIA3202549 (2023). [Google Scholar]
  • 30.Mokayed, H., Quan, T. Z., Alkhaled, L. & Sivakumar, V. Real-time human detection and counting system using deep learning computer vision techniques. Artif. Intell. Appl.1(4), 221–229. 10.47852/bonviewAIA2202391 (2023). [Google Scholar]
  • 31.Hassani, R., Boumehraz, M. & Hamzi, M. ECG signal classification based on combined CNN features and optimised support vector machine. Electroteh. Electron. Autom.72(2), 75–82. 10.46904/eea.23.72.2.1108008 (2024). [Google Scholar]
  • 32.Ramasamy, M. & Elangovan, M. Development of optimized cascaded LSTM with Seq2seqNet and transformer net for aspect-based sentiment analysis framework. Web Intell.23(3), 295–318. 10.3233/WEB-230096 (2025). [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

The datasets used and/or analysed during the current study available from the corresponding author on reasonable request.


Articles from Scientific Reports are provided here courtesy of Nature Publishing Group

RESOURCES