Skip to main content
Data in Brief logoLink to Data in Brief
. 2025 Mar 4;59:111440. doi: 10.1016/j.dib.2025.111440

FallVision: A benchmark video dataset for fall detection

Nakiba Nuren Rahman a,1, Abu Bakar Siddique Mahi a,1, Durjoy Mistry a, Shah Murtaza Rashid Al Masud a, Aloke Kumar Saha a, Rashik Rahman a,, Md Rajibul Islam b,
PMCID: PMC11950752  PMID: 40160526

Abstract

This article presents a comprehensive video dataset curated specifically for fall detection research, comprising categorized fall and no-fall videos. The dataset encompasses three primary categories of falls: falls from a bed, chair, and standing position. Initially collected as raw footage, these videos were subsequently processed to produce landmark videos, both with and without a background.

Recorded using handheld devices such as mobile phones and digital cameras, the dataset was sourced from voluntary participants, ensuring ethical compliance and informed consent. The dataset holds significant value for advancing fall detection algorithms, offering a robust platform for algorithm development and testing.

Fall detection systems are of paramount importance, particularly in scenarios where individuals are alone and unable to regain their footing post-fall or in cases where elderly individuals experience medical emergencies resulting in falls requiring immediate assistance. Leveraging this dataset, researchers can explore a plethora of techniques, including computer vision and deep learning, to devise and refine fall detection systems. Given its accessibility to researchers, this video dataset can be used in the advancement of fall detection technology to enhance safety measures for vulnerable populations.

Keywords: Fall detection, Video analysis, Fall classification, Human fall, Machine learning, Computer vision, Deep learning, Video dataset


Specifications Table

Subject Applied Machine Learning, Computer Vision.
Specific subject area Humans, especially individuals encountering medical emergencies. Videos depicting falls are gathered and converted into landmark videos, aiming to detect human falls by utilizing up to 17 key points within the processed footage.
Type of data The dataset consists of two types of data: videos in MP4 format and CSV files that track keypoints in each frame of the videos.
Data collection The dataset contains a number of videos obtained from voluntary participants who recorded the fall videos utilizing handheld devices. These raw videos are categorized into i) a number of fall videos and ii) a number of no-fall videos. The videos are categorized into videos of 58 volunteers falling from bed, falling from chair, and falling from standing positions.
Data source location Department of Computer Science and Engineering, University of Asia Pacific, Dhaka, Bangladesh.
Data accessibility Repository name: Fall Vision: A Benchmark Video Dataset for Advancing Fall Detection Technology
Data identification number: 10.7910/DVN/75QPKK
Direct URL to data: https://dataverse.harvard.edu/dataset.xhtml?persistentId=doi:10.7910/DVN/75QPKK
Instructions for accessing these data: This dataset has been made publicly available at the Harvard Dataverse repository for the purpose of any kind of academic, research, or instructional objective.

1. Value of the Data

  • The dataset comprises an extensive collection of videos, delineating two distinct classes. This comprehensive dataset functions to develop and test models and systems tailored for the detection of human falls, particularly within contexts associated with medical emergencies.

  • The fall video dataset is designed with a specific emphasis on advancing automated fall detection technology. It offers the opportunity to develop computer-assisted models, conduct comparative analyses, and refine systems aimed at detecting human falls more accurately, particularly in the context of individuals encountering medical emergencies.

  • This dataset can be utilized for building comprehensive and automated human fall detection systems involving deep learning techniques.

  • The dataset comprises videos recorded to closely simulate real-world human fall scenarios, ensuring alignment with practical conditions. These videos feature various backgrounds, enhancing the dataset's capacity to train machine learning models with greater accuracy and robustness.

  • This dataset includes human falls in three different scenarios: from a bed, from a chair, and from a standing position. These scenarios are represented by volunteers who show a variety of falling patterns simulating practical real-world situations which are not intricately covered by other existing datasets.

  • The dataset focuses on the keypoint-based fall detection, using 17 landmarks or key points of the human body extracted from the frames of the videos. This, in turn, makes certain that the detection process is unbiased toward the subjects' age, gender, or profession.

  • The dataset features CSV files for each video, which include the X and Y positions of keypoints along with their confidence scores. These files are intended to capture the movement patterns, focusing on the spatial positions of the keypoints.

2. Background

Human falls are common in people of different ages due to facing illness and various health issues. More commonly, people with cardiac disease tend to face sudden falls leading to requiring emergency support. Human falls are more frequent among elderly people, which can result in fatalities from hip or bone fractures to major physical injuries in vulnerable states. Studies have shown that around 28–35 % of elderly people experience at least one fall a year that may culminate in moderate to severe injuries. For this reason, falling has become a major problem for this age demographic [5]. On account of this, automatic fall detection systems are crucial for people who might not be able to call for assistance after falling [4]. To detect human falls, there are several fall detection systems, but they lack accuracy as they exhibit a considerable amount of false positives that cause a substantial number of false alarms [3]. Effective and accurate fall detection systems require advanced algorithms to be trained on comprehensive datasets, the lack of which results in difficulty differentiating between real and fake falls.

In order to accurately identify real falls, video datasets offer enriched contextual cues about the surrounding layout, interactions between different body parts, and the progression of events leading up to a fall, compared to images. Falls entail certain motion patterns; therefore, videos are better equipped to capture these patterns than still images. Furthermore, videos containing sequences of frames can help reduce false positives and distinguish a real fall from other irregular movements. While there have been previous works on video datasets [4,5,7], they often lack a focus on human falls from different positions. Another work [8], introduced a dataset consisting of data from healthy young individuals performing different activities including falls where the sensor positions in the dataset were based on right-handed individuals leading to absence of the broader population representation accurately and bias in the data collection process.

The proposed dataset aims to address this gap by focusing specifically on human falls from different states, with the goal of successfully distinguishing real falls from anomalous movements, especially in the case of elderly individuals. This video dataset stands out from existing fall detection datasets by providing a more comprehensive collection of fall videos recorded under diverse conditions.

In contrast to many existing datasets that might lack detailed annotations, this dataset consists of transformed landmarked videos annotated with 17 specified key points. This approach provides variety in human positions during falls along with comprehensive anatomical landmarks determining precise analysis. Furthermore, the dataset contributes to easing the training and evaluation of machine learning models by providing high-quality videos and detailed annotations. The following is a comprehensive survey of the existing fall detection datasets, as determined by their limitations:

Other datasets Limitation Our dataset with corresponding solution
UP-fall [8] biased toward right-handed individuals inclusion of unbiased video data that encompasses both left- and right-handed individuals.
Dataset of human actions [4] absence of a human fall from a lying position addition of comprehensive video data of human falls from a lying position (fall from bed)
absence of a human fall from a sitting position addition of comprehensive video data of human falls from a sitting position (fall from chair)
KFall [5] absence of specifications for falls from different positions consisting of human falls from three different positions: standing, chair or sitting, and bed or lying covering postures from different viewpoints.
Dataset of human actions [4], KFall [5] absence of keypoints and skeletons in human poses inclusion of 17 key points and skeletons of human poses

Utilizing videos with detailed annotations in a variety of situations, this dataset's main objective is to ensure precise differentiation between fall and non-fall incidents and speed progress in human fall detection using robust video analysis.

This dataset can be used by researchers to improve the accuracy of fall detection algorithms and better the emergency response for those who are at risk of falling.

3. Data Description

Addressing critical challenges in emergency response, healthcare, and technology development, a comprehensive dataset of human fall videos can provide rich and diverse video content for advancing human fall detection. The primary objectives of a fall detection system include assisting elderly individuals in health emergencies, reducing false emergencies, and fostering advancements in fall detection technology for the benefit of people of all ages.

Falls are a major cause of injury-related deaths among the elderly, and effective detection can reduce complications and injuries. One of the previous works in this field [1], included multi-visual modalities offering benefits such as obfuscated facial features and improved performance in low-light conditions and addressing real-world considerations such as varied lighting, continuous activities of daily living, and camera placement [1]. Another work [2] proposed a novel, efficient fall detection method based on future frame prediction, introducing a fall score based on the error between predicted and real frames to distinguish fall events during testing and achieve better performance in frame prediction based on surveillance videos of human falls. In concern with detecting falls, [6] introduced a system to distinguish between real fall activities and activities that seem similar to original falls by assessing human joint points as motion indicators. [7] Presented a dataset consisting of normal and anomalous situations such as falls and paralytic events to utilize in the process of classifying movements. A table below demonstrates comparison of our dataset to the existing fall datasets, considering factors such as devices, subject, camera view etc.

Dataset Device Subject Camera View Frame rate and resolution
Multi Visual Modality Fall Detection Dataset [1] Hikvision IP network camera, additional cameras to work as attachments to smartphones, vision-based sensors,
thermal cameras
Age and gender of the subject is not mentioned Camera mounted at the middle of the ceiling capturing a comprehensive view 20 frames per second (fps) with a resolution of 704 × 408 pixels
8.7 fps with a resolution of 1440 × 1080 pixels
Dataset of human actions [4] Exact mobile devices or camera models are unmentioned Age and gender of the subject is not mentioned Specific angles or views are not detailed Information is not available
KFall [5] Synchronized video camera, wearable inertial sensors, specific details are unmentioned Healthy young adults primarily including male, age is not mentioned Synchronized video captures including various angles, specific details are unmentioned Synchronized video is captured at a maximum frame rate of 90 Hz, resolution is not explicitly mentioned
UP-fall [8] Wearable sensors, electroencephalography headset, infrared sensors, specific details are unmentioned Healthy young adults including male and female, age is not mentioned Specific details are unmentioned 18 fps approximately, resolution is not explicitly mentioned
UR-fall [9] Single CCD cameras, multiple cameras, specialized omni-directional cameras, and stereo-pair cameras Five healthy volunteers over the age of 26 Angular field of view of 57 degrees horizontally and 43 degrees vertically using a Kinect sensor 30 frames per second (fps) with a resolution of
640 × 480 pixels
Our Dataset Mobile devices such as Google pixel 6a, Samsung A32, iphone 11 pro max and the Optoma SC26B webcam 58 healthy young adults including male and female in the age range of 22-27 years Front, back, left, right, floor, and ceiling—all the views were covered strategically to make the data inclusive A minimum frame rate of 30 fps with a video resolution of 720p and 1080p

The proposed video dataset consists of an enriched collection of fall videos captured from different backgrounds, including falls from various states, for elderly individuals facing health emergencies. Along with assisting elderly adults, the dataset has the potential to reduce the risk of health emergencies for people in general, including all ages as well by promptly detecting human falls, leading to medical urgency.

This dataset contains 11,732 human fall videos obtained from voluntary participants and recorded using handheld devices. These videos are categorized into 6002 fall videos and 5730 no-fall videos, where half of the videos are raw and the other half are masked (landmarked). Among the 6002 fall videos, there are 1974 videos of falls from beds (987 raw videos & 987 masked videos), 1986 videos of falls from chairs (993 raw videos & 993 masked videos), and 2042 videos of falls from a standing position (1021 raw videos & 1021 masked videos). On the other hand, among the 5730 no-fall videos, there are 1792 videos of no-fall from bed (896 raw videos & 896 masked videos), 1918 videos of no-fall from the chair (959 raw videos & 959 masked videos), and 2020 videos of no-fall from standing position (1010 raw videos & 1010 masked videos). Each video is accompanied by a CSV file that contains the X and Y positions of keypoints, along with their corresponding confidence scores, for each frame of the video. These files provide direct access to the key points, allowing for detailed analysis. It is worth noting that about 6–7 % (776 videos) of the total videos involve female participants. Though the proportion is not high, it does not impact the accuracy of human fall detection in our study. The detection process depends on the 17 landmark points of the human body extracted from the frames of the video, as these are the main features for the detection of falls and are independent of the subjects' age, gender, and profession etc.

This video dataset is available at the Harvard Dataverse public repository. These videos are categorized into different folders, as depicted in Fig. 1.

Fig. 1.

Fig 1:

Dataset folder format in the data repository.

4. Experimental Design, Materials and Methods

The videos of this dataset are in MP4 format, which was captured using handheld devices. A considerable number of 58 volunteers volunteered to record the raw videos based on a set of instructions. Precisely, the videos have been recorded using the Optoma SC26B webcam which has the capability to produce 4K Ultra HD video at 30 frames per second (fps) with a 120-degree ultra-wide angle and a dual-channel microphone. As for mobile devices, we have used the iPhone 11 Pro Max, which has the capability of producing 4K video at 24/30/60fps and 1080p video at 30/60/120/240fps with stereo sound recording; the Samsung A32, which can produce 1080p video at 30fps; and the Google Pixel 6a, which could produce 4K video at 30/60fps and 1080p video at 30/60/120/240fps.

All the subsets of the dataset consist of video data corresponding to similar characteristics. Precisely, the videos have a resolution of 720p and 1080p and a 30 fps (frame per second) rate. There are three types of videos in the dataset: fall from bed, chair, and standing position. After collecting the videos from the volunteers, the videos were masked (landmarked) using Yolo Pose, which is a single shot method that learns in one shot. Yolo Pose functions in two steps: person detection and keypoint localization. YOLOv7 pose, an extension of the one-shot pose detector, YOLO-Pose, is a multi-person keypoint detector. It is trained on the COCO dataset, which has 17 landmark topologies, as shown in Figs. 2, 3, and 4. The 17 key points represent specific human body landmarks. Precisely, there are 5 facial landmarks, including the nose, left eye, right eye, left ear and right ear. The upper body is represented by 6 landmarks: left shoulder, right shoulder, left elbow, right elbow, left wrist and right wrist. Similarly, the lower body includes 6 landmarks: left hip, right hip, left knee, right knee, left ankle and right ankle. These landmarks comprehensively cover the human body making a total of 17 keypoints as shown in Fig. 2. It is implemented in PyTorch, with the option of easy customization. With 17 key points, YOLOv7 pose detection runs for all frames. Two distinct figures, one for the no fall state and the other for the fall state, are presented below. These figures primarily consist of example screenshots of the video frames from the corresponding raw and masked (landmarked) videos for three different positions. Here, raw and masked videos, respectively, refer to the original video and the converted video with a skeleton that holds the landmark points of the human body representing pose estimation.

Fig. 2.

Fig 2:

YOLOv7 Pose with 17 landmark points.

Fig. 3.

Fig 3:

YOLOv7 Pose with 17 landmark points for no fall frames.

Fig. 4.

Fig 4:

YOLOv7 Pose with 17 landmark points for fall frames.

To detect keypoints in videos, video frames are passed as input, producing annotated images with keypoints and skeletons along with fps (frames per second) as output.

The method for keypoint detection is designed to process an input video frame to detect and annotate keypoints and skeletons of human poses. It also calculates the frame rate of this processing. The primary steps involve preparing single frames from a video, performing pose estimation, and annotating the detected key points on the frame.

First, a single frame from a video is taken as input by creating a duplicate of the input frame to work with, ensuring the original frame remains unchanged. Next, the frame size is adjusted to fit a specific format using a method called letterbox resizing, which maintains the aspect ratio. Another copy of the resized frame is then created for further processing. Subsequently, the image data is changed into a 4-dimensional tensor format, which is required for neural network models to process. Consequently, the tensor is converted into a format compatible with the PyTorch library. Then, the tensor is transferred to the computational device to perform the calculations.

For performing pose estimation, certain computational features are temporarily disabled to speed up the process. First, the start time is recorded. Then, a pre-trained model is used to predict keypoints and skeletons from the frame, and the end time is recorded. Based on the recorded time taken, the processing speed (fps) is calculated. To filter detections, a technique called non-maximum suppression is used to eliminate redundant detections, keeping the most relevant data. Next, the model's output is converted into a format to represent keypoints, and the processed image is transformed back into a format suitable for display, from tensor to a regular image format.

Finally, to annotate keypoints, each detected set of keypoints is iterated over, and the skeleton and keypoints are drawn on the image. The annotated image and the frame rate are returned as output.

This process ensures that each video frame is processed efficiently and accurately, providing visual insights into human poses and the system's performance speed.

Limitations

While this dataset provides videos representing a variety of practical falling patterns with valuable insights, there are some limitations that must be acknowledged.

  • In our dataset, specific scientific methods were not employed during data collection. However, the videos were carefully sourced from raw footage captured using handheld devices such as mobile phones and digital cameras, with voluntary participants who provided informed consent. Although we didn't apply advanced scientific methods during collection, the dataset is carefully structured to support model training and testing, especially for fall detection algorithms. The videos were processed using a single-shot method called YOLO Pose, which works in two steps: person detection and keypoint localization, to create landmarked videos. While no advanced scientific methods were used initially, the dataset offers a strong foundation for future research, where such methods can be applied to further enhance its effectiveness.

  • The dataset includes 58 healthy young volunteers aged 22 to 27 years, with approximately 6–7 % of the data featuring female volunteers. While we acknowledge the underrepresentation of females and older participants, it is important to note that fall detection is based on 17 body landmark points, making the methodology largely independent of age, gender, or profession. However, the representativeness of the sample itself remains a limitation.

Ethics Statement

All contributors involved in the creation of this video dataset participated voluntarily with complete awareness and consent. Video recordings of individuals experiencing falls were captured using smartphone cameras. To ensure safety, volunteers were instructed to fall onto a padded surface, such as a cotton mat, placed on the ground. Consequently, no physical harm was incurred during the recording process of the videos. In consequence, there are no ethical concerns regarding the health and safety of participants.

CRediT Author Statement

Abu Bakar Siddique Mahi: Methodology, Validation, Investigation, Data Curation, Writing-Original Draft, Visualization. Nakiba Nuren Rahman: Validation, Investigation, Writing-Original Draft, Visualization. Durjoy Mistry: Supervision, Validation, Project Administration. Shah Murtaza Rashid Al Masud: Supervision, Writing-Review, Project Administration. Aloke Kumar Saha: Supervision, Writing-Review, Project Administration. Rashik Rahman: Methodology, Writing-Review and Editing, Supervision, Validation, Project Administration. Md. Rajibul Islam: Writing-Review and Editing, Project Administration.

Acknowledgments

Acknowledgments

We extend our sincere appreciation and gratitude to the 58 dedicated members of the data collection team specifically listed on the following page below:

fallvolunters.netlify.app

Their invaluable contribution lies in diligently capturing fall videos and generously volunteering their time and efforts solely for data collection purposes.

Declaration of Competing Interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this data article.

Contributor Information

Rashik Rahman, Email: rashikrahman@uap-bd.edu.

Md. Rajibul Islam, Email: md.rajibul.islam@bubt.edu.bd.

Data Availability

References

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement


Articles from Data in Brief are provided here courtesy of Elsevier

RESOURCES