Abstract
Attention Deficit Hyperactivity Disorder (ADHD) is a prevalent neurodevelopmental disorder characterized by inattention, hyperactivity, and impulsivity. Current diagnostic methods rely primarily on subjective clinical evaluations, which are prone to bias. Neurophysiological techniques such as electroencephalography (EEG), eye tracking, and electrodermal activity (EDA) offer promising objective alternatives; however, their adoption is limited by the scarcity of large, public, multimodal datasets. To address this gap, we introduce the BALLADEER ADHD Dataset, a comprehensive multimodal resource that integrates simultaneous EEG, eye-tracking, and physiological signals from children and adolescents with ADHD and neurotypical controls. Data were collected through carefully designed cognitive tasks aimed at eliciting neurophysiological responses related to attentional control, response inhibition, and cognitive flexibility-key domains affected in ADHD. The dataset facilitates the development of machine learning models for ADHD classification and biomarker discovery through cross-modal analyses of EEG, eye movements, and autonomic nervous system activity. By publicly releasing this dataset, we aim to enhance transparency, reproducibility, and innovation in computational neuroscience and ADHD research.
Subject terms: Diagnostic markers, Neurological disorders
Background & Summary
Attention Deficit Hyperactivity Disorder (ADHD) is one of the most prevalent neurodevelopmental disorders, affecting approximately 5-7% of children and adolescents1,2. It is characterized by persistent patterns of inattention, hyperactivity, and impulsivity, which interfere with academic, occupational, and social functioning3–5. The disorder presents a complex etiology involving genetic, environmental, and neurobiological factors, with growing evidence pointing toward altered brain connectivity and dysfunctions in executive control networks6–8. Despite its high prevalence and impact, ADHD remains a challenging disorder to diagnose and manage, largely due to its heterogeneity and the overlap of symptoms with other psychiatric and developmental conditions9.
Traditional ADHD diagnosis relies on clinical evaluations using specific standardized scales10. In this context, technological advancements in neurophysiological assessment are transforming the study and evaluation of ADHD. However, these assessments are subjective and often prone to bias, necessitating the development of objective, neurobiological markers to improve diagnostic accuracy and treatment monitoring11–13.
In recent years, neurophysiological methods, especially electroencephalography (EEG), have gained attention as potential tools for ADHD assessment14. EEG allows for the non-invasive recording of brain activity, providing valuable insights into neural oscillations and functional connectivity patterns associated with attentional control and cognitive processing14,15. Studies have shown that children with ADHD often exhibit increased theta power, decreased beta power, and altered theta/beta ratios, suggesting an imbalance in cortical excitability and inhibition14,16.
In addition to EEG, eye-tracking and physiological measures offer complementary insights into ADHD. Eye-tracking studies have demonstrated that children with ADHD exhibit atypical oculomotor control during executive function tasks, including impaired inhibitory control and altered gaze patterns that correlate with symptom severity6. Electrodermal activity and heart rate variability reflect autonomic nervous system changes commonly observed in ADHD, with research showing altered stress reactivity compared to neurotypical controls. Integrating EEG with eye-tracking sand physiological signals enables cross-modal investigation of how attentional lapses manifest simultaneously in brain activity, oculomotor behavior, and physiological. This is an approach that remains underexplored due to the scarcity of publicly available multimodal datasets .
Despite the growing research using EEG to better understand and classify ADHD, the field is still constrained by several critical limitations. One of the most prominent issues is the lack of large-scale, publicly available, and multimodal datasets, which hinders both the reproducibility of findings and the development of generalizable machine learning models. In particular, most studies rely on small, often private datasets that limit robust statistical analysis and fail to capture the full heterogeneity of ADHD presentations. To contextualize these limitations, we next review the most relevant EEG-based datasets used in ADHD research, outlining their characteristics and the current gaps that underscore the need for more comprehensive open-access resources.
During our search, we only found a total of two publicly available datasets. The first one17, hosted by the National Brain Mapping Laboratory of Iran, includes EEG recordings from 61 ADHD subjects and 60 neurotypical ones. The data were collected using standard EEG procedures, capturing various brainwave patterns relevant to ADHD diagnosis. While this database is a valuable resource for researchers, its relatively small sample size limits the generalization of findings derived from it. The second public dataset18 contains a total of 144 EEG recordings, with 100 from individuals diagnosed with ADHD and 44 neurotypical ones. Of the 100 ADHD subjects, 52 were diagnosed with the inattention (ADD) subtype and 48 with the combined subtype (ADHD-C). To the best of our knowledge, this is the only dataset that differentiates between ADHD subtypes. The data are highly preprocessed, including filtering between 0.5 and 20 Hz, which limits the EEG spectrum, as it typically ranges from 0.5 to 60 Hz.
Unfortunately, the remaining datasets are not publicly available, hindering the reproducibility of the results and potentially introducing biases that other researchers are not able to detect. For example, the dataset presented in19 consists of EEG recordings from 20 ADHD subjects and 20 neurotypical ones. The data were collected at high frequencies to capture detailed brainwave patterns. Although the authors present a CNN-based model that achieves 88% accuracy, the small sample size and lack of diversity are key limitations when it comes to the robustness and generalizability of the results. Similarly, the authors of20 train a Convolutional Neural Network (CNN)-based model on a dataset composed by 50 ADHD subjects and 57 neurotypical ones.
The dataset presented in21 was obtained by recording EEGs from a total of 16 participants during photic stimuli. The main purpose of the study was to reveal the most effective channel and recording status for ADHD diagnosis, showing that the highest accuracy was obtained by training a Long Short-Term Memory (LSTM) model on the “Fp1,F7” channel. In22 a total 107 subjects participated, including 50 children with ADHD and 57 neurotypical ones. The authors analyze both deep learning features derived from a Convolutional Neural Network (CNN) model and 13 manually engineered brain network features. The dataset presented in23 is among the largest collections, composed by data from 181 patients with ADHD and 147 neurotypical ones gathered over the course of two years. The dataset from24 is only composed by EEG data from 40 individuals, but discerns between ADHD-C and ADD using DSM-5 criteria. In25, they focus on studying ADHD in adults, gathering data from 34 ADHD patients and 45 neurotypical ones. Finally, in26 the authors perform ADHD classification on EEG data from 15 ADHD patients and 18 neurotypical ones.
There are also a few datasets that are not publicly available but can be provided by the authors on demand. For example,27 curate a dataset containing data from 55 ADHD patients and 45 neurotypical subjects to study the correlation between the theta/beta ratio and the CPT-3 test used to measure attention in ADHD subjects. In28 analyze and compare the EEG data recording during the sleep of 19 ADHD patients and 29 healthy controls.
Although the amount of previous studies shows the potential for analyzing EEG data regarding ADHD, we were able to find several major limitations:
The first and most important one is the lack of publicly available datasets29. Specially when developing new predictive methods, it is essential to evaluate them in a quality dataset that acts as a benchmark so that they can be properly compared. Also, they should be evaluated in several datasets to test their generalization capabilities, as there are many factors that can have a considerable effect on the data, such as the age of the patient or its comorbidities. On the opposite end of the spectrum, using private datasets impedes the reproducibility of the results, which is essential to ensure that the results have been properly obtained.
The sample size of most datasets is very small. This poses two problems: (i) training and evaluating predictive models on a small sample size greatly hinders its generalizability and robustness and (ii) deep learning models, widely used in the previous works, benefit from learning from data at large scale, therefore we could expect a large increase in performance as larger datasets are provided.
As pointed out by29, as ADHD is a heterogeneous disorder, it would be highly desirable to provide data from different modalities together with EEG, specially from wearable technologies and electrocardiogram (ECG). Very few works have focused on this despite its great potential due to the low cost of recording with wearable devices.
To address these limitations, we present the BALLADEER (A Big dAta anaLytical pLatform for the diagnosis and treatment of Attention deficit hyperactivity Disorder featuring ExtendEd Reality) ADHD Dataset, a multimodal dataset that integrates EEG recordings, eye tracking, and physiological signals (electrodermal activity (EDA) and heart rate (HR)) from children and adolescents with ADHD and neurotypical controls. This dataset extends the previous work published in Heliyon30 which presented only the Attention Salckline game. The current release includes the complete project data with Attention Robots and Cognifit task with all the associated EGG, electrodermal activity gathered across all experiments.
The dataset was collected using a controlled experimental protocol incorporating several cognitive and attentional tasks, including Attention Slackline, Attention Robots, CogniFit, and Nesplora, as well as physiological variables registers, designed to assess different aspects of attentional processing and executive function. These tasks were selected to elicit neurophysiological responses associated with sustained attention, response inhibition, and cognitive flexibility, key domains that are often impaired in people with ADHD.
Our dataset aims to facilitate advances in computational neuroscience and artificial intelligence applications for ADHD research. By providing high-quality multimodal neurophysiological data, it enables the training and validation of machine learning models for ADHD classification, as well as the development of new digital biomarkers for objective assessment. In addition, the dataset enables cross-modal analyses, exploring how EEG dynamics interact with eye tracking and autonomic nervous system activity to determine attentional performance.
In the following sections, we provide a comprehensive overview of the methodological framework used for data acquisition, preprocessing, and quality control. Methods section details the methodology used for dataset creation, including participant recruitment, experimental design, and data collection procedures. In Data Records section, we describe the characteristics and structure of the dataset, highlighting its key features and potential applications. Finally, in Technical Validation section we summarize the technical validation procedures used to ensure signal quality and data reliability across sessions, including device calibration and monitoring.
Methods
In this section, we describe the methodological approach used to extract and process data for the presented dataset. The data acquisition process was designed to ensure accuracy, consistency, and high-quality multimodal recordings across different experimental conditions. We also detail the recruitment criteria, ethical considerations, and experimental setup, including the division of participants into control and test groups. Furthermore, we outline the sequence of data collection sessions, specifying the cognitive tasks performed and the biometric modalities recorded. Additionally, we describe the hardware and software architecture implemented to seamlessly integrate data from multiple sources, ensuring synchronization and efficient storage.
The following subsections provide an in-depth explanation of the participant demographics, experimental methodology, and technical infrastructure supporting the dataset.
Population and Methodology
For the compilation of the dataset presented in this study, participants were recruited, encompassing children and adolescents aged between 6 and 18 years. Prior to initiating the data collection phase, ethical approval was secured from the Ethical Committees of the University of Alicante (UA-2023-05-17_1) and the Alicante Institute for Health and Biomedical Research (ISABIAL; PI2022-139). Furthermore, informed consent forms were provided to all participants and signed prior to their participation in the study. Signing the form was a mandatory requirement for inclusion, either by the participants themselves or, in the case of minors, by their legal guardians. Parents and legal guardians provided consent for their children’s participation, data collection, and public sharing of their children’s data in an anonymized form. The data-sharing permission was included in the ethics committee approvals (UA-2023-05-17 1; ISABIAL PI2022-139).
The data collection period started in June 2022 and was completed in December 2024. In total, data were collected from 164 participants, comprising 66 females and 98 males. These participants have been allocated into two distinct groups:
Control Group: This group consists of 102 neurotypical subjects, including 52 females and 50 males.
Experimental Group: Comprising 62 subjects diagnosed with ADHD, this group includes 14 females and 48 males.
Patients with ADHD were referred from a specialized neurodevelopmental disorders unit within the Child and Adolescent Mental Health Service at the General University Hospital Dr. Balmis in Alicante (Spain). All diagnoses were established by experienced clinicians in accordance with the DSM-5 (Diagnostic and Statistical Manual of Mental Disorders, Fifth Edition) criteria. The diagnostic process included clinical interviews with patients, their families, and teachers, as well as the administration of standardized assessment protocols. Healthy control participants were recruited from the general population using snowball sampling techniques. Individuals with any severe condition that could interfere with the correct comprehension of the evaluation tasks or impair sensorimotor functioning were excluded from the study. Additionally, referring clinicians were instructed to avoid including patients presenting with severe comorbid symptomatology (e.g., anxiety, depression, or severe learning disorders), and a complete clinical history interview was conducted to confirm the absence of major psychological or neurological problems in the participants of the study.
While all ADHD diagnoses were established following DSM-5 criteria, ADHD subtype classification was not systematically recorded as part of our data collection protocol. This limits subtype-specific analyses, though future work could explore neurophysiological subtype markers through post-hoc clinical review or correlation with behavioral assessments from the evaluation battery.
Regarding the age demographics of the participants, the overall mean age was recorded at 12.4 years. Further breakdown by groups and gender (used here to denote biological sex) revealed the following mean ages: 13.3 years for females in the control group, 12 years for males in the control group, 12.6 years for females in the test group, and 11.9 years for males in the test group. In this manuscript and in the database, the variable labeled -“gender”- denotes biological sex (female/male) as recorded in the clinical record.
The sample encompassed a broad age range (6-18 years), covering key developmental stages associated with substantial neurophysiological and cognitive changes. This diversity in age allowed for a more comprehensive representation of the ADHD population and enhanced the ecological validity and generalizability of the results included in the dataset.
The psychology team in the project prepared the assessment protocol which was composed of two sessions with both pencil-and-paper and computerized tests. The evaluation tools selected for this study are considered gold-standard instruments for the assessment of ADHD symptomatology. They are widely used in both clinical and research settings and have demonstrated excellent psychometric properties. Moreover, the evaluation protocol combined traditional paper and pencil methodologies with computerized tasks inspired by classical neuropsychological paradigms, allowing for a comprehensive assessment of cognitive and behavioral functioning. Each employed measure included in the evaluation protocol is described in detail in the corresponding section later in the manuscript. In addition, the patient’s computer file was opened, anonymizing the data throughout the process to ensure the confidentiality of the participants. Study participants attended the assessment sessions with their parents. For each participant, the data gathering process was performed in two different sessions, each one of them performed in different days.
Participants were instructed to discontinue their medication from the day before the assessment, ensuring at least a 24-hour washout period to minimize potential acute pharmacological influences on task performance. Compliance with this procedure was monitored by the clinical team Fig. 1.
Fig. 1.
Stacked histograms of participant ages by group and gender.
More specifically, these sessions consisted of the following tests (those where physiological/biometric data were recorded are marked with an asterisk):
- First session:
- Interview between psychologist and participant, in which sociodemographic and clinical data are collected, as well as data on academic performance.
- CARAS-R31: a difference perception test administered according to the standardized protocol. This test has a total duration of 3 minutes.
- CogniFit (see Environment 3: CogniFit section), which evaluates a complete cognitive profile that includes perception and attention aptitudes. The IT team proceeds to set up the IT devices, such as the Emotiv EEG evaluation case and the Empatica psychophysiological recording bracelet.
- Attention Robots (see Environment 2: Attention Robots section). A gamifying test which assesses attention and perception processes.
- Second session:
- Attention Slackline (see Environment 1: Attention Slackline section). A gamifying test which assesses attention and perception processes.
- d2-R32: an attention assessment test administered according to the standardized protocol.
- Nesplora: neuropsychological assessment with virtual reality33. The team of computer experts adjusts the CGX EEG signal evaluation helmet, the Empatica psychophysiological signal evaluation wristband and the Tobii Eyetracker device. The computerized neuropsychological assessment test of attentional processes AULA (for participants aged 6 to 16 years) or AQUARIUM (for participants aged 17 to 18 years) is administered. Both tests are similar, since they evaluate the same variables and attentional dimensions, and are part of the Nesplora virtual environment applied by means of a Virtual Reality system.
- MATRICES34: a computerized intelligence assessment administered according to the standardized protocol.
Hardware Environment and Software Architecture
With the aim of gathering data from different devices, the Software and Hardware architecture shown in Fig. 2 was implemented. It is worth noting that the architecture was defined in order to be used by current and future games of the Attention series, as well as by any other third-party software. The architecture was implemented in a backend responsible of gathering and storing all the produced data, thus making this task independent of the data production environment. The backend exposes a comprehensive RESTful Application Programming Interface (API) that can be called by any software to start and stop gathering data, and to automatically store it at the end of this process. The start and stop endpoints calls are issued concurrently to all devices so EEG and psychophysiology begin and end at the same time, sharing a common time base. At each in game event, the game sent an event message to the backend with timestamps referenced to the same session start used to open the device streams. No hardware triggers are injected directly into the CGX data, so alignment is needed, and it must be performed after recording by matching each event timestamp to the nearest EEG sample time using the shared session start and the EEG sampling rate.
Fig. 2.
Software and Hardware architecture.
In the following, the different components of the architecture will be presented in detail. Note that the components are shown as images, while the connections between them are shown by means of arrows pointing in the direction of the dataflow:
Data Production Environment: This was the main data production application, developed using Unity for the Attention Games. This component was responsible for communicating with the backend to start and stop the recordings of biometrical data. Moreover, it sent events to the backend when they were produced in the game. For instance, in Attention Slackline, an event was produced when the player reacts (or fails to) to the appearance of a flag by pressing a button, or when a flag was shown and no button was pressed.
Tobii 5 Eye-tracker: As aforementioned, the game will gather whether or not the player was looking at several gaze-enabled elements. In this case, this was the only device connected directly to the game (instead of being connected to the backend). That is because Tobii provides an Software Development Kit (SDK) for Unity and, by using it, we were able to know whether the player was looking at a specific element of the game (the flag in our case) instead of having to analyze the exact position of the player’s gaze on the screen by means of (x, y) coordinates that would depend on the display resolution.
Attention Python Flask Backend: As stated above, this was the software component in charge of connecting all the biometrical devices (apart from the eye-tracker) and to get and store data from them. To do that, this component was implemented using Flask, a Python framework to create RESTful APIs. Therefore, a game-independent API was provided to allow all Attention games to interact with it in a straightforward manner. Indeed, the games only had to invoke the APIs methods for starting and stopping data recording, as well as the methods for registering a game event. Besides, this component was responsible for connecting the APIs of Empatica, Emotiv and CGX to retrive and transform the data produced by these devices. Finally, a visual web interface was also provided in order monitor the current status of each biometrical device.
Empatica Electrodermal Activity (EDA) and Heart Rate (HR) Wristband: In order to get several biometrical indicators from the players in a non-invasive manner, we used an EmbracePlus Wristband. This device provides us with the following data from the user, which could help practitioners better understand his/her status during the sessions: heart rate, inter-beat interval, blood volume pulse, electrodermal activity, skin temperature and movement (3-axis accelerometer).
CGX & Emotiv EEG Headsets: While the performance of the sessions, we also gather information about the players’ brain activity by using an encephalography (EEG) devices. More specifically, it was decided to use CGX Quick-32r EEG, a high-density dry EEG wireless system with LED impedance measurement since our target players were children and teenagers possibly suffering to ADHD. Thus, by using a dry EEG, the time required to place the device on the player’s head was considerably lower than if we had used a wet EEG device. In our case, the headset can be placed in 2-5 minutes (depending on the players’ hair or skin sensibility) and no re-hydration was needed during the sessions. Regarding the obtained data, raw data from 30 EEG channels was recorded with a sampling rate of 500 samples per second. CGX was primarily used with Attention Salckline. Therefore, by using this data, we will be able to train Deep Learning models able to detect players suffering from ADHD in the future [32]. It is worth noting that the Attention Python Flask Backend was ready to work not only with the CGX EEG, but also with other lower-cost devices from Emotiv such as Epoc X and Epoc+. These devices feature 14 saline-based wet electrodes recording at a sampling frequency that can range from 64 to 256 samples per second. Both devices share the same electrode layout and sampling options and are treated as a single device family. The Emotiv headsets were used with the Attention Robots and Cognifit tasks.
Synology 40TB Network-Attached Storage (NAS): Once all the biometrical data from the devices was gathered, it was combined with the events generated by the games and stored in a NAS. Afterwards, the backend transfers all this data through an encrypted SSH connection to this NAS which was only connected to the backend for security reasons. Hence, we were able to query this data using a dashboard.
Environment 1: Attention Slackline
“Attention Slackline” represents the first endeavor within this research to collect biometric data. Narrative-wise, the game is situated in a snow-clad valley, where the protagonist’s objective is to save a friend stranded on a distant mountain, marked by an orange flag. This rescue involves traversing between mountains using a slackline. Importantly, the narrative avoids mentioning any perils, aligning with the objective of maintaining a child-friendly atmosphere. Consequently, the character is depicted as never falling off the slackline.
Gameplay mechanics are structured such that the player consistently observes a target flag,(top-left) and presses the designated button when the pattern matches a flag appearing on the destination mountain where the friend is presumed to be located. Flags are displayed at predetermined times and remain on screen for a duration of 2 seconds. The stimulus schedule continues regardless of any responses, and flags are not replaced upon response. Successful gameplay requires the player to press a button when the patterns on the two flags match, which activates an animation that shows the character moving a little faster across the slackline (as depicted in Fig. 3). Conversely, pressing the button when the flags display different patterns or no flag is present activates a slow animation but the character remains in a fixed position. Each level of the game has a duration of five minutes, culminating in a motivational message that remains positive irrespective of the game’s outcome. Further details regarding this game are discussed in30.
Fig. 3.

Player walking fast after selecting the correct flag in Attention Slackline.
As far as biometrical data gathering is concerned, this environment will collect the following signals, recorded simultaneously:
Gaze Monitoring: The system captures and logs whether the player is viewing the flag at any given moment and records the specific timestamps of such occurrences.
EEG: Continuous recording of cerebral activity is maintained throughout the gameplay duration.
EDA and HR Monitoring: The player is equipped with a wristband that measures physiological indicators of arousal and cardiac function, respectively, to assess levels of nervousness.
Gameplay Logging: The game recorded the level onset and flag-spawn times in seconds from the start of the level. Level startup coincides with the start of the backend session. Events are linked to the EEG by converting the elapsed seconds into sample indices using the EEG sampling rate.
Environment 2: Attention Robots
The second game developed for this research, titled “Attention Robots," involves a task where players are confronted with an array of robots, each characterized by distinct features. The objective is for the player to accurately identify and select robots that meet specific criteria, such as possessing exactly two arms, regardless of their placement (either both on one side or one on each side), and maintaining an upright posture. This exercise aims to evaluate the player’s focus and precision in observing detailed attributes as per the established rules.
As depicted in Fig. 4, the game environment consists of a large grid comprising 14 rows and 47 columns, each slot filled with a robot variant. This setup provides a robust framework for players to exercise and evaluate their attention and decision-making skills. Each gameplay round requires players to identify the correct robots within a 20-second timeframe per line. Progression to the next line is indicated by a specific auditory signal and the highlighting of the new line. Robots that are correctly identified are highlighted in red, and players have the opportunity to correct any selections by re-clicking on the robots. The session concludes with a clear visual distinction between correctly and incorrectly marked robots. Details about this game can be found in35.
Fig. 4.
Player during Attention Robots game.
In terms of biometric data collection, the approach mirrors that of the first game, “Attention Slackline” with EEG and psychophysiology recorded simultaneously. However, the key difference lies in tracking the specific robot the player is observing during gameplay.
Environment 3: CogniFit
For the third environment, we used CogniFit36,a digital cognitive assessment platform. The platform administered a series of interactive exercises designed to evaluate cognitive skills including perception, attention, coordination, and problem-solving (see Fig. 5 to see an example). These exercises are dynamically adjusted to match the user’s performance, making them suitable for a wide range of ages and cognitive needs.
Fig. 5.

CogniFit exercise in progress.
During CogniFit sessions, the IT team configured and monitored the Emotiv EEG headset and the Empatica psychophysiological recording wristband. Because Cognifit is a third-party application without compatible API or event stream, we could not integrate task events or screen coordinates into our acquisition backend. For Cognifit task we recorded EEG and biometrical data. The eye-tracking was not collected, and the only available logs are the session level metadata.
Environment 4: Nesplora
Nesplora Aula is a neuropsychological assessment tool based on virtual reality, designed to measure sustained attention, impulsivity, and hyperactivity in children and adolescents aged 6 to 16. Participants aged 17 to 18 completed the AQUARIUM version, which evaluates the same attentional dimensions in an age-appropriate virtual environment. The test takes place in an immersive environment that simulates a classroom setting, where the child must complete continuous performance tasks (CPT) while facing realistic distractions such as background noise, moving classmates, or teacher interactions. During the assessment, objective data was recorded, including reaction times, number of correct and incorrect responses, and head movements, providing a detailed profile of the child’s attentional performance. Nesplora Aula has been scientifically validated and includes normative studies that allow for comparisons with reference populations. Its use in clinical and educational settings helps complement ADHD diagnosis and supports the development of personalized intervention strategies.
Data Records
The dataset is available at Figshare37. This section provides an overview of the data sources used for constructing the presented dataset, detailing the hardware and software used in the acquisition process. This dataset includes the full project data. The subset corresponding to the Slackline game alone was described in our previous publication30. The present release extends that work by adding the Robots and CogniFit tasks and increasing participant, session and modality coverage.
The dataset integrates multiple modalities, including EEG, eye-tracking, and physiological signals, captured using a diverse set of devices. Within each session, the available modalities were recorded simultaneously and share a common task-derived start and stop time alignment. The collection contains only recording while doing the tasks, no dedicated resting state EEG is included, and tutorial periods preceding task were not recorded. Please note that the dataset contains raw EEG data. We acknowledge that saccadic eye movements during tasks like Attention Slackline may introduce artifacts in the EEG. It is recommended that users implement preprocessing methods to remove these artifacts as needed. Since no offline preprocessing (filtering, re-referencing, artifact rejection, or epoching) was applied beyond hardware-level signal conditioning. For Emotiv, EmotivPRO applies a digital fifth-order Sinc filter, notch filters. For CGX, minimal hardware filtering is applied as documented in Methods. Researchers should implement their own preprocessing pipelines suited to their analytical objectives.
Table 1 summarizes the availability of recordings across different data acquisition environments and devices. Notably, the Attention Slackline environment includes three distinct levels, each generating separate datasets, as described in30. Additionally, non-biometric data, such as game logs, were collected to provide contextual information on participants’ interactions during the experimental tasks. In the following subsection, we describe the file structure, naming conventions, and variable definitions for all data files. Counts reflect unique participants with at least one valid recording per task and device. Differences in totals occur because not all participants used or could not tolerate specific devices (for example, headset fit or discomfort), and some sessions had temporary device failures. Counts are not additive across environments.
Table 1.
Composition of the dataset, relating data acquisition environments, hardware devices and number of participants.
| Epoc+ | Epoc X | CGX | EmbracePlus | Tobii | Non-biometric | |
|---|---|---|---|---|---|---|
| Slackline (lvl. 1) | — | — | 140 | 99 | 154 | 154 |
| Slackline (lvl. 6) | — | — | 140 | 99 | 153 | 153 |
| Slackline (lvl. 11) | — | — | 140 | 100 | 151 | 151 |
| Robots | 59 | 73 | — | 95 | 145 | 150 |
| CogniFit | 59 | 70 | — | 84 | — | — |
Data Structure
This section focuses on the storage of data gathered from the different tests performed by the participants, which are stored on the NAS, as mentioned above. Figure 6 illustrates the folder structure and storage hierarchy.
Fig. 6.
File structure of each participant.
To maintain a clear and consistent organization, the naming format shown in Fig. 6 was adopted. This format helps identify each file and its contents using key variables. Below, we explain each variable used in the hierarchy and file names.
UserID: A unique identifier assigned to the user.
Level: Indicator representing which level is being used, in cases where the game has multiple levels.
UnixSessionDate: The directory name representing the date in Unix format when it was created.
SessionDate: The date in the format YYYY-MM-DDTHH-M-S+GMT that represents when the data was collected.
DeviceModel: The model of the Emotiv headset used. Both models have the same number and location of electrodes. The models may include Epoc X and Plus.
DeviceId: The serial number of the Emotiv headset used.
Each file is structured to provide specific insights, and the variables within where selected to capture data from the activities. In the following tables, detailed explanations of each variable are provided.
The EEG data captured using the CGX device is stored in a CSV named [UserID]_EEG_CGX_[SessionDate].csv file. The specific format and structure of the data, including timestamps, EEG channel readings, and other relevant metrics, are describe in Table 2.
Table 2.
EEG data captured from CGX EEG device.
| Column | Description |
|---|---|
| Timestamps | Data capture timestamp in seconds (represent the time elapsed since the computer started). |
| AF7, Fpz ... AF8 | Voltage in microvolts captured in the corresponding EEG channel. |
| A2, ExG 1, ExG 2 | Voltage in microvolts in additional or external channels. |
| ACC32, ACC33, ACC34 | Acceleration measured in milligrams along the X, Y, Z axes, respectively. |
| Packet Counter, TRIGGER | Packet counter and trigger values for synchronized events. |
For the EEG data from the Emotiv headset Table 3 contains EEG data captured during the session using an Emotiv device, providing detailed records of the brain’s electrical activity for the participants.
Table 3.
EEG data captured using Emotiv device.
| Column | Description |
|---|---|
| EEG.Counter | EEG sample counter. |
| EEG.Interpolated | Indicates if the sample was interpolated. |
| EEG.[Channel] | EEG data from different channels (AF3, F7, F3, etc.). |
| EEG.RawCq | Raw EEG signal quality. |
| EEG.Battery | Device battery status. |
| EEG.BatteryPercent | Percentage of remaining battery on the device. |
| EEG.MarkerHardware | Hardware marker used during EEG data collection. |
| CQ.[Channel] | Signal quality in each specific channel (AF3, F7, etc.). |
| EQ.SampleRateQuality | Quality of the sampling rate. |
| EQ.OVERALL | Overall equipment quality. |
| EQ.[Channel] | Specific quality of each EEG channel. |
| MOT.[Data] | Movement data captured by MEMS sensors. |
| PM.[Metric] | Mental performance metrics such as attention and relaxation. |
| POW.[Channel].[Frequency] | Spectral power in different frequency bands. |
The file [UserID]_TAGS_[SessionDate].csv, as presented in Table 4, contains event tags marked during the session. This data is primarily utilized for the activity Attention: Slackline.
Table 4.
Event tags recorded during the session.
| Column | Description |
|---|---|
| Timestamp | Timestamp of when the event occurred. |
| label | Label that identifies the type of event or marker. |
| value | Additional details about the event, structured in JSON format, including reaction times, whether the subject reacted, whether the reaction was correct, and other data related to the test. |
| port | Identifier of the endpoint or interface that generated or captured the tag. |
The data stored in [UserID]_EYE_TRACKING_DATA_[SessionDate].csv contains eye-tracking data recorded during Attention: Robots session (Table 5). This file captures which robots were observed by the participant.
Table 5.
Eye tracker recorded during Attention Robots.
| Column | Description |
|---|---|
| timeChecked | The time (in seconds) at which it was recorded that the robot was being observed. |
| looked_col | The column number of the robot that was observed. |
| looked_row | The row number of the robot that was observed. |
The file users_demographics.json contains demographic and general information about the participants. This data includes key attributes such as age, gender (used here to denote biological sex: ‘1’ = male, ‘2’ = female), handedness, ADHD diagnosis status (categorized as yes, no, or undetermined), and the assigned experimental group as determined by the project’s team of psychologists (see Table 6).
Table 6.
Data description containing participant demographics and ADHD diagnosis status.
| Column | Description |
|---|---|
| username | Unique identifier for the user. |
| age | Age of the user. |
| gender | User’s biological sex (’1’ for male, ’2’ for female). |
| group | Group assignment of the participant, as determined by the project’s team of psychologists. |
| diagnosed | Indicates whether the participant has an ADHD diagnosis. ‘yes’ means the participant has been diagnosed with ADHD, ‘no’ means they have no diagnosis, and ‘undetermined’ refers to cases where the participant has no formal diagnosis but project experts suspect they may have ADHD. |
The slackline_flags_data.json file contains comprehensive information regarding the visual flags presented during the Slackline activity. For each Slackline level, the file includes an array of flag details that not only describes the type of each flag but also specifies the exact time at which it is scheduled to appear. In other words, for every level, additional information is provided listing all the flags and their corresponding spawn times, thereby enabling a fine-grained analysis of the stimuli presented during the task Table 7.
Table 7.
Slackline levels 1,6 and 11 flag information.
| Column | Description |
|---|---|
| level | Slackline level. |
| flags | Array of objects describing flags in the level. |
| flag_type | Type of flag, represented by name (circle, square, rhombus, doubleCircle). |
| flag_spawn_time | Time (in seconds) elapsed from the start of the level. |
Technical Validation
Ensuring data quality is a critical step in any research or analysis, as the data directly affects the results. This process helps us ensure that the data meet the necessary standards and that this one is accurate, complete, and suitable for the analysis it also by reducing errors and potential biases.
To achieve the highest quality of data collection, all devices, including EmbracePlus wristband, EEG headsets (CGX and Emotiv), and the Tobii Eye Tracker, undergo regular maintenance and preparation before each session. This process involves cleaning, charging and checking for software updates, all in accordance with the manufacturer’s instructions. Additionally, before each trial, we ensure that the entire system, from the device to the data storage platform, is functioning correctly. These precautions minimize the risk of data loss or interruptions during the experiment.
Throughout the experiment and between activities, we continually monitor the devices to ensure everything remains in proper working order. We also check in with participants to confirm that the equipment, such as the wristband, headsets remains comfortable and does not cause discomfort that could affect the data collection process. In case any issues arise, we promptly seek a solution to ensure that the participant’s well-being is never compromised.
To ensure that each device is ready for data collection, we follow a detailed preparation process tailored to the specific requirements of each piece of equipment. Below, we outline the steps taken for each device to guarantee accurate and reliable data throughout the experiment.
To ensure the highest data quality from the EmbracePlus wristband, we use the CareLab application, which provides a user-friendly interface to guide us throw the placement and correct connection of the device. Following the step-by-step instructions we make sure that the wristband is positioned properly on the participant. Once the wristband is in place, the app has real-time indicators that tell if the device is functioning correctly and that data recording has begun.
Throughout the session, we regularly monitor these indicators to ensure data is being captured without interruptions. At the end of each session, we verify that the wristband successfully syncs the collected data to the cloud, ensuring no data is lost.
With the EEG data from the Emotiv headsets, the Emotiv desktop software is required for setup and monitoring. This application also has a step-by-step for fitting the headset. Once the headset is placed on the participant, the application displays for each sensor contact, based on the impedance, an indicator shown as gray for no connection, yellow for poor connection, and green for a strong one. The setup cannot proceed until all sensors show a green status, ensuring a proper contact.
In the final step, the application displays the quality of the EEG signal from each sensor. This step is crucial as it helps monitor the strength of the signal. To ensure optimal data quality, we aim for a 100% signal quality, with an acceptable threshold being no lower than 90%. Once this level is maintained, the setup is complete. The last step is to connect the headset via our API to begin the data recording process.
The process for the CGX headset follows a similar procedure of the Emotiv headset. After confirming that the device is fully charged and ready, we place the headset on the participant. Each sensor on the CGX headset has an integrated light system that reflects impedance-based contact quality in real time. The CGX acquisition software provides a color-coded impedance map, with red indicating poor contact and green indicating optimal contact. Once all sensors display green, indicating proper contact, we proceed to verify that the headset connects correctly with our system. After confirming the connection, we proceed with the rest of the activity.
For the Tobii eye tracker, the first step is to position the monitor at the appropriate height for the participant. Once aligned, we begin with the calibration process through the Tobii software, which confirms that the participant is seated at the correct distance from the screen.
The calibration involves displaying a series of dots on the screen that the participant must look with their eyes until each dot “explodes”. This step ensures that the eye tracker is accurately calibrated to the participant.
Finally, to assess the quality of the data generated directly from our activities, we implement a thorough validation process. Once the session is completed, a local backup of the generated data is stored on the local computer. The data is then sent to our server via our API, which stores it on our NAS. After each activity, a verification process is performed both locally and remotely to ensure the data has been successfully uploaded. In the event of an error that could affect data collection or quality, we evaluate the possibility of repeating that specific activity.
When the session ends we perform a basic analysis step that generates visualizations of the collected data. These visualization allow us to easily asses the dataset quality without altering the original data. This step provides a quick, visual overview of the data before moving forward with a more detailed analysis. It is important to note that formal cross-device equivalence testing between CGX and Emotiv was not conducted. The dataset is released to enable such work. Analyses should account for device family explicitly rather than pooling signals across families.
Limitations
Despite all the measures taken during the sessions, working with children, some of them diagnosed with ADHD, as well as using various types of technological devices, presents unique challenges. These factors may introduce occasional errors or unexpected issues during activities.
One challenge we have encountered is the loss of eye tracking accuracy during the activities, particularly during the Attention Robots test. As shown in Fig. 7, we present the case of a participant who achieved consistent eye-tracking throughout the test (Fig. 7a), compared to another participant who, due to various factors such as movement or looking away from the Tobii device, experienced intermittent tracking loss (Fig. 7b).
Fig. 7.
Eye-tracking heatmaps for Attention Robots. Each cell corresponds to one robot in grid (Y-axis: row index, X-axis: column index). Color encodes the number of gaze-check events recorded at that position during the session.(a) shows stable tracking and (b) shows intermittent tracking loss. A shared color scale is used.
We also encounter cases where participants wearing EEG headsets perform actions that affect the signal quality. These actions can include abrupt head movements or touching the sensors. While we strive to maintain control over the protocol, such situations are difficult to prevent entirely, especially when working with children.
In addition, during the Slackline attention task, the target flag and the mountain flags appear in different parts of the screen, so participants may shift their gaze back and forth. These eye movements can introduce artifacts in the EEG signals and may influence some analyses (for example, by adding eye-related activity and affecting the apparent distribution of activity over posterior electrodes). Since the EEG is shared as raw data, users should apply preprocessing and artifact-handling steps that fit their planned analyses.
No dedicated resting-state EEG was collected. Recordings start at task onset and stop at task end, and tutorial or instruction periods were not recorded.
In Fig. 8, we present a recording that includes some of the cases mentioned earlier, where abrupt head movements or sensor interference impacted the EEG sensor level signal quality. In the red box, an artifact can be observed, likely caused by a sudden movement, while in the blue box shows a visible artifact produced by electrode displacement, in this case affecting the P3 electrode channel.
Fig. 8.
Example of a filtered EEG signal, the blue box highlights an instance of interference affecting one of the sensors, and the red box indicates an artifact likely caused by abrupt movement.
Beyond the technical challenges encountered during data acquisition, the dataset presents some limitations that warrant discussion. First, the sample size, while larger than that of many existing neurophysiological datasets on ADHD, remains relatively small for training robust deep learning models or for capturing the full heterogeneity of ADHD presentations. Furthermore, the participant sample is geographically homogeneous, drawn from a single region, which may limit the generalizability of the findings to other populations. The dataset does not include a long-term study, which precludes the study of ADHD symptom trajectories and changes in neurophysiological markers over time. Additionally, we did not conduct formal equivalence studies between the CGX and Emotiv EEG systems, meaning that potential device-specific effects must be considered when analyzing data from different sessions or comparing findings from different tasks.
Despite these limitations, the dataset offers substantial value for clinical research. The multimodal nature of our data represents a first step toward new methods to assist healthcare professionals. However, clinical translation will require rigorous validation studies in larger and more diverse populations, as well as comparison with existing clinical standards.
The BALLADEER ADHD dataset opens several avenues for future research. First, it provides a platform for developing and evaluating machine learning algorithms that classify ADHD based on neurophysiological signatures, with particular potential for exploring how different modalities contribute complementary diagnostic information. Second, researchers can investigate whether specific patterns in EEG signals, eye movement characteristics, and task responses could aid in early diagnosis. Finally, simultaneous recording from multiple devices allows for cross-modal analysis, such as examining how visible distractions in eye-tracking data correspond to changes in EEG signals or electrodermal activity.
Acknowledgements
This work has been co-funded by the BALLADEER (PROMETEO /2021/088) project, a Big Data analytical platform for the diagnosis and treatment of Attention Deficit Hyperactivity Disorder (ADHD) featuring extended reality, funded by the Conselleria de Innovación, Universidades, Ciencia y Sociedad Digital (Generalitat Valenciana); the IAEAV project (INREIA/2024/176) funded by the Conselleria de Innovación, Industria, Comercio y Turismo (Generalitat Valenciana); the BALIDA-AA project (CIPROM/2024/13), funded by Conselleria de Educación, Cultura, Universidades y Empleo (Generalitat Valenciana); the KOSMOS-UA project (PID2024-155363OB-C43), funded by Spanish Ministry of Science and Innovation; the ENIA Chair of Artificial Intelligence (TSI-100927-2023-6), the AgroVAL (TSI-100122-2024-10), Sophia (TSI-100130-2024-10) and European mobility for efficient planning and new business opportunities (TSI100121-2024-10) projects, funded by the Recovery, Transformation and Resilience Plan from the European Union Next Generation through the Ministry for Digital Transformation and the Civil Service; and Grant RED2022-134656-T by MCIN/AEI/10.13039/501100011033.
Author contributions
Juan Trujillo and Rosario Ferrer-Cascales: Conceptualization, Methodology, Software, Validation, Formal analysis, Data curation, Writing - original draft, Project administration, Supervision. Miguel A. Teruel and Nicolás Ruiz-Robledillo: Conceptualization, Methodology, Software, Validation, Formal analysis, Data curation, Writing - original draft, Supervision. Javier Sanchis, Sandra García-Ponsoda, Alejandro Panagiotidis-Arrizabalaga, Natalia Albaladejo-Blázquez, Ángela Martí-nez-Nicolás, Jorge García-Carrasco, Alejandro Reina, Ana Lavalle, Alejandro Maté and Borja Costa-López: Methodology, Software, Validation, Data curation, Writing - original draft.
Data availability
The dataset supporting the findings of this study is publicly available at Figshare37 (10.6084/m9.figshare.28676042). The repository contains the full multimodal dataset in JSON and CSV formats organized by participant and session, including EEG (CGX, Emotiv), eye-tracking, EDA, heart-rate, game event logs (tags and metadata) and participant demographics. All data shared are anonymized. Further details about file structure and variable names are provided in the repository and in the manuscript Data Records section.
Code availability
No custom code was generated for this dataset, but specialized software tools were used. EEG signals were captured using the Emotiv API (via the Emotiv SDK) and the CGX Acquisition Software, while games were developed in Unity. In addition, physiological measurements of the embraceplus wristbands were recorded using the Carelab Platform by Empatica. The following section provides a concise overview of each tool and its role in the study. No custom code was generated for signal preprocessing or analysis. The only custom software developed was a Python Flask backend API for coordinating multi-device data acquisition and file transfer, which does not modify recorded signals. • Emotiv API (Emotiv SDK / EmotivPRO Platform): The Emotiv API, provided as part of the Emotiv SDK (associated with the EmotivPRO platform), was used to acquire and process EEG data during the Robots and Cognifit sessions. • CGX Acquisition Software: CGX Acquisition Software from CGX Systems is a high-end EEG recording and analysis platform designed for both research and clinical applications. Provides robust real-time data acquisition, advanced signal processing, and synchronized event marking to ensure high-fidelity neurophysiological data.• Unity: Is a widely used game development engine that serves as the basis for designing and deploying the games. Its robust real-time rendering and interactive capabilities allowed the creation of immersive environments and precise control over stimulus presentation and data integration. • Empatica Care: A cloud-based platform developed by Empatica for continuous remote monitoring of the EmbracePlus wristband. The platform integrates data such as pulse rate, electrodermal activity, and temperature, enabling real-time tracking and analytics of key biomarkers.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Polanczyk, G., de Lima, M. S., Horta, B. L., Biederman, J. & Rohde, L. A. The worldwide prevalence of adhd: A systematic review and metaregression analysis. The American Journal of Psychiatry164(6), 942–948, 10.1176/ajp.2007.164.6.942 (2007). [DOI] [PubMed] [Google Scholar]
- 2.Salari, N. et al. The global prevalence of adhd in children and adolescents: a systematic review and meta-analysis. Italian Journal of Pediatrics49(1), 48, 10.1186/s13052-023-01456-1 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Arnett, A. B., Antúnez, M., Zeanah, C., Fox, N. A., Nelson, C. A., Physical and neurophysiological maturation associated with adhd among previously institutionalized children: a randomized controlled trial, Journal of Child Psychology and Psychiatry 10.1111/jcpp.14110 (2025). [DOI] [PMC free article] [PubMed]
- 4.A. P. Association, Diagnostic and Statistical Manual of Mental Disorders, 5th Edition, American Psychiatric Publishing, Arlington, VA, 10.1176/appi.books.9780890425596(2013).
- 5.A. P. Association, Diagnostic and Statistical Manual of Mental Disorders, 5th Edition, American Psychiatric Publishing, Arlington, VA, 10.1176/appi.books.9780890425787(2022).
- 6.Castellanos, F. X. et al. Executive function oculomotor tasks in girls with adhd. Journal of the American Academy of Child & Adolescent Psychiatry39(5), 644–650, 10.1097/00004583-200005000-00019 (2000). [DOI] [PubMed] [Google Scholar]
- 7.Johnson, D. E. et al. Growth parameters help predict neurologic competence in profoundly deprived, institutionalized children in romania. Pediatric Research45, 126 (1999). [Google Scholar]
- 8.Navarro-Soria, I. et al. Consequences of confinement due to covid-19 in spain on anxiety, sleep and executive functioning of children and adolescents with adhd, Sustainability 13(5) 10.3390/su13052487 (2021).
- 9.Overgaard, K. R. et al. Functional impairment related to adhd from preschool to school age. Journal of Attention Disorders29(3), 220–230, 10.1177/10870547241301179 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.A. A. of Pediatrics, Clinical practice guideline for the diagnosis, evaluation, and treatment of attention-deficit/hyperactivity disorder in children and adolescents, Pediatrics 128, 1-16 (2011) [DOI] [PMC free article] [PubMed]
- 11.Cortese, S. The neurobiology and genetics of attention-deficit/hyperactivity disorder (adhd): what every clinician should know. European Journal of Paediatric Neurology16(5), 422–433, 10.1016/j.ejpn.2012.01.009 (2012). [DOI] [PubMed] [Google Scholar]
- 12.Scassellati, C., Bonvicini, C., Faraone, S. V. & Gennarelli, M. Biomarkers and attention-deficit/hyperactivity disorder: A systematic review and meta-analyses. Journal of the American Academy of Child & Adolescent Psychiatry51(10), 1003–1019.e20, 10.1016/j.jaac.2012.08.015 (2012). [DOI] [PubMed] [Google Scholar]
- 13.Lee, S.-Y., Wang, L.-J., Yen, C.-F., Identification of diagnostic and therapeutic biomarkers for attention-deficit/hyperactivity disorder, The Kaohsiung Journal of Medical Sciences e12931, 10.1002/kjm2.12931 (2025) [DOI] [PMC free article] [PubMed]
- 14.Joy, R. C., George, S. T., Rajan, A. A. & Subathra, M. Detection of adhd from eeg signals using different entropy measures and ann. Clinical EEG and Neuroscience53(1), 12–23, 10.1177/15500594211036788 (2022). [DOI] [PubMed] [Google Scholar]
- 15.Snyder, S. M., Rugino, T. A., Hornig, M. & Stein, M. A. Integration of an eeg biomarker with a clinician’s adhd evaluation. Brain and Behavior5(4), e00330, 10.1002/brb3.330 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Slater, J. et al. Can electroencephalography (eeg) identify adhd subtypes? a systematic review. Neuroscience & Biobehavioral Reviews 139, 104752, 10.1016/j.neubiorev.2022.104752 (2022) [DOI] [PubMed]
- 17.Motie Nasrabadi, A., Allahverdy, A., Samavati, M., Mohammadi, M. R., Eeg data for adhd / control children 10.21227/rzfh-zn36 (2020).
- 18.Vahid, A., Bluschke, A., Roessner, V., Stober, S. & Beste, C. Deep learning based on event-related eeg differentiates children with adhd from healthy controls. Journal of clinical medicine8(7), 1055, 10.3390/jcm8071055 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Dubreuil-Vall, L., Ruffini, G. & Camprodon, J. A. Deep learning convolutional neural networks discriminate adult adhd from healthy individuals on the basis of event-related spectral eeg. Frontiers in neuroscience14, 251, 10.3389/fnins.2020.00251 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Chen, H., Song, Y., Li, X., Use of deep learning to detect personalized spatial-frequency abnormalities in eegs of children with adhd, Journal of neural engineering 16(6) 066046. 10.1088/1741-2552/ab3a0a (2019) [DOI] [PubMed]
- 21.Tosun, M. Effects of spectral features of eeg signals recorded with different channels and recording statuses on adhd classification with deep learning. Physical and Engineering Sciences in Medicine44(3), 693–702, 10.1007/s13246-021-01018-x (2021). [DOI] [PubMed] [Google Scholar]
- 22.Chen, H., Song, Y. & Li, X. A deep learning framework for identifying children with adhd using an eeg-based brain network. Neurocomputing356, 83–96, 10.1016/j.neucom.2019.04.058 (2019). [Google Scholar]
- 23.Müller, A. et al. Eeg/erp-based biomarker/neuroalgorithms in adults with adhd: Development, reliability, and application in clinical practice. The World Journal of Biological Psychiatry10.1080/15622975.2019.1605198 (2020). [DOI] [PubMed]
- 24.Ahmadi, A., Kashefi, M., Shahrokhi, H. & Nazari, M. A. Computer aided diagnosis system using deep convolutional neural networks for adhd subtypes. Biomedical Signal Processing and Control63, 102227, 10.1016/j.bspc.2020.102227 (2021). [Google Scholar]
- 25.Kim, S. et al. Machine-learning-based diagnosis of drug-naive adult patients with attention-deficit hyperactivity disorder using mismatch negativity. Translational Psychiatry11(1), 484, 10.1038/s41398-021-01604-3 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Cura, O. K., Akan, A. & Atli, S. K. Detection of attention deficit hyperactivity disorder based on eeg feature maps and deep learning. Biocybernetics and Biomedical Engineering44(3), 450–460, 10.1016/j.bbe.2024.07.003 (2024). [Google Scholar]
- 27.Wang, T.-S., Wang, S.-S., Wang, C.-L. & Wong, S.-B. Theta/beta ratio in eeg correlated with attentional capacity assessed by conners continuous performance test in children with adhd. Frontiers in Psychiatry14, 1305397, 10.3389/fpsyt.2023.1305397 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Vojnits, B. et al. Mobile sleep eeg suggests delayed brain maturation in adolescents with adhd: A focus on oscillatory spindle frequency. Research in Developmental Disabilities146, 104693, 10.1016/j.ridd.2024.104693 (2024). [DOI] [PubMed] [Google Scholar]
- 29.Loh, H. W. et al. Automated detection of adhd: Current trends and future perspective. Computers in Biology and Medicine146, 105525, 10.1016/j.compbiomed.2022.105525 (2022). [DOI] [PubMed] [Google Scholar]
- 30.Teruel, M. A. et al. Measuring Attention of ADHD Patients by means of a Computer Game featuring Biometrical Data Gathering. Heliyon10(5), e26555, 10.1016/j.heliyon.2024.e26555 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Thurstone, L. L., Granizo, M. Y., Luque, T., CARAS-R: test de percepción de diferencias: revisado, TEA, (2012).
- 32.Brickenkamp, R., Schmidt-Atzert, L., Liepmann, D., Test d2-Revision: Aufmerksamkeits-und Konzentrationstest, 1, Hogrefe Göttingen (2010).
- 33.Castilla, N., Higuera-Trujillo, J. L., Llinares, C., Virtual reality-based study assessing the impact of lighting on attention in university classrooms. Journal of Building Engineering 108902 (2024).
- 34.Sánchez-Sánchez, F., Santamaría, P., Abad, F. J., MATRICERS: Test de Inteligencia General, TEA, (2015).
- 35.Panagiotidis-Arrizabalaga, A. et al. Comparing Desktop, Virtual and Augmented Reality Gaming Environments for ADHD Attention Measurement. International Journal of Human–Computer Interaction 1-21, 10.1080/10447318.2025.2581259 (2025).
- 36.Siberski, J. et al. Computer-Based Cognitive Training for Individuals With Intellectual and Developmental Disabilities. American Journal of Alzheimer’s Disease & Other Dementias30(1), 41–48, 10.1177/1533317514539376 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Trujillo, J. et al. A Multimodal Dataset for Neurophysiological and AI Applications, figshare. 10.6084/m9.figshare.28676042 (2025). [DOI] [PMC free article] [PubMed]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The dataset supporting the findings of this study is publicly available at Figshare37 (10.6084/m9.figshare.28676042). The repository contains the full multimodal dataset in JSON and CSV formats organized by participant and session, including EEG (CGX, Emotiv), eye-tracking, EDA, heart-rate, game event logs (tags and metadata) and participant demographics. All data shared are anonymized. Further details about file structure and variable names are provided in the repository and in the manuscript Data Records section.
No custom code was generated for this dataset, but specialized software tools were used. EEG signals were captured using the Emotiv API (via the Emotiv SDK) and the CGX Acquisition Software, while games were developed in Unity. In addition, physiological measurements of the embraceplus wristbands were recorded using the Carelab Platform by Empatica. The following section provides a concise overview of each tool and its role in the study. No custom code was generated for signal preprocessing or analysis. The only custom software developed was a Python Flask backend API for coordinating multi-device data acquisition and file transfer, which does not modify recorded signals. • Emotiv API (Emotiv SDK / EmotivPRO Platform): The Emotiv API, provided as part of the Emotiv SDK (associated with the EmotivPRO platform), was used to acquire and process EEG data during the Robots and Cognifit sessions. • CGX Acquisition Software: CGX Acquisition Software from CGX Systems is a high-end EEG recording and analysis platform designed for both research and clinical applications. Provides robust real-time data acquisition, advanced signal processing, and synchronized event marking to ensure high-fidelity neurophysiological data.• Unity: Is a widely used game development engine that serves as the basis for designing and deploying the games. Its robust real-time rendering and interactive capabilities allowed the creation of immersive environments and precise control over stimulus presentation and data integration. • Empatica Care: A cloud-based platform developed by Empatica for continuous remote monitoring of the EmbracePlus wristband. The platform integrates data such as pulse rate, electrodermal activity, and temperature, enabling real-time tracking and analytics of key biomarkers.






