Skip to main content
Digital Biomarkers logoLink to Digital Biomarkers
. 2026 Jul 21;10(1):178–190. doi: 10.1159/000553327

Consensus-Based Definitions for Vocal Biomarkers: The International VOCAL Initiative

Mégane Pizzimenti a, Ayush Kalia b,c, Jamie A Toghranegar b, Mohamed Ebraheem b, Nicholas Cummins d, Satrajit S Ghosh e, James T Anibal f,g, Rhoda Au h,i,j,k,l,m, Arian Azarang n, Ruth H Bahr o, Sybille Barvaux a, Steven D Bedrick p, Hugo Botha q, Oita C Coleman r, Abir Elbeji a, Lampros C Kourtis s, Anaïs Rameau t, Jaskanwal Deep Singh Sara u,v, Stephanie W Watts b, Daria Hemmerling w, Jiri Mekyska x,y, Marisha L Speights z, Jean-Christophe Bélisle-Pipon A, Yael E Bensoussan b, Guy Fagherazzi a,✉,; on behalf of the Bridge2AI Voice Consortium; the eVoiceNet COST Action (CA24128)
PMCID: PMC13585421  PMID: 42757077

Abstract

Introduction

Voice-based health technologies are growing rapidly, but they lack standardized terminology, which hinders interdisciplinary collaboration, research quality, and clinical translation. The objective of this work was to develop universally accepted definitions in the rapidly evolving field of vocal biomarkers, as part of the VOCAL (Vocal Biomarker Guidelines for Ontology, Classification, Application, and Logistics) initiative, a structured, international consensus-based framework that aims to provide standards and guidelines.

Methods

VOCAL is a rigorous, international, multistage consensus-building study conducted in 2024–2025. It is a multi-institutional collaboration between representatives from the Bridge2AI-Voice Consortium (North America) and the eVoiceNet Network (European Union), involving a group of 24 international experts in medicine, clinical research, speech and language, audio signal processing, statistics, methodology, regulation, and ethics. VOCAL’s iterative process involved 5 rounds of review and feedback and an in-person workshop at the 2025 Bridge2AI Voice Symposium.

Results

Consensus-based definitions for vocal biomarkers were developed, spanning from broad concepts to domain-specific measures. A hierarchical continuum model of vocal biomarkers was established. We first distinguished between the concepts of vocal measures and vocal biomarkers. We then defined terms from broad, overarching concepts (level 0: biomarker, digital biomarker, vocal biomarker) to more specific physiological and cognitive domains (level 1: cardio-respiratory acoustic; level 2: voice; level 3: speech/articulatory; level 4: cognitive/language, including linguistic and paralinguistic subtypes).

Conclusion

This work provides a shared vocabulary that is essential for fostering communication through interdisciplinary collaboration, improving the quality and efficiency of research and development, and ensuring the ethical, reliable, and scalable deployment of future voice-based health technologies. It lays foundational groundwork for upcoming guidelines and standards, which are crucial for advancing the field of vocal biomarkers into widespread clinical utility.

Keywords: Vocal biomarker, Digital biomarker, Digital health, Artificial intelligence, Signal processing

Introduction

In recent years, the concept of vocal biomarkers has garnered attention for its potential in detecting, monitoring, and managing health conditions through the analysis of voice-based technologies [1, 2]. Vocal biomarkers fall under the broader category of digital biomarkers, which are objective and quantifiable health indicators derived from data collected via digital devices such as wearables and sensors. These noninvasive digital biomarkers promise to transform clinical trials and healthcare delivery by enabling continuous, objective, and patient-centered data collection [3]. Because “digital biomarker” remains an evolving term without a standardized international definition, as highlighted by Macías Alonso et al. [4], the field lacks a common conceptual foundation. This ambiguity is particularly pronounced for vocal biomarkers, which sit at the intersection of multiple complex physiological systems, further motivating the specialized framework proposed in this manuscript. As work in this field evolves, multiple terms such as “voice biomarkers,” “vocal biomarkers,” “speech biomarkers,” “voice AI” among others are used sometimes interchangeably, without a clear unified understanding of their definitions. As the field is multidisciplinary at its core, involving clinicians, speech pathologists, acoustic experts and data scientists, finding a common framework for definitions becomes crucial.

Physiological Systems Contributing to Vocal Biomarkers

Speech and voice production depend on the coordinated interaction of several physiological and cognitive subsystems. The respiratory system generates the airflow and aerodynamic energy required to initiate vocalization. The phonatory system then converts this airflow into acoustic energy through vibration of the vocal folds, producing the primary voice source. The articulatory system subsequently shapes this sound via movements of the tongue, lips, jaw, and velum, transforming the raw signal into intelligible speech. Beyond these peripheral mechanisms, higher-level cognitive and linguistic processes guide speech planning, lexical selection, syntactic structuring, and the modulation of prosody and communicative intent. These subsystems are captured in established theoretical frameworks, including the source-filter model [5], which links respiration, phonation, and vocal tract filtering, as well as psycholinguistic models [6] describing stages from conceptualization to formulation and articulation. Neurocomputational models such as DIVA/GODIVA further integrate motor control and cognitive-linguistic planning within a unified architecture of speech production [7]. Together, these models provide a scientific foundation for interpreting vocal biomarkers as features emerging from distinct yet interconnected physiological and cognitive domains. This perspective directly motivates the hierarchical continuum introduced below, spanning levels 1–4.

Once validated, vocal biomarkers will be unique among digital biomarkers in that they derive from the interplay of multiple physiological systems (respiratory, phonatory, articulatory, and cognitive) captured through audio. Their integration into healthcare reflects a broader movement toward real-time and remote health sensing [8, 9].

Voice and Speech Affecting Conditions

Because these subsystems contribute differently to speech and voice production, disruptions at specific physiological or cognitive levels can give rise to characteristic vocal signatures. As a result, vocal biomarkers have demonstrated potential across a wide range of health conditions and clinical outcomes, depending on which subsystem, or combination of subsystems, is affected.

In neurodegenerative and neurological disorders, speech and language impairments are common and may precede overt motor symptoms; for example, in Parkinson’s disease, features such as reduced loudness, monotone speech, imprecise articulation, and altered rhythm have shown promise for early detection and disease monitoring [10]. In Alzheimer’s disease, lexical-semantic and acoustic measures demonstrate high diagnostic performance for mild cognitive impairment, correlate with amyloid-β status and hippocampal volume, and predict disease progression, supporting their use in early detection and longitudinal monitoring [11].

Vocal biomarkers have also been explored in psychiatric and affective conditions, where individuals with depression often exhibit slower speech, reduced intensity, and flattened prosody that correlate with symptom severity [12]. Beyond disorders of the central nervous system, voice-based measures have shown potential in cardiopulmonary conditions such as chronic heart failure, offering a noninvasive and cost-effective approach for diagnosis, risk stratification, and telemonitoring [13]. Vocal biomarkers have further been applied to lifestyle and metabolic contexts, including the identification of smoking status from ecological audio recordings [14] and the screening or monitoring of type 2 diabetes [15]. Together, these examples illustrate the breadth of potential clinical applications of vocal biomarkers rather than providing an exhaustive review of the field.

Barriers to Clinical Implementation

Despite this rapid growth, a fundamental barrier to clinical implementation is the common assumption that any digitally captured voice metric constitutes a vocal biomarker, regardless of its validation status. To bridge this gap, we must draw a parallel with traditional biomarkers (e.g., imaging measures for Alzheimer’s disease), which achieved clinical gold-standard status only through extensive, multi-year validation. Currently, the path to validated digital biomarkers is poorly defined as guidelines from authorities such as the FDA or EMA remain in their nascent stages.

This regulatory and scientific vacuum is further exacerbated by a persistent challenge: the absence of universally accepted definitions for terms like “voice biomarkers,” “vocal biomarkers,” and “speech biomarkers,” among others. These terms are frequently used interchangeably in the literature [1620], yet they reflect features originating from distinct physiological and cognitive systems. A review of existing research reveals wide variance in definitions, often leading to nearly as many interpretations as there are articles on the topic. Such terminological ambiguity presents a significant barrier to scientific progress. It hampers collaboration across disciplines, complicates data sharing, and impedes the development of standardized methodologies and best practices [21]. Without a common language, the field risks fragmentation and inefficiency, ultimately undermining scientific advancement, regulatory assessment, clinical translation, and trustworthiness [22, 23]. To clarify the terminology used in this manuscript, we distinguish several commonly conflated concepts. The acoustic signal refers to the raw audio. A feature is any measurable descriptor extracted from that signal, and a measure is its quantified instance within a specific protocol. A biomarker refers to a feature (or combination of features) that has been shown to be associated with a biological, physiological, or clinical state within a defined context of use.

This challenge is further reinforced by interrelated risks that threaten the field’s coherence and scalability. The widespread use of AI, proprietary algorithms, and closed datasets has created a “black box” problem, where limited transparency restricts reproducibility and erodes trust among researchers, clinicians, and regulators. In parallel, inconsistent terminologies and disciplinary silos exacerbate confusion, highlighting the need for a shared, evolving vocabulary. The technical landscape adds further complexity as voice data are often massive, heterogeneous, and collected under diverse conditions, making integration across studies difficult without adherence to FAIR (Findable, Accessible, Interoperable, Reusable) principles [24]. These principles aim to ensure that data and associated metadata are easily located (findable), openly available under appropriate conditions (accessible), structured to allow combination and reuse across systems (interoperable), and documented to support reproducibility and secondary analyses (reusable). Applying FAIR to digital vocal biomarkers facilitates data sharing, reproducibility, and collaborative research, which are critical for developing reliable and scalable clinical applications.

Voice data-specific regulatory frameworks remain fragmented and difficult to navigate, particularly across regions, posing significant hurdles to clinical validation and implementation [24]. Additionally, concerns about privacy, data ownership, and ethical use of sensitive voice data persist, alongside practical challenges such as ensuring sustained patient engagement in digital trials. A recent survey of stakeholders in voice AI development emphasized the need for ethically sourced, diverse, and transparent data practices, calling for trustworthiness not just as a technical feature but as a guiding norm [25].

The lack of consensus also impairs the development of methodological infrastructure. While master protocols are now standard in clinical trials and other biomarker fields, vocal biomarker research lacks unified frameworks, resulting in heterogeneous data collection, feature extraction, and analysis methods [4]. This absence complicates validation and clinical scalability. These converging barriers underscore the need for community-driven standards, transparent methodologies, and harmonized regulatory pathways to unlock the full potential of vocal biomarkers. Establishing a unified framework for vocal biomarkers is therefore essential. Standardization will provide a shared vocabulary that facilitates interdisciplinary collaboration, supports data interoperability, and enables reproducibility of research findings. This foundation is critical for advancing vocal biomarker research from exploratory studies to reliable clinical applications.

Objective of This Work

To address this challenge, this paper proposes a set of foundational definitions as the first step of the VOCAL (Vocal Biomarker Guidelines for Ontology, Classification, Application, and Logistics) initiative, a structured, international framework for harmonizing practices and guiding the field of vocal biomarkers. The aim was to develop a shared and controlled vocabulary that promotes clarity and consistency across disciplines, facilitates communication among stakeholders and the public, and guides the development of best practices for the field.

Methods

Consensus-Building Process

To develop clear and widely applicable definitions of vocal biomarkers, a structured, multistage consensus-building process was employed. This approach was meticulously designed to incorporate diverse perspectives from interdisciplinary stakeholders, ensuring transparency and rigor throughout the development of the framework. The iterative nature of the process enabled continuous refinement of the proposed definitions, resulting in a robust, consensus-based outcome.

Participants

The consensus process involved 3 distinct groups of participants, ensuring broad representation and a comprehensive range of expertise:

  • Internal team (n = 7): this group comprised core members directly involved in the initial drafting of definitions and the coordination of the entire consensus process. They were leaders (Y.B., G.F.) and representatives of 2 large international consortia on vocal biomarkers: the Bridge2AI Voice Consortium and the eVoiceNet network. Their expertise spanned critical areas, including laryngology, speech-language pathology, digital epidemiology, AI, research methodology, and bioethics. This diverse composition ensured a strong foundational understanding of both clinical and technical aspects of vocal biomarkers.

  • Selected group of experts (n = 5): this group consisted of international academic researchers with highly specialized expertise in speech analysis, clinical research, and the development and commercialization of vocal biomarkers. Their focused input provided in-depth technical and methodological critique, refining the precision and scientific accuracy of the definitions.

  • Broader expert group (n = 18): this extended panel comprised a diverse array of researchers, clinicians, voice AI researchers, and stakeholders from a wide range of relevant disciplines (Clinical and Medical Sciences, Speech and Language Sciences, Engineering and Technical Fields, Statistics and Methodology, Regulation, Ethics, and Translation). This group provided a comprehensive review of the definitions from various practical and theoretical standpoints, ensuring their applicability across the broader scientific and clinical landscape.

Iterative Process

The process was organized into 5 iterative rounds of review and feedback, systematically designed to progressively refine the definitions and achieve strong consensus among all participants (shown in Fig. 1; online suppl. Table 1; for all online suppl. material, see https://doi.org/10.1159/000553327).

  • Initial draft and first round of review: the internal team initiated the process by drafting preliminary definitions of vocal biomarkers. A comprehensive review informed this of existing literature and terminology. Following this, team members individually reviewed and provided comments on the draft. A subsequent group meeting was held to discuss and consolidate all feedback, leading to the first set of revisions (revision #1).

  • Second round of review: the revised definitions (revision #1) were then shared with both the internal team and the selected group of experts. Participants provided written feedback, which was then discussed in an online workshop. This workshop focused on identifying points of agreement and divergence, and the collective input led to a second set of revisions (revision #2).

  • Third round of review: the updated definitions (revision #2) were again reviewed by the internal team and the selected experts. A second online workshop was convened to refine terminology further and address any remaining ambiguities or concerns. This meticulous discussion resulted in a third revision (revision #3), which brought the definitions closer to a final consensus.

  • Fourth round of review: the extensively revised definitions (revision #3) were presented to the broader expert group for review and finalization. An in-person workshop was conducted at the Voice AI Conference in Tampa, April 2025, during which participants formally voted by a show of hands on their agreement or disagreement with each presented definition. Any definition receiving disagreement from 25% or more of the participants was subjected to further discussion and revision. This iterative refinement continued until a consensus was achieved for all definitions.

  • Fifth round of review and finalization: feedback from the workshop was further integrated into a final set of definitions developed by the internal team. Following this final round, all co-authors reviewed and approved the definitions and the overarching framework presented in this paper.

Fig. 1.

The process included an initial internal team draft and first review, two rounds of review with selected experts, a consensus workshop with a broader expert group, and a final review and approval. The number of participants at each stage is indicated (n).

Flowchart of the iterative process to establish the definitions of key terms of the VOCAL initiative.

This rigorous process ensured that all concerns were thoroughly addressed through collective discussion before the definitions were finalized.

Results

To capture the layered complexity of vocal biomarkers while maintaining clarity and usability, the definitions were structured across different levels of granularity, forming a continuum model. This hierarchical approach enables both a broad conceptual understanding and precise categorization based on the underlying physiological and cognitive systems involved in voice production.

  • Level 0 represents broad, overarching terms in the biomarker domain. These foundational definitions are intended to frame the scope of discussion and provide shared reference points across diverse disciplines and technologies.

  • Levels 1 to 4 provide more specific and structured definitions of terms that fall under the umbrella of vocal biomarkers. These are categorized based on the stage or domain of vocal production and analysis, ranging from the earliest physiological processes to higher-level cognitive functions (shown in Fig. 2, 3).

Fig. 2.

The pyramid illustrates how vocal biomarkers relate to underlying physiological and cognitive subsystems. The layered design reflects conceptual relationships among subsystems rather than quantitative or hierarchical differences. The structure represents the increasing integration of processes involved in generating vocal signals, from foundational cardio-respiratory acoustics (Level 1) and phonatory activity (Level 2), through articulatory modulation (Level 3), to cognitive–linguistic contributions (Level 4). Biomarkers may span multiple levels depending on context. Corresponding anatomical regions for each subsystem are shown on the left.

Hierarchical continuum of vocal biomarkers: from voice production to high-level cognition.

Fig. 3.

Vocal biomarkers can be categorized across four levels based on their physiological and cognitive origins. Level 0: Biomarker, Digital Biomarker, Vocal biomarker represent broad, overarching terms. Level 1: Cardio-Respiratory Acoustic Biomarkers (CRA) relate to acoustic signals associated with respiration and cardiac function (e.g., breathing sounds, cough acoustics). Level 2: Voice Biomarkers pertain to features of phonation and vocal fold vibration (e.g., harmonic-to-noise ratio, cepstral peak prominence, fundamental frequency). Level 3: Speech/Articulatory Biomarkers involve articulatory processes and speech motor control (e.g., articulation rate, speech rate, articulatory precision, speech prosody and rhythm). Level 4: Cognitive/Language Biomarkers reflect the content of speech, including lexical access, word pairing, syntax, and paralinguistic elements such as emotion, cultural cues, and dialects (e.g., vocabulary, syntax). Created in https://BioRender.com

Vocal biomarkers levels 0–4.

It is crucial to understand that these categories are not mutually exclusive and often overlap, reflecting the interconnected nature of human voice production. While all stages of vocal production are ultimately under neural control (from motor execution to high-level language planning), our classification focuses on the primary physiological subsystem manifesting the signal. For clarity, we use vocal as an umbrella term covering all phenomena of human voice production, including voice-acoustic and speech-related measures. In this manuscript, voice is used in this inclusive sense, and speech-based markers are treated as a subset of vocal markers. Furthermore, extracted features can be interpreted as different types of biomarkers, depending on the specific context in which they are applied. For instance, fundamental frequency (F0) may serve as a phonatory marker of dysphonia in vocal fold paralysis (level 2), whereas in psychiatric contexts, the same feature may reflect paralinguistic or emotional states (level 4).

Level 0 Definitions

Table 1 establishes the foundational terminology for the entire framework, ensuring that readers from diverse backgrounds, including clinicians, engineers, researchers, and regulators, have a common understanding of “biomarker,” “digital Biomarker,” and “vocal biomarker” before delving into more specific categories (shown in online supp Fig. S1). This clarity is paramount for interdisciplinary discourse and regulatory alignment.

Table 1.

Level 0 definitions

Term Definition Explanatory notes/examples Source
Biomarker A measurable biological substance or characteristic whose detection or evaluation provides information about the function or state of a system or organ in the body. It can be used to understand normal or altered biological processes and may serve roles in risk prediction, screening, diagnosis, monitoring, treatment response measurement, or prognosis There are various types of biomarkers, including physiological, serological, imaging, and genetic ones. For instance, HbA1c serves as a blood biomarker for diabetes. Another example of a biomarker is the presence of the BRCA gene, which indicates a genetic predisposition to certain types of breast cancer. In line with regulatory definitions, a biomarker is formally recognized only after undergoing appropriate validation (analytical and clinical). However, within the research and innovation context, especially in emerging areas such as voice-based health technologies, the term “biomarker” is commonly used pre-validation to denote a candidate indicator under investigation FDA-NIH BEST Resource (Biomarkers, EndpointS, and other Tools Resource)
Digital biomarker A biomarker derived from data collected using a digital device such as a wearable, portable, or implantable device For example, a heart rate measured from a smartwatch is a digital biomarker, and an ECG measured by a wearable is a digital biomarker. A digital biomarker is a digital measure that can be used as a digital endpoint in a clinical study. However, not all digital measures are digital biomarkers See Macias Alonso et al. (2024) for an overview of current definitions
Digital measures are simply quantified outputs derived from digital data sources; only those that have demonstrated an association with a biological, physiological, or clinical state within a defined context of us qualify as digital biomarkers
Vocal biomarker (overarching term, includes voice, speech, linguistic, paralinguistic, and prosody) A vocal biomarker is an overarching term for defining biomarkers derived from audio recordings of the voice, speech, linguistic, and paralinguistic systems. These are collected through acoustic data recording and can be analyzed using traditional acoustic analysis or digital signal processing and AI methods. This overarching term includes biomarkers that are impacted by the process of creating speech from our lungs and cardio-respiratory system to our larynx, resonators, and brain, which helps with lexical access and influences the content of our speech It is important to understand that these should not represent mutually exclusive categories as they can overlap and should be interpreted more as a continuum. Moreover, extracted features can be interpreted as different types of biomarkers, depending on the context Proposed by the VOCAL consensus
Example: F0 can be used as a voice biomarker in dysphonia or voice disorders and can also be considered in mental health screening for depression

Level 1–4 Definitions

Table 2 breaks down the complex domain of vocal biomarkers into actionable, physiologically grounded categories. These levels are organized according to the primary physiological subsystem involved in the production of the signal.

Table 2.

Proposed consensus-based definitions for vocal biomarker levels 1–4 (VOCAL initiative)

Term Definition Examples Illustrations
CRA biomarkers (level 1) CRA biomarkers relate to acoustic signals pertaining to respiration and cardiac function. They represent the earliest stage in the speech production pathway and reflect physiological processes that influence vocalization Breathing sounds, cough acoustics, and pauses caused by breath support A cough with low intensity may reflect reduced respiratory drive or weakened supportive muscles. Acoustic analysis of coughs or breathing may indicate respiratory illness or muscular disorders. Frequent pauses in speech may reflect dyspnea or reduced breath control
Voice biomarkers (level 2) Voice biomarkers relate to features from the source of phonation: the lungs pushing air through the vocal folds, which vibrate to generate sound, later shaped by the vocal tract (throat, mouth, nose) HNR, CPP, fundamental frequency (F0) In laryngitis, vocal fold inflammation can cause a rough or hoarse voice quality, measurable using CPP. Parkinson’s disease can affect prosody, leading to monotonous speech with reduced pitch variability (F0 range) and inflection
Speech/articulatory biomarkers (level 3) Speech/articulatory biomarkers relate to articulatory processes and speech motor control. They reflect the coordination and functioning of articulators (tongue, lips, soft palate, etc.), which are regulated by neurological control Articulation rate, speech rate, syllable precision, consonant clarity In ALS, tongue weakness may impair the articulation of syllables like “Ka” or “Ga.” Variations in rhythm and timing can indicate altered motor coordination or speech effort
Cognitive/language biomarkers (level 4) Cognitive/language biomarkers relate to the content and meaning of speech, reflecting higher-level cognitive functions, language processing, and emotional or social communication cues Linguistic biomarkers refer to the content of our speech, our lexical access, the words we choose in a sentence, and how we pair words together to form sentences. They can also include syntax, grammar, semantics, and vocabulary concepts. Paralinguistic biomarkers include non-verbal communication elements, such as conveying emotion, cultural cues, and dialects Linguistic illustration: reduced vocabulary diversity in Alzheimer’s disease or syntactic simplification in aphasia; paralinguistic illustration: intonation patterns signaling questions or emotional states (e.g., rising pitch) may reflect cognitive or affective modulation in conditions like autism spectrum disorder

Level 1: Cardio-Respiratory Acoustic

This level focuses on the physiological “power supply” of the voice. It encompasses measures related to respiratory support and breath control, while also considering the cardiac system’s indirect influence on vocal stability through hemodynamic effects on vocal fold hydration and the modulation of respiratory rhythms.

Level 2: Voice/Phonatory

These definitions target the laryngeal source, focusing on the acoustic properties of phonation. These phonatory or laryngeal measures describe the sound generated by vocal fold vibration (e.g., jitter, shimmer) before it is modified by the vocal tract.

Level 3: Speech/Articulatory

This category involves supraglottic and articulatory movements, including the tongue, lips, and jaw, that transform the laryngeal source into identifiable phonemes and speech patterns.

Level 4: Cognitive/Language

This final level addresses the high-level neural processes governing communication. It includes lexical access and speech content, as well as paralinguistic and emotional subtypes that reflect a speaker’s cognitive and psychological state. By providing clear definitions, examples, and illustrations, it serves as a practical guide for researchers and clinicians to accurately classify and interpret vocal features, thereby reducing ambiguity and facilitating standardized research design and reporting.

Discussion

This study presents the first stage of the VOCAL (Vocal Biomarker Guidelines for Ontology, Classification, Application, and Logistics) initiative, offering foundational definitions to address terminological fragmentation in the field of vocal biomarkers. By introducing a layered, physiologically grounded vocabulary organized across 5 hierarchical levels, the framework establishes a shared reference system that facilitates interdisciplinary collaboration, supports regulatory alignment, and promotes scientific rigor. In doing so, it fills a critical gap in the biomedical and digital health landscape, where vocal biomarkers are increasingly positioned as scalable, noninvasive tools for early detection and monitoring of health conditions.

Rethinking Voice Data and Analytics through a Unified Lexicon

This study’s central contribution is the reframing of “vocal biomarkers” not as a monolithic or interchangeable term but as a continuum of overlapping physiological and cognitive signals. The framework avoids rigid boundaries, opting instead for a layered structure that mirrors the interconnected nature of voice production, encompassing breath and muscle control as well as language and emotion. This continuum model provides researchers and clinicians with a shared yet flexible vocabulary that accommodates scientific precision and real-world variability. Although the levels may overlap depending on context and speech tasks, this does not introduce ambiguity. Rather, it requires explicit documentation of context, task, and feature annotation, thereby clarifying the physiological or cognitive origin of each signal.

What makes this framework novel is its ability to capture complexity without collapsing it. The five-level structure, from level 0’s foundational definitions to Level 4’s cognitive and language biomarkers, differentiates vocal signals not only by signal type but also by the physiological or cognitive systems they reflect. Level 1 captures cardio-respiratory acoustic signals (breath, cough), level 2 captures laryngeal voice-source features (pitch, harmonicity), level 3 focuses on articulatory modulation by the vocal tract, and level 4 encompasses cognitive and language-related control (lexical choice, prosody, social communication). This hierarchy mirrors the flow from source to modulation to message, providing both biological and interpretive justification for the levels. Anchoring the hierarchy in established subsystems of voice production (respiration, phonation, articulation, cognitive-linguistic control) affords direct mechanistic attribution of features, which strengthens interpretability and facilitates clinical and regulatory evaluation.

The framework also contributes to addressing AI’s “black box” problem. While LLMs and other AI systems can extract complex features from voice data, their outputs often lack clinical interpretability. By mapping features to defined levels, the framework enables physiologically and cognitively meaningful interpretation, even when features originate from automated models. LLMs thus serve as extraction and annotation tools within a structure that maintains transparency and regulatory relevance.

Addressing Fragmentation and Supporting Clinical Uptake

This layered approach also acts as a conceptual correction to the field’s fragmentation. The multistage review process highlighted that the term “vocal biomarkers” is frequently used without explicit specification of the vocal subsystem or interpretive level, reflecting a key contributor to the field’s fragmentation. This imprecision limits interoperability, obscures reproducibility, and undermines clinical translation and regulatory review. The proposed definitions address this gap by creating clear referents across disciplines while allowing interpretive flexibility. Overlaps between levels become opportunities for explicit annotation rather than sources of confusion, improving transparency and comparability across studies.

Beyond classification, the framework may contribute to shaping emerging normative and regulatory implications. By standardizing language in a fast-evolving and commercially saturated space, it can enhance transparency and challenge the opacity of proprietary systems that often treat voice as an undifferentiated input. A common lexicon, co-produced by experts across laryngology, AI, neuroscience, ethics, and clinical research, fosters transparency and enables scrutiny of what is measured, how, and why. In this way, the framework functions not only as a taxonomic scaffold but also as a potential epistemic and ethical guide, helping realign the field around grounded and shared understandings of voice and health as vocal biomarkers move toward clinical integration.

Limitations

While this framework represents a significant step toward definitional coherence in vocal biomarker research, several limitations must be acknowledged. First, the process, although interdisciplinary and international, remains in its early stages and primarily involved researchers embedded in academic or consortium networks. The perspectives of patients, regulators, and public health decision-makers from underrepresented regions were absent, despite their relevance when voice data intersect with trust, access, and cultural specificity. Furthermore, although the VOCAL initiative brings together experts from across Europe and the USA, it does not yet capture the full global diversity of voice. The current framework is grounded largely in Western linguistic and clinical traditions. This represents a limitation, as vocal expression, paralinguistic norms, and culturally embedded interpretations of what constitutes a “disordered” voice can vary substantially across regions and languages. As a result, the applicability of the proposed definitions to non-Western languages and culturally diverse populations remains to be empirically examined.

Second, the framework is intentionally definitional and does not yet provide guidance on study design, feature extraction protocols, or validation metrics. While it clarifies what is being measured, it does not resolve how measurements should be standardized, interpreted, or integrated into clinical workflows. Questions regarding longitudinal stability, generalizability across languages and populations, and device dependence remain outside the scope of this initial effort. Third, the flexible continuum approach may present operational challenges. Nonetheless, the requirement for explicit contextual documentation enhances clarity and supports clinical and regulatory use.

Finally, the consensus process does not constitute empirical validation. Whether these definitions improve reproducibility, regulatory acceptance, or interdisciplinary collaboration remains to be demonstrated. The framework should therefore be viewed as provisional and evolving.

Contribution to Vocal Biomarkers Development and Clinical Uptake

The definitional framework proposed in this paper addresses a longstanding barrier in vocal biomarker research: the absence of precise and consistent terminology. By structuring vocal biomarkers along a physiologically grounded continuum, it provides a common language for describing and interpreting voice-based health data, laying the groundwork for adoption and clinical relevance.

Standardization acts as a catalyst for trust and uptake across the healthcare ecosystem. Clear definitions enable clinicians to interpret findings, encourage patient trust, and allow regulators to evaluate technologies more efficiently. This creates a reinforcing cycle in which clarity builds trust and trust supports adoption. By linking each feature to a defined level, the framework allows outputs from LLMs or other AI models to be interpreted in a clinically meaningful way, reducing the “black box” problem and enabling regulatory scrutiny. This supports reproducibility, ethical governance, and cross-study comparison. Shared terminology enhances communication for patients and clinicians, while for payers and regulators it reduces uncertainty and facilitates alignment with existing biomarker approval pathways.

Beyond building clinical and regulatory trust, the framework also brings important efficiency gains at a time when healthcare systems face growing resource constraints. A shared vocabulary reduces the administrative and technical effort needed for data integration and collaboration across institutions. In clinical practice, clearer definitions support quicker decisions and smoother triage, helping ease the pressure on specialized services. For researchers, improved interoperability limits redundant data collection and makes it easier to reuse existing datasets. Economically, these efficiencies lower the barriers to adopting vocal-AI tools and reduce long-term implementation costs, supporting the sustainable use of vocal biomarkers for large-scale population health monitoring.

Future Research

This framework offers a foundation rather than an endpoint. Future research should test the utility of the continuum model across use cases, assessing whether it improves consistency in feature labeling, interpretability, and reproducibility. Comparative studies may clarify its impact on transparency and cross-study compatibility. Methodological extensions are needed, including reporting standards aligned with the levels, protocols for labeling multi-level features, and guidance for ambiguous cases. Embedding the framework into shared methodological infrastructures would further support interoperability.

Cultural, linguistic, and ethical dimensions must also be considered. Vocal features carry social and contextual meaning beyond physiology, and research must examine how definitions perform across diverse populations while avoiding bias. Participatory approaches will be critical to ensure equitable design and deployment. To address the current geographical bias, future phases of the VOCAL initiative should explicitly involve stakeholders from the Global South. Establishing partnerships with research and clinical networks in Africa, Asia, and Latin America will be essential to validate and adapt the framework across a broader range of languages, paralinguistic norms, and phonetic structures. Targeted studies will be needed to understand how cultural specificities in communication styles may influence the mapping of articulatory (level 3) and cognitive-linguistic (level 4) biomarkers. Engaging with regulatory bodies beyond the EMA and FDA will likewise be important to ensure that emerging standards for vocal biomarkers are equitable and globally relevant.

Finally, regulatory alignment remains essential. Mapping the framework to existing evidentiary categories such as those defined in the FDA’s Biomarker Qualification Program and the EMA’s Qualification of Novel Methodologies, which distinguish analytical validation, clinical validation, and context-of-use evidence, can help establish vocal biomarkers as clinically valid and ethically deployable tools [26, 27].

Conclusion

This paper presents the first structured, consensus-based definitional framework for vocal biomarkers, addressing a critical need for clarity, consistency, and interoperability in a rapidly evolving field. Developed through a rigorous, multistage process involving international experts across disciplines, the framework offers a shared vocabulary reflecting the physiological and cognitive dimensions of vocal signal generation, modulation and emission. By anchoring vocal biomarkers in a layered, interpretable structure, this work supports interdisciplinary collaboration, ethical development, and clinical translation. It also provides a structured way to interpret AI-derived features, including those from LLMs, ensuring that clinical relevance is maintained even in automated analyses. It helps unify fragmented research efforts, strengthens reproducibility, and facilitates regulatory alignment. As voice increasingly emerges as a source of health-relevant data, a common language is indispensable. This framework offers a path toward trustworthy, patient-centered, and clinically impactful voice-based health technologies.

Acknowledgments

We thank Claire Bortolotto for her contributions during selected meetings of the first-round review. We also acknowledge the contributions of the members of the Bridge2AI Voice Consortium and the eVoiceNet COST Action (CA24128) who contributed to discussions and activities supporting this work. A full list of collaborators is provided in online supplementary material.

Statement of Ethics

This work consisted of an expert consensus-building process involving academic and professional collaborators and did not involve data collection. According to local and national guidelines, formal ethics approval and written informed consent were not required for this type of activity. Participation in the consensus process was voluntary.

Conflict of Interest Statement

The authors have no conflict of interest to declare.

Funding Sources

This article is based upon work from COST Action eVoiceNet (CA24128), supported by COST (European Cooperation in Science and Technology), and the Bridge2AI-Voice initiative, the Precision Public Health Grand Challenge of the Bridge2AI Program funded by the NIH Common Fund (Award No. OT2OD032720-01S1). Funding supported networking activities, consortium meetings, and collaborative activities related to this work. The authors were responsible for the study design, analysis, interpretation of results, manuscript preparation, and decision to publish.

Author Contributions

M.P., Y.E.B., and G.F. conceptualized and designed the consensus-building process and coordinated the study and led the drafting of the initial definitions. M.P., A.K., J.T., M.E., N.C., S.S.G., J.T.A., R.A., A.A., R.H.B., S.D.B., H.B., O.C., A.E., L.C.K., A.R., J.D.S., S.W.W., D.H., J.M., M.L.S., J.C.B.P., Y.E.B., and G.F. contributed to the methodological framework, contributed to the review, discussion, and refinement of the definitions, and revised the manuscript and approved the final version. Y.E.B. and G.F. supervised the process.

Funding Statement

This article is based upon work from COST Action eVoiceNet (CA24128), supported by COST (European Cooperation in Science and Technology), and the Bridge2AI-Voice initiative, the Precision Public Health Grand Challenge of the Bridge2AI Program funded by the NIH Common Fund (Award No. OT2OD032720-01S1). Funding supported networking activities, consortium meetings, and collaborative activities related to this work. The authors were responsible for the study design, analysis, interpretation of results, manuscript preparation, and decision to publish.

Data Availability Statement

All data generated or analyzed during this study are included in this published article and its online supplementary information files. No additional raw data were generated or are available due to the nature of the consensus process involving expert participants. Further inquiries can be directed to the corresponding author.

Supplementary Material.

Supplementary Material.

References

  • 1. Fagherazzi G, Fischer A, Ismael M, Despotovic V. Voice for health: the use of vocal biomarkers from research to clinical practice. Digit Biomark. 2021;5(1):78–88. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2. Bensoussan Y, Elemento O, Rameau A. Voice as an AI biomarker of Health- introducing audiomics. JAMA Otolaryngol Neck Surg. 2024;150(4):283–4. [DOI] [PubMed] [Google Scholar]
  • 3. Fagherazzi G, Bensoussan Y. The imperative of voice data collection in clinical trials. Digit Biomark. 2024;8(1):207–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4. Macias Alonso AK, Hirt J, Woelfle T, Janiaud P, Hemkens LG. Definitions of digital biomarkers: a systematic mapping of the biomedical literature. BMJ Health Care Inform. 2024;31(1):e100914. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5. Tokuda I. The source-filter theory of speech. In: Aronoff M, editor. Oxford research encyclopedia of linguistics: Oxford University Press; 2021. [Google Scholar]
  • 6. Baker E, Croot K, McLeod S, Paul R. Psycholinguistic models of speech development and their application to clinical practice. J Speech Lang Hear Res. 2001;44(3):685–702. [DOI] [PubMed] [Google Scholar]
  • 7. Miller HE, Guenther FH. Modelling speech motor programming and apraxia of speech in the DIVA/GODIVA neurocomputational framework. Aphasiology. 2021;35(4):424–41. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8. Kalia A, Boyer M, Fagherazzi G, Bélisle-Pipon JC, Bensoussan Y. Master protocols in vocal biomarker development to reduce variability and advance clinical precision: a narrative review. Front Digit Health. 2025;7:1619183. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9. Patel R, Price N, Bahr R, Bedrick S, Bensoussan Y, Bélisle-Pipon JC, et al. Summary of keynote speeches from the 2024 voice AI symposium, presented by the Bridge2AI-Voice consortium. Front Digit Health. 2024;6:1484503. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10. Roland V, Huet K, Harmegnies B, Piccaluga M, Verhaegen C, Delvaux V. Vowel production: a potential speech biomarker for early detection of dysarthria in Parkinson’s disease. Front Psychol. 2023;14:1129830. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. Hajjar I, Okafor M, Choi JD, Moore E 2nd, Abrol A, Calhoun VD, et al. Development of digital voice biomarkers and associations with cognition, cerebrospinal biomarkers, and neural representation in early Alzheimer’s disease. Alzheimers Dement. 2023;15(1):e12393. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12. Mundt JC, Vogel AP, Feltner DE, Lenderking WR. Vocal acoustic biomarkers of depression severity and treatment response. Biol Psychiatry. 2012;72(7):580–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. Bauser M, Kraus F, Koehler F, Rak K, Pryss R, Weiß C, et al. Voice assessment and vocal biomarkers in heart failure: a systematic review. Circ Heart Fail. 2025;18(8):e012303. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Ayadi H, Elbéji A, Despotovic V, Fagherazzi G. Digital vocal biomarker of smoking status using ecological audio recordings: results from the colive voice study. Digit Biomark. 2024;8(1):159–70. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15. Elbeji A, Aguayo G, Fischer A, Fagherazzi G. Identification of vocal biomarkers for screening diabetes and monitoring health of people with diabetes: preliminary results from the COLIVE voice study. Diabetes Technol Ther. 2022;24(Suppl 1):A224. [Google Scholar]
  • 16. Watase T, Omiya Y, Tokuno S. Severity classification using dynamic time warping-based voice biomarkers for patients with COVID-19: feasibility cross-sectional study. JMIR Biomed Eng. 2023;8:e50924. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. Lin YC, Yan HT, Lin CH, Chang HH. Identifying and estimating frailty phenotypes by vocal biomarkers: cross-sectional study. J Med Internet Res. 2024;26:e58466. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18. Gaur S, Kalani P, Mohan M. Harmonic-to-noise ratio as speech biomarker for fatigue: K-nearest neighbour machine learning algorithm. Med J Armed Forces India. 2024;80(Suppl 1):S120–6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. Sterner B. Explaining ambiguity in scientific language. Synthese. 2022;200(5):354. [Google Scholar]
  • 20. Livieri G, Mangina E, Protopapadakis ED, Panayiotou AG. The gaps and challenges in digital health technology use as perceived by patients: a scoping review and narrative meta-synthesis. Front Digit Health. 2025;7:1474956. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Key factors behind IT implementation failures in mid-sized healthcare settings - CapMinds. 2025. https://www.capminds.com/blog/key-factors-behind-it-implementation-failures-in-mid-sized-healthcare-settings/
  • 22. Babrak LM, Menetski J, Rebhan M, Nisato G, Zinggeler M, Brasier N, et al. Traditional and digital biomarkers: two worlds apart? Digit Biomark. 2019;3(2):92–102. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23. Schmidt J, Schutte NM, Buttigieg S, Novillo-Ortiz D, Sutherland E, Anderson M, et al. Mapping the regulatory landscape for artificial intelligence in health within the European Union. Npj Digit Med. 2024;7(1):229. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Polat G, Grzybowski A. Shaping the future of healthcare: ethical clinical challenges and pathways to trustworthy AI. J Clin Med. 2025;14(5):1605. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25. Bélisle-Pipon JC, Powell M, English R, Malo MF, Ravitsky V, Bridge2AI–Voice Consortium, et al. Stakeholder perspectives on ethical and trustworthy voice AI in health care. Digit Health. 2024;10:20552076241260407. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Jadoenathmisier KD, Gardarsdottir H, Mol PGM, Pasmooij AMG. Insights from the European Medicines Agency on digital health technology derived endpoints. Drug Discov Today. 2025;30(6):104388. [DOI] [PubMed] [Google Scholar]
  • 27. European Medicines Agency . Questions and answers: qualification of digital technology-based methodologies to support approval of medicinal products (EMA/219860/2020); 2020. [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Data Availability Statement

All data generated or analyzed during this study are included in this published article and its online supplementary information files. No additional raw data were generated or are available due to the nature of the consensus process involving expert participants. Further inquiries can be directed to the corresponding author.


Articles from Digital Biomarkers are provided here courtesy of Karger Publishers

RESOURCES