Abstract
Purpose
The purpose of this study was to develop best practices for conducting cognitive debriefing interviews with pediatric populations by drawing on the currently available literature and insights from experts in the health-related quality of life research community.
Methods
A scoping review of the literature was conducted to identify existing recommendations, considerations, and methods for conducting cognitive debriefing interviews with pediatric populations. Findings from the review informed the development of a draft set of best practices, which were subsequently reviewed and refined through a two‑round modified Delphi process. The Delphi panel was composed of experts (i.e., researchers) in patient-reported outcome instrument development and evaluation with experience conducting interviews with pediatric populations.
Results
Thirty-five articles or guidance documents were included in the final scoping review, contributing insights that were used to develop a draft set of 17 best practices. Following the two rounds of review by the Delphi panel, the final set of 17 best practices, which reflects panel consensus, addresses: developing the interview guide, evaluating the characteristics of the instrument to be debriefed, and interview conduct.
Conclusion
The best practices described in this article provide evidence‑based guidance that can help to standardize and strengthen the rigor of cognitive debriefing interviews with pediatric populations while potentially also improving the experience for participants. Broader implementation across drug development programs may enhance the reliability of measurement, improve the quality of pediatric participant input, and promote more patient‑centered decision‑making.
Keywords: Cognitive debriefing, Pediatric patients, Content validity, Best practice
Introduction
The United States (US) Food and Drug Administration’s (FDA) Patient-Focused Drug Development (PFDD) guidance encourages medical product developers to have patients self-report their symptoms and/or impacts on health-related quality of life (HRQoL) via patient-reported outcome (PRO) instruments when appropriate to do so [1–4]. When asking patients (pediatric and adult) to self-report on their symptoms and HRQoL, it is essential to use a PRO instrument that demonstrates content validity, i.e., one that is relevant to the patients’ lived experiences, is easy to understand, and is easy to complete [3]. This reflects a broader emphasis on patient experience data, which play a critical role not only in evaluating treatment benefit–risk profiles but also in informing clinical guidelines and shared decision-making [5].
Recommendations for developing a PRO instrument that demonstrates content validity include using two distinct qualitative research methodologies: concept elicitation and cognitive debriefing [1–3, 6, 7]. While PRO instrument developers can rely on numerous publications describing best practices for establishing the content validity of assessments to be administered in adult populations, guidance addressing these methods in pediatric populations (i.e., children ≤ 11 years and adolescents aged 12–17 years) is comparatively sparse. In 2010, Bevans et al. emphasized the importance of using cognitive interviews to assess children’s understanding of any instruments they are meant to complete [8]. In 2013, Arbuckle and Abetz-Webb and the ISPOR PRO Good Research Practices for the Assessment of Children and Adolescents Task Force reaffirmed this position in their recommendations for working with pediatric populations when developing PRO instruments and conducting research designed to inform regulatory decision-making and support label claims [9, 10].
Although the FDA PFDD guidance cites both Arbuckle and Abetz-Webb and ISPOR as useful resources for researchers working with pediatric populations, the guidelines do not specify how those recommendations should be implemented in regard to cognitive debriefing [3, 4].
While some aspects of the existing best practices for cognitive debriefing interviewing in health outcomes research among adults can be adapted successfully for use with pediatric populations [2, 3, 7], it remains essential to use developmentally appropriate methods when evaluating PRO instruments. Too often researchers rely on instruments and approaches that miss this important consideration. This includes dismissing the ability of young children, in particular, to provide valuable feedback. Fortunately, recent research from Gale et al., which focused on addressing some of the methodological challenges associated with engaging young children (6–7 years) in cognitive interviewing, provides new evidence to support the feasibility of this practice [11].
As the state of health measurement continues to evolve and we learn more about how patients experience different conditions and interact with the instruments intended to measure their lived experiences, our efforts to ensure that the methods we use to evaluate those instruments must also evolve. To help in that process of evolution, we designed and executed a scoping review and modified Delphi process to develop best practices for conducting cognitive debriefing interviews with pediatric populations, drawing on the most currently available literature and insights of experts from the health-economics and outcome research (HEOR) community.
Methods
The study design included a scoping review followed by a 2-round modified Delphi process (Fig. 1). The study team was composed of four researchers with expertise in qualitative methods and health-related quality of life measurement.
Fig. 1.

Overview of study design. A graphic representation of the study design which includes a scoping review, followed by draft best practices, two rounds of a modified Delphi, and revisions to finalize best practices
Scoping review
The goal of the scoping review was to examine the peer-reviewed and gray literature to identify existing recommendations, considerations, and methods used in conducting cognitive debriefing interviews with pediatric populations to inform a draft set of best practices for the Delphi panel to review. Database-specific search strings were developed, refined, and implemented in PubMed (using Medical Subject Heading (MeSH) terms when available) and Embase (Fig. 2). The methods and results were reported following the Preferred Reporting Items for Systematic reviews and Meta-Analyses extension for Scoping Reviews (PRISMA-ScR) Checklist.
Fig. 2.

Search strings used in the PubMed and Embase databases. PubMed and Embase search strings
Additional articles were identified through hand searching (i.e., scanning the reference lists of relevant articles to identify additional relevant studies) and desk searching (i.e., keyword searches in Google/Google Scholar). This approach allowed the study team to go beyond study-based articles and inquire into regulatory or professional guidance documents, such as those published by the US FDA, European Medical Association (EMA), and ISPOR.
Records review and data extraction
All records were reviewed in two rounds. Prior to reviewing the records, the study team established a set of inclusion criteria to guide the review and selection process (Fig. 3). The same criteria were used throughout the review process. Interrater reliability was not calculated as each record was screened and reviewed independently by a single reviewer. To promote consistency in record selection, reviewers agreed on inclusion criteria prior to reviewing any records and held regular team discussions to address and resolve any uncertainties.
Fig. 3.

Search result record inclusion criteria. List of inclusion criteria applied during the scoping review of the literature
Three members of our study team (LB, EB, MO’C) participated in the title/abstract screening, full text review, and data extraction process. Titles and abstracts were screened and those that did not meet the inclusion criteria were excluded (MS Excel). Those that met the inclusion criteria were reviewed in full and either included for data extraction or excluded as they did not meet the inclusion criteria.
During the screening and review process, a reason for exclusion—with respect to the inclusion criteria—was recorded for each article not selected for further review or data extraction. While some articles may have been excluded for multiple reasons, only one reason was recorded. A hierarchy of exclusion criteria was used to ensure excluded articles were coded in a standard approach.
Data extraction focused on identifying practices, suggestions, concerns, and solutions related to conducting cognitive debriefing interviews with children and/or adolescents. Data were extracted into a shared document (MS Word) and, using an inductive approach, we identified key themes (e.g., practices described by multiple articles or guidance documents). Those themes were used to draft an initial set of best practices for cognitive debriefing with pediatric populations.
Modified Delphi method
Once the best practices were drafted, we used a modified Delphi method to assess and confirm the relevance and importance of each practice and refine the language of each practice. The Delphi method uses a panel of experts to develop, over multiple rounds of input, an output that represents the consensus of the group [12, 13]. In the health sciences, the Delphi method can be used for various purposes, including instrument development and validation, identifying research priorities, setting clinical standards, and, in the case of this study, setting practice guidelines [14].
To recruit a target of 10 panelists, we invited 18 individuals with expertise conducting HRQoL research and/or PRO instrument development with pediatric populations to participate in the panel. “Expertise” was defined as having published or presented work on the topic. Potential panelists were identified using the research team’s professional network and by reviewing author lists from key publications and presentations related to cognitive debriefing and/or PRO instrument development with pediatric populations.
The modified Delphi consisted of two rounds (Round 1 [R1] and Round 2 [R2]) of asynchronous feedback. For each round, panelists were provided with a guidance document that included an overview of the study and its progress and the draft best practices along with a link to the online questionnaire. Panelists were asked to review the best practices and provide feedback (e.g., the extent to which they agreed with the proposed statement, the importance of the practice) through a series of open- and closed-ended questions. Panelists were given two weeks to complete each questionnaire. Following each round of data collection, we reviewed the panelists’ responses.
Prior to data collection, we defined consensus as agreement among ≥ 75% of the panelists. This threshold, which represents a majority of the panelists and aligns with other published studies on the Delphi method in the health sciences, was applied across both rounds of data collection [15]. Further, if the panel came to consensus on a practice’s description, it was considered final. If the panel did not reach consensus, we reviewed the questionnaire results, discussed suggested edits, and revised the practice’s description.
Results
Scoping review
Search results
Between the two databases, we identified a total of 494 abstracts. After removing duplicates, screening titles and abstracts, and reading full-text, ultimately 35 records were included in the scoping review (Fig. 4).
Fig. 4.

PRISMA diagram. Graphical representation of the process to identify and evaluate records for the scoping review
Practices identified in the literature
We identified several practices used by researchers according to the published and gray literature, which we grouped into two non-mutually exclusive categories: developmental considerations and methodological considerations (Table 1). We also identified how these considerations align with current regulatory guidance.
Table 1.
Articles with insights into the developmental and methodological considerations for conducting cognitive debriefing interviews with pediatric populations
| Article content | Citation |
|---|---|
| Developmental considerations | Aldhouse et al. [16] |
| Arbuckle and Abetz-Webb [10] | |
| Bevans et al. [8] | |
| Coombes et al. [17] | |
| Gabes et al. [18] | |
| Gale et al. [11] | |
| Halstead et al. [19] | |
| Hwang et al. [20] | |
| Jacobson et al. [21] | |
| Kamath et al. [22] | |
| Kramer and Schwartz [23] | |
| Romano et al. [24] | |
| Tomlinson et al. [25] | |
| Tucker et al. [26] | |
| Turner-Bowker et al. [27] | |
| Willis et al. [28] | |
| Zigler et al. [29] | |
| Methodological considerations | Aldhouse et al. [16] |
| Anthony et al. [30] | |
| Arbuckle and Abetz-Webb [10] | |
| Bevans et al. [8] | |
| Carlton [31] | |
| Cella et al. [32] | |
| Coombes et al. [17] | |
| Gale et al. [11] | |
| Halstead et al. [19] | |
| Hwang et al. [20] | |
| Irwin et al. [33] | |
| Jacobson et al. [21] | |
| Kamath et al. [22] | |
| Kramer and Schwartz [23] | |
| O’Sullivan et al. [34] | |
| Propp et al. [35] | |
| Rams et al. [36] | |
| Romano et al. [24] | |
| Sarda et al. [37] | |
| Taylor et al. [38] | |
| Tomlinson et al. [25] | |
| Truninger et al. [39] | |
| Tucker et al. [26] | |
| Turner-Bowker et al. [27] | |
| Webber et al. [40] | |
| Zigler et al. [29] | |
| Zizi et al. [41] |
Developmental considerations
The scoping review indicated that, prior to conducting cognitive debriefing with pediatric patients, researchers must consider the patients’ ability to meaningfully participate and provide feedback on the relevance, comprehensibility, and/or comprehensiveness of an instrument [10, 11, 16, 28]. For example, several articles noted the importance of considering the cognitive and reading/literacy capabilities of the intended sample population before determining whether a think-aloud approach (i.e., an exercise during which a participant reads aloud all aspects of an instrument) is appropriate, because both cognition and literacy pertain to age-related development and potential health-related developmental delays [8, 10, 18, 19, 29]. To manage these issues, two of the articles [21, 24] considered participants’ reading ability prior to their cognitive debriefing interviews. Five studies, taking into account that their younger participants may want or need support with reading, allowed for participants to have the questionnaire read aloud to them (verbatim) while completing the measure [20, 25, 31, 34, 42]. Irwin et al. (2009) offered breaks during interviews to accommodate the children participating (none ultimately took a break) [33].
In addition to challenges with literacy, pediatric patients may struggle with conceptualizing time, complicating their ability to accurately reflect on an instrument’s recall period. Developmentally, pediatric populations, particularly those under age 8, may lack the cognitive ability to fully comprehend a recall period that asks them to think about a set amount of time (e.g., the last 7 days) [17, 22]. For example, Turner-Bowker, et. al. found that while their sample of pediatric patients seemed to be interpreting recall periods as intended, they were not consistently applying them when answering items. [27] Although this was attributed to a lack of attention on behalf of the participants, the finding points to a need to consider whether the entirety of the task (i.e., reporting on the acceptability of a treatment during a specific timeframe) was too cognitively challenging.
Methodological considerations
The articles included in the review also outlined methodological considerations for conducting cognitive debriefing interviews with pediatric populations. Commonly, research teams took care with the structure and language of their interview guides. Arbuckle and Abetz-Webb (2013) noted the importance of asking unique questions throughout verbal probing, as asking repetitive questions could lead to participants changing their answers and increase the likelihood of social desirability bias in the data [10]. They also highlighted the importance of clear, concrete questions noting that hypothetical questions could be cognitively challenging for some pediatric patients to answer.
To enhance comfort during the interviews and assist their sample’s understanding of the cognitive debriefing process, Kramer and Schwartz (2017) used practice items allowing their participants to practice the technique before reviewing the instrument in question [23]. They also permitted their interviewers to “customize questions” as necessary throughout the verbal probing process. Kramer and Schwartz also enlisted “young adult” interviewers to facilitate their interviews and increase the comfort of the participants [23]. Tomlinson et al. (2019) intentionally started their interviews among children 8–11 with “direct, easy to answer questions” followed by more complex open-ended questions [25]. To support their respondents during interviews assessing the content validity of an instrument designed to measure the symptoms of COVID-19, Romano, et al. (2023) used reference cards with age appropriate definitions for each symptom they asked about [24]. Gale et al. (2025) used a number of approaches to tailor their interviews to their participants aged 6–7: they opened interviews with “information stories and videos about the research”; used practice questions to familiarize the participants with the process (much like Kramer and Schwartz); had the interviewer read each item aloud; and utilized props, such as a teddy bear to whom participants were asked to explain the item (rather than being asked to rephrase in their own words for the interviewer) and a spinner toy, to help keep the participants engaged [11].
Other articles also touched on the importance of using different approaches to engage patients over the course of an interview. Among the articles reviewed, six described using a combined approach of think-aloud and in-depth probing [17, 19, 20, 25, 36, 40]. One used only think-aloud [16], and eight used only verbal probing [24, 26, 29, 31, 32, 34, 35, 42]. The remaining articles did not state which approach was used. Four studies described debriefing instruments in sections or limiting the number of items to be debriefed in one interview to avoid overwhelming participants [24, 26, 39, 42]. Jacobson, et al. (2015) took a different approach by utilizing a retrospective debriefing approach instead of think-aloud [21]. In their study, respondents first completed a subset of items on paper, after which the interviewer briefly inquired into the respondent’s understanding of the item, decision-making processes for choosing among the response choices, and difficulties experienced while answering the questions. O’Sullivan et al. (2014) had participants (ages 8–18) evaluate the ease of completion for each item by using a Likert scale [34]. Bevans et al. (2010) suggested employing a multi-media approach (e.g., graphics, video, audio) to keep participants engaged [8], while Halstead et al. (2020) used a card sorting task to establish the response options for questions asking about symptom severity [19]. Carlton (2013) also used cards to facilitate their interviews: participants (children ages 5–9) were asked to rank questionnaire items by arranging cards on which the items were printed [31]. However, Carlton noted that 10 participants did not understand the card sorting/ranking exercise; as a result, the exercise was not completed for those participants. Tomlinson et al. (2019) used “an ungendered hand puppet to engage with respondents” along with “a board with a windowed frame” when reviewing written items, which helped participants (children, ages 4–7) to focus on just a single item at a time [25]. Finally, Gale et al.’s (2025) aforementioned use of the spinner toy was meant to make follow up probes more game-like as participants would spin the toy to determine which probe would be asked next [11].
In addition to finding ways to make cognitive debriefing interviews more engaging through data collection methodologies, carefully crafted questions, support for the participant, and gamification, a handful of articles also described the use of technology designed specifically to appeal to younger participants. A usability study from Anthony et al. (2024) examined an ePRO platform with participants ages 8–17 years old and found that the integration of developmentally responsive design (e.g., child’s choice of colors and avatar) in pediatric ePRO platforms fostered participants’ sense of engagement and motivation to complete instruments [30]. Related to the use of technology to engage young respondents, two articles described the need to present an instrument to participants as it would be seen during the clinical trial [36, 40], suggesting that instruments that will be used on an ePRO platform should be debriefed on said platform. Additionally, Taylor et al. (2015)[38] did not specify whether they debriefed the digital version of their target instrument, but noted that confusion with skip patterns could have been eliminated if the instrument had been administered electronically as it would be during a trial.
Lastly, the articles addressed the essential decision about who is the best reporter for studies of pediatric populations (self-report vs. parent/observer report vs. combined report), noting that this may change over time and from one study to the next. Arbuckle and Abetz-Webb state this determination should be made based on several considerations, including the concepts being measured, the age of the target population, and the disease area being studied [10]. In the articles reviewed, seven included self-report from children ages 8–11 years old [16, 17, 22, 27, 29, 35, 40], seven included self-report from children ages 12–17 [16–18, 22, 27, 29, 40], three included observer report [16, 17, 27], and three included combined-report [19, 37, 41].
Regulatory guidance
Regulatory guidance on cognitive debriefing in pediatric populations remains limited, with documents from the US FDA, EMA, and the UK’s Medicines and Healthcare Products Regulatory Agency (MHRA) either not addressing cognitive debriefing directly or lacking specific methodological direction. Among available guidance, the US FDA’s PFDD Guidance 3 comes closest to addressing pediatric evaluation by emphasizing that instruments should use age-appropriate vocabulary (“age relevant”) and recall periods that are understandable, both in terms of the word choice and the concepts of interest [3]. However, it does not provide concrete recommendations for conducting cognitive debriefing, instead referring researchers to existing methodological literature (e.g., Arbuckle and Abetz-Webb [10], Bevans et al. [8], Matza et al. [9], and Papadopolous et al. [43]) and Section VI of the PFDD Guidance 2, which broadly addresses potential barriers to self-report and ways to overcome them.
PFDD Guidance 2 highlights practical considerations on how to manage challenges with self-report that may arise when collecting data from children, including children’s limited attention spans and the use of engagement strategies such as drawing or props, as well as providing guidance on caregiver involvement and minimizing their influence during interviews [2]. This is similar to PFDD Guidance 1, which addresses the collection of patient input and suggests “eliciting patient experience through play and drawings” when working with children [1]. Guidance from the EMA is even briefer than that of the US FDA and directs researchers to Matza et al., while acknowledging the potential role of caregivers in data collection with children, and “using creative and age related approaches” such as pictures instead of words, for children not yet reading [44]. The UK’s MHRA does not provide any information on pediatric cognitive debriefing.
Draft best practices
Guided by the findings from the scoping review, we drafted a set of 17 statements of best practices for conducting cognitive debriefing interviews with pediatric populations. Seven of the practices were grouped within the developmental domain (i.e., related to target participants’ developmental capabilities), while the remaining 10 were grouped within the methodological domain.
Modified Delphi method and best practices for conducting cognitive debriefing interviews with pediatric populations
Panel characteristics
The Delphi panel included 10 United States- and Europe-based researchers with expertise in outcomes research and/or PRO instrument development with pediatric populations. The panelists self-identified as working in academia (n = 6) and consulting (n = 4); one panelist was a former colleague of three of the authors. All 10 panelists participated in R1 and eight returned for R2. Two did not respond to invitations to take part in R2, despite multiple contact attempts.
Consensus results by round
The panel reached consensus on the appropriateness of three of the 17 draft best practices reviewed in R1. For the 14 that did not reach consensus, the suggested revisions were generally minor. After incorporating the panel’s feedback (e.g., updating language to clarify), we reorganized the revised 14 best practices into three domains: developing the interview guide (n = 3), characteristics of the instrument being debriefed (n = 4), and interview conduct (n = 7).
The panel reached consensus on the appropriateness of seven of the 14 draft best practices reviewed in R2. Among those for which the panel did not reach consensus, four underwent minor updates to align with panel suggestions. The remaining three were not revised as the panelists’ suggested changes were unclear (e.g., no specific edits were recommended), pertained to another practice (and thus were already covered), or were related to instrument development (e.g., drafting items for an instrument) instead of conducting cognitive debriefing interviews.
Best practices for conducting cognitive debriefing interviews with pediatric populations
The final 17 best practices for conducting cognitive debriefing interviews with pediatric populations are organized into three categories representing various aspects of conducting cognitive debriefing interviews with pediatric participants: developing the interview guide, characteristics of the instrument being debriefed, and interview conduct. Within each category, the practices are not listed in order of importance, but rather with a view to taking the researcher through a logical set of steps in a process.
Developing the interview guide
Best Practice 1: When feasible, researchers should avoid hybrid concept elicitation and cognitive debriefing interviews to minimize cognitive strain associated with switching tasks. If unavoidable, researchers should ensure the interview is as brief and engaging as possible.
Best Practice 2: The interview guide should include clear, concise, and concrete questions. This includes avoiding repetitive and hypothetical questions as much as possible. If questions need to be repeated, explain why it is necessary and important. If hypothetical questions are necessary (e.g., they are dictated by the PRO instrument), they should align with the comprehension and developmental levels of the sample population.
Best Practice 3: Researchers should identify any items in the target instrument that may be complex (e.g., presents a hypothetical situation, uses advanced language) and prepare a simplified backup description(s) that can be used to explain the item in a uniform, age-appropriate way to maintain rapport, participation, and engagement.
Best Practice 4: When possible, the interview guide should be pilot tested with a small number of potential participants (e.g., children or adolescents with the target health condition) to evaluate its appropriateness.
Characteristics of the instrument being debriefed
Best Practice 5: Researchers should review the target instrument(s) before debriefing to proactively identify any elements that may present a challenge for the intended pediatric study sample, including reading level, layout, and use of images or graphics. Developing a study-specific, standardized approach for how to address each potentially challenging element can be helpful.
Best Practice 6: Researchers should consider the cognitive challenge associated with conceptualizing and reporting on the health concept(s) of interest over the period of time specified in the recall period.
Best Practice 7: Researchers should assess the instrument’s recall period to determine if it is appropriately brief for the target study sample but still aligns with the concept(s) being measured.
Best Practice 8: Researchers should consider whether the target pediatric study sample will have the ability to reflect on both the number and labelling of response options for each item within the PRO instrument being debriefed.
Interview conduct
Best Practice 9: Researchers should use alternative approaches to boost engagement over the course of the interview (e.g., debriefing in sets of items, rather than item by item; using drawing, images, graphics, video, and/or adaptative/assistive technology when necessary).
Best Practice 10: When using the think-aloud method, researchers should maintain flexibility to account for different reading comprehension levels.
Best Practice 11: Researchers should reserve the think-aloud for pediatric participants who can read sufficiently on their own. Interviews should be flexible to allow the interviewer to introduce, or if necessary, stop using the think-aloud method depending on the individual participant.
Best Practice 12: Researchers should allow for the instrument to be read aloud/administered by an interviewer if the pediatric participant chooses or requires it (i.e., the instrument’s reading level is too advanced). While independent reading skills are not necessarily required for pediatric populations to self-report via PRO measures, the cognitive debriefing interview provides a good opportunity to determine what level of support may be needed in order to implement the instrument effectively.
Best Practice 13: Researchers should use verbal probing questions with all pediatric participants regardless of whether the participant is able to complete the think-aloud. This approach helps to confirm whether each item and its response options are relevant and understood as intended.
Best Practice 14: Researchers should limit the number of items to debrief at one time. One way to achieve this is by debriefing items in small sets (e.g., 3–4 items) with breaks provided between sets.
Best Practice 15: Researchers should debrief instruments in the expected mode(s) of administration. This includes computer-based formats (ePRO), which can be displayed as screenshots if the ePRO itself is not available for the debriefing study.
Best Practice 16: Researchers should allow pediatric participants to briefly practice the think-aloud method using a practice item (unrelated to the items being debriefed) before debriefing the target instrument(s).
Best Practice 17: Researchers should explore not only whether the participant can recall the instrument's specific timeframe (i.e., recall period), but whether they are able to do it consistently when completing the instrument being debriefed.
Discussion
Although PFDD guidance encourages including children’s perspectives in drug development and recent studies have shed light on how children understand and describe their health [28, 45], there has been limited direction on how to effectively gather pediatric input when evaluating the content validity of PRO instruments. Cognitive debriefing is a complex task that both examines the elements of an instrument and delves into the cognitive processes of the potential respondent. Adapting this task for children and adolescents requires a thoughtful approach from researchers. The 17 best practices described herein offer clear, evidence-based methods—informed by the literature and endorsed by experts—to help researchers go beyond the obvious, because, as pointed out by one of our panelists, “reading level is just one component.”
As evidenced by the scoping review, cognitive debriefing interviews with pediatric populations are often subject to common pitfalls, including interview guides that use language that is neither age- nor developmentally-appropriate [10, 11, 16, 18–21, 23–25, 28, 29, 31, 33, 34]; interviews that are too long [24, 26, 33, 39], too cognitively burdensome [11, 23–25], and not engaging [8, 11, 19, 24–26, 31, 33, 39]; and the use of data collection methods that are unfamiliar or challenging for young participants [17, 22, 27]. The best practices described above aim to help researchers avoid these pitfalls by engaging in careful planning and building in opportunities for flexibility.
By outlining concrete strategies that build on the existing evidence [7, 9, 10] and account for developmental variation, literacy levels, and the unique challenges of pediatric self-report, these practices help ensure that data collected from pediatric populations are both meaningful and reliable. This strengthens the ability of researchers to gather the evidence required to support content validity for pediatric PRO instruments—an area where regulators are increasingly focused and expect rigorous methods. Moreover, as pediatric drug development continues to expand and regulatory agencies increasingly emphasize the inclusion of younger patients in the drug development process worldwide, these practices can provide researchers and sponsors with a consistent and defensible approach for demonstrating that PRO instruments truly reflect the experiences of pediatric patients and are easy for them to complete. Ultimately, this can enhance the quality of data, submitted for regulatory review and inform patient-centered decision-making [46–49]. As such, these best practices are important to both industry and regulatory agencies because they offer an evidence based, methodologically sound framework for engaging pediatric participants in cognitive debriefing research.
It is also important to consider how these best practices will be disseminated and integrated into instrument development efforts. Although this work focuses specifically on the process of cognitive debriefing, that process does not occur in isolation. Concept elicitation interviews, instrument development and selection are critical precursors to cognitive debriefing, and if pediatric perspectives are not incorporated during those earlier stages, conducting effective cognitive debriefing becomes far more challenging. Ensuring collaboration with instrument developers—and encouraging adoption of these best practices from the outset—will help create instruments that are better aligned with the needs and abilities of pediatric populations.
This is especially true when working with pediatric populations with rare diseases. Methods employed for cognitive debriefing studies in rare disease pediatric populations need to be thoughtfully chosen and highly relevant to the patients’ lived experiences, developmental stages, and functional abilities. When sample sizes for all phases of PRO development and testing are small, the best practices for cognitive debriefing will introduce the necessary scientific rigor to support inclusion of PRO instruments in studies to develop much-needed new therapies.
The current work can also be expanded by the development of clear guidance on how to operationalize these best practices. This guidance might include formal steps for assessing cognitive burden and/or reading level, pilot-testing interview guides, adapting these best practices for children who are younger than eight or neurodiverse, and addressing how they might be implemented in cross-cultural validation studies. This work could also be expanded into adjacent domains, such as clinical practice or the development of pediatric patient-reported experience measures. Finally, future iterations of these best practices would benefit from the input of pediatric patients, their caregivers, and/or their clinicians. The current best practices were developed without direct input from these groups as the intent was to ensure the best practices reflected expertise in the development and content validation of patient-reported outcome instruments.
The development of these best practices was subject to some limitations. In terms of the scoping review, by restricting the search parameters for the scoping review to English-only records, it is possible that the review is missing insights from other languages. Similarly, the inaccessibility of certain records identified through the database search may have resulted in the exclusion of potentially informative articles. In terms of the modified Delphi, panelist attrition from Round 1 to Round 2 may have constrained further refinement of the best practices. Although the volume of feedback provided in Round 2 was relatively minor, and therefore the attrition is unlikely to have meaningfully influenced the final recommendations, this remains a potential limitation. Finally, the composition of the panel itself may have influenced the final set of best practices; although efforts were made to invite panelists with diverse expertise and viewpoints, selection bias may have occurred.
The methods used to develop these best practices have several strengths. The scoping review was essential for mapping the available evidence, identifying key concepts, and highlighting gaps in the literature regarding the use of cognitive debriefing to develop or validate PRO instruments for pediatric populations. The Delphi panel strengthened the development of the best practices by enabling structured, expert consensus. The two rounds supported iterative refinement of the best practices, improving their specificity, clarity, and agreement among panelists.
Together, these best practices offer practical, evidence-based research standards for strengthening the rigor of cognitive debriefing interviews with pediatric populations. As health measurement advances and our understanding of pediatric patients’ lived experiences deepens, the methods used to evaluate the appropriateness of instruments must evolve in parallel. Integrating these best practices into cognitive debriefing studies will help researchers conduct rigorous studies that have considered and adapted for the unique needs of their pediatric study sample. Broader adoption across drug development programs will further advance the field by promoting more reliable measurement, more meaningful pediatric input, and more patient centered decision-making.
Acknowledgements
We would like to thank the following individuals for their valuable contributions to this work: Linda Abetz-Webb, Eline Alons, Rob Arbuckle, Jill Carlton, Victoria Gale, Jeanne Landgraf, Miriam Linver, Tessa Peasgood, Zabin Patel-Syed, and Laura Waldman. Their time, interest, and insights are greatly appreciated.
Author contributions
Lynne Broderick, Meaghan O’Connor, and Elizabeth Brennan contributed to the study conception and design. Material preparation, data collection and analysis were performed by Lynne Broderick, Meaghan O’Connor, Elizabeth Brennan, and Martha Bayliss. The first draft of the manuscript was written by Lynne Broderick and all authors commented on previous versions of the manuscript. All authors read and approved the final manuscript. Lynne Broderick: Conceptualization, Methodology, Validation, Formal analysis, Investigation, Resources, Writing—Original draft, Writing—Review and editing, Visualization, Project administration, Funding acquisition Meaghan O’Connor: Methodology, Validation, Formal analysis, Investigation, Writing–Review and editing Elizabeth Brennan: Formal analysis, Investigation, Writing—Review and editing Martha Bayliss: Formal analysis, Writing—Review and editing, Supervision.
Funding
This work was funded by IQVIA, Inc.
Data availability
No datasets were generated or analysed during the current study.
Declarations
Conflict of interest
The authors declare no competing interests.
Ethical approval
All individuals involved in the modified Delphi participated voluntarily in their capacity as experts. This study included no interventions, and participation posed no risk for Delphi panelists. As such, formal ethics review and approval were not required. Nevertheless, the study team followed good research practices.
Consent to participate
Informed consent was not required as this study was not human subjects research. Nevertheless, all participants in the modified Delphi were provided with a detailed overview of the study, its goals, and the expectations for panelists prior to their committing to participate.
Consent for publication
Not applicable.
Footnotes
Publisher's Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.U.S. Food and Drug Administration. (2020). Patient-Focused Drug Development: Collecting Comprehensive and Representative Input.
- 2.U.S. Food and Drug Administration. (2022). Patient-Focused Drug Development: Methods to Identify What Is Important to Patients.
- 3.U.S. Food and Drug Administration. (2025). Patient-Focused Drug Development: Selecting, Developing, or Modifying Fit-for-Purpose Clinical Outcome Assessments.
- 4.U.S. Food and Drug Administration. (2023). Patient-Focused Drug Development: Incorporating Clinical Outcome Assessments Into Endpoints For Regulatory Decision-Making.
- 5.Chan, E. K. H., Edwards, T. C., Haywood, K., Mikles, S. P., & Newton, L. (2019). Implementing patient-reported outcome measures in clinical practice: A companion guide to the ISOQOL user’s guide. Quality of Life Research,28(3), 621–627. 10.1007/s11136-018-2048-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Patrick, D. L., Burke, L. B., Gwaltney, C. J., Leidy, N. K., Martin, M. L., Molsen, E., et al. (2011). Content validity–establishing and reporting the evidence in newly developed patient-reported outcomes (PRO) instruments for medical product evaluation: ISPOR PRO good research practices task force report: Part 1–eliciting concepts for a new PRO instrument. Value Health J Int Soc Pharmacoeconomics Outcomes Res.,14(8), 967–977. 10.1016/j.jval.2011.06.014 [DOI] [PubMed] [Google Scholar]
- 7.Patrick, D. L., Burke, L. B., Gwaltney, C. J., Leidy, N. K., Martin, M. L., Molsen, E., et al. (2011). Content validity–establishing and reporting the evidence in newly developed patient-reported outcomes (PRO) instruments for medical product evaluation: ISPOR PRO Good Research Practices Task Force report: Part 2–assessing respondent understanding. Value Health,14(8), 978–988. 10.1016/j.jval.2011.06.013 [DOI] [PubMed] [Google Scholar]
- 8.Bevans, K. B., Riley, A. W., Moon, J., & Forrest, C. B. (2010). Conceptual and methodological advances in child-reported outcomes measurement. Expert Review of Pharmacoeconomics and Outcomes Research,10(4), 385–396. 10.1586/erp.10.52 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Matza, L. S., Patrick, D. L., Riley, A. W., Alexander, J. J., Rajmil, L., Pleil, A. M., et al. (2013). Pediatric patient-reported outcome instruments for research to support medical product labeling: Report of the ISPOR PRO good research practices for the assessment of children and adolescents task force. Value Health,16(4), 461–479. 10.1016/j.jval.2013.04.004 [DOI] [PubMed] [Google Scholar]
- 10.Arbuckle, R., & Abetz-Webb, L. (2013). “Not just little adults”: Qualitative methods to support the development of pediatric patient-reported outcomes. The Patient.,6(3), 143–159. 10.1007/s40271-013-0022-3 [DOI] [PubMed] [Google Scholar]
- 11.Gale, V., Powell, P. A., & Carlton, J. (2025). Young children (6–7 years) can meaningfully participate in cognitive interviews assessing comprehensibility in health-related quality of life domains: A qualitative study. Quality of Life Research. 10.1007/s11136-025-03940-z [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Hsu, C. C., & Sandford, B. A. (2007). The Delphi technique: Making sense of consensus. Practical Assessment, Research, and Evaluation,12(1), 10. 10.7275/PDZ9-TH90 [DOI] [Google Scholar]
- 13.Khodyakov, D., Grant, S., Kroger, J., & Bauman, M. (2023). RAND methodological guidance for conducting and critically appraising Delphi panels. RAND Corporation. [cited 2026 Jan 27]. Available from: https://www.rand.org/pubs/tools/TLA3082-1.html, 10.7249/TLA3082-1.
- 14.Junger, S. (2023). Delphi studies in the health sciences: Epistemic potentials and challenges. Delphi methods in the social and health sciences (1st ed, pp. 51–74). Springer. [Google Scholar]
- 15.Diamond, I. R., Grant, R. C., Feldman, B. M., Pencharz, P. B., Ling, S. C., Moore, A. M., et al. (2014). Defining consensus: A systematic review recommends methodologic criteria for reporting of Delphi studies. Journal of Clinical Epidemiology,67(4), 401–409. 10.1016/j.jclinepi.2013.12.002 [DOI] [PubMed] [Google Scholar]
- 16.Aldhouse, N. V. J., Kitchen, H., Johnson, C., Marshall, C., Pegram, H., Pease, S., et al. (2022). Key measurement concepts and appropriate clinical outcome assessments in pediatric achondroplasia clinical trials. Orphanet Journal of Rare Diseases,17(1), 182. 10.1186/s13023-022-02333-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Coombes, L., Braybrook, D., Harðardóttir, D., Scott, H. M., Bristowe, K., Ellis-Smith, C., et al. (2024). Cognitive testing of the Children’s Palliative Outcome Scale (C-POS) with children, young people and their parents/carers. Palliative Medicine,38(6), 644–659. 10.1177/02692163241248735 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Gabes, M., Ragamin, A., Baker, A., Kann, G., Donhauser, T., Gabes, D., et al. (2022). Content validity of RECAP in Dutch, English and German to measure eczema control in young people with atopic eczema: cognitive interview study. Exp Dermatol. 48th Annual Meeting of the Arbeitsgemeinschaft Dermatologische Forschung, ADF. Virtual. 31(2):e40–1. Located at: Embase; 637795118. 10.1111/exd.14511. [DOI]
- 19.Halstead, P., Arbuckle, R., Marshall, C., Zimmerman, B., Bolton, K., & Gelotte, C. (2020). Development and content validity testing of patient-reported outcome items for children to self-assess symptoms of the common cold. The Patient.,13(2), 235–250. 10.1007/s40271-019-00404-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Hwang, M., Zebracki, K., Vogel, L. C., Mulcahey, M. J., & Varni, J. W. (2020). Development of the pediatric quality of life inventory™ spinal cord injury (PedsQL™ SCI) module: Qualitative methods. Spinal Cord.,58(10), 1134–1142. 10.1038/s41393-020-0450-6 [DOI] [PubMed] [Google Scholar]
- 21.Jacobson, C. J. J., Kashikar-Zuck, S., Farrell, J., Barnett, K., Goldschneider, K., Dampier, C., et al. (2015). Qualitative evaluation of pediatric pain behavior, quality, and intensity item candidates and the PROMIS pain domain framework in children with chronic pain. The Journal of Pain,16(12), 1243–1255. 10.1016/j.jpain.2015.08.007 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Kamath, B. M., Abetz-Webb, L., Kennedy, C., Hepburn, B., Gauthier, M., Johnson, N., et al. (2018). Development of a novel tool to assess the impact of itching in pediatric cholestasis. The patient.,11(1), 69–82. 10.1007/s40271-017-0266-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Kramer, J. M., & Schwartz, A. (2017). Refining the pediatric evaluation of disability inventory-patient-reported outcome (PEDI-PRO) item candidates: Interpretation of a self-reported outcome measure of functional performance by young people with neurodevelopmental disabilities. Developmental Medicine and Child Neurology,59(10), 1083–1088. 10.1111/dmcn.13482 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Romano, C., Mayorga, M., Ruiz-Guiñazú, J., Trudel, G. C., Fehnel, S., McQuarrie, K., et al. (2023). Development of patient- and observer-reported outcome measures to assess COVID-19 signs and symptoms in children and adolescents. Journal of Patient-Reported Outcomes,7(1), 7. 10.1186/s41687-023-00542-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Tomlinson, D., Hyslop, S., Stein, E., Spiegler, B., Vettese, E., Kuczynski, S., et al. (2019). Development of mini-SSPedi for children 4–7 years of age receiving cancer treatments. BMC Cancer,19(1), 32. 10.1186/s12885-018-5210-z [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Tucker, C. A., Bevans, K. B., Teneralli, R. E., Smith, A. W., Bowles, H. R., & Forrest, C. B. (2014). Self-reported pediatric measures of physical activity, sedentary behavior, and strength impact for PROMIS: Item development. Pediatric Physical Therapy,26(4), 385–392. 10.1097/PEP.0000000000000074 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Turner-Bowker, D. M., An Haack, K., Krohe, M., Yaworsky, A., Vivas, N., Kelly, M., et al. (2020). Development and content validation of the pediatric oral medicines acceptability questionnaires (P-OMAQ): Patient-reported and caregiver-reported outcome measures. Journal of Patient-Reported Outcomes,4(1), 80. Located at: Embase; 643890578. 10.1186/s41687-020-00246-1 [DOI] [PMC free article] [PubMed]
- 28.Willis, J., Zeratkaar, D., Ten Hove, J., Rosenbaum, P., & Ronen, G. M. (2021). Engaging the voices of children: A scoping review of how children and adolescents are involved in the development of quality-of-life-related measures. Value Health,24(4), 556–567. 10.1016/j.jval.2020.11.007 [DOI] [PubMed] [Google Scholar]
- 29.Zigler, C. K., Ardalan, K., Lane, S., Schollaert, K. L., & Torok, K. S. (2020). A novel patient-reported outcome for paediatric localized scleroderma: A qualitative assessment of content validity. British Journal of Dermatology,182(3), 625–635. 10.1111/bjd.18512 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Anthony, S. J., Pol, S. J., Selkirk, E. K., Matthiesen, A., Klaassen, R. J., Manase, D., et al. (2024). User-centered design and usability of voxe as a pediatric electronic patient-reported outcome measure platform: Mixed methods evaluation study. JMIR Human Factors,19(11), Article e57984. 10.2196/57984 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Carlton, J. (2013). Developing the draft descriptive system for the child amblyopia treatment questionnaire (CAT-Qol): A mixed methods study. Health and Quality of Life Outcomes,22(11), 174. 10.1186/1477-7525-11-174 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Cella, D., Blackwell, C. K., & Wakschlag, L. S. (2022). Bringing PROMIS to early childhood: Introduction and qualitative methods for the development of early childhood parent report instruments. Journal of Pediatric Psychology,47(5), 500–509. 10.1093/jpepsy/jsac027 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Irwin, D. E., Gross, H. E., Stucky, B. D., Thissen, D., DeWitt, E., Lai, J., et al. (2012). Development of six PROMIS pediatrics proxy-report item banks. Health and Quality of Life Outcomes,10(1), Article 22. 10.1186/1477-7525-10-22 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.O’Sullivan, C., Dupuis, L. L., Gibson, P., Johnston, D. L., Baggott, C., Portwine, C., et al. (2014). Refinement of the symptom screening in pediatrics tool (SSPedi). British Journal of Cancer,111(7), 1262–1268. 10.1038/bjc.2014.445 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Propp, R., McAdam, L., Davis, A. M., Salbach, N. M., Weir, S., Encisa, C., et al. (2019). Development and content validation of the muscular dystrophy child health index of life with disabilities questionnaire for children with Duchenne muscular dystrophy. Developmental Medicine and Child Neurology,61(1), 75–81. 10.1111/dmcn.13977 [DOI] [PubMed] [Google Scholar]
- 36.Rams, A., Baldasaro, J., Bunod, L., Delbecque, L., Strzok, S., Meunier, J., et al. (2024). Assessing itch severity: Content validity and psychometric properties of a patient-reported pruritus numeric rating scale in atopic dermatitis. Advances in Therapy,41(4), 1512–1525. 10.1007/s12325-024-02802-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Sarda, S. P., De La Cruz, M., Flood, E. M., Vanya, M., Hwang, D. G., Ta, C. N., et al. (2019). Content validity of a novel patient-reported and observer-reported outcomes assessment to evaluate ocular symptoms associated with infectious conjunctivitis in both adult and pediatric populations. Health and Quality of Life Outcomes,17(1), 163. 10.1186/s12955-019-1223-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Taylor, R. M., Fern, L. A., Solanki, A., Hooker, L., Carluccio, A., Pye, J., et al. (2015). Development and validation of the BRIGHTLIGHT Survey, a patient-reported experience measure for young people with cancer. Health and Quality of Life Outcomes,28(13), 107. 10.1186/s12955-015-0312-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Truninger, M. I., Werner, H., Landolt, M. A., Hahn, A., Hennermann, J. B., Lagler, F. B., et al. (2024). The PompeQoL questionnaire: Development and validation of a new measure for children and adolescents with Pompe disease. Journal of Inherited Metabolic Disease.,47, 1348–1362. 10.1002/jimd.12777 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Webber, A., Randhawa, S., Felizzi, F., Soos, M., Arbuckle, R., O’Brien, P., et al. (2023). The amblyopia quality of life (AmbQoL): Development and content validation of a novel health-related quality of life instrument for use in adult and pediatric amblyopia populations. Ophthalmology and Therapy,12(2), 1281–1313. Located at: Embase; 2021766087. 10.1007/s40123-023-00668-2 [DOI] [PMC free article] [PubMed]
- 41.Zizzi, C. E., Luebbe, E., Mongiovi, P., Hunter, M., Dilek, N., Garland, C., et al. (2021). The spinal muscular atrophy health index: A novel outcome for measuring how a patient feels and functions. Muscle and Nerve,63(6), 837–844. 10.1002/mus.27223 [DOI] [PubMed] [Google Scholar]
- 42.Irwin, D. E., Varni, J. W., Yeatts, K., & DeWalt, D. A. (2009). Cognitive interviewing methodology in the development of a pediatric item bank: A patient reported outcomes measurement information system (PROMIS) study. Health and Quality of Life Outcomes,23(7), 3. 10.1186/1477-7525-7-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Papadopoulos, E. J., Bush, E. N., Eremenco, S., & Coons, S. J. (2020). Why reinvent the wheel? Use or modification of existing clinical outcome assessment tools in medical product development. Value in Health: The Journal of the International Society for Pharmacoeconomics and Outcomes Research,23(2), 151–153. 10.1016/j.jval.2019.09.2745 [DOI] [PubMed] [Google Scholar]
- 44.Committee for Medicinal Products for Human Use. (2016). Appendix 2 to the guideline on the evaluation of anticancer medicinal products in man: The use of patient-reported outcome (PRO) measures in oncology studies [Internet]. European Medicines Agency. Available from: https://www.ema.europa.eu/en/documents/other/appendix-2-guideline-evaluation-anticancer-medicinal-products-man_en.pdf.
- 45.Kroh, J., Tuppat, J., Gentile, R., & Reichelt, H. (2023). How do children rate their health? An investigation of considered health dimensions, health factors, and assessment strategies. Child Indicators Research,16(6), 2545–2580. 10.1007/s12187-023-10066-6 [DOI] [Google Scholar]
- 46.Pharmaceuticals and Medical Devices Agency. (2025). Initiatives to Promote Pediatric Drug Development. Report PMDA/CPE Notification No. 1618. Available from: https://www.pmda.go.jp/files/000274940.pdf.
- 47.U.S. Food and Drug Administration. (2025). Interested Parties Meeting: Implementation of the Best Pharmaceuticals for Children Act and Pediatric Research Equity Act. Available from: https://www.fda.gov/news-events/fda-meetings-conferences-and-workshops/interested-parties-meeting-implementation-best-pharmaceuticals-children-act-and-pediatric-research.
- 48.Pagano, A., Green, D. J., Goldman, J. L., & Deshmukh, A. (2026). Promises, pitfalls, and paths forward for the Pediatric Research Equity Act. JAMA Health Forum,7(5), Article e260993. 10.1001/jamahealthforum.2026.0993 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.European Medicines Agency. (2026). Procedural advice on paediatric applications: Guidance for applicants. Apr. Report EMA/672643/2017 Rev. 151. Available from: https://www.ema.europa.eu/en/documents/regulatory-procedural-guideline/procedural-advice-paediatric-applications_en.pdf.
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
No datasets were generated or analysed during the current study.
