Skip to main content
JAMA Network logoLink to JAMA Network
. 2026 Aug 20;9(8):e2630074. doi: 10.1001/jamanetworkopen.2026.30074

Reliability of Physical Examination Findings in Youths Diagnosed With Pneumonia

Shubhada Hooli 1, Ron Reeder 2, Lauren Cutler 3, Laura F Sartori 4, Geoff Capraro 5, Amy Y Cheng 6, Allison Cator 7,8, Matthew J Lipshaw 7,8, Lilliam Ambroggio 9,10, Chris A Rees 11, Son H McLaren 12, Justin Moher 13, Leah Tzimenatos 14, Patrick S Walsh 15, Chari D Larsen 16, Richard M Ruddy 17, Samir S Shah 18,19, Nathan Kuppermann 20,21, Todd A Florin 22,23,, for the Pediatric Emergency Care Applied Research Network (PECARN) PedCAPS Investigators
PMCID: PMC13494721  PMID: 42623068

This cohort study investigates the interrater reliability of physical examination findings among children and adolescents in emergency departments diagnosed with community-acquired pneumonia.

Key Points

Question

What is the interrater reliability of physical examination findings in youths diagnosed with community-acquired pneumonia?

Findings

In this cohort study among 252 youths with community-acquired pneumonia, individual respiratory auscultatory findings (such as crackles or decreased breath sounds) had limited interrater reliability. Wheezing and retractions had the highest reliability.

Meaning

These results suggest that auscultatory examination findings of crackles or decreased breath sounds alone are not sufficiently reliable to diagnose community-acquired pneumonia in youths.

Abstract

Importance

Community-acquired pneumonia (CAP) accounts for nearly 2 million pediatric outpatient and 375 000 emergency department (ED) visits annually in the US. Guidelines recommend relying on physical examination findings, not imaging, to diagnose CAP in youths who can be treated as outpatients.

Objective

To determine the interrater reliability (IRR) of physical examination findings in youths diagnosed with CAP in EDs.

Design, Setting, and Participants

This was a planned analysis from an ongoing prospective cohort study (pediatric CAP severity [PedCAPS]). Youths aged 3 months to 17 years with CAP were recruited at 7 academic pediatric EDs within the US from August 1, 2023, until May 24, 2025; participants had signs of lower respiratory tract infections, fever within 48 hours, and pneumonia on chest radiography, if performed. Youths with chronic pulmonary diseases (except asthma), sickle cell disease, immunodeficiency, cardiac disease, neurological disorders affecting respiration, and aspiration pneumonia were excluded, as were those hospitalized within the preceding 30 days or transferred from other EDs or hospitals.

Main Outcomes and Measures

Two examiners evaluated the same patient within 60 minutes of each other and independently recorded their findings. IRR of physical examination findings was reported by raw agreement and Fleiss κ. A lower bound of the 95% CI of 0.4 for κ was considered acceptable reliability.

Results

Among 252 youths with paired physical examinations (median [IQR] age, 5.7 [3.4-8.8] years; 127 female [50.4%]), the most frequent comorbidity was asthma (56 youths [22.2%]). In the overall study population, no physical examination finding met predefined significance for IRR. Wheezing (κ = 0.50; 95% CI, 0.39-0.62) and retractions (κ = 0.49; 95% CI, 0.37-0.60) had the highest IRR. In subanalyses of 124 youths discharged home and 128 youths who were hospitalized, IRRs of physical examinations were similar between the 2 groups.

Conclusions and Relevance

In this study, individual auscultation findings, such as decreased breath sounds, crackles, or rhonchi, did not demonstrate sufficient reliability to be used alone for diagnosis.

Introduction

Acute respiratory illnesses are the most common reason youths present for acute care.1 Community-acquired pneumonia (CAP) is one of the most frequent respiratory infections and accounts for up to 2 million outpatient visits and 375 000 emergency department (ED) visits annually in the US.2

Given the challenges in distinguishing CAP from other respiratory illnesses in youths, emergency medicine clinicians often obtain chest radiographs (CXRs) in those with suspected CAP.3,4 However, the Pediatric Infectious Diseases Society/Infectious Diseases Society of America recommends that CAP be diagnosed clinically (ie, without radiographic imaging) in youths who do not have comorbidities or prolonged symptoms and who are treated as outpatients.5 The clinical diagnosis of CAP often relies on several physical examination findings, such as fever, cough, crackles or rales, decreased breath sounds, tachypnea, and increased work of breathing. However, the subjective nature of physical examination findings can lead to misdiagnosis and subsequent undertreatment or overtreatment of CAP.6

The interrater reliability (IRR) of common examination findings in youths with acute respiratory illnesses, including CAP, is not well described. Additionally, it has largely been examined only in wheezing illnesses, such as asthma; from auscultation recordings; or in single-center studies.7,8,9,10,11,12,13 To overcome these limitations we analyzed data from a prospective, multicenter study to evaluate the agreement and IRR of respiratory physical examination findings in youths diagnosed with CAP in emergency care settings.

Methods

Study Design

This cohort study aimed to determine the IRR of physical examination findings using data from an ongoing prospective, multicenter cohort study to derive and validate a pediatric CAP severity (PedCAPS) clinical prediction rule14 within the Pediatric Emergency Care Applied Research Network (PECARN).15 Details of the parent study protocol have been described previously.14 This study was conducted in accordance with the Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) reporting guideline. The University of Utah Institutional Review Board (IRB), acting as a single IRB with all sites in reliance, approved this study (IRB_00159139). The reviewing IRB granted a waiver of informed consent because the study met 45 CFR §46.116 requirements by posing minimal risk and to avoid enrollment bias.

Study Setting and Population

We recruited youths aged 3 months through 17 years diagnosed with CAP at 7 academic pediatric EDs within the US from August 1, 2023, until May 24, 2025. These sites participated in the derivation phase of the parent study. We included youths diagnosed with CAP by the ED clinician, exhibiting signs or symptoms consistent with acute lower respiratory infections, fever within 48 hours of presentation, and findings suggestive of pneumonia on CXR, if performed. Exclusion criteria included chronic pulmonary disease (eg, bronchopulmonary dysplasia, tracheostomy, or oxygen dependence), sickle cell disease, immunodeficiency (eg, active chemotherapy or chronic steroid use), clinically substantial cardiac disease (eg, single-ventricle physiology), and muscular dystrophy or other neurological disorders affecting respiration. Given that our primary disease of interest was CAP, we excluded youths with aspiration pneumonia or those previously enrolled or hospitalized within 30 days prior to the ED visit, which would put them at greater risk for recurrent or hospital-acquired pneumonia. We also excluded youths transferred from other EDs or hospitals.

Study Procedures

ED clinicians of youths who were eligible for inclusion in the study were approached to complete a standardized case report form (CRF). ED clinicians included attending physicians, advanced practice clinicians, or pediatric emergency medicine fellows. Each site had a robust pediatric emergency medicine fellowship program with at least 25 faculty members.

The CRF detailed patient symptoms, comorbidities, and physical examination findings. For this planned analysis, our goal was to collect IRR examinations on 5% to 10% of our original 2000-patient enrollment goal. For the subset of participants included in this analysis, a second ED clinician independently conducted a physical examination of the patient and completed an identical CRF within 60 minutes of the first clinician’s examination, with no interventions provided in between. To reflect actual clinical practice conditions, we did not instruct clinicians on how to assess physical examination findings. Time of completion and knowledge of chest imaging findings were documented on both CRFs. Clinicians were masked to the other evaluator’s CRF but may have reviewed the patient’s imaging or discussed their examination findings as part of routine care coordination. Study research coordinators approached families and asked them to self-report patient demographics. When this was not possible, coordinators extracted information from the electronic health record. This information consisted of demographics, including race and ethnicity and sex, as well as additional clinical information, such as vital signs, imaging results, and disposition. Data were entered into a REDCap database.16

Variables

Physical examination findings were primarily binary and included altered mental status, capillary refill time (≤2 seconds vs ≥3 seconds), chest retractions or accessory muscle use, and grunting. Auscultatory findings (ie, wheezing, diminished or absent breath sounds, crackles or rales, and rhonchi or coarse breath sounds) were all categorized as not present, focal, or diffuse. As a subanalysis, multicategory breath sound variables were transformed to dichotomous variables indicating the presence or absence of a specific finding. Other variables of interest included age, demographic characteristics, comorbidities, vital signs, chest radiography findings (eg, atelectasis, consolidation, and peribronchial thickening), disposition, and discharge diagnoses. Race and ethnicity were assessed in the study to ensure that the study population was representative and to assess differences in IRR given potential differences in ability to assess findings based on varied skin tones. In the data collection tool, race options included American Indian or Alaska Native, Asian, Black or African American, Native Hawaiian or Other Pacific Islander, White, unknown, other, and not reported. Ethnicity options included Hispanic or Latino, not Hispanic or Latino, unknown, other, and not reported. Racial minority groups were defined as consisting of youths whose parents or caregivers stated that their race was American Indian or Alaska Native, Asian, Black or African American, Native Hawaiian or Other Pacific Islander, or other race.

Statistical Analysis

We summarized participant demographics and clinical characteristics using counts and percentages for categorical variables and median and IQR for continuous variables. The incidence of each physical examination finding was calculated as the proportion of observations (individual examinations) in which the finding was present. Because participants underwent up to 2 examinations, each contributed up to 2 observations. The number of patients enrolled by site is reported in eTable 1 in Supplement 1. To assess for selection bias, we compared participant enrollment by site and compared demographics and clinical characteristics of our analytical cohort with those of youths without a second physical examination in the pedCAPS derivation cohort. We compared categorical variables using Fisher exact tests (with Monte Carlo approximation for tables larger than 2 × 2), and continuous variables were compared using the Wilcoxon rank-sum test. Our threshold for determining a difference was a 2-sided P value < .05. These findings are reported in eTable 1 in Supplement 1.

The proportion of total agreement (ie, raw agreement) was calculated for each examination finding. Interrater reliability was evaluated using Fleiss κ to account for agreement greater than what would be expected by chance given the involvement of multiple raters.17 Fleiss κ and associated 95% CIs were computed using the MAGREE macro version 3.918 in SAS version 9.4 (SAS Institute) and presented in a graphical plot of IRR (Figure 1). As a sensitivity analysis, we present a Gwet agreement coefficient (AC1)19 analysis in eTable 2 in Supplement 1.

Figure 1. Forest Plot of Interrater Reliability (IRR) and Agreement on Physical Examination Findings in Study Population.

Data table and forest plot of Fleiss kappa and raw agreement by exam finding. Left panel table with four columns and nine rows including one header row. Column headers, left to right: Physical examination finding; Frequency in population, No. positive observations slash total No. observations percent a; Fleiss kappa (95% C I); Raw agreement. Row entries top to bottom: Altered mental status, 18 slash 502 (3.6), 0.31 (0.02 to 0.60), 0.95. Capillary refill time, 20 slash 502 (4.0), 0.27 (0.00 to 0.54), 0.94. Decreased breath sounds, 252 slash 500 (50.4), 0.11 (0.02 to 0.20), 0.46. Fine crackles or rales, 228 slash 502 (45.4), 0.21 (0.11 to 0.31), 0.55. Grunting, 29 slash 502 (5.8), 0.38 (0.14 to 0.61), 0.93. Retractions or accessory muscle use, 160 slash 504 (31.7), 0.49 (0.37 to 0.60), 0.78. Rhonchi or coarse breath sounds, 189 slash 500 (37.8), 0.18 (0.08 to 0.28), 0.56. Wheezing, 110 slash 500 (22.0), 0.50 (0.39 to 0.62), 0.82. Right panel forest plot aligned to the same eight findings, with horizontal axis labeled I R R (95% C I) ranging from zero to 1.00 with ticks at 0, 0.25, 0.50, 0.75, and 1.00. A vertical dotted reference line near 0.40. For each row, a dark teal circle with a horizontal whisker for Fleiss kappa and its 95 percent confidence interval, and a dark teal triangle for raw agreement. Legend in the upper right: circle labeled kappa; triangle labeled Raw agreement.

The overall study population was 252 youths.

aEach examiner’s assessment was considered a distinct observation; therefore, the total number of observations equaled twice the number of patients with paired examinations performed.

Interrater reliability was also evaluated in 4 subgroups by age (<4 years vs ≥4 years), race (White vs racial minority groups), disposition (home vs hospitalization), rater knowledge of CXR findings, and CXR-confirmed pneumonia. We chose to assess the IRR of capillary refill time among White youths compared with youths in racial minority groups because there are concerns that capillary refill time may be misjudged in individuals with darker skin tones.20,21 Our full rationale for subgroup analyses is presented in eTable 3 in Supplement 1.

Fleiss κ is reported on a scale from 0 to 1, where 0 indicates poor agreement and 1 indicates near-perfect agreement. Agreement was qualitatively judged as poor (<0.20), fair (0.21-0.40), moderate (0.41-0.60), substantial (0.61-0.80), and near perfect (0.81-1.00).22 For the purposes of statistical significance and power calculations, we used asymptotic variance to calculate a κ threshold of 0.7, with the lower limit of the 95% CI of 0.4 or greater, given that this represents at least moderate agreement.7,23 When calculating a target sample size, we determined that to achieve a κ of 0.7 with a lower limit of the 95% CI of 0.4, assuming the prevalence of a finding is at least 5%, 122 paired examinations of unique patients were required.19,22,23 To further evaluate sample size adequacy, we conducted simulation analyses using observed distributions of wheezing and retractions. Fleiss κ was then calculated for wheezing and retractions under scenarios with 100 and 200 additional second physical exams. Each scenario was repeated over 1000 simulations. These simulations suggested that increasing the number of paired examinations was unlikely to meaningfully change results.

Results

We enrolled 252 youths in the IRR component of the PedCAPS study (median [IQR] age, 5.7 [3.4-8.8] years; 127 female [50.4%]; 53 African American or Black [21.0%], 1 American Indian or Alaska Native [0.4%], 12 Asian [4.8%], 143 White [56.7%], and 28 other race or ethnicity [11.1%]; 64 Hispanic or Latino [26.0%]), and the most frequent comorbidity was asthma (56 youths [22.2%]) (Table). A parent or guardian self-reported the race and ethnicity of 151 participants (59.9%). The median (IQR) initial oxygen saturation (Spo2) was 95% (93%-97%). One-half of participants were hospitalized (128 youths [50.8%]), among whom 9 participants (3.6%) were admitted to an intensive care unit; 119 were admitted to an acute care ward (47.2%) (Table). Nearly all patients (249 patients [98.8%]) had both clinical and radiographic diagnoses of pneumonia, with the remaining few being diagnosed clinically without radiographs. The frequency of abnormal physical examination findings ranged from 18 observations (3.6%) with altered mental status to 252 observations (50.4%) with decreased breath sounds (Figure 1). Our study population, had a higher proportion of White participants (56.7% vs 2035 participants [46.5%]), as well as higher frequency of hospitalization (47.2% vs 1627 participants [37.2%]) and wheezing (59 participants [23.6%] vs 760 participants [17.4%]) compared with 4375 youths without a second physical examination in the pedCAPS derivation cohort (eTable 1 in Supplement 1). In contrast, proportions of participants classified as other race (11.1% vs 898 participants [20.5%]) or Hispanic or Latino ethnicity (26.0% vs 1573 participants [36.3%]) were lower in our analytical cohort.

Table. Patient Demographics, Characteristics, and Outcomes.

Variable Patients, No. (%) (N = 252)
Age, median (IQR), y 5.7 (3.4-8.8)
Sex
Female 127 (50.4)
Male 125 (49.6)
Race
African American or Black 53 (21.0)
American Indian or Alaska Native 1 (0.4)
Asian 12 (4.8)
White 143 (56.7)
>1 Race 9 (3.6)
Othera 28 (11.1)
Unknown or not reported 6 (2.4)
Ethnicity
Hispanic or Latino 64 (26.0)
Not Hispanic or Latino 182 (74.0)
Comorbidities
Asthma 56 (22.2)
Premature birth (<37 wk) 7 (2.8)
Genetic or metabolic 6 (2.4)
Cardiovascular 3 (1.2)
Neurologic 2 (0.8)
Initial oxygen saturation, median (IQR), % 95.0 (93.0-97.0)
Initial respiratory rate, median (IQR), breaths/min 32.0 (24.0-44.0)
Disposition from ED
Outpatient or home 124 (49.2)
Hospital ward 119 (47.2)
Intensive care unit 9 (3.6)

Abbreviation: ED, emergency department.

a

Any race not otherwise specified.

No physical examination findings demonstrated substantial IRR (κ>0.7 (Figure 1). The highest IRR was observed for wheezing (κ = 0.50; 95% CI, 0.39-0.62) and retractions (κ = 0.49; 95% CI, 0.37-0.60). When multicategory breath sound variables were dichotomized to reflect presence or absence rather than location (focal vs diffuse), the IRR of wheezing improved (κ = 0.58; 95% CI, 0.46-0.70). IRRs of altered mental status, capillary refill time, and grunting were less precise, with 95% CIs as wide as 0.5, likely reflecting low prevalence (18 of 502 observations [3.6%], 20 of 502 observations [4.0%], and 29 of 502 observations [5.8%], respectively). When these examination findings were noted to be present by at least 1 examiner, raw agreement ranged from 3 of 17 examiner pairs (17.6%) for capillary refill time to 6 of 23 examiner pairs (26.1%) for grunting. However, overall raw agreement of altered mental status, capillary refill time, and grunting was 93.2% to 95.2% consistent with the known limitation of κ in low prevalence conditions. Sensitivity analysis using Gwet AC1 yielded higher agreement estimates than Fleiss κ (eTable 2 in Supplement 1), with near-perfect agreement for altered mental status, prolonged capillary refill time, and grunting and substantial agreement for wheezing and retractions or accessory muscle.

Subgroup analyses stratified by disposition are shown in Figure 2. However, subgroup analyses by age, race and ethnicity, and knowledge of CXR results were underpowered. Nearly all patients had their CXRs reviewed by at least 1 evaluator before their physical examination was documented (222 of 251 patients). There was no difference in the IRR of physical examination findings between 118 patients for whom both evaluators and 104 patients for whom 1 evaluator knew the CXR results; however, neither met our prespecified sample size of 122 paired physical examinations. There were no significant differences in the IRR of physical examination findings between groups of 124 youths discharged from the ED and 128 youths who were admitted. The IRR of wheezing among youths who were hospitalized was modestly higher compared with the overall study population (κ = 0.57; 95% CI, 0.43-0.72 vs κ = 0.50; 95% CI, 0.39-0.62) (Figure 2). The IRR of grunting among discharged patients was higher but did not reach the predetermined cutoff (κ = 0.66; 95% CI, 0.21-1.00) and was observed in only 6 of 246 observations (2.4%) (Figure 2).

Figure 2. Forest Plot of Interrater Reliability (IRR) and Agreement on Physical Examination Findings in Youths Hospitalized and Discharged.

Data figure with table and forest plot of Fleiss kappa and raw agreement. Two-panel data figure. Left panel: table with four columns and nine data rows, plus one header row. Column headers, left to right: Physical examination finding; Frequency in population, No. positive observations slash total No. observations percent; Fleiss kappa, 95 percent C I; Raw agreement. Each finding has two stacked lines of values. Row 1, Altered mental status: 13 slash 256, five point one; kappa zero point four three, C I zero point zero eight to zero point seven eight; raw agreement zero point nine five. Second line: 5 slash 246, two point zero; kappa minus zero point zero two, C I minus zero point zero four to minus zero point zero zero; raw agreement zero point nine six. Row 2, Capillary refill time: 17 slash 256, six point six; kappa zero point three one, C I zero point zero zero to zero point six one; raw agreement zero point nine one. Second line: 3 slash 246, one point two; kappa minus zero point zero one, C I minus zero point zero three to zero point zero zero; raw agreement zero point nine eight. Row 3, Decreased breath sounds: 147 slash 254, fifty seven point nine; kappa zero point zero one, C I minus zero point one one to zero point one three; raw agreement zero point three seven. Second line: 105 slash 246, forty two point seven; kappa zero point two zero, C I zero point zero six to zero point three four; raw agreement zero point five five. Row 4, Fine crackles or rales: 119 slash 256, forty six point five; kappa zero point one three, C I minus zero point zero two to zero point two seven; raw agreement zero point four nine. Second line: 109 slash 246, forty four point three; kappa zero point two nine, C I zero point one five to zero point four four; raw agreement zero point six one. Row 5, Grunting: 23 slash 256, nine point zero; kappa zero point two eight, C I zero point zero two to zero point five five; raw agreement zero point eight eight. Second line: 6 slash 246, two point four; kappa zero point six six, C I zero point two one to one point one zero; raw agreement zero point nine eight. Row 6, Retractions or accessory muscle use: 120 slash 256, forty six point nine; kappa zero point five zero, C I zero point three five to zero point six five; raw agreement zero point seven five. Second line: 40 slash 248, sixteen point one; kappa zero point two eight, C I zero point zero seven to zero point five zero; raw agreement zero point eight one. Row 7, Rhonchi or coarse breath sounds: 121 slash 254, forty seven point six; kappa zero point one nine, C I zero point zero five to zero point three two; raw agreement zero point five. Second line: 68 slash 246, twenty seven point six; kappa zero point one one, C I minus zero point zero four to zero point two six; raw agreement zero point six one. Row 8, Wheezing: 66 slash 254, twenty six point zero; kappa zero point five seven, C I zero point four three to zero point seven two; raw agreement zero point eight three. Second line: 44 slash 246, seventeen point nine; kappa zero point three nine, C I zero point two two to zero point five seven; raw agreement zero point eight one. Right panel: forest plot with horizontal axis labeled I R R, 95 percent C I, ranging from minus zero point two five to one point zero zero, with ticks at minus zero point two five, zero, zero point two five, zero point five, zero point seven five, and one point zero zero. A vertical dotted reference line at approximately zero point four. Legend in upper right: blue circle k hospitalized; blue triangle raw agreement hospitalized; dark gray circle k discharged; dark gray triangle raw agreement discharged. For each finding, circles have horizontal C I lines; triangles appear as single points near the right side of the axis.

There were 128 youths hospitalized and 124 youths discharged from the emergency department.

aEach examiner’s assessment was considered a distinct observation; therefore, the total number of observations equaled twice the number of patients with paired examinations performed.

Discussion

In this multicenter, prospective cohort study of 252 youths with CAP presenting for emergency care, no individual physical examination finding demonstrated moderate or strong IRR across the study population when assessed using Fleiss κ. When we evaluated with Gwet AC1 as an alternative measure of reliability, we found that key auscultatory findings typically used to diagnose CAP,6 such as decreased breath sounds and crackles, remained inconsistently identified between different ED clinicians. Current guidelines, including those from the Infectious Diseases Society of America, recommend that clinicians rely on physical examination findings to diagnose pediatric CAP in outpatient settings.5 However, the substantial variability in how these findings are recognized across clinicians suggests that use of auscultatory findings alone is not a reliable method for diagnosing CAP in youths.

This is one of the few published studies that have focused on the IRR of physical examination findings in youths diagnosed with pneumonia.7 Previous studies evaluated the reliability of respiratory findings in populations with wheezing illnesses,9 including bronchiolitis11 and asthma8,13 as well as suspected pneumonia.7 Outcomes for this prior work were often composite clinical scores24 rather than individual examination findings, and few, if any, of these studies included features thought to be specific to the diagnosis of pneumonia, such as crackles or rales. While the asthma literature has demonstrated moderate or high interrater reliability for clinical respiratory scores used to monitor patients and guide care pathways,13 similar evaluations are largely lacking for populations without asthma. Notably, we were unable to determine whether the IRR of these examination findings varied across key subgroups, such as younger children, due to limited sample size. This is important given that the degree of chest wall ossification and adipose deposition, which vary with age, could influence the IRR of findings, such as retractions and auscultation findings.

Diagnosing pediatric CAP in outpatient, urgent care, and emergency settings often relies on assessing physical examination findings. One single-center study7 reported the IRR of physical examination findings in youths with clinically suspected CAP. The study included 128 youths with suspected CAP who underwent paired physical examinations; it found moderate IRRs for tachypnea, crackles or rales, and decreased breath sounds. Our study augments these findings in a larger, more diverse pediatric population with a definitive rather than suspected diagnosis of CAP.7 Given that this study was conducted in 7 distinct sites across the US, this analysis may also be more generalizable. Notably, physical examination findings typically used to diagnose CAP, such as crackles or rales and decreased breath sounds, had limited IRRs in the single-center analysis7 and in our multicenter study.

A known limitation of Fleiss κ is its sensitivity to low category prevalence. We considered additional measures of concordance, such as prevalence and bias-adjusted κ25 or Gwet AC1.19 These measures use different assumptions about chance agreement and may be difficult to interpret in the setting of very rare findings, where observed agreement is largely driven by concordance on the absence of the sign. For CAP diagnosis, the clinically meaningful question is whether raters agreed on the presence of a sign, not its absence. Given that our primary interest is rater agreement on signs indicative of CAP, we believe that raw agreement tables and Fleiss κ provide the most transparent and interpretable summary of the data.

We conducted simulations to ensure adequate sample size. However, uncommon but prognostically important physical examination findings, such as altered mental status, delayed capillary refill time, and grunting, demonstrated high raw agreement between raters but low IRRs when the finding was present. This reflects a known limitation of the κ statistic in low-prevalence scenarios, where even near-perfect agreement in the absence of findings does not translate into high κ scores. Our sensitivity analysis with Gwet AC1 demonstrated high agreement, likely due to large agreement regarding the absence of these findings. Despite its infrequency, altered mental status has been shown in multiple large-scale prospective studies conducted in low- and middle-income countries to be strongly predictive of poor outcomes, such as mortality and hypoxemia.26,27,28 Therefore, the presence of altered mental status and grunting should prompt heightened clinical concern and further evaluation, despite the low IRRs.

We hypothesized that there may be differences in IRRs based on whether a patient was hospitalized (due to disease severity), age (due to differences in body size and chest wall compliance), race (particularly with regard to capillary refill time, which is difficult to assess reliably in individuals with darker skin tones),20,21 and knowledge of CXR findings (due to knowledge of where in the lungs findings may be present). We found no significant difference in IRRs of physical examination findings between youths hospitalized and those discharged from the ED, suggesting that disease severity did not change reliability. Sample size limitations prevented us from conclusively making these distinctions within other subpopulations, such as by age group, race, or evaluator knowledge of radiographic findings.

Our study has several important implications. Current guidelines in the US encourage relying on patient history and clinical features, such as auscultation, to diagnose CAP in youths in outpatient settings given that CXR results infrequently influence management decisions and may not be feasible to obtain.5 However, our findings confirm that IRRs for these examination features are inadequate, such that different clinicians often reach different conclusions when examining the same patient with CAP. This variability undermines confidence in auscultation as a primary diagnostic tool in CAP and challenges the notion of its accuracy. Given that two-thirds of youths with CAP experience antibiotic-associated adverse events, improving diagnostic accuracy (which would limit downstream use of antibiotics when suspicion for CAP is lower) would have substantial benefits for patients and clinicians.29 These challenges highlight the importance of exploring new diagnostic technologies and the reproducibility of their findings. Examples of such innovations include electronic stethoscopes,30 machine learning approaches to interpret breath sounds,31 the most suitable imaging modalities,32 and the potential use of biospecimen biomarkers of disease.33,34

Limitations

This study has several limitations. Raters did not receive standardized training on terminology for describing breath sounds, which may have contributed to variability in the interpretation of physical examination findings. However, this represents clinical practice, thereby improving the generalizability of the study findings. Sample size precluded us from examining site-level differences. We did not consistently record the amount of time that passed between rater examinations; however, all examiners were instructed to conduct their evaluations within 60 minutes of each other per study protocol. An intervention, such as oxygen or albuterol in a patient with asthma, could have caused differences in rater-observed wheezing or retractions; however, we attempted to ensure that no interventions were provided between examinations. Nonetheless, it is unlikely that CAP treatment, such as antibiotics or respiratory support, would lead to the resolution of, or a change in, auscultatory findings (eg, crackles or rales) within 60 minutes. An individual rater’s knowledge of CXR results may have introduced anchoring bias, artificially inflating agreement (or disagreement, if 2 evaluators had discordant knowledge) on auscultatory findings. We attempted to overcome this limitation by conducting a subanalysis limited to youths who underwent chest radiography; however, this analysis was not sufficiently powered. Although our analytical cohort was not fully representative of the overall pedCAPS population with respect to race, ethnicity, wheezing prevalence, or disposition, we do not believe these differences meaningfully affected interrater reliability. Prior studies demonstrated high interrater reliability for scoring tools in youths with asthma; therefore, any bias introduced by the higher prevalence of wheezing in our cohort would be expected to inflate, rather than diminish IRR estimates of physical examination findings.13

Conclusions

This cohort study’s findings augment prior evidence that auscultation of decreased breath sounds, crackles, or rhonchi are insufficiently reliable and may lead to differing diagnoses of CAP between clinicians. Future studies should include validation of examination-based scoring systems used to guide management of pediatric CAP and diseases with diagnostic overlap35 and development of novel tools to help diagnose CAP across diverse care settings.

Supplement 1.

eTable 1. Comparison of Demographics, Clinical Characteristics, and Outcomes of Analytical Cohort With Overall PedCAPS Derivation Cohort

eTable 2. Gwet AC1 of Physical Examination Findings

eTable 3. Subgroup Analyses and Rationale

Supplement 2.

Nonauthor Collaborators. Pediatric Emergency Care Applied Research Network (PECARN) PedCAPS Investigators

Supplement 3.

Data Sharing Statement

References

  • 1.McDermott KW, Stocks C, Freeman WJ. Statistical Brief #242: Overview of Pediatric Emergency Department Visits, 2015. Agency for Healthcare Research and Quality; 2018. Accessed July 10, 2026. https://www.ncbi.nlm.nih.gov/books/NBK526418/ [PubMed] [Google Scholar]
  • 2.Neuman MI, Shah SS, Shapiro DJ, Hersh AL. Emergency department management of childhood pneumonia in the United States prior to publication of national guidelines. Acad Emerg Med. 2013;20(3):240-246. doi: 10.1111/acem.12088 [DOI] [PubMed] [Google Scholar]
  • 3.Florin TA, French B, Zorc JJ, Alpern ER, Shah SS. Variation in emergency department diagnostic testing and disposition outcomes in pneumonia. Pediatrics. 2013;132(2):237-244. doi: 10.1542/peds.2013-0179 [DOI] [PubMed] [Google Scholar]
  • 4.Parikh K, Hall M, Blaschke AJ, et al. Aggregate and hospital-level impact of national guidelines on diagnostic resource utilization for children with pneumonia at children’s hospitals. J Hosp Med. 2016;11(5):317-323. doi: 10.1002/jhm.2534 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Bradley JS, Byington CL, Shah SS, et al. ; Pediatric Infectious Diseases Society and the Infectious Diseases Society of America . The management of community-acquired pneumonia in infants and children older than 3 months of age: clinical practice guidelines by the Pediatric Infectious Diseases Society and the Infectious Diseases Society of America. Clin Infect Dis. 2011;53(7):e25-e76. doi: 10.1093/cid/cir531 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Shah SN, Bachur RG, Simel DL, Neuman MI. Does this child have pneumonia: the rational clinical examination systematic review. JAMA. 2017;318(5):462-471. doi: 10.1001/jama.2017.9039 [DOI] [PubMed] [Google Scholar]
  • 7.Florin TA, Ambroggio L, Brokamp C, et al. Reliability of examination findings in suspected community-acquired pneumonia. Pediatrics. 2017;140(3):e20170310. doi: 10.1542/peds.2017-0310 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Angelilli ML, Thomas R. Inter-rater evaluation of a clinical scoring system in children with asthma. Ann Allergy Asthma Immunol. 2002;88(2):209-214. doi: 10.1016/S1081-1206(10)61998-9 [DOI] [PubMed] [Google Scholar]
  • 9.Bekhof J, Reimink R, Bartels IM, Eggink H, Brand PLP. Large observer variation of clinical assessment of dyspnoeic wheezing children. Arch Dis Child. 2015;100(7):649-653. doi: 10.1136/archdischild-2014-307143 [DOI] [PubMed] [Google Scholar]
  • 10.Ramgopal S, Saper JK, Rudloff JR, et al. Interrater reliability of pediatric respiratory auscultation findings. Hosp Pediatr. 2025;15(9):e431-e435. doi: 10.1542/hpeds.2025-008510 [DOI] [PubMed] [Google Scholar]
  • 11.Gajdos V, Beydon N, Bommenel L, et al. Inter-observer agreement between physicians, nurses, and respiratory therapists for respiratory clinical evaluation in bronchiolitis. Pediatr Pulmonol. 2009;44(8):754-762. doi: 10.1002/ppul.21016 [DOI] [PubMed] [Google Scholar]
  • 12.Hay AD, Wilson A, Fahey T, Peters TJ. The inter-observer agreement of examining pre-school children with acute cough: a nested study. BMC Fam Pract. 2004;5:4. doi: 10.1186/1471-2296-5-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.McLaughlin P, Banuelos RC, Camp EA, Kancharla V, Sampayo EM. The Clinical Respiratory Score: investigating the reliability of an asthma scoring tool across a multidisciplinary team. J Asthma. 2022;59(10):1915-1922. doi: 10.1080/02770903.2021.1978481 [DOI] [PubMed] [Google Scholar]
  • 14.Florin TA, Reeder R, Ambroggio L, et al. ; Pediatric Emergency Care Applied Research Network (PECARN) PedCAPS Investigators . Derivation and validation of the Pediatric Community-Acquired Pneumonia Severity (PedCAPS) score: A prospective cohort study. J Hosp Med. 2026;21(5):590-598. doi: 10.1002/jhm.70220 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Pediatric Emergency Care Applied Research Network . The Pediatric Emergency Care Applied Research Network (PECARN): rationale, development, and first steps. Acad Emerg Med. 2003;10(6):661-668. doi: 10.1111/j.1553-2712.2003.tb00053.x [DOI] [PubMed] [Google Scholar]
  • 16.Harris PA, Taylor R, Thielke R, Payne J, Gonzalez N, Conde JG. Research electronic data capture (REDCap)—a metadata-driven methodology and workflow process for providing translational research informatics support. J Biomed Inform. 2009;42(2):377-381. doi: 10.1016/j.jbi.2008.08.010 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Zapf A, Castell S, Morawietz L, Karch A. Measuring inter-rater reliability for nominal data—which coefficients and confidence intervals are appropriate? BMC Med Res Methodol. 2016;16(1):93. doi: 10.1186/s12874-016-0200-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.SAS Institute Inc . Sample 25006: compute estimates and tests of agreement among multiple raters. Accessed October 6, 2025. https://support.sas.com/kb/25/006.html
  • 19.Vach W, Gerke O. Gwet’s AC1 is not a substitute for Cohen’s kappa—a comparison of basic properties. MethodsX. 2023;10:102212. doi: 10.1016/j.mex.2023.102212 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Polfer EM, Zimmerman RM, Tefera E, Katz RD, Higgins JP, Means KR Jr. The effect of skin pigmentation on determination of limb ischemia. J Hand Surg Am. 2018;43(1):24-32.e1. doi: 10.1016/j.jhsa.2017.09.002 [DOI] [PubMed] [Google Scholar]
  • 21.Gakwaya RB, Zonfrillo MR, Ellison AM, Holmes JF, Kuppermann N, Ruest SM. Racial differences in the identification of seat belt signs among pediatric motor vehicle crash occupants. Pediatr Emerg Care. 2025;41(11):e178-e181. doi: 10.1097/PEC.0000000000003462 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Landis JR, Koch GG. The measurement of observer agreement for categorical data. Biometrics. 1977;33(1):159-174. doi: 10.2307/2529310 [DOI] [PubMed] [Google Scholar]
  • 23.Vanbelle S. Asymptotic variability of (multilevel) multirater kappa coefficients. Stat Methods Med Res. 2019;28(10-11):3012-3026. doi: 10.1177/0962280218794733 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Balaguer M, Alejandre C, Vila D, et al. Bronchiolitis score of Sant Joan de Déu: BROSJOD score, validation and usefulness. Pediatr Pulmonol. 2017;52(4):533-539. doi: 10.1002/ppul.23546 [DOI] [PubMed] [Google Scholar]
  • 25.Hoehler FK. Bias and prevalence effects on kappa viewed in terms of sensitivity and specificity. J Clin Epidemiol. 2000;53(5):499-503. doi: 10.1016/S0895-4356(99)00174-2 [DOI] [PubMed] [Google Scholar]
  • 26.Hooli S, Colbourn T, Lufesi N, et al. Predicting hospitalised paediatric pneumonia mortality risk: an external validation of RISC and mRISC, and local tool development (RISC-Malawi) from Malawi. PLoS One. 2016;11(12):e0168126. doi: 10.1371/journal.pone.0168126 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Gordon B, Nyiro JU, Nair H, et al. External validation of pediatric pneumonia and bronchiolitis risk scores to predict mortality in children hospitalized in Kenya: a retrospective cohort study. J Infect Dis. 2026;233(1):e230-e238. doi: 10.1093/infdis/jiaf377 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Schuh HB, Hooli S, Ahmed S, et al. Clinical hypoxemia score for outpatient child pneumonia care lacking pulse oximetry in Africa and South Asia. Front Pediatr. 2023;11:1233532. doi: 10.3389/fped.2023.1233532 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Williams DJ, Creech CB, Walter EB, et al. ; The DMID 14-0079 Study Team . Short- vs standard-course outpatient antibiotic therapy for community-acquired pneumonia in children: the SCOUT-CAP randomized clinical trial. JAMA Pediatr. 2022;176(3):253-261. doi: 10.1001/jamapediatrics.2021.5547 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Sola J, Braun F, Muntane E, et al. Towards an unsupervised device for the diagnosis of childhood pneumonia in low resource settings: automatic segmentation of respiratory sounds. Annu Int Conf IEEE Eng Med Biol Soc. 2016;2016:283-286. doi: 10.1109/EMBC.2016.7590695 [DOI] [PubMed] [Google Scholar]
  • 31.Park JS, Park SY, Moon JW, Kim K, Suh DI. Artificial intelligence models for pediatric lung sound analysis: systematic review and meta-analysis. J Med Internet Res. 2025;27:e66491. doi: 10.2196/66491 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Balk DS, Lee C, Schafer J, et al. Lung ultrasound compared to chest x-ray for diagnosis of pediatric pneumonia: a meta-analysis. Pediatr Pulmonol. 2018;53(8):1130-1139. doi: 10.1002/ppul.24020 [DOI] [PubMed] [Google Scholar]
  • 33.Ambroggio L, Florin TA, Williamson K, et al. Urine metabolites of suspected community-acquired pneumonia. J Infect Dis. 2025;232(2):370-380. doi: 10.1093/infdis/jiaf072 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Florin TA, Ambroggio L, Brokamp C, et al. Biomarkers and disease severity in children with community-acquired pneumonia. Pediatrics. 2020;145(6):e20193728. doi: 10.1542/peds.2019-3728 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Sheikh Z, Potter E, Li Y, et al. ; PROMISE Investigators . Validity of clinical severity scores for respiratory syncytial virus: a systematic review. J Infect Dis. 2024;229(suppl 1):S8-S17. doi: 10.1093/infdis/jiad436 [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplement 1.

eTable 1. Comparison of Demographics, Clinical Characteristics, and Outcomes of Analytical Cohort With Overall PedCAPS Derivation Cohort

eTable 2. Gwet AC1 of Physical Examination Findings

eTable 3. Subgroup Analyses and Rationale

Supplement 2.

Nonauthor Collaborators. Pediatric Emergency Care Applied Research Network (PECARN) PedCAPS Investigators

Supplement 3.

Data Sharing Statement


Articles from JAMA Network Open are provided here courtesy of American Medical Association

RESOURCES