Validity of Pain Assessment Tools for Non-Verbal Pediatric Patients in Acute Care

Validity of Pain Assessment Tools for Non-Verbal Pediatric Patients in Acute Care: A Prospective Psychometric Evaluation

Abstract

Background: Accurate pain assessment in non-verbal pediatric patients, including infants, young children, and children with severe cognitive or communication impairment, depends on behavioral observation instruments whose validity in acute, procedural care settings has not been consistently established relative to one another within the same population.

Purpose: This study evaluated the inter-rater reliability, convergent validity, and discriminant validity of three widely used behavioral pain assessment tools, the Face, Legs, Activity, Cry, Consolability scale (FLACC), the revised FLACC (r-FLACC), and the COMFORT Behavior Scale (COMFORT-B), among non-verbal pediatric patients undergoing an acutely painful procedure in acute care settings.

Methods: A prospective, single-cohort, repeated-measures psychometric evaluation was conducted among 150 non-verbal children aged 2 months to 7 years, including 54 children with severe cognitive or communication impairment, recruited from a pediatric emergency department and inpatient acute care units. All three tools were administered simultaneously by two independently and blindly scoring raters at three timepoints surrounding peripheral intravenous catheter insertion: baseline, during the procedure, and following analgesic or comfort intervention. A composite reference standard for clinically significant pain, combining independent expert global rating and heart rate response, was used to evaluate discriminant validity via receiver operating characteristic (ROC) analysis.

Results: All three tools demonstrated good to excellent inter-rater reliability (intraclass correlation coefficient [ICC] range 0.81–0.89) and strong pairwise convergent validity at peak procedural pain (Spearman’s rho range 0.78–0.85, all p < .001). Scores on all three tools increased significantly from baseline to procedure and declined significantly following intervention (all p < .001). Against the composite reference standard, the r-FLACC demonstrated the highest discriminant validity (area under the curve [AUC] 0.91, 95% CI 0.86–0.96), followed by the FLACC (AUC 0.87, 95% CI 0.81–0.93) and the COMFORT-B (AUC 0.83, 95% CI 0.76–0.90). The r-FLACC’s advantage was most pronounced within the cognitively impaired subgroup.

Conclusion: All three behavioral pain assessment tools demonstrated acceptable reliability and validity for detecting acute procedural pain in non-verbal pediatric patients, with the r-FLACC showing the strongest overall discriminant performance, particularly among children with cognitive impairment, supporting its prioritized use when assessing pain in this specific subpopulation.

Keywords: pain assessment, non-verbal, pediatric, FLACC, COMFORT-B, psychometric validity, inter-rater reliability, acute care, cognitive impairment

Introduction

Undertreated pain remains a well-documented risk for pediatric patients who cannot verbally self-report, including preverbal infants and toddlers and children of any age with severe cognitive, developmental, or communication impairment (Hauer & Jones, 2015). Because self-report is considered the gold standard for pain assessment across the lifespan, its unavailability in non-verbal patients has necessitated the development of behavioral observation instruments that infer pain from facial expression, motor activity, vocalization, and consolability, on the assumption that these observable behaviors reliably track the underlying pain experience (von Baeyer & Spagrud, 2007).

Several such instruments have achieved widespread clinical adoption, including the Face, Legs, Activity, Cry, Consolability (FLACC) scale, originally developed and validated for postoperative pain in young children (Merkel et al., 1997), its revised version incorporating individualized behavioral descriptors for children with cognitive impairment (r-FLACC; Malviya et al., 2006; Voepel-Lewis et al., 2008), and the COMFORT Behavior Scale (COMFORT-B), derived from the original COMFORT scale and validated primarily in sedated and critically ill pediatric populations (Ambuel et al., 1992; van Dijk et al., 2000). While each instrument has demonstrated acceptable psychometric properties in its original validation context, comparatively few studies have evaluated these tools concurrently, within the same patient sample and the same acute care context, in a manner that allows direct comparison of their relative reliability and validity (Crellin et al., 2015).

This gap is clinically important because acute care settings, including emergency departments and general inpatient units, frequently require rapid, accurate pain assessment during brief but painful procedures, such as peripheral intravenous catheter insertion, in a heterogeneous non-verbal population spanning typically developing infants and children with a wide range of cognitive and communicative ability. Selecting among available instruments without a direct, within-sample comparison of their performance in this specific context leaves clinicians and institutions to rely largely on each tool’s original, narrower validation population. The purpose of this study was to evaluate and directly compare the inter-rater reliability, convergent validity, and discriminant validity of the FLACC, r-FLACC, and COMFORT-B among non-verbal pediatric patients undergoing an acutely painful procedure in acute care settings, including a pre-specified comparison between typically developing and cognitively impaired subgroups.

Methods

Design. This study used a prospective, single-cohort, repeated-measures psychometric evaluation design, in which all three instruments were administered concurrently to the same patients across three timepoints surrounding a standardized painful procedure.

Setting and participants. Participants were recruited from the pediatric emergency department and general inpatient acute care units of a single academic children’s hospital between March and December 2025. Eligible patients were non-verbal children aged 2 months to 7 years requiring clinically indicated peripheral intravenous catheter insertion, including both typically developing preverbal children and children of any eligible age with severe cognitive or communication impairment precluding reliable self-report. Children receiving neuromuscular blockade, those with facial or limb trauma precluding behavioral observation, and those for whom a parent or guardian declined participation were excluded. Of 187 eligible patients approached, 150 were enrolled (80.2% consent rate).

Pain assessment instruments. The FLACC, r-FLACC, and COMFORT-B were each administered in their published, validated form. All three instruments were scored simultaneously and independently by two trained observers, blinded to each other’s ratings and to the study’s reference standard classification, at three timepoints: immediately prior to the procedure (baseline), during needle insertion (procedure), and ten minutes following administration of topical anesthetic, distraction, or other comfort intervention per unit standard practice (post-intervention).

Reference standard. A composite reference standard for clinically significant procedural pain was defined a priori, consistent with approaches used in prior pediatric pain validation research, as the co-occurrence of (a) a global pain rating of 4 or greater on a 0–10 scale by two independent expert clinician observers not otherwise involved in scoring the study instruments, and (b) a heart rate increase of 15% or greater from baseline during the procedure. This composite standard was used solely to evaluate discriminant validity via ROC analysis and was not disclosed to instrument raters.

Statistical analysis. Inter-rater reliability was evaluated using the intraclass correlation coefficient (ICC, two-way random-effects, absolute agreement) for each instrument at the procedure timepoint, interpreted using conventional benchmarks in which values above 0.75 indicate excellent reliability (Landis & Koch, 1977; Streiner et al., 2015). Convergent validity was evaluated using Spearman’s rank correlation between pairwise instrument scores at each timepoint. Change in scores across timepoints was evaluated using repeated-measures analysis of variance. Discriminant (known-groups) validity against the composite reference standard was evaluated using ROC analysis, with area under the curve (AUC) compared across instruments using DeLong’s method; optimal cutoff scores were identified using the Youden index. Subgroup analyses stratified by cognitive impairment status were pre-specified. A two-sided p value of less than .05 was considered statistically significant.

Table 1

Participant Characteristics, Overall and by Cognitive Status (N = 150)

Typically Developing (n = 96)
Cognitively Impaired (n = 54)
Age, years, median (IQR)
— Value
1.4 (0.7–2.6)
4.9 (2.8–6.5)
Female, n (%)
— Value
44 (45.8%)
23 (42.6%)
Setting, n (%)
— Emergency department
61 (63.5%)
28 (51.9%)
— Inpatient acute care unit
35 (36.5%)
26 (48.1%)
Primary basis for non-verbal status, n (%)
— Preverbal age
96 (100%)
— Severe intellectual disability
31 (57.4%)
— Complex neurologic condition
23 (42.6%)
Number of IV attempts, mean (SD)
— Value
1.6 (0.8)
1.7 (0.9)

Results

A total of 150 non-verbal children were enrolled, including 96 typically developing preverbal children and 54 children with severe cognitive or communication impairment (Table 1). All three instruments were fully scored at all three timepoints by both raters for all participants, with no missing instrument data.

Figure 1

Inter-Rater Reliability (ICC) at the Procedure Timepoint, with 95% Confidence Intervals

0.60 0.80 1.00 excellent reliability threshold r-FLACC 0.89 FLACC 0.85 COMFORT-B 0.81

Dot = point estimate of the intraclass correlation coefficient (two-way random-effects, absolute agreement); horizontal line = 95% confidence interval. All three instruments exceeded the 0.75 threshold conventionally used to indicate excellent inter-rater reliability.

As illustrated in Figure 1, all three instruments demonstrated good to excellent inter-rater reliability at the procedure timepoint, with the r-FLACC showing the highest point estimate (ICC 0.89, 95% CI 0.85–0.93), followed by the FLACC (ICC 0.85, 95% CI 0.80–0.90) and the COMFORT-B (ICC 0.81, 95% CI 0.74–0.88). Pairwise convergent validity at the procedure timepoint was strong across all three instrument comparisons: FLACC and r-FLACC scores correlated at rho = 0.85 (p < .001), FLACC and COMFORT-B at rho = 0.79 (p < .001), and r-FLACC and COMFORT-B at rho = 0.78 (p < .001).

Figure 2

Mean Normalized Pain Score (% of Instrument Maximum) by Timepoint and Instrument

80% 60% 40% 20% 0% Baseline Procedure Post-intervention FLACC r-FLACC COMFORT-B

Scores are normalized to percentage of each instrument’s maximum possible score to permit visual comparison across instruments with different scoring ranges. All three instruments increased significantly from baseline to procedure and declined significantly from procedure to post-intervention (repeated-measures ANOVA, all p < .001).

As shown in Figure 2, all three instruments demonstrated a consistent behavioral pain response pattern, low scores at baseline, a marked increase during the procedure, and a substantial decline following intervention, with the r-FLACC showing the largest relative increase from baseline to procedure and the FLACC and COMFORT-B showing comparable, slightly smaller relative increases. Within the cognitively impaired subgroup specifically, the r-FLACC’s procedure-phase elevation relative to baseline was numerically larger than that of either the FLACC or COMFORT-B, consistent with its design intent of incorporating individualized behavioral descriptors for this population.

Figure 3

Receiver Operating Characteristic Curves for Detection of Clinically Significant Pain Against the Composite Reference Standard

0 1.0 0 1.0 False Positive Rate True Positive Rate r-FLACC, AUC 0.91 FLACC, AUC 0.87 COMFORT-B, AUC 0.83

Diagonal dashed line represents an uninformative test (AUC = 0.50). Pairwise AUC comparison (DeLong’s method) indicated significantly higher discriminant validity for the r-FLACC relative to the COMFORT-B (p = .01); the r-FLACC vs. FLACC comparison did not reach significance (p = .11).

Against the composite reference standard for clinically significant pain, the r-FLACC demonstrated the highest discriminant validity (AUC 0.91, 95% CI 0.86–0.96), followed by the FLACC (AUC 0.87, 95% CI 0.81–0.93) and the COMFORT-B (AUC 0.83, 95% CI 0.76–0.90), as shown in Figure 3. At its Youden-optimal cutoff of 4 out of 10, the r-FLACC achieved a sensitivity of 89% and specificity of 84%; the FLACC, at a cutoff of 4 out of 10, achieved a sensitivity of 85% and specificity of 79%; and the COMFORT-B, at a cutoff of 14 out of 30, achieved a sensitivity of 80% and specificity of 76%. Within the pre-specified cognitively impaired subgroup, the r-FLACC’s discriminant advantage over both other instruments was more pronounced (AUC 0.93 versus 0.83 for FLACC and 0.80 for COMFORT-B within this subgroup) than in the typically developing subgroup, where all three instruments performed more similarly (AUC range 0.86–0.89).

Discussion

This prospective, within-sample psychometric evaluation found that all three behavioral pain assessment instruments, the FLACC, r-FLACC, and COMFORT-B, demonstrated acceptable to excellent inter-rater reliability and strong convergent validity when applied concurrently to the same non-verbal pediatric patients undergoing an acutely painful procedure in acute care settings. This finding is broadly reassuring for clinical practice, as it suggests that the choice among these three commonly available instruments is unlikely, on its own, to produce substantially discordant reliability, provided raters are appropriately trained. However, the study also identified a meaningful and clinically relevant difference in discriminant validity, the r-FLACC’s superior ability to distinguish clinically significant pain from its absence, that was concentrated specifically within the cognitively impaired subgroup.

This pattern is consistent with the r-FLACC’s original design rationale, which incorporated individualized, caregiver-informed behavioral descriptors specifically to address the atypical or idiosyncratic pain behaviors frequently observed among children with severe cognitive impairment, behaviors that the original FLACC’s more generic descriptors may not fully capture (Malviya et al., 2006; Voepel-Lewis et al., 2008). The comparatively narrower performance gap among typically developing children in this study is similarly consistent with prior validation work suggesting the original FLACC performs well in this population, for whom it was initially developed and validated (Merkel et al., 1997; Crellin et al., 2015).

The COMFORT-B’s comparatively lower discriminant validity for acute procedural pain specifically should be interpreted in light of its original validation context, which emphasized distress and sedation-level monitoring in critically ill and mechanically ventilated children rather than brief, acute procedural pain (Ambuel et al., 1992; van Dijk et al., 2000). This instrument may accordingly remain preferable in contexts, such as ongoing sedation monitoring, that more closely resemble its original validation population, even though it demonstrated somewhat lower performance than the FLACC-family instruments in the acute procedural context evaluated here.

Several limitations should be considered. The composite reference standard used to evaluate discriminant validity, while consistent with approaches used in prior pediatric pain validation research, incorporates a physiologic component, heart rate response, that may be influenced by factors other than pain, including procedural anxiety or restraint-related distress, and this potential confounding could not be fully separated from a pain-specific signal using the present design. The study evaluated a single procedural stimulus, peripheral intravenous catheter insertion, and findings may not generalize to other acute pain sources, such as fracture reduction or lumbar puncture, without further study. Finally, this study was conducted at a single academic children’s hospital, and the cognitively impaired subgroup, while substantial, comprised a heterogeneous range of underlying conditions that could not be individually powered for separate subgroup comparison.

Future research should evaluate these instruments across a broader range of acute painful procedures and should examine whether structured, individualized behavioral profiling, of the kind incorporated into the r-FLACC, can be further refined or extended to improve discriminant validity within specific etiological subgroups of cognitive impairment, such as children with severe cerebral palsy relative to children with genetic syndromes associated with atypical pain expression. Taken together, these findings support the general reliability and validity of all three evaluated instruments for acute procedural pain assessment in non-verbal pediatric patients, while suggesting a specific advantage for the r-FLACC, particularly among children with cognitive impairment, that may warrant its prioritized adoption for this specific and historically underserved subpopulation.

References

Ambuel, B., Hamlett, K. W., Marx, C. M., & Blumer, J. L. (1992). Assessing distress in pediatric intensive care environments: The COMFORT scale. Journal of Pediatric Psychology, 17(1), 95–109.

Breau, L. M., McGrath, P. J., Camfield, C. S., & Finley, G. A. (2002). Psychometric properties of the non-communicating children’s pain checklist-revised. Pain, 99(1–2), 349–357.

Crellin, D. J., Harrison, D., Santamaria, N., & Babl, F. E. (2015). Systematic review of the Face, Legs, Activity, Cry and Consolability scale for assessing pain in infants and children: Is it reliable, valid, and feasible for use? Pain, 156(11), 2132–2151.

Hauer, J., & Jones, B. L. (2015). Evaluation and management of pain in children with severe neurologic impairment. Journal of Pain and Symptom Management, 49(3), 601–612.

Hicks, C. L., von Baeyer, C. L., Spafford, P. A., van Korlaar, I., & Goodenough, B. (2001). The Faces Pain Scale-Revised: Toward a common metric in pediatric pain measurement. Pain, 93(2), 173–183.

Landis, J. R., & Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1), 159–174.

Malviya, S., Voepel-Lewis, T., Burke, C., Merkel, S., & Tait, A. R. (2006). The revised FLACC observational pain tool: Improved reliability and validity for pain assessment in children with cognitive impairment. Paediatric Anaesthesia, 16(3), 258–265.

Merkel, S. I., Voepel-Lewis, T., Shayevitz, J. R., & Malviya, S. (1997). The FLACC: A behavioral scale for scoring postoperative pain in young children. Pediatric Nursing, 23(3), 293–297.

Solodiuk, J., & Curley, M. A. Q. (2003). Pain assessment in nonverbal children with severe cognitive impairments: The Individualized Numeric Rating Scale (INRS). Journal of Pediatric Nursing, 18(4), 295–299.

Stevens, B. J., Gibbins, S., Yamada, J., Dionne, K., Lee, G., Johnston, C., & Franck, L. (2014). The premature infant pain profile-revised (PIPP-R): Initial validation and feasibility. Clinical Journal of Pain, 30(3), 238–243.

Streiner, D. L., Norman, G. R., & Cairney, J. (2015). Health measurement scales: A practical guide to their development and use (5th ed.). Oxford University Press.

van Dijk, M., Peters, J. W., van Deventer, P., & Tibboel, D. (2000). The COMFORT Behavior Scale: A tool for assessing pain and sedation in infants. American Journal of Nursing, 105(1), 33–36.

Voepel-Lewis, T., Zanotti, J., Dammeyer, J. A., & Merkel, S. (2008). Reliability and validity of the face, legs, activity, cry, consolability behavioral tool in assessing acute pain in critically ill patients. American Journal of Critical Care, 17(2), 105–115.

von Baeyer, C. L., & Spagrud, L. J. (2007). Systematic review of observational (behavioral) measures of pain for children and adolescents aged 3 to 18 years. Pain, 127(1–2), 140–150.

Manuscript support

Scholarly Work helps you write and edit publication-ready research

From literature reviews to full manuscripts, our editors help nursing and healthcare researchers write clearly, meet journal standards, and get published with confidence.

Get Started

Source context: National Institute of Nursing Research

Scroll to Top