Triage Accuracy and Decision-Making Consistency Among Emergency Department Nurses: A Vignette-Based Inter-Rater Reliability Study
Abstract
Background: Triage acuity assignment is among the highest-consequence decisions an emergency department nurse makes within the first minutes of a patient encounter, yet the consistency of that decision, both between nurses and against an expert-defined standard, is not routinely measured in most emergency departments despite its direct relevance to patient safety.
Purpose: This study evaluated inter-rater reliability and accuracy of Emergency Severity Index (ESI) triage assignment among emergency department nurses using a standardized set of clinical vignettes, and examined the association of nurse experience level and ESI-specific training with triage accuracy.
Methods: Sixty-eight emergency department nurses across four hospitals independently assigned an ESI acuity level (1–5) to each of 30 standardized clinical vignettes, developed and assigned a gold-standard ESI level by an expert panel of emergency physicians and triage-certified nurse educators through structured consensus. Inter-rater reliability among nurses was assessed using Fleiss’ kappa; accuracy against the gold standard was assessed using weighted kappa, exact agreement percentage, and rates of under-triage and over-triage. Subgroup comparisons by nurse experience level and ESI training status were conducted.
Results: Overall inter-rater agreement among nurses was substantial (Fleiss’ kappa 0.68, 95% CI 0.63–0.73). Agreement with the expert gold standard was similarly substantial (weighted kappa 0.71, 95% CI 0.66–0.76), with an exact agreement rate of 74.3%, an under-triage rate of 14.2%, and an over-triage rate of 11.5%. Disagreement was concentrated at the ESI level 2/3 boundary, which accounted for 61.4% of all under-triage errors. Expert nurses (>10 years of emergency experience) had significantly higher exact agreement (81.7%) and lower under-triage rate (8.9%) than novice nurses with less than 2 years of experience (64.2% exact agreement, 21.6% under-triage; p < .001). Nurses who had completed formal ESI-specific training had significantly higher weighted kappa against the gold standard than those who had not (0.77 vs. 0.61, p < .001).
Conclusion: Emergency department nurses demonstrated substantial but imperfect inter-rater reliability and accuracy in ESI triage assignment, with disagreement concentrated at the clinically consequential level 2/3 boundary and with experience and formal ESI training each independently associated with higher accuracy, supporting targeted, boundary-specific training as a strategy for improving triage decision-making consistency.
Keywords: triage accuracy, Emergency Severity Index, inter-rater reliability, emergency nursing, clinical decision-making, Fleiss’ kappa, under-triage, vignette-based assessment
Introduction
Triage acuity assignment, the process by which an emergency department nurse determines the urgency of a patient’s condition and the resources it is likely to require, is among the highest-consequence clinical judgments made within an emergency department, directly shaping how quickly a patient is seen and which diagnostic and treatment resources are mobilized (Fernandes et al., 2005). The Emergency Severity Index (ESI), a five-level triage acuity scale incorporating both illness severity and anticipated resource need, is among the most widely adopted triage instruments internationally and has demonstrated reasonable reliability and predictive validity across a substantial body of prior research (Gilboy et al., 2011; Wuerz et al., 2000).
Despite this general validation, prior research examining inter-rater reliability of ESI assignment has consistently identified a specific area of vulnerability: the boundary between ESI level 2, indicating a high-risk situation warranting immediate physician evaluation, and ESI level 3, indicating a stable but resource-intensive presentation, has been repeatedly documented as the most difficult acuity distinction for triage nurses to apply consistently, with disagreement at this boundary carrying particular patient safety significance given the clinical difference between these two acuity categories (Tanabe et al., 2004; Worster et al., 2004). More recent multicenter research has continued to document meaningful, though generally moderate to substantial, variability in ESI accuracy and reliability across different institutions, triage nurse experience levels, and training backgrounds (Mistry et al., 2018; Jordi et al., 2015).
Given the direct patient safety relevance of triage consistency, and the continued evolution of emergency department staffing, training, and nurse experience distribution across institutions, ongoing, institution-level assessment of triage reliability and accuracy, rather than reliance on the general validation literature alone, remains valuable for identifying specific, addressable gaps in local triage practice. The purpose of this study was to evaluate inter-rater reliability and accuracy of ESI triage assignment among emergency department nurses using a standardized set of clinical vignettes, and to examine the association of nurse experience level and ESI-specific training with triage accuracy.
Methods
Design. This study used a vignette-based inter-rater reliability and diagnostic agreement design, in which all participating nurses independently assigned a triage acuity level to the same standardized set of clinical vignettes, allowing both between-nurse agreement and agreement against an expert-defined gold standard to be calculated from a controlled, common stimulus set rather than from variable, real-world patient encounters.
Vignette development and gold standard. Thirty clinical vignettes were developed by the research team from de-identified, anonymized emergency department presentations, each including a chief complaint, relevant history, vital signs, and general appearance description consistent with the information available to a nurse at the point of triage. Vignettes were selected to span all five ESI acuity levels and to include a deliberately higher proportion of presentations near the clinically ambiguous ESI level 2/3 boundary, consistent with this boundary’s documented vulnerability in prior literature. A gold-standard ESI level for each vignette was established by an expert panel comprising three emergency physicians and two triage-certified nurse educators, who independently assigned an ESI level to each vignette and then resolved any initial disagreement through structured group discussion until full panel consensus was reached.
Participants. Sixty-eight emergency department registered nurses across four hospitals within a single regional health system participated. Eligible nurses had a minimum of six months of triage experience. Participants were categorized by experience level as novice (less than 2 years of emergency nursing experience), experienced (2 to 10 years), or expert (more than 10 years), and by whether they had completed formal, standardized ESI-specific training beyond general orientation.
Procedure. Each participant independently reviewed and assigned an ESI acuity level (1 through 5) to each of the 30 vignettes, presented in randomized order, without access to the gold-standard classification or to other participants’ responses.
Statistical analysis. Inter-rater reliability among the 68 participating nurses was assessed using Fleiss’ kappa, appropriate for evaluating agreement among more than two raters assigning cases to categorical, ordered levels (Fleiss, 1971). Accuracy against the expert gold standard was assessed using weighted kappa, exact agreement percentage, under-triage rate (assigned ESI level indicating lower acuity than the gold standard), and over-triage rate (assigned ESI level indicating higher acuity than the gold standard). Kappa values were interpreted using conventional benchmarks in which values of 0.61 to 0.80 indicate substantial agreement and values above 0.80 indicate near-perfect agreement (Landis & Koch, 1977). Subgroup comparisons of exact agreement and under-triage rate by experience level were conducted using chi-square tests; weighted kappa comparison by ESI training status was conducted using a bootstrap-based comparison of dependent kappa coefficients. A two-sided p value of less than .05 was considered statistically significant.
Table 1
Participant Characteristics (N = 68)
Results
Among 68 participating nurses (Table 1) completing 30 vignettes each (2,040 total triage assignments), overall inter-rater agreement was substantial (Fleiss’ kappa 0.68, 95% CI 0.63–0.73). Agreement against the expert-derived gold standard was similarly substantial (weighted kappa 0.71, 95% CI 0.66–0.76), with an overall exact agreement rate of 74.3%. Under-triage occurred in 14.2% of assignments and over-triage in 11.5%. The distribution of nurse-assigned versus gold-standard ESI levels is shown in Figure 1.
Figure 1
Confusion Matrix: Nurse-Assigned ESI Level vs. Expert Gold-Standard ESI Level (% of All 2,040 Assignments)
Diagonal cells (darkest shading) represent exact agreement between nurse-assigned and gold-standard ESI level. Off-diagonal cells below the diagonal (nurse assigned a numerically higher ESI level, indicating lower acuity) represent under-triage; cells above the diagonal represent over-triage. The largest single disagreement cell is ESI 3 assigned by nurses when the gold standard was ESI 2 (5.9%), consistent with the level 2/3 boundary as the primary source of triage error.
Disagreement was concentrated at the ESI level 2/3 boundary specifically: the single most common error, nurses assigning ESI 3 to a vignette with a gold-standard classification of ESI 2, accounted for 5.9% of all assignments and 61.4% of all under-triage errors overall. Exact agreement and under-triage rate differed significantly by nurse experience level, as shown in Figure 2.
Figure 2
Exact Agreement and Under-Triage Rate with Gold Standard, by Nurse Experience Level
Both exact agreement and under-triage rate differed significantly across experience levels (chi-square test, both p < .001), with expert nurses showing the highest accuracy and lowest under-triage rate.
Weighted kappa against the gold standard also differed significantly by ESI training status, as shown in Figure 3: nurses who had completed formal ESI-specific training achieved significantly higher agreement with the gold standard than nurses who had not (weighted kappa 0.77 vs. 0.61, p < .001).
Figure 3
Weighted Kappa Against Gold Standard, Overall and by Subgroup
Dots represent weighted kappa point estimates comparing nurse-assigned to gold-standard ESI level; horizontal lines represent 95% confidence intervals. Weighted kappa comparison between ESI-trained and non-trained subgroups was statistically significant (p < .001).
Discussion
This vignette-based inter-rater reliability study found that emergency department nurses demonstrated substantial, but clinically imperfect, agreement both with each other and with an expert-derived gold standard in Emergency Severity Index triage assignment, with an overall weighted kappa of 0.71, consistent with prior multicenter ESI reliability research reporting kappa values generally in the substantial range (Mistry et al., 2018; Jordi et al., 2015). The specific concentration of disagreement at the ESI level 2/3 boundary, accounting for the majority of under-triage errors in this study, directly replicates a pattern identified in earlier ESI reliability research conducted with earlier versions of the instrument, suggesting that this specific boundary has remained a persistent source of triage difficulty across ESI versions and practice eras rather than reflecting a limitation specific to an earlier tool iteration (Tanabe et al., 2004; Worster et al., 2004).
The clinical significance of this concentration at the level 2/3 boundary should not be understated: under-triage from level 2 to level 3 delays a patient who genuinely requires immediate physician evaluation into a lower-priority queue, representing precisely the kind of triage error most directly linked to preventable delay in time-sensitive care. This finding suggests that generic ESI training addressing the instrument broadly may be less effective than training specifically structured around this documented boundary, using deliberately challenging practice cases similar to those developed for this study’s vignette set.
Both nurse experience and formal ESI-specific training were independently associated with higher triage accuracy, with expert nurses achieving the highest observed weighted kappa (0.79) and lowest under-triage rate (8.9%) among the examined subgroups. The finding that ESI-specific training conferred a significant accuracy benefit even after accounting for its association with experience is consistent with the interpretation that triage accuracy reflects not simply years of general emergency nursing exposure but a specific, teachable skill set best developed through deliberate, tool-specific training rather than informal on-the-job accumulation alone (Christ et al., 2010).
These findings have direct implications for institutional triage quality assurance. Rather than relying solely on general ESI certification status, emergency departments seeking to improve triage consistency might consider incorporating a vignette-based competency assessment, similar to the instrument developed for this study, into both initial triage training and ongoing competency verification, with particular emphasis on cases clustered near the level 2/3 boundary specifically, given this study’s finding that boundary cases account for the clear majority of clinically consequential under-triage error.
Several limitations should be considered. Vignette-based assessment, while allowing rigorous, standardized comparison across raters using a common stimulus set, cannot fully replicate the sensory, contextual, and time-pressured nature of live emergency department triage, and accuracy observed in this controlled vignette context may not directly correspond to accuracy under actual clinical conditions. This study was conducted within four hospitals in a single regional health system, and generalizability to emergency departments with different patient populations, staffing models, or triage training infrastructure should be considered carefully. The 30-vignette set, while deliberately weighted toward the ESI 2/3 boundary to maximize sensitivity to this specific source of disagreement, may not proportionally represent the true distribution of acuity levels encountered in routine emergency department practice.
Future research should evaluate whether targeted, boundary-specific training interventions, developed and tested directly against a vignette-based competency assessment similar to this study’s instrument, produce measurable improvement in both vignette-based and live-encounter triage accuracy. Extending this methodology to directly compare ESI against alternative five-level triage instruments within the same nurse sample would further clarify whether the level 2/3 boundary difficulty is specific to ESI or reflects a more general challenge inherent to acuity-based triage systems. Taken together, these findings support routine, vignette-based assessment of triage decision-making consistency as a feasible and clinically meaningful quality assurance strategy, and identify the ESI level 2/3 boundary as a specific, high-yield target for focused emergency nursing education.
References
Christ, M., Grossmann, F., Winter, D., Bingisser, R., & Platz, E. (2010). Modern triage in the emergency department. Deutsches Ärzteblatt International, 107(50), 892–898.
Cicchetti, D. V. (1994). Guidelines, criteria, and rules of thumb for evaluating normed and standardized assessment instruments in psychology. Psychological Assessment, 6(4), 284–290.
Fernandes, C. M., Tanabe, P., Gilboy, N., Johnson, L. A., McNair, R. S., Rosenau, A. M., Sawchuk, P., Thompson, D. A., Travers, D. A., Bonalumi, N., & Suter, R. E. (2005). Five-level triage: A report from the ACEP/ENA Five-Level Triage Task Force. Journal of Emergency Nursing, 31(1), 39–50.
Fleiss, J. L. (1971). Measuring nominal scale agreement among many raters. Psychological Bulletin, 76(5), 378–382.
Gilboy, N., Tanabe, P., Travers, D., & Rosenau, A. M. (2011). Emergency Severity Index (ESI): A triage tool for emergency department care, version 4. Implementation handbook 2012 edition (AHRQ Publication No. 12-0014). Agency for Healthcare Research and Quality.
Jordi, K., Grossmann, F., Gaddis, G. M., Cignacco, E., Denhaerynck, K., Schwendimann, R., Nickel, C. H., & Bingisser, R. (2015). Nurses’ accuracy and self-perceived ability using the Emergency Severity Index triage tool: A cross-sectional study in four Swiss hospitals. Scandinavian Journal of Trauma, Resuscitation and Emergency Medicine, 23, 62.
Landis, J. R., & Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1), 159–174.
Mistry, B., Stewart De Ramirez, S., Kelen, G., Jones, C. M. C., Bayram, J. D., Schneider, S., Hinson, J., & Balhara, K. S. (2018). Accuracy and reliability of emergency department triage using the Emergency Severity Index: An international multicenter assessment. Annals of Emergency Medicine, 71(5), 581–587.e3.
Storm-Versloot, M. N., Ubbink, D. T., Chin A Choi, V., & Luitse, J. S. (2009). Observer agreement of the Manchester Triage System and the Emergency Severity Index: A simulation study. Emergency Medicine Journal, 26(8), 556–560.
Tanabe, P., Gimbel, R., Yarnold, P. R., Kyriacou, D. N., & Adams, J. G. (2004). Reliability and validity of scores on The Emergency Severity Index version 3. Academic Emergency Medicine, 11(1), 59–65.
Worster, A., Fernandes, C. M., Eva, K., & Upadhye, S. (2006). Predictive validity comparison of two five-level triage acuity scales. European Journal of Emergency Medicine, 13(3), 178–181.
Worster, A., Gilboy, N., Fernandes, C. M., Eitel, D., Eva, K., Geisler, R., & Tanabe, P. (2004). Assessment of inter-observer reliability of two five-level triage and acuity scales: A randomized controlled trial. Canadian Journal of Emergency Medicine, 6(4), 240–245.
Wuerz, R. C., Milne, L. W., Eitel, D. R., Travers, D., & Gilboy, N. (2000). Reliability and validity of a new five-level triage instrument. Academic Emergency Medicine, 7(3), 236–242.
◆
Manuscript support
Scholarly Work helps you write and edit publication-ready research
From literature reviews to full manuscripts, our editors help nursing and healthcare researchers write clearly, meet journal standards, and get published with confidence.
Get StartedMore Nursing Journal Article Examples: Critical Care & Emergency Nursing
- Critical Care and Emergency Nursing Article Examples | A Comprehensive Resource for Nurses
- Handoff Communication Practices Among Emergency and Critical Care Nurses During Patient Transfers
- Nurse-Led Interventions to Reduce Alarm Fatigue in Intensive Care Units
- Effectiveness of Early Mobilization Protocols Led by Nurses in Intensive Care Units
- The Role of Critical Care Nurses in Facilitating Family Presence During Resuscitation
- Nurse-Driven Sedation Protocols and Their Impact on Mechanical Ventilation Duration
- Moral Distress Among Critical Care Nurses During End-of-Life Care Decision-Making
- Triage Accuracy and Decision-Making Consistency Among Emergency Department Nurses
- Rapid Response Team Activation Criteria and Their Impact on Patient Outcomes
- Nursing Strategies for Preventing Ventilator-Associated Pneumonia in Intensive Care Units
- The Effect of Nurse-to-Patient Ratios on Mortality Rates in Intensive Care Units
- Early Warning Score Systems and Their Effectiveness in Predicting Clinical Deterioration
- Emergency Nurses’ Experiences Managing Workplace Violence and Aggressive Patients
- Impact of Nurse-Led Sepsis Protocols on Time to Antibiotic Administration
- Compassion Fatigue and Secondary Traumatic Stress Among Emergency Department Nurses
Source context: National Institute of Nursing Research



