Abstract
This study aimed to evaluate intrarater reliability for nasopharyngoscopy ratings of velopharyngeal closure across two contexts: ratings made in a clinical setting and ratings made in a research setting using standardized rating definitions.
Five speech-language pathologists participated in this study. Two assessments of intrarater reliability were performed: between ratings from clinical reports and re-ratings made in a research setting and videos re-rated twice in a research setting. Nasopharyngoscopy ratings analyzed included velopharyngeal closure pattern, total percent closure, and extent of velar movement. Reliability was calculated using Cohen's kappa and the intraclass correlation coefficient (ICC).
Intrarater reliability between clinical reports and re-ratings made in a research setting varied for closure pattern from weak (
= .24) to perfect (
= 1.00), for total percent closure from poor (ICC = -.09) to good (ICC = .81), and for extent of velar movement from moderate (ICC = .57) to good (ICC = .80). Compared to intrarater reliability between clinical and research settings, intrarater reliability for videos re-rated twice in a research setting was higher for all raters. For closure pattern, intrarater reliability within a research setting ranged from substantial (
= .68) to perfect (
= 1.00), for total percent closure from moderate (ICC = .70) to excellent (ICC = .99), and for extent of velar movement from good (ICC = .80) to excellent (ICC = .91).
Findings suggest that nasopharyngoscopy ratings made in a clinical setting may be systematically different than ratings made in a research setting; thus, research results that depend on nasopharyngoscopy ratings may not be generalizable to clinical care. This may be caused by a variety of patient-specific and practice-specific factors, in addition to a lack of standardization. Findings suggest that high reliability for nasopharyngoscopy ratings can be achieved if raters receive consistent training and are following the same rating scale definitions.