CAPE-V: What It Measures
Consensus Auditory-Perceptual Evaluation of Voice
Quick answer
The CAPE-V measures the auditory-perceptual quality of a voice. A clinician listens to sustained vowels, standard sentences, and conversational speech, then rates overall severity, roughness, breathiness, strain, pitch, and loudness on 100 mm visual analog scales, producing a standardized description of dysphonia that can be compared across clinicians and over time.
Voice quality is inherently perceptual, and the CAPE-V is the profession's attempt to make that perception reliable. It standardizes the tasks, the attributes, and the rating scale so that 'moderately breathy' means something similar in two different clinics.
CAPE-V at a glance
- Full name
- Consensus Auditory-Perceptual Evaluation of Voice
- Publisher
- ASHA (Division 3 consensus protocol)
- Domain
- Voice
- Age range
- Adolescents and adults (adaptable to children)
- Administration time
- About 10–15 minutes
- Format
- Auditory-perceptual rating protocol using 100 mm visual analog scales
- Score types
- Visual analog scale ratings (0–100 mm) per attribute, Consistency ratings (consistent vs. intermittent)
What the CAPE-V measures
- Overall severity of dysphonia
- Roughness — perceived irregularity in the voicing source
- Breathiness — audible air escape during phonation
- Strain — perceived effort or hyperfunction
- Pitch — deviation from expected pitch for age and sex
- Loudness — deviation from expected loudness
Subtests and sections
- Sustained vowels
- /a/ and /i/ held 3–5 seconds each, repeated three times, to hear the source signal without articulation.
- Six standard sentences
- Phonetically loaded sentences targeting all voiced content, hard glottal attacks, nasals, and voiceless consonants.
- Running speech
- At least 20 seconds of conversational speech about a neutral topic.
- Visual analog rating
- Each attribute marked on a 100 mm line and reported as a number out of 100 with a consistency notation.
CAPE-V vs. GRBAS
GRBAS rates grade, roughness, breathiness, asthenia, and strain on a four-point ordinal scale (0–3). It is quick and internationally used, but the coarse scale limits sensitivity to small changes. The CAPE-V uses continuous 100 mm visual analog scales, adds pitch and loudness, and specifies the elicitation tasks, which makes it more sensitive for outcome measurement.
If a question asks which tool best documents subtle improvement after six weeks of voice therapy, the continuous scale is the better answer.
Building a complete voice evaluation
A defensible voice evaluation layers four kinds of data: perceptual (CAPE-V), acoustic (fundamental frequency, cepstral peak prominence, perturbation measures), aerodynamic (maximum phonation time, s/z ratio, airflow), and patient-reported outcome (Voice Handicap Index or V-RQOL). Laryngeal imaging from ENT sits underneath all of it as the diagnostic foundation.
No single layer is sufficient. Exam items often present a plan missing one layer and ask what the SLP should add.
Strengths
- Standardized tasks and scales improve inter-rater reliability over free description.
- Free to use and widely adopted, so results transfer across settings.
- Sensitive enough to document change from voice therapy or surgery.
- Captures dimensions that acoustic and aerodynamic measures alone do not.
Limitations
- Perceptual judgment still varies with rater training and experience.
- Not a diagnosis — it describes quality but says nothing about laryngeal pathology.
- Must be paired with laryngeal imaging before initiating voice therapy for an undiagnosed dysphonia.
- No normative cut scores; interpretation is descriptive and comparative.
When clinicians use it
- Baseline and outcome documentation across a voice therapy episode.
- Communicating voice quality precisely to otolaryngology colleagues.
- Complementing acoustic (jitter, shimmer, CPPS) and aerodynamic (MPT, subglottal pressure) measures.
- Training students in reliable perceptual judgment.
What the Praxis 5331 asks about the CAPE-V
- Know that laryngeal examination by an ENT is required before voice therapy for an undiagnosed voice disorder — a very common keyed answer.
- Match attributes to pathophysiology: breathiness with glottal incompetence, strain with hyperfunction, roughness with irregular vibration.
- Distinguish perceptual (CAPE-V, GRBAS), acoustic, aerodynamic, and imaging measures.
Practice assessment questions with rationales
Praxis Path pairs this reference material with scenario questions, spaced-repetition flashcards, and a timed 132-question mock exam.
FAQ
What does the CAPE-V measure?
Auditory-perceptual voice quality across six attributes: overall severity, roughness, breathiness, strain, pitch, and loudness, rated on 100 mm visual analog scales.
What tasks are used in the CAPE-V?
Sustained vowels /a/ and /i/, six standardized sentences, and at least 20 seconds of running conversational speech.
What is the difference between CAPE-V and GRBAS?
GRBAS uses a four-point ordinal scale with five attributes; the CAPE-V uses continuous visual analog scales, adds pitch and loudness, and standardizes the elicitation tasks.
Can an SLP start voice therapy based on a CAPE-V alone?
No. A laryngeal examination by an otolaryngologist is required to rule out pathology before initiating voice therapy for an undiagnosed dysphonia.
Related assessments
- CELF-5 — Clinical Evaluation of Language Fundamentals, Fifth Edition
- GFTA-3 — Goldman-Fristoe Test of Articulation, Third Edition
- PLS-5 — Preschool Language Scales, Fifth Edition
- WAB-R — Western Aphasia Battery–Revised
- BDAE-3 — Boston Diagnostic Aphasia Examination, Third Edition
- Browse the full assessment reference library