1. Click on "Start" to set all values on default.
2. Upload one or multiple mp3 or wav audio file(s) with speech or voice examples of your choice.
3. optional: Select the noise gate setting (VAD), to avoid artifacts in pauses in speech: noise gate off (quiet parts will be analysed) or
noise gate on (pauses in speech will not be analysed).
4. optional: If you have uploaded multiple audio files to analyze them in a batch:
5. Click on "Analyse", to start the analyzing process.
6. optional: View data in the interactive plots.
7. Click on "Show JS Arrays" or "Export CSV", to show the recorded data as JavaScript arrays or to export them as a CSV/Excel file.
Goldstandard der Stimmgüte in Spontansprache; misst das Hervortreten der Harmonischen im Quefrency-Bereich über der Regressionsgeraden. Werte > 12–15 dB stehen für eine klare Stimme, niedrige Werte für Heiserkeit/Dysphonie.
SPR (Speaker's Power Ratio)
Variabel (dB)
Singing / Speaker's Power Ratio; logarithmische Energierate zwischen 2–4 kHz und 0–2 kHz. Quantifiziert die Tragfähigkeit, Durchsetzungsfähigkeit und Resonanz der Sprechstimme.
HNR (Harmonics-to-Noise)
0 bis ~30 dB
Verhältnis harmonischer Stimmbandenergie zu turbulentem Rauschen. Hohe Werte signalisieren saubere Phonation, niedrige Werte Behauchtheit/Aperiodizität.
Jitter (lokal)
0 bis ~15 %
Zyklus-zu-Zyklus-Schwankung der Grundschwingungsperiode. Wichtiges Maß für Stimmstabilität (Normwert bei gehaltener Phonation typisch < 1.04 %).
Shimmer (lokal)
0 bis ~25 %
Relative Schwankung der Spitzenamplitude von Periode zu Periode. Indikator für glottalen Schluss und Ermüdung (Norm typisch < 3.81 %).
Delta C (ΔC)
ca. 30 bis 90 ms
Standardabweichung der Dauern konsonantischer Intervalle nach Ramus et al. (1999). Hohe Werte kennzeichnen akzentzählende Sprachen oder präzise Konsonantencluster.
Delta V (ΔV)
ca. 30 bis 80 ms
Standardabweichung der vokalischen Intervalle. Zeigt den Längenkontrast zwischen betonten und reduzierten Vokalen an.
Vokalanteil (%V)
30 bis 65 %
Prozentualer Anteil vokalischer Segmente an der Gesamtsprachzeit. Höher in silben-/morazählenden Sprachen (Spanisch, Japanisch), niedriger im Deutschen/Englischen.
VarcoV / VarcoC
Variabel
Variationskoeffizienten der Vokal- und Konsonantendauern ($\text{SD}/\text{Mean} \times 100$); normalisiert Rhythmusmaße gegen das Sprechtempo.
nPVI-V
ca. 30 bis 80
Normalized Pairwise Variability Index aufeinanderfolgender Vokaldauern; quantifiziert den lokalen Rhythmuswechsel zwischen starken und schwachen Silben.
Artikulationsrate
Silben / s
Silbenanzahl geteilt durch die reine Nettosprechzeit (Pausen herausgerechnet).
Sprechrate (Gesamt)
Silben / s
Silbenanzahl bezogen auf die Bruttodauer inklusive Pausen.
F0 Mean & Range
Hz & Halbtöne
Mittlere Grundfrequenz sowie die perzeptiv normierte Intonationsbreite ($12 \cdot \log_2(F_{0,\text{max}} / F_{0,\text{min}})$) in musikalischen Halbtönen.
Prosodische Kontur
Kategorie
Form des F0-Makroverlaufs via Regression: Rising (steigend), Falling (fallend), Convex (Bogen/Peak), Concave (Tal) oder Flat (monoton).
Spectral Centroid
Hz
Spektraler Helligkeitsschwerpunkt; korreliert mit subjektiver Stimmschärfe und Klangfarbe.
Spectral Tilt (85% Roll-off)
Hz
Frequenzgrenze, unterhalb derer 85 % der spektralen Gesamtenergie liegen. Beschreibt den Energieabfall zu hohen Frequenzen.
AM-Silben-Peak
0.5 bis 16 Hz
Dominante Modulationsfrequenz der Amplitudenhüllkurve; bildet die Silbenpulsrate (typisch 3–6 Hz) ab.
Pausen & Sprechzeitanteil
Anzahl, s, %
Erfasst Sprechpausen (> 180 ms) und den prozentualen Sprachaktivitätsanteil (Fluency).
Gold standard of voice quality in continuous speech; measures the prominence of the fundamental cepstral peak over linear regression. Values > 12–15 dB indicate a healthy, clear voice; low values reflect dysphonia/breathiness.
SPR /(Speaker's Power Ratio)
Variable (dB)
Singing / Speaker's Power Ratio; logarithmic energy ratio between 2–4 kHz and 0–2 kHz. Quantifies vocal projection, resonance, and penetration.
HNR (Harmonics-to-Noise)
0 to ~30 dB
Harmonics-to-Noise Ratio; relative harmonic glottal energy compared to aspiration noise. High values indicate clear phonation.
Jitter (local)
0 to ~15 %
Cycle-to-cycle duration perturbation of vocal fold vibration. Measures pitch stability (norm in sustained vowels typically < 1.04 %).
Shimmer (local)
0 to ~25 %
Relative cycle-to-cycle amplitude perturbation. Reflects glottal closure and vocal fatigue (norm typically < 3.81 %).
Delta C (ΔC)
approx. 30 to 90 ms
Standard deviation of consonantal intervals (Ramus et al., 1999). Higher in stress-timed languages with complex consonant clusters.
Delta V (ΔV)
approx. 30 to 80 ms
Standard deviation of vocalic intervals. Quantifies length contrasts between stressed and reduced vowels.
Vocalic Ratio (%V)
30 to 65 %
Percentage of vocalic intervals relative to total speech time. Higher in syllable-/mora-timed languages (Spanish, Japanese).
VarcoV / VarcoC
Variable
Variation coefficients of vowel and consonant durations ($\text{SD}/\text{Mean} \times 100$); normalizes rhythm metrics against speech rate.
nPVI-V
approx. 30 to 80
Normalized Pairwise Variability Index of consecutive vowel durations; quantifies local rhythmic alternation between strong and weak syllables.
Articulation Rate
Syllables / s
Number of syllables divided by net speech duration (pauses excluded).
Speech Rate
Syllables / s
Total number of syllables divided by gross audio duration including pauses.
F0 Mean & Range
Hz & Semitones
Average fundamental frequency and perceptually scaled pitch span ($12 \cdot \log_2(F_{0,\text{max}} / F_{0,\text{min}})$) in semitones.
Prosodic Contour
Category
F0 macro-trajectory classified via regression: Rising, Falling, Convex (peak), Concave (valley), or Flat.
Spectral Centroid
Hz
Spectral center of gravity; correlates with perceived timbre brightness and sharpness.
Spectral Tilt (85% Roll-off)
Hz
Frequency below which 85% of total spectral energy is contained; characterizes spectral decay.
AM Syllable Peak
0.5 to 16 Hz
Dominant modulation frequency of the amplitude envelope; captures the syllabic pulse rate (typically 3–6 Hz).
Pause & Speech Ratio
Count, s, %
Quantifies silent intervals (> 180 ms) and proportional voiced time (speaking fluency).
Select one or more audio files (wav, mp3) and click "Analyse".
Analyse Audio... Initialize Audio Signal Processing...
0%
SInES Speech & Voice Visualizer
1. F0 & Intonation Contour
2. Consonant (C) & Vowel Segments (V) on the Envelope Curve
Legend:■ Vowel Intervals (V) | ■ Consonant Intervals (C) | Dark gray line = RMS Level Envelope
Click to jump to the corresponding audio playback time position.
3. Ramus et al. (1999) Speech Rhythm Map (%V vs. ΔC)
Typology according to Ramus et al. (1999): Red = Stress-timed (German, English) | Blue = Syllable-timed (Spanish, French) | Green = Mora-timed (Japanese) | Golden Dot = Currently Selected File.
4. Amplitude Modulation Spectrum (0.5 – 16 Hz)
Modulation frequency (0.5 – 16 Hz): The red marker indicates the dominant syllable rate (peak modulation rate in the 2–8 Hz band).
Interactive Navigation: Clicking on the top graphs (1 and 2) jumps to the corresponding audio playback time position.