Intelligibility across two ASR families
Semantic WER on the same 500 generated clips per system · lower is better
Whisper-large-v3
wav2vec2 LV-60K
0.0%
3.0%
6.0%
9.0%
12.1%
15.1%
Inflect-Micro-v2
Whisper 1.93%
wav2vec2 LV 4.81%
Inflect-Nano-v2
Whisper 2.48%
wav2vec2 LV 5.67%
KittenTTS Nano · Bruno
Whisper 0.97%
wav2vec2 LV 3.55%
KittenTTS Nano · Hugo
Whisper 1.15%
wav2vec2 LV 3.63%
Piper · Ryan Low
Whisper 2.33%
wav2vec2 LV 6.70%
Piper · Danny Low
Whisper 1.81%
wav2vec2 LV 5.31%
Supertonic 3 · James · 3-step
Whisper 692.82%
wav2vec2 LV 13.47%
Supertonic 3 · James · 8-step
Whisper 622.52%
wav2vec2 LV 3.76%
Large disagreement identifies ASR sensitivity rather than a hidden aggregate score; all hypotheses and per-clip errors ship in the raw reports.