Predicted quality versus model footprint
UTMOS22 on 500 matched unseen prompts · complete deployable weight footprint
16
32
64
128
256
512
4.12
4.21
4.29
4.38
4.46
1
2
3
4
5
MODEL FAMILY
UTMOS22
MB
1
Inflect-Micro-v2
4.395
37.5
2
Inflect-Nano-v2
4.386
16.0
3
Supertonic 3 · 8-step
4.295
398.1
4
Piper Low · 2-voice mean
4.242
63.1
5
KittenTTS Nano · 2-voice mean
4.204
56.8
Model weight footprint (MB, logarithmic scale)
Predicted MOS · higher is better
UTMOS22 is an automated predictor, not human MOS. Family means use equal weighting across the named voices.
Supertonic 3-step: 2.471 UTMOS22 · below plotted range.
Supertonic UTMOS uses the James voice; family means use equal voice weighting.