Predicted quality versus model footprint UTMOS22 on 500 matched unseen prompts · complete deployable weight footprint 16 32 64 128 256 512 4.12 4.21 4.29 4.38 4.46 1 2 3 4 5 MODEL FAMILY UTMOS22 MB 1 Inflect-Micro-v2 4.395 37.5 2 Inflect-Nano-v2 4.386 16.0 3 Supertonic 3 · 8-step 4.295 398.1 4 Piper Low · 2-voice mean 4.242 63.1 5 KittenTTS Nano · 2-voice mean 4.204 56.8 Model weight footprint (MB, logarithmic scale) Predicted MOS · higher is better UTMOS22 is an automated predictor, not human MOS. Family means use equal weighting across the named voices. Supertonic 3-step: 2.471 UTMOS22 · below plotted range. Supertonic UTMOS uses the James voice; family means use equal voice weighting.