Chatterbox-Turbo TTS β€” GGUF (ggml)

GGUF / ggml conversion of ResembleAI/chatterbox-turbo for use with CrispStrobe/CrispASR.

Chatterbox-Turbo is a distilled 350M-parameter TTS pipeline: GPT-2 tokenizer + AR text-to-speech model + meanflow S3Gen (2-step CFM, vs 10 for base Chatterbox) + HiFTGenerator vocoder. Distributed under MIT license.

Two GGUF files are needed: the T3 model (text to speech tokens) and the S3Gen model (speech tokens to audio).

Updates

2026-06-21 β€” tokenizer fix (re-uploaded T3 files). The T3 GGUFs previously embedded only the 50257-token base GPT-2 vocab, while the T3 text embedding is 50276 β€” the 19 extra ids are the turbo emotion/style control tokens. The files have been re-uploaded with the full 50276-token tokenizer (base vocab + added_tokens.json), so they are now internally consistent and load cleanly on strict loaders (CrispASR β‰₯ v0.8.1). The weights are unchanged (byte-for-byte), so this is a tokenizer-only update.

Emotion / style tags. You can drive prosody by putting any of these bracketed tags in the input text (CrispASR β‰₯ v0.8.1 emits them as their special token id):

[laugh] [chuckle] [sigh] [gasp] [cough] [groan] [sniff] [shush] [clear throat]
[whispering] [angry] [happy] [crying] [fear] [surprised] [sarcastic] [dramatic]
[narration] [advertisement]

Example: [laugh] Thank you so much! prepends a laugh to the line. (Effect strength varies per tag and prompt.)

Files

File Size Notes
chatterbox-turbo-t3-f16.gguf 964 MB T3 GPT-2 AR model (24L, 1024D)
chatterbox-turbo-t3-q8_0.gguf 628 MB Quantized T3, recommended deployment default
chatterbox-turbo-t3-q4_k.gguf 457 MB Smaller T3 quant for memory-constrained use
chatterbox-turbo-s3gen-f16.gguf 628 MB S3Gen encoder + meanflow CFM + HiFT vocoder
chatterbox-turbo-s3gen-q8_0.gguf 350 MB Quantized S3Gen, recommended deployment default
chatterbox-turbo-s3gen-q4_k.gguf 244 MB Smaller S3Gen quant for memory-constrained use

Encoder attention/FFN weights are stored at F32 precision for quality. Vocoder weights (conv_pre, resblocks, conv_post, source fusion, F0 predictor) are F32.

Quick start

# 1. Build CrispASR
git clone https://github.com/CrispStrobe/CrispASR
cd CrispASR
cmake -B build -DCMAKE_BUILD_TYPE=Release -DBUILD_SHARED_LIBS=OFF
cmake --build build -j --target chatterbox

# 2. Pull both model files (Q8_0 recommended)
huggingface-cli download cstr/chatterbox-turbo-GGUF chatterbox-turbo-t3-q8_0.gguf --local-dir .
huggingface-cli download cstr/chatterbox-turbo-GGUF chatterbox-turbo-s3gen-q8_0.gguf --local-dir .

# 3. Synthesise (C API β€” CLI adapter in progress)
# See test programs in SESSION_HANDOVER.md for usage examples

Architecture

Text -> GPT-2 BPE tokenizer (50276 tokens: 50257 base + 19 emotion/style tags)
     -> T3 GPT-2 AR (24 layers, 1024D, 16 heads, learned pos emb, SwiGLU)
     -> 25 Hz speech tokens (6561 codebook)
     -> UpsampleConformerEncoder (6 pre + 4 post upsample, 512D, 8 heads, rel-pos attn)
        -> Upsample1D: nearest-neighbor 2x + Conv1d(512,512,k=5) + Linear + LayerNorm + xscale
     -> 80-channel mel spectrogram (50 Hz)
     -> Meanflow CFM denoiser (2 Euler steps, linear schedule, no CFG)
        UNet1D: 1 down + 12 mid + 1 up blocks, 256 ch, 4 transformer blocks each
     -> HiFTGenerator vocoder (F0 predictor + SineGen + 3x ConvTranspose1d + iSTFT)
     -> 24 kHz mono WAV

Key differences from base Chatterbox

Feature Base Chatterbox Chatterbox-Turbo
T3 architecture Llama (30L, 520M) GPT-2 Medium (24L, 350M)
T3 tokenizer Character (704 tokens) BPE (50276 tokens, incl. 19 emotion tags)
CFM steps 10 (cosine schedule) 2 (linear, meanflow distilled)
CFG Yes (rate=0.7) No (distilled)
Total params ~520M ~350M

Quality verification

ASR roundtrip using same speech tokens as Python reference:

Metric Value
ASR output (moonshine-base) "Hello world" (correct)
Language detection confidence 0.939
encoder_out RMS 0.4602 (exact match to Python)
matrix_bd (rel-pos scores) h0[0,0] 24.70 (matches Python to 2dp)

Conversion

# From HuggingFace model (requires chatterbox-tts pip package):
python models/convert-chatterbox-to-gguf.py \
  --input ResembleAI/chatterbox-turbo \
  --output-dir /path/to/output \
  --variant turbo

Related models

Provenance and EU AI Act Art. 53 note

  • Upstream model: ResembleAI/chatterbox-turbo β€” published by ResembleAI.
  • Upstream licence: mit. This repository redistributes under the same terms; it grants no rights the upstream licence does not.
  • What was done here: format conversion and/or quantisation only (GGUF/GGML). No training, no fine-tuning, no merging, no distillation, no change to architecture, vocabulary or capability. Only the numeric representation of the upstream weights differs.
  • Training data: documented β€” where it is documented at all β€” by the upstream provider; see the upstream model card. No training data was used, added or selected by this repository. No training-content summary was found on the upstream model card at the time of writing; that documentation gap is upstream's and is not filled here.
  • Provider status: under Regulation (EU) 2024/1689 the upstream authors remain the provider of this model. Converting the serialisation format does not make this repository the provider of a new general-purpose AI model, and no such claim is made. Questions about training content, copyright policy or model capability belong upstream.
Downloads last month
1,680
GGUF
Model size
0.3B params
Architecture
chatterbox-s3gen
Hardware compatibility
Log In to add your hardware

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for cstr/chatterbox-turbo-GGUF

Quantized
(11)
this model