--- license: mit language: - de library_name: transformers pipeline_tag: automatic-speech-recognition tags: - moonshine - streaming - asr - german --- # Moonshine Streaming Small — German German streaming speech recognition, 112.9M parameters. Same architecture as [moonshine-ai/moonshine-streaming-small](https://huggingface.co/moonshine-ai/moonshine-streaming-small), trained for German with a 12,288-entry German tokenizer. The [Tiny model](https://huggingface.co/moonshine-ai/moonshine-streaming-tiny- de) is a quarter of the size. Moonshine Streaming pairs a 50 Hz time-domain audio frontend with a sliding-window Transformer encoder, so it transcribes incrementally rather than waiting for an utterance to finish. It is intended for on-device use on edge-class hardware. ## Checkpoint identity This repository is a conversion of one specific training checkpoint, recorded here because the weights behind a language move as later stages win: | | | |---|---| | Checkpoint | `de12k_small_stageB_best.safetensors` | | Stage | B (fine-tune from Stage A weights) | | Architecture | `spindlier_prime_adapted` | | Tokenizer | `tokenizer_de12k.json`, vocab 12,288 | | Snapshot taken | 2026-08-24 | | Parameters | 112.9M | If you need reproducibility, pin the revision of this repository rather than tracking `main`. ## Usage ```bash pip install --upgrade transformers datasets[audio] ``` ```python from transformers import MoonshineStreamingForConditionalGeneration, AutoProcessor import torch model = MoonshineStreamingForConditionalGeneration.from_pretrained( "moonshine-ai/moonshine-streaming-small-de" ).eval() processor = AutoProcessor.from_pretrained("moonshine-ai/moonshine-streaming-small-de") inputs = processor(audio, return_tensors="pt", sampling_rate=16000) # Cap the output length. Like other seq2seq ASR models this one can fall into a # repetition loop, and short or noisy clips are where it happens. seq_lens = inputs.attention_mask.sum(dim=-1) max_new_tokens = int((seq_lens * 6.5 / 16000).max().item()) + 2 generated = model.generate(**inputs, max_new_tokens=max_new_tokens) print(processor.batch_decode(generated, skip_special_tokens=True)[0]) ``` **Pass the `attention_mask`.** The encoder applies its per-layer sliding windows only when it is given one; called without a mask it attends over the whole utterance instead, which is a different model from the one that was trained. The processor returns the mask, so the snippet above is the safe form. The processor also pads audio to a whole number of 80-sample frames, which the frontend requires. ## Architecture | | | |---|---| | Encoder | 10 layers, width 620, 8 heads, sliding windows (16, 4) on the first two and last two layers and (16, 0) between | | Decoder | 10 layers, width 512, 8 heads, RoPE over 32 of each head's 64 dimensions | | Frontend | 50 Hz features, CMVN, asinh compression, two causal stride-2 convolutions | | Adapter | learned absolute positional embeddings, then a projection from 620 to 512 | The lookahead layers give roughly 80 ms of lookahead; the intermediate layers have none. ## Training data Trained on a large-scale **automatically labeled** German corpus, plus a much smaller human-labeled read-speech set: - **Podcast crawl**, roughly 103,000 hours, pseudo-labeled. - **Track A read speech**, roughly 3,700 hours, human-transcribed (Common Voice, Multilingual LibriSpeech, FLEURS and VoxPopuli). The crawled transcripts are **pseudo-labels**: they were produced by running a Whisper-family teacher model over crawled audio, not by human transcription. The model therefore inherits the teacher's error modes, including its handling of proper nouns, numerals and code-switching. No human-verified transcript was used for the bulk of training. ## Evaluation German is scored on **word error rate** (WER), after the usual case and punctuation normalization. Mandarin and Japanese in this model family are instead scored on no-space CER, because they are written without spaces; every other language, this one included, uses WER. `suite_de` is FLEURS German and Multilingual LibriSpeech German. Both are read speech, so neither panel measures spontaneous or conversational German. ### Seeded 400-utterance sample, batch 1 Batch 1 is the honest number for deployment. Batched evaluation zero-pads short clips up to the longest in the batch, and that trailing silence flatters the model. | Panel | WER | |---|---:| | `fleurs_de` | 7.19 | | `mls_de` | 7.88 | | **macro** | **7.534** | ### This repository against the training checkpoint These weights were converted from the `neo` training checkpoint, and the conversion was checked by measurement rather than inspection: same seeded sample, same batch size, same normalizer. A conversion that loads and emits plausible text can still have a permuted weight mapping, which only a score catches. | | `fleurs_de` | `mls_de` | macro | |---|---:|---:|---:| | Training checkpoint | 7.19 | 7.88 | 7.534 | | Same checkpoint, same stopping rule | 7.19 | 7.51 | 7.350 | | This repository | 7.19 | 7.51 | 7.350 | 400/400 and 395/400 transcripts are byte-identical. The middle row is the comparison that matters. `neo`'s decoder also stops when it sees a repeating token pattern, and `transformers` does not, so the top row is measured under a different stopping rule than this repository can use. Rescoring the checkpoint without that heuristic gives 7.350 against this repository's 7.350: the same number to three decimals. The gap in the top row is that heuristic, not the conversion. ### The quantized build we ship The `.ort` package served to the Moonshine deployment library is quantized to int8 from these same weights, and scores 7.530 against 7.350 for the float checkpoint on the same sample under the same stopping rule -- a difference of +0.181, which is inside the noise of a 400-clip sample and should not be read as the quantized build being better or worse. That build is a different artifact from this repository, which is float32. ## Limitations - **Machine-labeled training data.** See above; the model reproduces its teacher's mistakes as well as its strengths. - **Repetition loops on short clips.** Like other seq2seq ASR models this one can fall into a repetition loop, and short or noisy clips are where it happens. Cap the output length, as the usage snippet does. - **Evaluated on 2 panels only.** No evaluation of telephony, children's speech, heavy dialect, or noisy far-field conditions. ## Out-of-scope use Not intended for non-consensual surveillance, speaker identification, or high-stakes decisions. ## License MIT.