Instructions to use lemuralabs/lemura-arabic-asr-lite with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NeMo
How to use lemuralabs/lemura-arabic-asr-lite with NeMo:
import nemo.collections.asr as nemo_asr asr_model = nemo_asr.models.ASRModel.from_pretrained("lemuralabs/lemura-arabic-asr-lite") transcriptions = asr_model.transcribe(["file.wav"]) - Notebooks
- Google Colab
- Kaggle
lemura-arabic-asr-lite
Lemura Labs' Arabic-first speech-recognition model โ leaderboard-grade at a fraction of the size.
Arabic accuracy, without the GPU bill.
Overview ยท Benchmarks ยท Efficiency ยท Inference ยท Live Demo ยท Sibling model
What it is
lemura-arabic-asr-lite is a compact multi-dialect Arabic ASR model built for accuracy and efficiency. It is a FastConformer-CTC acoustic model adapted in-house from the NVIDIA FastConformer foundation โ the contribution is the adaptation (the Arabic data curriculum and dialect coverage), not the foundation.
- Small and fast โ ~115M parameters; runs comfortably on CPU and in real time, no GPU required.
- Dialect-aware โ fine-tuned on ~2,900 hours of Arabic spanning MSA and the Gulf, Egyptian, Levantine and Maghrebi dialect groups, not MSA-only.
- Robust on real audio โ strongest on broadcast, conversational and Gulf/MSA speech.
- Open and simple โ a single
.nemofile, loadable in a few lines with NVIDIA NeMo.
The result is leaderboard-grade Arabic ASR at edge scale โ #2 of 36 systems on the Open Universal Arabic ASR Leaderboard, ahead of every 2Bโ30B audio-LLM evaluated, at ~115M parameters.
Model summary
| Model | lemura-arabic-asr-lite โ compact multi-dialect Arabic ASR |
| Task | Automatic speech recognition (audio โ text) |
| Approach | Discriminative ASR โ FastConformer encoder + CTC decoder |
| Training | adapted from the NVIDIA FastConformer foundation; ~2,900 h Arabic fine-tuning across 5 dialect groups |
| Total parameters | ~115,000,000 (~0.12B) |
| Checkpoint size | ~0.4 GB (single .nemo file) |
| Audio input | 16 kHz mono (auto-resampled) |
| Languages | Arabic โ MSA + Gulf / Egyptian / Levantine / Maghrebi |
| Runtime | NVIDIA NeMo โ CPU ยท GPU ยท real-time |
| License | CC-BY-4.0 |
Benchmarks
Arabic dialectal ASR is hard โ heavily dialectal, conversational, code-switched speech is the frontier for every system. On the Open Universal Arabic ASR Leaderboard, lemura-arabic-asr-lite ranks #2 of 36 systems with an average WER of 25.08 % โ ahead of every model evaluated except one, including systems 17ร to 250ร larger.
Open Universal Arabic ASR Leaderboard โ full standings
Per-dataset WER % across all six leaderboard test sets. Lower is better; Avg WER is the ranking metric. Our rows were produced with the official leaderboard code (same normalizer, same WER metric). Ours in bold.
| # | Model | Params | Avg WER | SADA | CV-18 | MASC-clean | MASC-noisy | MGB-2 | Casablanca |
|---|---|---|---|---|---|---|---|---|---|
| 1 | lemuralabs/lemura-arabic-asr-lite (Ours) | 0.12B | 25.08 | 37.28 | 9.74 | 7.27 | 23.65 | 14.33 | 58.24 |
| 2 | CohereLabs/cohere-transcribe-arabic-07-2026 | ~2B | 25.87 | 37.47 | 5.82 | 19.60 | 27.07 | 15.54 | 49.71 |
| 3 | omnilingual-asr/omniASR_LLM_7B | 7B | 28.32 | 41.61 | 8.75 | 19.69 | 29.29 | 14.13 | 56.46 |
| 4 | omnilingual-asr/omniASR_LLM_3B | 3B | 29.96 | 46.18 | 9.15 | 19.90 | 30.03 | 14.22 | 60.27 |
| 5 | omnilingual-asr/omniASR_LLM_1B | 1B | 29.96 | 43.84 | 9.55 | 20.03 | 30.26 | 15.34 | 60.68 |
| 6 | CohereLabs/cohere-transcribe-03-2026 | ~2B | 30.67 | 60.11 | 8.17 | 8.66 | 19.01 | 25.33 | 62.71 |
| 7 | Qwen/Qwen3-Omni-30B-A3B-Instruct | 30B | 30.71 | 44.82 | 11.46 | 21.47 | 30.85 | 13.09 | 62.55 |
| 8 | nvidia-conformer-ctc-large-arabic (lm) | 0.6B | 32.91 | 44.52 | 8.80 | 23.74 | 34.29 | 17.20 | 68.90 |
| 9 | omnilingual-asr/omniASR_LLM_300M | 0.3B | 32.96 | 51.38 | 12.03 | 20.66 | 32.45 | 16.58 | 64.64 |
| 10 | google/gemma-4-E4B-it | 4B | 32.98 | 43.40 | 19.65 | 24.86 | 33.59 | 17.72 | 58.63 |
| 11 | Qwen/Qwen3-ASR-1.7B | 1.7B | 33.36 | 45.53 | 16.90 | 24.37 | 34.29 | 16.57 | 64.47 |
| 12 | mistralai/Voxtral-Small-24B-2507 | 24B | 34.47 | 50.82 | 15.25 | 23.96 | 34.43 | 16.03 | 66.30 |
| 13 | nvidia-conformer-ctc-large-arabic (greedy) | 0.6B | 34.74 | 47.26 | 10.60 | 24.12 | 35.64 | 19.69 | 71.13 |
| 14 | google/gemma-4-E2B-it | 2B | 35.87 | 46.23 | 23.76 | 27.47 | 36.15 | 20.72 | 60.87 |
| 15 | openai/whisper-large-v3 | 1.5B | 36.86 | 55.96 | 17.83 | 24.66 | 34.63 | 16.26 | 71.81 |
| 16 | omnilingual-asr/omniASR_CTC_3B | 3B | 37.78 | 69.85 | 14.19 | 21.48 | 34.60 | 18.96 | 67.58 |
| 17 | omnilingual-asr/omniASR_CTC_7B | 7B | 38.12 | 72.69 | 12.47 | 21.08 | 35.04 | 20.43 | 67.02 |
| 18 | facebook/seamless-m4t-v2-large | 2.3B | 38.16 | 62.52 | 21.70 | 25.04 | 33.24 | 20.23 | 66.25 |
| 19 | omnilingual-asr/omniASR_CTC_1B | 1B | 39.29 | 71.42 | 17.55 | 22.76 | 35.73 | 19.96 | 68.32 |
| 20 | openai/whisper-large-v3-turbo | 0.8B | 40.05 | 60.36 | 25.73 | 25.51 | 37.16 | 17.75 | 73.79 |
| 21 | openai/whisper-large-v2 | 1.5B | 40.20 | 57.46 | 21.77 | 27.25 | 38.55 | 25.17 | 71.01 |
| 22 | Qwen/Qwen3-ASR-0.6B | 0.6B | 42.19 | 53.75 | 28.28 | 31.34 | 42.63 | 25.45 | 71.68 |
| 23 | openai/whisper-large | 1.5B | 42.57 | 63.24 | 26.04 | 28.89 | 40.79 | 24.28 | 72.18 |
| 24 | mistralai/Voxtral-Mini-3B-2507 | 3B | 42.58 | 63.65 | 22.12 | 28.37 | 41.27 | 22.56 | 77.52 |
| 25 | asafaya/hubert-large-arabic-transcribe | 0.3B | 45.50 | 67.82 | 8.01 | 32.94 | 50.16 | 37.51 | 76.53 |
| 26 | openai/whisper-medium | 0.8B | 45.57 | 67.71 | 28.07 | 29.99 | 42.91 | 29.32 | 75.44 |
| 27 | nvidia-Parakeet-ctc-1.1b-concat | 1.1B | 46.54 | 70.70 | 26.34 | 30.49 | 45.95 | 24.94 | 80.80 |
| 28 | omnilingual-asr/omniASR_CTC_300M | 0.3B | 46.65 | 78.11 | 27.90 | 28.40 | 43.26 | 26.85 | 75.35 |
| 29 | nvidia-Parakeet-ctc-1.1b-universal | 1.1B | 51.96 | 73.58 | 40.01 | 36.16 | 50.03 | 30.68 | 81.30 |
| 30 | microsoft/VibeVoice-ASR | โ | 52.99 | 69.83 | 44.25 | 32.95 | 52.43 | 25.10 | 93.37 |
| 31 | facebook/mms-1b-all | 1B | 54.54 | 77.48 | 26.52 | 38.82 | 57.33 | 39.16 | 87.95 |
| 32 | openai/whisper-small | 0.24B | 55.13 | 78.02 | 24.18 | 35.93 | 56.36 | 48.64 | 87.64 |
| 33 | whitefox123/w2v-bert-2.0-arabic-4 | 0.6B | 58.13 | 87.34 | 41.79 | 37.82 | 53.28 | 40.66 | 87.88 |
| 34 | jonatasgrosman/wav2vec2-large-xlsr-53-arabic | 0.3B | 60.98 | 86.82 | 23.00 | 42.75 | 64.27 | 56.29 | 92.72 |
| 35 | speechbrain/asr-wav2vec2-commonvoice-14-ar | 0.1B | 65.74 | 88.54 | 29.17 | 49.10 | 69.57 | 64.37 | 93.68 |
Bold = our model and our best-in-column results. lemura-arabic-asr-lite is the single best system on MASC-clean (7.27) and MASC-noisy (23.65) of every model evaluated, and the highest-ranked system below 1B parameters by a wide margin. Casablanca (Moroccan Darija) is the hardest set for every system.
Competitor rows are the published Open Universal Arabic ASR Leaderboard standings, reproduced for context; they are not our measurements.
Efficiency
The systems around us on the leaderboard are large generative audio-LLMs (2โ30B parameters) that need GPUs. lemura-arabic-asr-lite reaches the same accuracy tier with a ~115M-parameter CTC model:
| lemura-arabic-asr-lite | Typical top-5 systems | |
|---|---|---|
| Parameters | ~115M | 2B โ 30B |
| Hardware | CPU or GPU | GPU |
| Latency | Real-time | Seconds / clip |
| Footprint | ~0.4 GB | 4 โ 60 GB |
| Avg WER | 25.08 | 24.78 โ 30.71 |
That makes it practical for on-device, low-cost and high-throughput Arabic transcription where the larger models are impractical.
Inference
import nemo.collections.asr as nemo_asr
model = nemo_asr.models.ASRModel.restore_from("asr_final.nemo")
print(model.transcribe(["audio.wav"])) # 16 kHz mono
Try it live, no install: lemura-arabic-asr demo
Languages, dialects and tasks
| Languages | Arabic |
| Dialects | MSA, Gulf/Khaleeji, Egyptian, Levantine, Maghrebi |
| Tasks | Transcription (audio โ text) |
| Audio | 16 kHz mono, auto-resampled |
Intended use and limitations
Intended: transcribing Arabic speech across dialects โ call centres, voice notes, media captioning, voice interfaces, and any deployment where CPU-only or real-time operation matters.
Limitations:
- Maghrebi/Darija (Casablanca, 58.24 WER) remains hard, as it does for every system on the leaderboard.
- Very noisy or far-field audio degrades accuracy.
- CTC output has no built-in punctuation or diacritic restoration.
- May reflect biases present in the training corpora.
Related model
For a larger, generative alternative with broader dialect fine-tuning, see lemura-arabic-asr-qwen3 (1.7B, audio-LLM). Note its published WER is an in-domain figure and is not directly comparable to the zero-shot leaderboard numbers above.
License
CC-BY-4.0.
Citation
@misc{lemura_arabic_asr_lite_2026,
title = {lemura-arabic-asr-lite: Compact Multi-Dialect Arabic Speech Recognition},
author = {Lemura Labs},
year = {2026},
url = {https://huggingface.co/lemuralabs/lemura-arabic-asr-lite}
}
About Lemura Labs
Arabic-first, efficiency-first speech intelligence.
Lemura Labs builds compact, deployable speech and language models โ accuracy at a size and cost that works outside the datacentre.
- Downloads last month
- 52
Space using lemuralabs/lemura-arabic-asr-lite 1
Evaluation results
- Average WER on Open Universal Arabic ASR Leaderboard (average of 6 sets)self-reported25.080
- WER on MASC (clean)self-reported7.270
- WER on Common Voice 18 (Arabic)self-reported9.740
- WER on MGB-2self-reported14.330
- WER on MASC (noisy)self-reported23.650
- WER on SADAself-reported37.280
- WER on Casablancaself-reported58.240