lemura-arabic-asr-lite

Lemura Labs' Arabic-first speech-recognition model โ€” leaderboard-grade at a fraction of the size.

Arabic accuracy, without the GPU bill.

License Task Format Params Open-AR-ASR Efficiency Runtime

Overview ยท Benchmarks ยท Efficiency ยท Inference ยท Live Demo ยท Sibling model


What it is

lemura-arabic-asr-lite is a compact multi-dialect Arabic ASR model built for accuracy and efficiency. It is a FastConformer-CTC acoustic model adapted in-house from the NVIDIA FastConformer foundation โ€” the contribution is the adaptation (the Arabic data curriculum and dialect coverage), not the foundation.

  • Small and fast โ€” ~115M parameters; runs comfortably on CPU and in real time, no GPU required.
  • Dialect-aware โ€” fine-tuned on ~2,900 hours of Arabic spanning MSA and the Gulf, Egyptian, Levantine and Maghrebi dialect groups, not MSA-only.
  • Robust on real audio โ€” strongest on broadcast, conversational and Gulf/MSA speech.
  • Open and simple โ€” a single .nemo file, loadable in a few lines with NVIDIA NeMo.

The result is leaderboard-grade Arabic ASR at edge scale โ€” #2 of 36 systems on the Open Universal Arabic ASR Leaderboard, ahead of every 2Bโ€“30B audio-LLM evaluated, at ~115M parameters.

Model summary

Modellemura-arabic-asr-lite โ€” compact multi-dialect Arabic ASR
TaskAutomatic speech recognition (audio โ†’ text)
ApproachDiscriminative ASR โ€” FastConformer encoder + CTC decoder
Trainingadapted from the NVIDIA FastConformer foundation; ~2,900 h Arabic fine-tuning across 5 dialect groups
Total parameters~115,000,000 (~0.12B)
Checkpoint size~0.4 GB (single .nemo file)
Audio input16 kHz mono (auto-resampled)
LanguagesArabic โ€” MSA + Gulf / Egyptian / Levantine / Maghrebi
RuntimeNVIDIA NeMo โ€” CPU ยท GPU ยท real-time
LicenseCC-BY-4.0

Benchmarks

Arabic dialectal ASR is hard โ€” heavily dialectal, conversational, code-switched speech is the frontier for every system. On the Open Universal Arabic ASR Leaderboard, lemura-arabic-asr-lite ranks #2 of 36 systems with an average WER of 25.08 % โ€” ahead of every model evaluated except one, including systems 17ร— to 250ร— larger.

Open Universal Arabic ASR Leaderboard โ€” full standings

Per-dataset WER % across all six leaderboard test sets. Lower is better; Avg WER is the ranking metric. Our rows were produced with the official leaderboard code (same normalizer, same WER metric). Ours in bold.

# Model Params Avg WER SADA CV-18 MASC-clean MASC-noisy MGB-2 Casablanca
1 lemuralabs/lemura-arabic-asr-lite (Ours) 0.12B 25.08 37.28 9.74 7.27 23.65 14.33 58.24
2 CohereLabs/cohere-transcribe-arabic-07-2026 ~2B 25.87 37.47 5.82 19.60 27.07 15.54 49.71
3 omnilingual-asr/omniASR_LLM_7B 7B 28.32 41.61 8.75 19.69 29.29 14.13 56.46
4 omnilingual-asr/omniASR_LLM_3B 3B 29.96 46.18 9.15 19.90 30.03 14.22 60.27
5 omnilingual-asr/omniASR_LLM_1B 1B 29.96 43.84 9.55 20.03 30.26 15.34 60.68
6 CohereLabs/cohere-transcribe-03-2026 ~2B 30.67 60.11 8.17 8.66 19.01 25.33 62.71
7 Qwen/Qwen3-Omni-30B-A3B-Instruct 30B 30.71 44.82 11.46 21.47 30.85 13.09 62.55
8 nvidia-conformer-ctc-large-arabic (lm) 0.6B 32.91 44.52 8.80 23.74 34.29 17.20 68.90
9 omnilingual-asr/omniASR_LLM_300M 0.3B 32.96 51.38 12.03 20.66 32.45 16.58 64.64
10 google/gemma-4-E4B-it 4B 32.98 43.40 19.65 24.86 33.59 17.72 58.63
11 Qwen/Qwen3-ASR-1.7B 1.7B 33.36 45.53 16.90 24.37 34.29 16.57 64.47
12 mistralai/Voxtral-Small-24B-2507 24B 34.47 50.82 15.25 23.96 34.43 16.03 66.30
13 nvidia-conformer-ctc-large-arabic (greedy) 0.6B 34.74 47.26 10.60 24.12 35.64 19.69 71.13
14 google/gemma-4-E2B-it 2B 35.87 46.23 23.76 27.47 36.15 20.72 60.87
15 openai/whisper-large-v3 1.5B 36.86 55.96 17.83 24.66 34.63 16.26 71.81
16 omnilingual-asr/omniASR_CTC_3B 3B 37.78 69.85 14.19 21.48 34.60 18.96 67.58
17 omnilingual-asr/omniASR_CTC_7B 7B 38.12 72.69 12.47 21.08 35.04 20.43 67.02
18 facebook/seamless-m4t-v2-large 2.3B 38.16 62.52 21.70 25.04 33.24 20.23 66.25
19 omnilingual-asr/omniASR_CTC_1B 1B 39.29 71.42 17.55 22.76 35.73 19.96 68.32
20 openai/whisper-large-v3-turbo 0.8B 40.05 60.36 25.73 25.51 37.16 17.75 73.79
21 openai/whisper-large-v2 1.5B 40.20 57.46 21.77 27.25 38.55 25.17 71.01
22 Qwen/Qwen3-ASR-0.6B 0.6B 42.19 53.75 28.28 31.34 42.63 25.45 71.68
23 openai/whisper-large 1.5B 42.57 63.24 26.04 28.89 40.79 24.28 72.18
24 mistralai/Voxtral-Mini-3B-2507 3B 42.58 63.65 22.12 28.37 41.27 22.56 77.52
25 asafaya/hubert-large-arabic-transcribe 0.3B 45.50 67.82 8.01 32.94 50.16 37.51 76.53
26 openai/whisper-medium 0.8B 45.57 67.71 28.07 29.99 42.91 29.32 75.44
27 nvidia-Parakeet-ctc-1.1b-concat 1.1B 46.54 70.70 26.34 30.49 45.95 24.94 80.80
28 omnilingual-asr/omniASR_CTC_300M 0.3B 46.65 78.11 27.90 28.40 43.26 26.85 75.35
29 nvidia-Parakeet-ctc-1.1b-universal 1.1B 51.96 73.58 40.01 36.16 50.03 30.68 81.30
30 microsoft/VibeVoice-ASR โ€” 52.99 69.83 44.25 32.95 52.43 25.10 93.37
31 facebook/mms-1b-all 1B 54.54 77.48 26.52 38.82 57.33 39.16 87.95
32 openai/whisper-small 0.24B 55.13 78.02 24.18 35.93 56.36 48.64 87.64
33 whitefox123/w2v-bert-2.0-arabic-4 0.6B 58.13 87.34 41.79 37.82 53.28 40.66 87.88
34 jonatasgrosman/wav2vec2-large-xlsr-53-arabic 0.3B 60.98 86.82 23.00 42.75 64.27 56.29 92.72
35 speechbrain/asr-wav2vec2-commonvoice-14-ar 0.1B 65.74 88.54 29.17 49.10 69.57 64.37 93.68

Bold = our model and our best-in-column results. lemura-arabic-asr-lite is the single best system on MASC-clean (7.27) and MASC-noisy (23.65) of every model evaluated, and the highest-ranked system below 1B parameters by a wide margin. Casablanca (Moroccan Darija) is the hardest set for every system.

Competitor rows are the published Open Universal Arabic ASR Leaderboard standings, reproduced for context; they are not our measurements.

Efficiency

The systems around us on the leaderboard are large generative audio-LLMs (2โ€“30B parameters) that need GPUs. lemura-arabic-asr-lite reaches the same accuracy tier with a ~115M-parameter CTC model:

lemura-arabic-asr-lite Typical top-5 systems
Parameters ~115M 2B โ€“ 30B
Hardware CPU or GPU GPU
Latency Real-time Seconds / clip
Footprint ~0.4 GB 4 โ€“ 60 GB
Avg WER 25.08 24.78 โ€“ 30.71

That makes it practical for on-device, low-cost and high-throughput Arabic transcription where the larger models are impractical.

Inference

import nemo.collections.asr as nemo_asr

model = nemo_asr.models.ASRModel.restore_from("asr_final.nemo")
print(model.transcribe(["audio.wav"]))   # 16 kHz mono

Try it live, no install: lemura-arabic-asr demo

Languages, dialects and tasks

Languages Arabic
Dialects MSA, Gulf/Khaleeji, Egyptian, Levantine, Maghrebi
Tasks Transcription (audio โ†’ text)
Audio 16 kHz mono, auto-resampled

Intended use and limitations

Intended: transcribing Arabic speech across dialects โ€” call centres, voice notes, media captioning, voice interfaces, and any deployment where CPU-only or real-time operation matters.

Limitations:

  • Maghrebi/Darija (Casablanca, 58.24 WER) remains hard, as it does for every system on the leaderboard.
  • Very noisy or far-field audio degrades accuracy.
  • CTC output has no built-in punctuation or diacritic restoration.
  • May reflect biases present in the training corpora.

Related model

For a larger, generative alternative with broader dialect fine-tuning, see lemura-arabic-asr-qwen3 (1.7B, audio-LLM). Note its published WER is an in-domain figure and is not directly comparable to the zero-shot leaderboard numbers above.

License

CC-BY-4.0.

Citation

@misc{lemura_arabic_asr_lite_2026,
  title  = {lemura-arabic-asr-lite: Compact Multi-Dialect Arabic Speech Recognition},
  author = {Lemura Labs},
  year   = {2026},
  url    = {https://huggingface.co/lemuralabs/lemura-arabic-asr-lite}
}

About Lemura Labs

Arabic-first, efficiency-first speech intelligence.

Lemura Labs builds compact, deployable speech and language models โ€” accuracy at a size and cost that works outside the datacentre.

Downloads last month
52
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Space using lemuralabs/lemura-arabic-asr-lite 1

Evaluation results