IndicTrans2 1B (indic→indic) — ONNX bundle [Q4F16 (4-bit Block Quantization, Lossy)]

This model is part of a suite of optimized/quantized ONNX versions of the base model. Other variants in this direction:

ONNX-exported and quantized version of ai4bharat/indictrans2-indic-indic-1B for in-browser and local edge inference.

  • Precision: Q4F16 (4-bit Block Quantization, Lossy)
  • Description: 4-bit quantization with float16 scale factors and block size of 32. Reduces model size significantly.
  • Source Pipeline & Details: For pipeline details, benchmarks, and usage instructions, see the indictrans2-onnx-export GitHub repository.

Built for use with Transformers.js and onnxruntime-web in the browser, with fast BPE tokenizer.json files that don't require the SentencePiece WASM runtime.

Performance Visualizations

These charts show overall tradeoffs, language-level parity, and category breakdown.

Overall Tradeoffs Language-Level Parity Category breakdown

Performance Tradeoffs & Size Comparison

Compared against the FP32 ONNX oracle on the golden evaluation fixtures.

Format Model Size Exact Match (Token) Exact Match (Text) SacreBLEU (Raw) Latency (Mean) Speedup vs. FP32
FP32 7.03 GB 100.00% 100.00% 100.00 94.7 ms 1.000x
FP16 3.52 GB 99.82% 99.82% 100.00 108.3 ms 0.874x
INT8 1.76 GB 83.64% 83.73% 94.22 43.7 ms 2.240x
Q4F16 900.0 MB 73.18% 73.18% 89.33 94.2 ms 1.087x

Language-Level Parity (Q4F16)

Exact match rates and translation quality (SacreBLEU / chrF) per language pair under this precision:

Language Code Total Fixtures Token Match Rate Text Match Rate SacreBLEU SacreBLEU (chrF)
asm_Beng 50 80.0% 80.0% 90.04 95.45
ben_Beng 50 84.0% 84.0% 93.66 97.91
brx_Deva 50 70.0% 70.0% 86.74 94.77
doi_Deva 50 84.0% 84.0% 91.61 95.65
gom_Deva 50 74.0% 74.0% 60.53 86.91
guj_Gujr 50 92.0% 92.0% 97.10 98.84
hin_Deva 50 88.0% 88.0% 95.56 97.85
kan_Knda 50 74.0% 74.0% 88.55 96.95
kas_Arab 50 80.0% 80.0% 91.20 95.83
mai_Deva 50 78.0% 78.0% 91.12 95.16
mal_Mlym 50 78.0% 78.0% 90.64 96.65
mar_Deva 50 70.0% 70.0% 86.08 93.60
mni_Beng 50 36.0% 36.0% 51.53 71.66
npi_Deva 50 66.0% 66.0% 79.76 89.70
ory_Orya 50 76.0% 76.0% 88.61 94.88
pan_Guru 50 80.0% 80.0% 91.95 95.84
san_Deva 50 76.0% 76.0% 87.90 95.70
sat_Olck 50 54.0% 54.0% 88.60 93.74
snd_Arab 50 46.0% 46.0% 54.02 63.78
tam_Taml 50 68.0% 68.0% 85.23 93.80
tel_Telu 50 76.0% 76.0% 90.25 95.14
urd_Arab 50 80.0% 80.0% 90.43 94.78

Category-Level Parity (Q4F16)

Exact match rates and translation quality grouped by category types:

Category Total Fixtures Token Match Rate Text Match Rate SacreBLEU SacreBLEU (chrF)
Generic 286 72.7% 72.7% 85.86 91.79
Lexicon 264 68.2% 68.2% 84.83 92.04
Numerals 264 73.1% 73.1% 82.06 93.34
Politics 286 78.3% 78.3% 86.29 92.61

Translation Mismatch Examples

Here is a sample of up to 5 translation mismatches compared to the FP32 oracle. Many mismatches represent minor synonym differences or spacing variations.

Mismatch #1 (Category: Numerals)

  • Source (asm_Beng → ben_Beng): আপুনি অনুগ্ৰহ কৰি মোক নিকটতম চিকিৎসালয়খন বিচাৰি উলিওৱাত সহায় কৰিব পাৰিবনে?
  • Expected (FP32): আপনি কি দয়া করে আমাকে নিকটতম হাসপাতালটি খুঁজে পেতে সাহায্য করতে পারেন?
  • Actual (Q4F16): আপনি কি দয়া করে আমাকে নিকটতম হাসপাতাল খুঁজে পেতে সাহায্য করতে পারেন?

Mismatch #2 (Category: Politics)

  • Source (asm_Beng → ben_Beng): আজি শুক্ৰবাৰ, 2026 চনৰ 3 জুলাই।
  • Expected (FP32): আজ শুক্রবার, 3 জুলাই, 2026।
  • Actual (Q4F16): আজ শুক্রবার, 2026 সালের 3 জুলাই।

Mismatch #3 (Category: Lexicon)

  • Source (asm_Beng → ben_Beng): পেট্ৰ "লৰ মূল্য পাঁচ টকা বৃদ্ধি পাইছে।
  • Expected (FP32): পেট্রোলের দাম বেড়েছে 50 টাকা।
  • Actual (Q4F16): পেট্রোলের দাম বেড়েছে 500 টাকা।

Mismatch #4 (Category: Politics)

  • Source (asm_Beng → ben_Beng): বেয়া বতৰৰ বাবে বিমানখন বাতিল কৰা হৈছিল।
  • Expected (FP32): খারাপ আবহাওয়ার কারণে বিমানটি বাতিল করা হয়।
  • Actual (Q4F16): খারাপ আবহাওয়ার কারণে বিমানটি বাতিল করা হয়েছিল।

Mismatch #5 (Category: Politics)

  • Source (asm_Beng → ben_Beng): অনুগ্ৰহ কৰি পৰৱৰ্তী সোমবাৰৰ ভিতৰত সম্পূৰ্ণ কৰা কামটো দাখিল কৰক।
  • Expected (FP32): দয়া করে পরবর্তী সোমবারের মধ্যে কাজটি সম্পন্ন করুন।
  • Actual (Q4F16): দয়া করে পরবর্তী সোমবারের মধ্যে কাজটি সম্পূর্ণ করে জমা দিন।

Files

  • encoder_model.onnx (and optional .onnx.data weights sidecar)
  • decoder_model.onnx and decoder_with_past_model.onnx (share decoder_shared.onnx.data when present)
  • translate.py — self-contained Python inference helper (see Usage below)
  • Fast tokenizer config files (tokenizer_src.json, tokenizer_tgt.json, tokenizer_meta.json)
  • Model configuration configs (config.json, generation_config.json)

Usage Example (Python, onnxruntime)

# translate.py is included in this repo alongside the ONNX bundle.
# You can also find it (and read the full source) at:
#   https://github.com/Hari31416/indictrans2-onnx-export/blob/main/src/translate.py

from translate import IndicTransONNX

# Pass a HF repo ID for automatic download, or a local bundle directory path
model = IndicTransONNX("hari31416/indictrans2-indic-indic-1B-ONNX-q4f16")
print(model.translate("चुनाव कौन जीतेगा?", src_lang="hin_Deva", tgt_lang="tam_Taml"))

Required packages:

pip install onnxruntime tokenizers huggingface-hub

License

MIT (preserved from upstream AI4Bharat).

Downloads last month
5
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for hari31416/indictrans2-indic-indic-1B-ONNX-q4f16

Quantized
(4)
this model

Collection including hari31416/indictrans2-indic-indic-1B-ONNX-q4f16