IndicTrans2 1B (indic→indic) — ONNX bundle [Q4F16 (4-bit Block Quantization, Lossy)]
This model is part of a suite of optimized/quantized ONNX versions of the base model. Other variants in this direction:
- FP32 (Full Precision / Base):
hari31416/indictrans2-indic-indic-1B-ONNX- FP16 (Half Precision):
hari31416/indictrans2-indic-indic-1B-ONNX-fp16- INT8 (Dynamic Quantization):
hari31416/indictrans2-indic-indic-1B-ONNX-int8- Q4F16 (4-bit Block Quantization):
hari31416/indictrans2-indic-indic-1B-ONNX-q4f16(Current)
ONNX-exported and quantized version of ai4bharat/indictrans2-indic-indic-1B
for in-browser and local edge inference.
- Precision: Q4F16 (4-bit Block Quantization, Lossy)
- Description: 4-bit quantization with float16 scale factors and block size of 32. Reduces model size significantly.
- Source Pipeline & Details: For pipeline details, benchmarks, and usage instructions, see the indictrans2-onnx-export GitHub repository.
Built for use with Transformers.js and onnxruntime-web in the browser, with fast BPE tokenizer.json files that don't require the SentencePiece WASM runtime.
Performance Visualizations
These charts show overall tradeoffs, language-level parity, and category breakdown.
Performance Tradeoffs & Size Comparison
Compared against the FP32 ONNX oracle on the golden evaluation fixtures.
| Format | Model Size | Exact Match (Token) | Exact Match (Text) | SacreBLEU (Raw) | Latency (Mean) | Speedup vs. FP32 |
|---|---|---|---|---|---|---|
| FP32 | 7.03 GB | 100.00% | 100.00% | 100.00 | 94.7 ms | 1.000x |
| FP16 | 3.52 GB | 99.82% | 99.82% | 100.00 | 108.3 ms | 0.874x |
| INT8 | 1.76 GB | 83.64% | 83.73% | 94.22 | 43.7 ms | 2.240x |
| Q4F16 | 900.0 MB | 73.18% | 73.18% | 89.33 | 94.2 ms | 1.087x |
Language-Level Parity (Q4F16)
Exact match rates and translation quality (SacreBLEU / chrF) per language pair under this precision:
| Language Code | Total Fixtures | Token Match Rate | Text Match Rate | SacreBLEU | SacreBLEU (chrF) |
|---|---|---|---|---|---|
| asm_Beng | 50 | 80.0% | 80.0% | 90.04 | 95.45 |
| ben_Beng | 50 | 84.0% | 84.0% | 93.66 | 97.91 |
| brx_Deva | 50 | 70.0% | 70.0% | 86.74 | 94.77 |
| doi_Deva | 50 | 84.0% | 84.0% | 91.61 | 95.65 |
| gom_Deva | 50 | 74.0% | 74.0% | 60.53 | 86.91 |
| guj_Gujr | 50 | 92.0% | 92.0% | 97.10 | 98.84 |
| hin_Deva | 50 | 88.0% | 88.0% | 95.56 | 97.85 |
| kan_Knda | 50 | 74.0% | 74.0% | 88.55 | 96.95 |
| kas_Arab | 50 | 80.0% | 80.0% | 91.20 | 95.83 |
| mai_Deva | 50 | 78.0% | 78.0% | 91.12 | 95.16 |
| mal_Mlym | 50 | 78.0% | 78.0% | 90.64 | 96.65 |
| mar_Deva | 50 | 70.0% | 70.0% | 86.08 | 93.60 |
| mni_Beng | 50 | 36.0% | 36.0% | 51.53 | 71.66 |
| npi_Deva | 50 | 66.0% | 66.0% | 79.76 | 89.70 |
| ory_Orya | 50 | 76.0% | 76.0% | 88.61 | 94.88 |
| pan_Guru | 50 | 80.0% | 80.0% | 91.95 | 95.84 |
| san_Deva | 50 | 76.0% | 76.0% | 87.90 | 95.70 |
| sat_Olck | 50 | 54.0% | 54.0% | 88.60 | 93.74 |
| snd_Arab | 50 | 46.0% | 46.0% | 54.02 | 63.78 |
| tam_Taml | 50 | 68.0% | 68.0% | 85.23 | 93.80 |
| tel_Telu | 50 | 76.0% | 76.0% | 90.25 | 95.14 |
| urd_Arab | 50 | 80.0% | 80.0% | 90.43 | 94.78 |
Category-Level Parity (Q4F16)
Exact match rates and translation quality grouped by category types:
| Category | Total Fixtures | Token Match Rate | Text Match Rate | SacreBLEU | SacreBLEU (chrF) |
|---|---|---|---|---|---|
| Generic | 286 | 72.7% | 72.7% | 85.86 | 91.79 |
| Lexicon | 264 | 68.2% | 68.2% | 84.83 | 92.04 |
| Numerals | 264 | 73.1% | 73.1% | 82.06 | 93.34 |
| Politics | 286 | 78.3% | 78.3% | 86.29 | 92.61 |
Translation Mismatch Examples
Here is a sample of up to 5 translation mismatches compared to the FP32 oracle. Many mismatches represent minor synonym differences or spacing variations.
Mismatch #1 (Category: Numerals)
- Source (asm_Beng → ben_Beng):
আপুনি অনুগ্ৰহ কৰি মোক নিকটতম চিকিৎসালয়খন বিচাৰি উলিওৱাত সহায় কৰিব পাৰিবনে? - Expected (FP32):
আপনি কি দয়া করে আমাকে নিকটতম হাসপাতালটি খুঁজে পেতে সাহায্য করতে পারেন? - Actual (Q4F16):
আপনি কি দয়া করে আমাকে নিকটতম হাসপাতাল খুঁজে পেতে সাহায্য করতে পারেন?
Mismatch #2 (Category: Politics)
- Source (asm_Beng → ben_Beng):
আজি শুক্ৰবাৰ, 2026 চনৰ 3 জুলাই। - Expected (FP32):
আজ শুক্রবার, 3 জুলাই, 2026। - Actual (Q4F16):
আজ শুক্রবার, 2026 সালের 3 জুলাই।
Mismatch #3 (Category: Lexicon)
- Source (asm_Beng → ben_Beng):
পেট্ৰ "লৰ মূল্য পাঁচ টকা বৃদ্ধি পাইছে। - Expected (FP32):
পেট্রোলের দাম বেড়েছে 50 টাকা। - Actual (Q4F16):
পেট্রোলের দাম বেড়েছে 500 টাকা।
Mismatch #4 (Category: Politics)
- Source (asm_Beng → ben_Beng):
বেয়া বতৰৰ বাবে বিমানখন বাতিল কৰা হৈছিল। - Expected (FP32):
খারাপ আবহাওয়ার কারণে বিমানটি বাতিল করা হয়। - Actual (Q4F16):
খারাপ আবহাওয়ার কারণে বিমানটি বাতিল করা হয়েছিল।
Mismatch #5 (Category: Politics)
- Source (asm_Beng → ben_Beng):
অনুগ্ৰহ কৰি পৰৱৰ্তী সোমবাৰৰ ভিতৰত সম্পূৰ্ণ কৰা কামটো দাখিল কৰক। - Expected (FP32):
দয়া করে পরবর্তী সোমবারের মধ্যে কাজটি সম্পন্ন করুন। - Actual (Q4F16):
দয়া করে পরবর্তী সোমবারের মধ্যে কাজটি সম্পূর্ণ করে জমা দিন।
Files
encoder_model.onnx(and optional.onnx.dataweights sidecar)decoder_model.onnxanddecoder_with_past_model.onnx(sharedecoder_shared.onnx.datawhen present)translate.py— self-contained Python inference helper (see Usage below)- Fast tokenizer config files (
tokenizer_src.json,tokenizer_tgt.json,tokenizer_meta.json) - Model configuration configs (
config.json,generation_config.json)
Usage Example (Python, onnxruntime)
# translate.py is included in this repo alongside the ONNX bundle.
# You can also find it (and read the full source) at:
# https://github.com/Hari31416/indictrans2-onnx-export/blob/main/src/translate.py
from translate import IndicTransONNX
# Pass a HF repo ID for automatic download, or a local bundle directory path
model = IndicTransONNX("hari31416/indictrans2-indic-indic-1B-ONNX-q4f16")
print(model.translate("चुनाव कौन जीतेगा?", src_lang="hin_Deva", tgt_lang="tam_Taml"))
Required packages:
pip install onnxruntime tokenizers huggingface-hub
License
MIT (preserved from upstream AI4Bharat).
- Downloads last month
- 5
Model tree for hari31416/indictrans2-indic-indic-1B-ONNX-q4f16
Base model
ai4bharat/indictrans2-indic-indic-1B

