Text Generation
MLX
Safetensors
mistral
quantized
4bit
indian-languages
multilingual
apple-silicon
sarvam
conversational
4-bit precision
Instructions to use Jimmi42/sarvam-m-4bit-mlx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use Jimmi42/sarvam-m-4bit-mlx with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("Jimmi42/sarvam-m-4bit-mlx") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- MLX LM
How to use Jimmi42/sarvam-m-4bit-mlx with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "Jimmi42/sarvam-m-4bit-mlx"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "Jimmi42/sarvam-m-4bit-mlx" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Jimmi42/sarvam-m-4bit-mlx", "messages": [ {"role": "user", "content": "Hello"} ] }'
File size: 3,906 Bytes
f6c7c50 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 | # 🛠️ LM Studio Setup Guide for Sarvam-M 4-bit MLX
## 🔍 Problem Diagnosis
The "EOS token issue" where the model stops after a few words is caused by **incorrect prompt formatting**, not the EOS token itself.
## ✅ Solution: Proper Chat Format
### **Option 1: Use Chat Mode in LM Studio**
1. **Load the model** in LM Studio
2. **Switch to Chat mode** (not Playground mode)
3. **Set Chat Template** to "Custom" or "Mistral"
4. **Configure these settings:**
```
System Prompt: "You are a helpful assistant."
Chat Template Format:
<s>[SYSTEM_PROMPT]{system}[/SYSTEM_PROMPT][INST]{user}[/INST]
```
### **Option 2: Manual Prompt Format**
If using Playground mode, format your prompts like this:
**Simple Format:**
```
[INST] Your question here [/INST]
```
**With System Prompt:**
```
<s>[SYSTEM_PROMPT]You are a helpful assistant.[/SYSTEM_PROMPT][INST]Your question here[/INST]
```
### **Option 3: Thinking Mode (Advanced)**
For reasoning tasks, use:
```
<s>[SYSTEM_PROMPT]You are a helpful assistant. Think deeply before answering the user's question. Do the thinking inside <think>...</think> tags.[/SYSTEM_PROMPT][INST]Your question here[/INST]<think>
```
## 🎛️ Recommended LM Studio Settings
### **Generation Parameters:**
- **Max Tokens:** 512-1024
- **Temperature:** 0.7-0.8
- **Top P:** 0.9
- **Repetition Penalty:** 1.1
- **Context Length:** 4096
### **Stop Sequences:**
Add these stop sequences:
- `</s>`
- `[/INST]`
- `\n\nUser:`
- `\n\nHuman:`
### **MLX Settings:**
- ✅ Enable MLX acceleration
- ✅ Use GPU memory
- Set batch size to 1-4
## 🧪 Test Examples
### **Test 1: Basic Math**
```
Prompt: [INST] What is 2+2? Please explain your answer. [/INST]
Expected: The sum of 2 and 2 is **4**. [explanation follows]
```
### **Test 2: Reasoning**
```
Prompt: <s>[SYSTEM_PROMPT]Think before answering.[/SYSTEM_PROMPT][INST]Why is the sky blue?[/INST]<think>
Expected: <think>[reasoning]</think> The sky appears blue because...
```
### **Test 3: Hindi Language**
```
Prompt: [INST] भारत की राजधानी क्या है? [/INST]
Expected: भारत की राजधानी **नई दिल्ली** है...
```
## 🚨 Common Issues & Fixes
| Issue | Cause | Solution |
|-------|-------|----------|
| Empty responses | No chat template | Use `[INST]...[/INST]` format |
| Stops after few words | Wrong stop tokens | Remove `</s>` from stop sequences |
| Repeating text | Low repetition penalty | Increase to 1.1-1.2 |
| Slow responses | CPU inference | Enable MLX acceleration |
## 📝 Model Information
- **Format:** MLX 4-bit quantized
- **Languages:** English + 10 Indic languages
- **Context:** 4096 tokens
- **Based on:** Mistral Small architecture
- **Special Features:** Thinking mode, multi-language support
## 🔗 Working Example Commands
### **MLX-LM (Command Line):**
```bash
# Basic chat
python -m mlx_lm.generate --model Jimmi42/sarvam-m-4bit-mlx --prompt "[INST] Hello, how are you? [/INST]" --max-tokens 100
# With thinking
python -m mlx_lm.generate --model Jimmi42/sarvam-m-4bit-mlx --prompt "<s>[SYSTEM_PROMPT]Think deeply.[/SYSTEM_PROMPT][INST]Explain quantum physics[/INST]<think>" --max-tokens 200
```
### **Python Code:**
```python
from mlx_lm import load, generate
model, tokenizer = load('Jimmi42/sarvam-m-4bit-mlx')
# Format prompt correctly
messages = [{"role": "user", "content": "What is AI?"}]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
response = generate(model, tokenizer, prompt, max_tokens=100)
print(response)
```
## ✅ Success Checklist
- [ ] Model loads in LM Studio with MLX enabled
- [ ] Chat template is set to Custom/Mistral format
- [ ] Test prompt: `[INST] Hello [/INST]` generates response
- [ ] Stop sequences configured correctly
- [ ] Generation parameters optimized
- [ ] Multi-language capability tested |