somersetent 's Collections

3-bit Quantization

3-bit quantization in LLMs means compressing model weights from 16-bit precision down to an average of 3 bits per weight.