somersetent 's Collections

4-bit Quantization

4-bit in LLMs means compressing model weights from 16-bit to 4-bit precision. This cuts the model's memory size by about 75%.