2-bit quantization reduces the memory required to store each model weight from 16 bits down to 2 bits. This compresses the model size by roughly 8x.
-
huihui-ai/Huihui-DeepSeek-V4-Flash-0731-abliterated-GGUF
Text Generation • 284B • Updated • 370k • 162 -
MaziyarPanahi/Qwen2.5-Coder-0.5B-QwQ-draft-GGUF
Text Generation • 0.5B • Updated • 747 • 5 -
MaziyarPanahi/Qwen2.5-7B-HomerAnvita-NerdMix-GGUF
Text Generation • 8B • Updated • 135 • 1 -
legraphista/Yi-Coder-9B-Chat-IMat-GGUF
Text Generation • 9B • Updated • 737 • 2