openerotica/erotiquant3
Viewer • Updated • 5.12k • 47 • 8
[🏠Sentient Simulations] | [Discord] | [Patreon]
This repository contains a 4 bit GPTQ-quantized version of the ArliAI Llama 3.1 70B model using llm-compressor.
| Attribute | Value |
|---|---|
| Algorithm | GPTQ |
| Layers | Linear |
| Weight Scheme | W4A16 |
| Group Size | 128 |
| Calibration Dataset | openerotica/erotiquant3 |
| Calibration Sequence Length | 4096 |
| Calibration Samples | 512 |
The dataset was preprocessed with the following steps:
SYSTEM, USER, ASSISTANT).View the shell and python script used to quantize this model.
4 A40s with 300gb of ram was rented on runpod.
Quantization took approximately 11 hours with a total of $23.65 in compute costs. (And another $70 of me screwing up the quants like 10 times but anyways...)
Base model
ArliAI/Llama-3.1-70B-ArliAI-RPMax-v1.3