GPTQ post-training 4-bit quantization for LLMs: fit big models on small GPUs
Group-wise 4-bit quantization with ~2% perplexity loss, 4x memory reduction and 3-4x inference speedup, kernel backends (ExLlamaV2, Marlin, Triton), transformers/QLoRA integration, and benchmark tables
- What
- Group-wise 4-bit quantization with ~2% perplexity loss, 4x memory reduction and 3-4x inference speedup, kernel backends (ExLlamaV2, Marlin, Triton), transformers/QLoRA integration, and benchmark tables
- Cost
- Free
- Needs
- Python with pip and a CUDA GPU for real quantization (CPU-only works for reading the workflow, not for quantizing); HuggingFace access for pre-quantized models
- Install
- Copy the installer prompt below into your Muse — your agent does the rest.
Curated by Skill Harbor — @orchestra-research's GPTQ skill: the complete post-training 4-bit quantization workflow for LLMs — how group-wise quantization works (per-group scale/zero-point, Hessian-aware error minimization), the group-size trade-off table (32/128/256/1024 with memory, accuracy and speed guidance — 128 is the recommended default), three ready configs (standard 4-bit, high-accuracy 3-bit, max-accuracy 4-bit with small groups), kernel backends (ExLlamaV2 default, Marlin for Ampere+ GPUs, Triton for Linux), auto-gptq install and usage, loading pre-quantized models from HuggingFace (TheBloke's 1000+ GPTQ models), quantizing your own model with calibration data, transformers integration, QLoRA fine-tuning on a GPTQ base (a 70B model trainable on a single A100 80GB), multi-GPU deployment and CPU offloading patterns, and benchmark tables (4x memory reduction — Llama 2-70B 140GB → 35GB; 3.4-4.8x inference speedup; <2% perplexity degradation). Honest caveats: guidance plus pip-installable libraries — needs CUDA GPUs for real quantization; when to prefer AWQ or bitsandbytes is covered honestly (better accuracy on Ampere/Ada, simpler 8-bit integration); performance claims are the skill's, validate on your own hardware. MIT licensed. Skill Harbor never reviews the code, review it yourself before use. Discovered via skills.sh.
Version:
Install
Prerequisites: Python with pip and a CUDA GPU for real quantization (CPU-only works for reading the workflow, not for quantizing); HuggingFace access for pre-quantized models Install "GPTQ post-training 4-bit quantization for LLMs: fit big models on small GPUs" for me. It gives my agent @orchestra-research's GPTQ playbook: group-wise 4-bit quantization configs (standard, high-accuracy 3-bit, max-accuracy small-group), kernel backends (ExLlamaV2, Marlin, Triton), auto-gptq install, loading pre-quantized models, quantizing your own model with calibration data, transformers/QLoRA integration, multi-GPU and CPU-offload deployment patterns, and benchmark tables. MIT licensed. Repository: https://github.com/orchestra-research/ai-research-skills/blob/main/10-optimization/gptq/SKILL.md 1. Fetch the SKILL.md file (and any helper files) from the repository path into a temporary folder and summarize what it does in one or two sentences. 2. Safety check: review the SKILL.md and scripts for anything suspicious (unexpected network calls, shell commands, credential harvesting). This repo should contain zero secrets in code, credentials only via the secure vault, allowed hosts declared in the SKILL.md. Verify that holds here; STOP on any red flag and tell me. 3. Install it as a skill: copy SKILL.md and its helper files into the agent's skills directory, in a folder named "gptq". 4. Verify with no network calls: frontmatter valid, files in place. 5. Report what was installed, where, and what I still need to do myself (e.g. confirm a CUDA GPU is available and which path to take — load a pre-quantized model from HuggingFace or quantize my own model with calibration data; validate benchmark claims on my own hardware; pick the kernel backend for my GPU generation — Marlin needs Ampere+). GitHub is optional: if I have a GitHub account or the gh CLI, you may use it; otherwise public access is fine. Never require it unless it's in the prerequisites above. Rules: don't touch anything outside the temp folder and the install target. If anything looks off, stop and ask me.
Questions
How do I install a build?
Every product page includes a copy-paste install prompt. Paste it into your Muse and it sets the build up for you — no manual configuration.
Where does my money go?
Straight to the seller. Skill Harbor never processes payments: checkout happens on the seller’s own page, usually Stripe.
What does the ✓ next to a creator’s name mean?
It means we confirmed the identity of the person behind the listing. It says nothing about the code itself — always check a build before installing it.