GPTQ — post-training 4-bit quantization for LLMs with minimal accuracy loss
Compress LLMs to 4-bit with group-wise quantization — 4× memory reduction, <2% perplexity loss, 3-4× faster inference; AutoGPTQ guides, kernel backends (ExLlamaV2, Marlin, Triton), transformers integration, QLoRA fine-tuning
- What
- Compress LLMs to 4-bit with group-wise quantization — 4× memory reduction, <2% perplexity loss, 3-4× faster inference; AutoGPTQ guides, kernel backends (ExLlamaV2, Marlin, Triton), transformers integration, QLoRA fine-tuning
- Cost
- Free
- Needs
- an NVIDIA CUDA GPU and Python — the skill is a quantization workflow guide, not the software itself
- Install
- Copy the installer prompt below into your Muse — your agent does the rest.
Curated by Skill Harbor — @davila7's gptq skill: a practical guide to GPTQ post-training quantization for running large LLMs on limited GPU memory. Covers when to use GPTQ vs AWQ vs bitsandbytes, quick start (install AutoGPTQ, load pre-quantized models from HuggingFace, quantize your own model with calibration data), group-wise quantization mechanics and the group-size trade-off table, three configs (standard 4-bit recommended, 3-bit high-compression, 4-bit maximum accuracy), kernel backends (ExLlamaV2 default-fastest, Marlin for Ampere+, Triton Linux-only), direct transformers integration, QLoRA fine-tuning (a 70B model trainable on a single A100 80GB), performance benchmarks (memory reduction per model, tokens/sec, WikiText-2 perplexity degradation <2%), common patterns (multi-GPU, CPU offloading, batch inference), and how to find pre-quantized models. Honest caveats: targets CUDA GPUs (RTX 4090, A100 class) — limited CPU/Mac utility; AutoGPTQ and backend versions move fast; model quality claims reflect the benchmarks quoted, verify on your workload. MIT licensed (frontmatter and manifest agree). Skill Harbor never reviews the code, review it yourself before use. Discovered via skills.sh.
Version:
Install
Prerequisites: an NVIDIA CUDA GPU and Python — the skill is a quantization workflow guide, not the software itself Install "GPTQ — post-training 4-bit quantization for LLMs with minimal accuracy loss" for me. It gives my agent @davila7's GPTQ workflow: choose between GPTQ/AWQ/bitsandbytes, install AutoGPTQ, load pre-quantized models or quantize my own with calibration data, pick a quantization config (standard 4-bit, 3-bit high-compression, or 4-bit max accuracy) and kernel backend (ExLlamaV2, Marlin, Triton), integrate with transformers, run QLoRA fine-tuning, and apply the memory/inference benchmarks and deployment patterns (multi-GPU, CPU offloading, batch inference). MIT-licensed. Repository: https://github.com/davila7/claude-code-templates/blob/main/cli-tool/components/skills/ai-research/optimization-gptq/SKILL.md 1. Fetch the SKILL.md file (and any helper files) from the repository path into a temporary folder and summarize what it does in one or two sentences. 2. Safety check: review the SKILL.md and scripts for anything suspicious (unexpected network calls, shell commands, credential harvesting). This repo should contain zero secrets in code. Verify that holds here; STOP on any red flag and tell me. 3. Install it as a skill: copy SKILL.md and its helper files into the agent's skills directory, in a folder named "gptq". 4. Verify with no network calls: frontmatter valid, files in place. 5. Report what was installed, where, and what I still need to do myself (e.g. install auto-gptq/transformers myself and pick my target model; nothing else — it's a methodology). GitHub is optional: if I have a GitHub account or the gh CLI, you may use it; otherwise public access is fine. Never require it unless it's in the prerequisites above. Rules: don't touch anything outside the temp folder and the install target. If anything looks off, stop and ask me.
Questions
How do I install a build?
Every product page includes a copy-paste install prompt. Paste it into your Muse and it sets the build up for you — no manual configuration.
Where does my money go?
Straight to the seller. Skill Harbor never processes payments: checkout happens on the seller’s own page, usually Stripe.
What does the ✓ next to a creator’s name mean?
It means we confirmed the identity of the person behind the listing. It says nothing about the code itself — always check a build before installing it.