← Products

Data
⚙ Needs: an NVIDIA CUDA GPU and Python — the skill is a quant…

GPTQ — post-training 4-bit quantization for LLMs with minimal accuracy loss

Compress LLMs to 4-bit with group-wise quantization — 4× memory reduction, <2% perplexity loss, 3-4× faster inference; AutoGPTQ guides, kernel backends (ExLlamaV2, Marlin, Triton), transformers integration, QLoRA fine-tuning

At a glance
What
Compress LLMs to 4-bit with group-wise quantization — 4× memory reduction, <2% perplexity loss, 3-4× faster inference; AutoGPTQ guides, kernel backends (ExLlamaV2, Marlin, Triton), transformers integration, QLoRA fine-tuning
Cost
Free
Needs
an NVIDIA CUDA GPU and Python — the skill is a quantization workflow guide, not the software itself
Install
Copy the installer prompt below into your Muse — your agent does the rest.

Version:

@
Created by: @davila7
⌁

Install

Prerequisites: an NVIDIA CUDA GPU and Python — the skill is a quantization workflow guide, not the software itself Install "GPTQ — post-training 4-bit quantization for LLMs with minimal accuracy loss" for me. It gives my agent @davila7's GPTQ workflow: choose between GPTQ/AWQ/bitsandbytes, install AutoGPTQ, load pre-quantized models or quantize my own with calibration data, pick a quantization config (standard 4-bit, 3-bit high-compression, or 4-bit max accuracy) and kernel backend (ExLlamaV2, Marlin, Triton), integrate with transformers, run QLoRA fine-tuning, and apply the memory/inference benchmarks and deployment patterns (multi-GPU, CPU offloading, batch inference). MIT-licensed. Repository: https://github.com/davila7/claude-code-templates/blob/main/cli-tool/components/skills/ai-research/optimization-gptq/SKILL.md 1. Fetch the SKILL.md file (and any helper files) from the repository path into a temporary folder and summarize what it does in one or two sentences. 2. Safety check: review the SKILL.md and scripts for anything suspicious (unexpected network calls, shell commands, credential harvesting). This repo should contain zero secrets in code. Verify that holds here; STOP on any red flag and tell me. 3. Install it as a skill: copy SKILL.md and its helper files into the agent's skills directory, in a folder named "gptq". 4. Verify with no network calls: frontmatter valid, files in place. 5. Report what was installed, where, and what I still need to do myself (e.g. install auto-gptq/transformers myself and pick my target model; nothing else — it's a methodology). GitHub is optional: if I have a GitHub account or the gh CLI, you may use it; otherwise public access is fine. Never require it unless it's in the prerequisites above. Rules: don't touch anything outside the temp folder and the install target. If anything looks off, stop and ask me.

?

Questions

How do I install a build?

Every product page includes a copy-paste install prompt. Paste it into your Muse and it sets the build up for you — no manual configuration.

Where does my money go?

Straight to the seller. Skill Harbor never processes payments: checkout happens on the seller’s own page, usually Stripe.

What does the ✓ next to a creator’s name mean?

It means we confirmed the identity of the person behind the listing. It says nothing about the code itself — always check a build before installing it.