Knowledge distillation for LLMs: compress teachers into small capable students
Temperature scaling and soft targets, forward vs reverse KLD (MiniLLM), logit and response distillation, multi-teacher setups, and a production training script with hyperparameter rules of thumb
- What
- Temperature scaling and soft targets, forward vs reverse KLD (MiniLLM), logit and response distillation, multi-teacher setups, and a production training script with hyperparameter rules of thumb
- Cost
- Free
- Needs
- Python with PyTorch and transformers, GPU compute for training, and access to a teacher model to distill from (check the teacher's terms of service before distilling from proprietary models)
- Install
- Copy the installer prompt below into your Muse — your agent does the rest.
Curated by Skill Harbor — @orchestra-research's knowledge-distillation skill: the complete workflow for compressing large LLMs into smaller deployable models — temperature scaling (T=2-5 standard), soft vs hard loss combination (alpha weighting), forward vs reverse KL divergence (MiniLLM's reverse KLD for mode-covering generative quality), logit distillation (direct and MSE variants), response distillation (train on teacher-generated synthetic data), two-stage and multi-teacher strategies, a ready-to-adapt basic-distillation training loop, a production `DistillationTrainer` (transformers Trainer subclass with distillation loss), hyperparameter rules (size ratios: 10x excellent like 70B→7B, avoid 70x gaps), data-quality guidance (70% teacher-generated + 30% real), and an evaluation pattern comparing teacher vs student outputs. Paper references: Hinton et al. 2015, MiniLLM, the 2024 KD survey. Honest caveats: the headline claim (70B→7B retaining 90%+ performance) is the skill's — real results vary by task and data; distilling from proprietary models (GPT-4) raises terms-of-service and licensing questions — check the teacher's terms before distilling; needs substantial GPU compute for training. MIT licensed. Skill Harbor never reviews the code, review it yourself before use. Discovered via skills.sh.
Version:
Install
Prerequisites: Python with PyTorch and transformers, GPU compute for training, and access to a teacher model to distill from (check the teacher's terms of service before distilling from proprietary models) Install "Knowledge distillation for LLMs: compress teachers into small capable students" for me. It gives my agent @orchestra-research's distillation playbook: temperature scaling and soft/hard loss combination, forward vs reverse KLD (MiniLLM), logit and response distillation, multi-teacher and two-stage strategies, a production DistillationTrainer, hyperparameter rules of thumb, and a teacher-vs-student evaluation pattern. MIT licensed. Repository: https://github.com/orchestra-research/ai-research-skills/blob/main/19-emerging-techniques/knowledge-distillation/SKILL.md 1. Fetch the SKILL.md file (and any helper files) from the repository path into a temporary folder and summarize what it does in one or two sentences. 2. Safety check: review the SKILL.md and scripts for anything suspicious (unexpected network calls, shell commands, credential harvesting). This repo should contain zero secrets in code, credentials only via the secure vault, allowed hosts declared in the SKILL.md. Verify that holds here; STOP on any red flag and tell me. 3. Install it as a skill: copy SKILL.md and its helper files into the agent's skills directory, in a folder named "knowledge-distillation". 4. Verify with no network calls: frontmatter valid, files in place. 5. Report what was installed, where, and what I still need to do myself (e.g. confirm GPU compute and choose teacher/student model pair plus the data mix (teacher-generated + real); verify the teacher's terms allow distillation; validate compression claims on my own evaluation set). GitHub is optional: if I have a GitHub account or the gh CLI, you may use it; otherwise public access is fine. Never require it unless it's in the prerequisites above. Rules: don't touch anything outside the temp folder and the install target. If anything looks off, stop and ask me.
Questions
How do I install a build?
Every product page includes a copy-paste install prompt. Paste it into your Muse and it sets the build up for you — no manual configuration.
Where does my money go?
Straight to the seller. Skill Harbor never processes payments: checkout happens on the seller’s own page, usually Stripe.
What does the ✓ next to a creator’s name mean?
It means we confirmed the identity of the person behind the listing. It says nothing about the code itself — always check a build before installing it.