Megatron Bridge GPU memory tuning: OOM fixes, fragmentation and PEFT tricks
Reduce peak GPU memory in Megatron Bridge training — expandable segments, PEFT input re-gather, parallelism resizing, activation recompute, measured results on Llama3 70B
- What
- Reduce peak GPU memory in Megatron Bridge training — expandable segments, PEFT input re-gather, parallelism resizing, activation recompute, measured results on Llama3 70B
- Cost
- Free
- Needs
- an NVIDIA Megatron Bridge / NeMo distributed-training setup on multi-GPU hardware — the skill is methodology the agent follows, not software to install
- Install
- Copy the installer prompt below into your Muse — your agent does the rest.
Curated by Skill Harbor — @nvidia's GPU memory-tuning skill for Megatron Bridge: a decision-first playbook for training OOMs, starting from the finding that most OOMs are memory fragmentation, not raw capacity. The single most effective fix: `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True` (zero throughput cost). Then, in order: for LoRA with sequence parallelism, enable `sequence_parallel_input_regather` to stop retaining full gathered LoRA-A inputs; add selective activation recompute (`recompute_modules=[core_attn]`) with its ~16% utilization cost; resize parallelism (with warnings — doubling TP costs -28% throughput on Llama3 70B, raising PP costs ~6% and CPU offloading is blocked when PP > 1); the skill includes measured strategy comparisons on Llama3 70B SFT (32x H100 80GB), PEFT+SP input re-gather results (Qwen3-8B, Qwen3-30B-A3B, GPT-OSS-120B), compatibility constraints (expandable segments incompatible with `--use-nccl-ub`, CUDA-graph rules), code anchors, a failure-diagnosis table and a verification procedure with pytest commands. Honest caveats: deeply specialized — NVIDIA Megatron Bridge / NeMo distributed training on multi-GPU clusters only, not for single-GPU hobby training; references companion skills (activation recompute, FSDP) and docs in the same repo that are not bundled here. Apache-2.0 licensed. Skill Harbor never reviews the code, review it yourself before use. Discovered via skills.sh.
Version:
Install
Prerequisites: an NVIDIA Megatron Bridge / NeMo distributed-training setup on multi-GPU hardware — the skill is methodology the agent follows, not software to install Install "Megatron Bridge GPU memory tuning: OOM fixes, fragmentation and PEFT tricks" for me. It gives my agent @nvidia's GPU memory playbook for Megatron Bridge: the decision order for training OOMs (expandable segments first, then PEFT input re-gather for LoRA+SP, selective activation recompute, parallelism resizing with measured tradeoffs), compatibility constraints, code anchors, a failure-diagnosis table and verification procedures. Apache-2.0 licensed. Repository: https://github.com/nvidia/skills/blob/main/skills/nemo-mbridge-perf-memory-tuning/SKILL.md 1. Fetch the SKILL.md file (and any helper files) from the repository path into a temporary folder and summarize what it does in one or two sentences. 2. Safety check: review the SKILL.md and scripts for anything suspicious (unexpected network calls, shell commands, credential harvesting). Review the files for suspicious commands or credential handling; stop on any red flag and never assume the repository is safe. Tell me if anything looks off. 3. Install it as a skill: copy SKILL.md and its helper files into the agent's skills directory, in a folder named "nemo-mbridge-perf-memory-tuning". 4. Verify with no network calls: frontmatter valid, files in place. 5. Report what was installed, where, and what I still need to do myself (e.g. hand the agent your Megatron Bridge training config and OOM symptoms; note the companion skills activation-recompute and megatron-fsdp are referenced but not bundled here). GitHub is optional: if I have a GitHub account or the gh CLI, you may use it; otherwise public access is fine. Never require it unless it's in the prerequisites above. Rules: don't touch anything outside the temp folder and the install target. If anything looks off, stop and ask me.
Questions
How do I install a build?
Every product page includes a copy-paste install prompt. Paste it into your Muse and it sets the build up for you — no manual configuration.
Where does my money go?
Straight to the seller. Skill Harbor never processes payments: checkout happens on the seller’s own page, usually Stripe.
What does the ✓ next to a creator’s name mean?
It means we confirmed the identity of the person behind the listing. It says nothing about the code itself — always check a build before installing it.