Muse Glimmer on one Arc Pro B70
vLLM-XPU + DFlash recipe pushing Muse Glimmer 30B to ~90 tok/s on a single Intel Arc Pro B70; full benchmarks.
- What
- vLLM-XPU + DFlash recipe pushing Muse Glimmer 30B to ~90 tok/s on a single Intel Arc Pro B70; full benchmarks.
- Cost
- Free
- Needs
- Reproduce the Muse Glimmer B70 recipe.
- Install
- Copy the installer prompt below into your Muse — your agent does the rest.
Curated by Skill Harbor: a research recipe and benchmark suite that runs Meta's Muse Glimmer 30B on a single Intel Arc Pro B70 (32 GB) far above the public baseline. By switching from llama.cpp to vLLM-XPU with a GPTQ 4-bit target, XPU graphs and a quantized DFlash draft (20 speculative tokens), the author documents 89-101 tok/s greedy single-stream decode on completed-answer workloads (GSM8K, HumanEval) versus roughly 27-32 tok/s for the published llama.cpp recipes, plus a 131k-token native context profile and a concurrency sweep (840 tok/s aggregate at C96 in burst). Every number ships with its reproduction protocol, scripts, checksums and stated limits. By @mgaruccio (Mike Garuccio), listed here with credit to its creator. Honest caveats: this is a research profile, not a production-safety claim; a near-capacity forced-length six-request test produced two empty responses and that failure remains unresolved; you need the exact hardware (one Arc Pro B70, Linux with the xe driver, Docker, /dev/dri), this is not CUDA. Skill Harbor never reviews the code, review it yourself before use.
Version:
Install
Copy the install package below, then paste it into MuseCommunity-built. Skill Harbor doesn't audit code — review the source before installing.
Reproduce the Muse Glimmer B70 recipe. Prerequisites: ONE Intel Arc Pro B70 (32 GB); Linux with the xe driver; Docker; /dev/dri available. This is not CUDA. 1. Download the weights: `hf download mgaruccio/Muse-Glimmer-30B-GPTQ-Int4-sym-G128 --local-dir ./models/target` and `hf download mgaruccio/Muse-Glimmer-30B-assistant-GPTQ-Int4-sym-G128 --local-dir ./models/draft`, then compare the shard hashes against docs/checksums.md in the repo. 2. `export MODEL="$PWD/models/target"`, `export DRAFT="$PWD/models/draft"`, `export DFLASH_KV_MODE=none`. 3. Launch: `bash scripts/start-muse-vllm-dflash-c1-graph-draft-gptq.sh`, then `bash scripts/wait-vllm-health.sh 420 8000`. 4. Verify: `curl -sf http://127.0.0.1:8000/v1/models` (expect served id muse-glimmer-gptq). 5. Measure: `python3 scripts/vllm-dflash-share-suite.py 3 2048 share-suite.json`; the process exits non-zero on any unquoteable prompt, so do not publish a headline from a partial run. Note: Muse streams delta.reasoning then delta.content, so a client that only reads content looks idle until thinking finishes. See docs/concurrency.md for the C8/K4 multi-client profile, and docs/share-suite.md for the measurement protocol.
Saved to your recent installs. Find it anytime on /connect.
Questions
How do I install a build?
Every product page includes a copy-paste install prompt. Paste it into your Muse and it sets the build up for you — no manual configuration.
Where does my money go?
Straight to the seller. Skill Harbor never processes payments: checkout happens on the seller’s own page, usually Stripe.
What does the ✓ next to a creator’s name mean?
It means we confirmed the identity of the person behind the listing. It says nothing about the code itself — always check a build before installing it.