museglimmer-shim
OpenAI-compatible server for Muse Glimmer on Apple Silicon, via mlx_vlm; for local runners whose MLX build lags.
- What
- OpenAI-compatible server for Muse Glimmer on Apple Silicon, via mlx_vlm; for local runners whose MLX build lags.
- Cost
- Free
- Needs
- Serve Muse Glimmer from your Mac to any OpenAI-compatible client.
- Install
- Copy the installer prompt below into your Muse — your agent does the rest.
Curated by Skill Harbor: museglimmer-shim is a small OpenAI-compatible /v1/chat/completions server that runs Muse Glimmer on Apple Silicon. It exists because bundled MLX runtimes (like the one in LM Studio) lag behind the model: the mlx_vlm Python package on PyPI already supports Glimmer's architecture, so this project loads the model once, in-process, and serves it over a FastAPI app shaped exactly like the OpenAI Chat Completions API, including tool-calling translation, real vision input and streaming with heartbeat frames. Any OpenAI-compatible client keeps working unchanged, just pointed at a different host and port. By @GordonWei, listed here with credit to its creator. Honest caveats: Apple Silicon + Metal only, so no Linux and no Docker on macOS (no Metal GPU passthrough); a 30B model at roughly 13-15 tok/s means a large fixed system prompt routinely pushes single-request latency past 300-400 seconds, so budget client timeouts accordingly; one request at a time, a busy request gets an immediate 503 with Retry-After: 30 instead of queueing. Skill Harbor never reviews the code, review it yourself before use.
Version:
Install
Copy the install package below, then paste it into MuseCommunity-built. Skill Harbor doesn't audit code — review the source before installing.
Serve Muse Glimmer from your Mac to any OpenAI-compatible client. Prerequisites: an Apple Silicon Mac running macOS natively (not Docker, not Linux: Docker Desktop for Mac runs a Linux VM with no Metal GPU passthrough); Python 3; enough disk/RAM for the Muse Glimmer 30B 4-bit model. 1. Clone: `git clone https://github.com/GordonWei/museglimmer-shim.git` and cd into it. 2. `python3 -m venv venv && source venv/bin/activate && pip install -r requirements.txt` 3. Ad hoc, in the foreground: `uvicorn server:app --host 0.0.0.0 --port 8091`. As a persistent, auto-restarting macOS service: `./museglimmer-shim.sh start` (generates a launchd plist for this checkout's path; `status`, `logs`, `restart` and `stop` are also available). 4. Point any OpenAI-compatible client at http://localhost:8091. Endpoints: POST /v1/chat/completions, GET /v1/models, GET /health. 5. Choose the model with the MUSEGLIMMER_MODEL environment variable (default mlx-community/Muse-Glimmer-30B-4bit); swapping models needs only an env var and a restart. Notes: always pass `tools` in OpenAI's function-calling shape if the caller might want tool use, otherwise the model can emit garbled output; prefer stream: true, because some clients mistake a slow silent non-streaming request for a stalled connection and abort it.
Saved to your recent installs. Find it anytime on /connect.
Questions
How do I install a build?
Every product page includes a copy-paste install prompt. Paste it into your Muse and it sets the build up for you — no manual configuration.
Where does my money go?
Straight to the seller. Skill Harbor never processes payments: checkout happens on the seller’s own page, usually Stripe.
What does the ✓ next to a creator’s name mean?
It means we confirmed the identity of the person behind the listing. It says nothing about the code itself — always check a build before installing it.