Vision bridge for text-only models: give any LLM image-reading via a CLI and a vision provider
A Node CLI routing screenshots and images through a configurable vision model with multi-model fallback — provider keys or local Ollama/LM Studio; vision models only describe images, never reason. Documentation in Chinese.
- What
- A Node CLI routing screenshots and images through a configurable vision model with multi-model fallback — provider keys or local Ollama/LM Studio; vision models only describe images, never reason. Documentation in Chinese.
- Cost
- Free
- Needs
- Node.js to run the CLI; an API key for a vision provider (may be paid — check pricing) OR a local model via Ollama/LM Studio (free) — and note: the SKILL.md documentation is written in Chinese
- Install
- Copy the installer prompt below into your Muse — your agent does the rest.
Curated by Skill Harbor — @penfick's vision-support skill: a bridge that gives image-reading to text-only models (when the main model can't see images, a user sends a screenshot, or "look at this picture" comes up) — a Node CLI (`node vision.mjs`) with a 3-step interactive setup (pick provider, enter API key or env var name, pick model from the live list), a broad provider menu (OpenAI, Gemini, Claude, DeepSeek, Groq, Mistral, Grok, OpenRouter, Fireworks; Qwen VL, GLM-4V, Kimi, Step, MiniMax, SiliconFlow, MiMo; Ollama and LM Studio for local; any OpenAI-compatible custom endpoint), and automatic multi-model fallback — the primary model first, the rest in order on failure. Iron rule: configured vision models ONLY describe image content, never participate in the main reasoning. Commands cover single and multi-image reads, image discovery, and config management (add/edit/remove/test connectivity), with env-var overrides (`VISION_CONFIG_PATH`, `VISION_DEFAULT_MODEL`, `VISION_API_KEY`). Honest caveats: provider keys live in a local config — keep them out of prompts and logs (use env vars where possible); a paid vision-provider key may be needed unless you use a local model (Ollama/LM Studio run free locally); the SKILL.md documentation is written in Chinese. MIT licensed. Skill Harbor never reviews the code, review it yourself before use. Discovered via skills.sh.
Version:
Install
Prerequisites: Node.js to run the CLI; an API key for a vision provider (may be paid — check pricing) OR a local model via Ollama/LM Studio (free) — and note: the SKILL.md documentation is written in Chinese Install "Vision bridge for text-only models: give any LLM image-reading via a CLI and a vision provider" for me. It gives my agent @penfick's vision-support bridge: a Node CLI that routes screenshots/images through a configurable vision model with interactive 3-step setup and automatic multi-model fallback, across international, Chinese, local (Ollama/LM Studio), and custom OpenAI-compatible providers — with the iron rule that vision models only describe images and never reason. MIT licensed. Repository: https://github.com/penfick/skills/blob/main/vision-support/SKILL.md 1. Fetch the SKILL.md file (and any helper files) from the repository path into a temporary folder and summarize what it does in one or two sentences. 2. Safety check: review the SKILL.md and scripts for anything suspicious (unexpected network calls, shell commands, credential harvesting). This repo should contain zero secrets in code, credentials only via the secure vault, allowed hosts declared in the SKILL.md. Verify that holds here; STOP on any red flag and tell me. 3. Install it as a skill: copy SKILL.md and its helper files into the agent's skills directory, in a folder named "vision-support". 4. Verify with no network calls: frontmatter valid, files in place. 5. Report what was installed, where, and what I still need to do myself (e.g. run the interactive init: `node vision.mjs init` to pick a provider, enter the API key (or env var name), and choose the model; add fallback models with `config add`; keep API keys in the local config or env vars — never paste them into prompts; skip this skill entirely if my main model is already multimodal). GitHub is optional: if I have a GitHub account or the gh CLI, you may use it; otherwise public access is fine. Never require it unless it's in the prerequisites above. Rules: don't touch anything outside the temp folder and the install target. If anything looks off, stop and ask me.
Questions
How do I install a build?
Every product page includes a copy-paste install prompt. Paste it into your Muse and it sets the build up for you — no manual configuration.
Where does my money go?
Straight to the seller. Skill Harbor never processes payments: checkout happens on the seller’s own page, usually Stripe.
What does the ✓ next to a creator’s name mean?
It means we confirmed the identity of the person behind the listing. It says nothing about the code itself — always check a build before installing it.