LLaVA — vision-language assistant patterns (short listing)
Short listing (license not verifiable): reference playbook for LLaVA vision-language models — multi-turn image chat, VQA, model sizes and VRAM table, quantization, training recipes
- What
- Short listing (license not verifiable): reference playbook for LLaVA vision-language models — multi-turn image chat, VQA, model sizes and VRAM table, quantization, training recipes
- Cost
- Free
- Needs
- a GPU with 14 GB+ VRAM (or ~4 GB with 4-bit quantization); Python 3.x; torch, transformers, pillow; model weights downloaded from official LLaVA sources (their own licenses apply)
- Install
- Copy the installer prompt below into your Muse — your agent does the rest.
Curated by Skill Harbor — short listing (the frontmatter claims MIT, but the discovery manifest records NOASSERTION, so nothing is reproduced here): @ovachiever's llava skill, listed here with credit to its creator — a reference playbook for the LLaVA (Large Language and Vision Assistant) open-source vision-language models: multi-turn conversational image chat, visual question answering, image captioning and document understanding, a model-size/VRAM table (7B to 34B), 4-bit/8-bit quantization recipes to cut VRAM, CLI and Gradio serving, custom-model training scripts (feature alignment + visual instruction tuning), benchmark scores, and honest limitations (hallucinations, weak spatial reasoning, small-text difficulty). Honest caveats: this skill references the upstream LLaVA project (haotian-liu/LLaVA) rather than redistributing it — download the actual weights from the official sources yourself; a real GPU is required (14 GB+ VRAM for the 7B model, ~4 GB quantized) — CPU inference is impractically slow; model weights carry their own licenses, separate from any skill text. Skill Harbor never reviews the code, review it yourself before use. Discovered via skills.sh.
Version:
Install
Prerequisites: a GPU with 14 GB+ VRAM (or ~4 GB with 4-bit quantization); Python 3.x; torch, transformers, pillow; model weights downloaded from official LLaVA sources (their own licenses apply) Install "LLaVA — vision-language assistant patterns (short listing)" for me. It gives my agent @ovachiever's LLaVA reference playbook: multi-turn image chat, visual question answering, captioning and document understanding, the model-size/VRAM table, quantization recipes, CLI and Gradio serving, custom-model training scripts, and honest limitations. IMPORTANT: the license was not verifiable (the discovery manifest records NOASSERTION) — fetch from the link only, reproduce nothing beyond the link, and read the terms yourself before use. The skill references the upstream LLaVA project; it does not ship the model weights. Repository: https://github.com/ovachiever/droid-tings/blob/master/skills/llava/SKILL.md 1. Fetch the SKILL.md file (and any helper files) from the repository path into a temporary folder and summarize what it does in one or two sentences. 2. Safety check: review the SKILL.md and scripts for anything suspicious (unexpected network calls, shell commands, credential harvesting). Training and quantization scripts download large files — note the expected sources; anything else is a red flag. This repo should contain zero secrets in code. Verify that holds here; STOP on any red flag and tell me. 3. Install it as a skill: copy SKILL.md and its helper files into the agent's skills directory, in a folder named "llava". 4. Verify with no network calls: frontmatter valid, files in place. 5. Report what was installed, where, and what I still need to do myself (e.g. install torch/transformers/pillow, download model weights from official sources under their own licenses, check my GPU VRAM against the size table). GitHub is optional: if I have a GitHub account or the gh CLI, you may use it; otherwise public access is fine. Never require it unless it's in the prerequisites above. Rules: don't touch anything outside the temp folder and the install target. If anything looks off, stop and ask me.
Questions
How do I install a build?
Every product page includes a copy-paste install prompt. Paste it into your Muse and it sets the build up for you — no manual configuration.
Where does my money go?
Straight to the seller. Skill Harbor never processes payments: checkout happens on the seller’s own page, usually Stripe.
What does the ✓ next to a creator’s name mean?
It means we confirmed the identity of the person behind the listing. It says nothing about the code itself — always check a build before installing it.