Jetson LLM serving benchmarks: vLLM, llama.cpp, and Ollama with comparable JSON output
Runtime-matched wrapper scripts for vLLM, llama.cpp/GGUF, and Ollama emitting the same JSON envelope (TTFT, ITL, throughput), plus Jetson-specific reading guidance (bandwidth saturation, power modes, thermal warnings)
- What
- Runtime-matched wrapper scripts for vLLM, llama.cpp/GGUF, and Ollama emitting the same JSON envelope (TTFT, ITL, throughput), plus Jetson-specific reading guidance (bandwidth saturation, power modes, thermal warnings)
- Cost
- Free
- Needs
- a Jetson device hosting (or able to host) the model runtime — this skill benchmarks on-device; plus a running vLLM server for the vLLM path, or an Ollama daemon for the Ollama path
- Install
- Copy the installer prompt below into your Muse — your agent does the rest.
Curated by Skill Harbor — @nvidia's Jetson LLM benchmark skill: reproducible serving benchmarks on Jetson hardware with structured JSON an agent can diff. It picks the wrapper matching your runtime — `bench_vllm.sh` against a running OpenAI-compatible vLLM server (concurrency sweeps), `bench_llama_cpp.sh` through the NVIDIA-AI-IOT container for local GGUF models, `bench_ollama.sh` against an Ollama daemon via its REST API — and each emits the same envelope (model, SKU, generation, L4T, container, TTFT/ITL/TPOT/throughput p50/p99, config, warnings). Warnings fire on non-max power modes, >5% background GPU use, and `tegrastats` thermal throttling; SKU facts are populated live from the device, never guessed. Includes Jetson-specific reading guidance most LLMs don't know: on Orin Nano/NX, single-stream vs concurrency-8 throughput gaps signal memory-bandwidth saturation (shrink quantization before tuning), TTFT regressions after a JetPack upgrade are usually CUDA graph cache misses (re-warm), and Thor NVFP4 numbers never mix with Orin W4A16 without a `quant` column. Honest caveats: runs ON the Jetson device only; vLLM needs an already-running server (this skill benchmarks, it doesn't serve); Ollama numbers are single-stream and not comparable to vLLM sweeps; the llama.cpp path pulls a Docker container — tell the user before running it. Apache-2.0 licensed. Skill Harbor never reviews the code, review it yourself before use. Discovered via skills.sh.
Version:
Install
Prerequisites: a Jetson device hosting (or able to host) the model runtime — this skill benchmarks on-device; plus a running vLLM server for the vLLM path, or an Ollama daemon for the Ollama path Install "Jetson LLM serving benchmarks: vLLM, llama.cpp, and Ollama with comparable JSON output" for me. It gives my agent @nvidia's Jetson benchmarking workflow: pick the runtime-matched wrapper script (bench_vllm.sh, bench_llama_cpp.sh, bench_ollama.sh), always run with --help first, do a warmup pass before measured runs, compare the shared JSON envelope (TTFT/ITL/TPOT/throughput p50/p99 plus device facts populated live), and read results with Jetson-specific guidance (bandwidth saturation, power modes, thermal warnings). Apache-2.0 licensed. Repository: https://github.com/nvidia/skills/blob/main/skills/jetson-llm-benchmark/SKILL.md 1. Fetch the SKILL.md file (and any helper files) from the repository path into a temporary folder and summarize what it does in one or two sentences. 2. Safety check: review the SKILL.md and scripts for anything suspicious (unexpected network calls, shell commands, credential harvesting). This repo should contain zero secrets in code, credentials only via the secure vault, allowed hosts declared in the SKILL.md. Verify that holds here; STOP on any red flag and tell me. 3. Install it as a skill: copy SKILL.md and its helper files into the agent's skills directory, in a folder named "jetson-llm-benchmark". 4. Verify with no network calls: frontmatter valid, files in place. 5. Report what was installed, where, and what I still need to do myself (e.g. point the agent at the Jetson device and which runtime path to benchmark — vLLM server already running, Ollama daemon, or local GGUF file; put the device in MAXN before measuring; never invent device facts — the wrapper fills them in live; tell me before the llama.cpp path pulls its Docker container). GitHub is optional: if I have a GitHub account or the gh CLI, you may use it; otherwise public access is fine. Never require it unless it's in the prerequisites above. Rules: don't touch anything outside the temp folder and the install target. If anything looks off, stop and ask me.
Questions
How do I install a build?
Every product page includes a copy-paste install prompt. Paste it into your Muse and it sets the build up for you — no manual configuration.
Where does my money go?
Straight to the seller. Skill Harbor never processes payments: checkout happens on the seller’s own page, usually Stripe.
What does the ✓ next to a creator’s name mean?
It means we confirmed the identity of the person behind the listing. It says nothing about the code itself — always check a build before installing it.