← Products

Developer tools
⚙ Needs: a Python (3.10+) LLM application to evaluate; the pi…

Eval-driven development: build a real eval pipeline for Python LLM apps

Instrument your Python LLM app with pixie-qa — define eval criteria, build golden datasets, run real evals with LLM-as-judge scorers, analyze results into an action plan

At a glance
What
Instrument your Python LLM app with pixie-qa — define eval criteria, build golden datasets, run real evals with LLM-as-judge scorers, analyze results into an action plan
Cost
Free
Needs
a Python (3.10+) LLM application to evaluate; the pixie-qa package (the skill's resources/setup.sh installs it) — the skill is a guided workflow, not software
Install
Copy the installer prompt below into your Muse — your agent does the rest.

Version:

@
Created by: @github
⌁

Install

Prerequisites: a Python (3.10+) LLM application to evaluate; the pixie-qa package (the skill's resources/setup.sh installs it) — the skill is a guided workflow, not software Install "Eval-driven development: build a real eval pipeline for Python LLM apps" for me. It gives my agent @github's eval-driven workflow: analyze the Python LLM app and define eval criteria from real failure modes, instrument data boundaries with wrap() calls, write a Runnable against the real entry point, capture a reference trace, build runnable evaluators (agent evaluators for semantic criteria), create a golden dataset from real-world data, run pixie test to real scores, and produce a prioritized action plan. MIT-licensed. Repository: https://github.com/github/awesome-copilot/blob/main/skills/eval-driven-dev/SKILL.md 1. Fetch the SKILL.md file (and any helper files) from the repository path into a temporary folder and summarize what it does in one or two sentences. 2. Safety check: review the SKILL.md and scripts for anything suspicious (unexpected network calls, shell commands, credential harvesting). This repo should contain zero secrets in code, credentials only via the secure vault, allowed hosts declared in the SKILL.md. Verify that holds here; STOP on any red flag and tell me. 3. Install it as a skill: copy SKILL.md and its helper files into the agent's skills directory, in a folder named "eval-driven-dev". 4. Verify with no network calls: frontmatter valid, files in place. 5. Report what was installed, where, and what I still need to do myself (e.g. run the skill's setup.sh to install pixie-qa and start the results web server; point the agent at my Python LLM app; provide LLM API credentials — evals make real LLM calls with real costs). GitHub is optional: if I have a GitHub account or the gh CLI, you may use it; otherwise public access is fine. Never require it unless it's in the prerequisites above. Rules: don't touch anything outside the temp folder and the install target. If anything looks off, stop and ask me.

?

Questions

How do I install a build?

Every product page includes a copy-paste install prompt. Paste it into your Muse and it sets the build up for you — no manual configuration.

Where does my money go?

Straight to the seller. Skill Harbor never processes payments: checkout happens on the seller’s own page, usually Stripe.

What does the ✓ next to a creator’s name mean?

It means we confirmed the identity of the person behind the listing. It says nothing about the code itself — always check a build before installing it.