Eval Harness
Practice eval-driven development: define capability and regression evals before coding, grade, and track pass@k.
- What
- Practice eval-driven development: define capability and regression evals before coding, grade, and track pass@k.
- Cost
- Free
- Needs
- Use "Eval Harness" with your Muse.
- Install
- Copy the installer prompt below into your Muse — your agent does the rest.
Curated by Skill Harbor: a formal eval-driven development (EDD) framework for AI coding sessions. Evals are the unit tests of AI development: define capability evals (can the agent do something new) and regression evals (did a change break existing behavior) BEFORE implementation, then grade with code-based deterministic graders, model-based LLM-as-judge rubrics, or human review with risk levels. Metrics: pass@k (at least one success in k attempts) and pass^k (all k trials succeed, the higher bar for critical paths), with recommended thresholds (pass@3 >= 0.90 for capability, pass^3 = 1.00 for release-critical regressions). Includes the eval workflow (define, implement, evaluate, report), storage layout (.claude/evals/ definitions, logs, baselines), eval anti-patterns (overfitting prompts to eval examples, happy-path-only measurement, flaky graders in release gates), and best practices (define before coding, deterministic graders first, version evals with code). Note: the skill's example utilities describe a capsule and replay framework where candidate execution is disabled by design (no verified OS containment backend); static checks and receipts are not execution evidence. By @affaan-m, listed here with credit to its creator. From the affaan-m/ECC repository (MIT). Honest caveats: a method framework, not an auto-runner; some paths follow Claude Code conventions. Skill Harbor never reviews the code, review it yourself before use.
Version:
Install
Copy the install package below, then paste it into MuseCommunity-built. Skill Harbor doesn't audit code — review the source before installing.
Use "Eval Harness" with your Muse. Prerequisites: the Muse app (mobile or web). No installation, no API key. A method framework, not an auto-runner. 1. Open the skill: https://github.com/affaan-m/ECC/blob/main/skills/eval-harness/SKILL.md and copy the full SKILL.md text. 2. Paste it into a chat with Muse and add: "Set up eval-driven development for [feature or agent task]. Define capability evals and regression evals first, with pass@k targets." 3. Ask Muse to draft the eval definitions, pick grader types (code first, model grader for open-ended, human for risky), and produce the eval report format after each round. Tip: define evals before any implementation. Writing them after is just documenting what you already built. Safety: a skill is plain-text instructions; it runs nothing by itself. Never paste secrets into a chat, and review anything Muse proposes before it acts.
Saved to your recent installs. Find it anytime on /connect.
Questions
How do I install a build?
Every product page includes a copy-paste install prompt. Paste it into your Muse and it sets the build up for you — no manual configuration.
Where does my money go?
Straight to the seller. Skill Harbor never processes payments: checkout happens on the seller’s own page, usually Stripe.
What does the ✓ next to a creator’s name mean?
It means we confirmed the identity of the person behind the listing. It says nothing about the code itself — always check a build before installing it.