Agent Eval
Compare coding agents head-to-head on your own tasks with pass rate, cost, time, and consistency metrics.
- What
- Compare coding agents head-to-head on your own tasks with pass rate, cost, time, and consistency metrics.
- Cost
- Free
- Needs
- Use "Agent Eval" with your Muse.
- Install
- Copy the installer prompt below into your Muse — your agent does the rest.
Curated by Skill Harbor: a lightweight CLI tool plus method for comparing coding agents head-to-head on reproducible tasks you define in YAML (prompt, files to touch, judge criteria). Judges can be code-based and deterministic (pytest, build), pattern-based (grep), or model-based (LLM-as-judge). Each agent run gets its own git worktree for isolation, no Docker required, and the report shows pass rate, API cost, wall-clock time, and consistency across repeated runs. Best practices included: 3 to 5 real-workload tasks (not toy examples), 3+ trials per agent, pinned commits for reproducibility, at least one deterministic judge per task, and tracking cost alongside pass rate. By @affaan-m, listed here with credit to its creator. From the affaan-m/ECC repository (MIT). Note: the underlying CLI lives in the joaquinhuigomez/agent-eval repository linked by the skill. Honest caveats: developer-oriented, you need a git repo and the agent CLIs you want to compare installed locally. Skill Harbor never reviews the code, review it yourself before use.
Version:
Install
Copy the install package below, then paste it into MuseCommunity-built. Skill Harbor doesn't audit code — review the source before installing.
Use "Agent Eval" with your Muse. Prerequisites: a git repository with real tasks to test, and the coding-agent CLIs you want to compare installed on your machine (e.g. Claude Code, Aider, Codex). The Muse app (mobile or web) for the planning side. 1. Open the skill: https://github.com/affaan-m/ECC/blob/main/skills/agent-eval/SKILL.md and copy the full SKILL.md text. 2. Paste it into a chat with Muse and add: "Help me define 3 to 5 benchmark tasks for [my project] in the agent-eval YAML format. Repo at [path], pinned commit [sha]." 3. Install the agent-eval CLI from the repository linked in the skill (review the source first), write your task YAMLs, then run: agent-eval run --task tasks/<name>.yaml --agent claude-code --agent aider --runs 3 4. Ask Muse to help interpret the report table (pass rate, cost, time, consistency) and make the call. Tip: start with tasks from your real workload, not toy examples, and always include at least one deterministic judge (tests or build) per task. Safety: a skill is plain-text instructions; it runs nothing by itself. Review any command before running it, never paste secrets into a chat, and review anything Muse proposes before it acts.
Saved to your recent installs. Find it anytime on /connect.
Questions
How do I install a build?
Every product page includes a copy-paste install prompt. Paste it into your Muse and it sets the build up for you — no manual configuration.
Where does my money go?
Straight to the seller. Skill Harbor never processes payments: checkout happens on the seller’s own page, usually Stripe.
What does the ✓ next to a creator’s name mean?
It means we confirmed the identity of the person behind the listing. It says nothing about the code itself — always check a build before installing it.