Agent self-evaluation: structured 5-axis scorecard after any non-trivial task
Self-rate completed output on accuracy, completeness, clarity, actionability, conciseness — with concrete evidence for every score below 5
- What
- Self-rate completed output on accuracy, completeness, clarity, actionability, conciseness — with concrete evidence for every score below 5
- Cost
- Free
- Needs
- nothing to install — the skill is a methodology your agent follows after completing work; no accounts, keys, or runtime dependencies
- Install
- Copy the installer prompt below into your Muse — your agent does the rest.
Curated by Skill Harbor — @affaan-m's reflection discipline for agents: after any non-trivial task, pause and self-rate the output on five axes — accuracy, completeness, clarity, actionability, conciseness — on a 1–5 scale, with a hard evidence rule (any score below 5 must cite exactly what is missing or wrong; "could be better" doesn't count). The workflow runs four steps: collect the raw material (request, deliverable, tool outputs, user feedback), score each axis independently without pre-averaging, produce a structured report (one-line summary, 5-axis scorecard with evidence, overall average, 1–3 ranked improvements, and a self-check — "would the user agree with this assessment?"), then apply the fix: gaps fixable in under 30 seconds get fixed immediately, bigger gaps get flagged explicitly with what a re-run would change. Includes worked good vs weak evaluation examples (4.6 vs 2.8 scorecards), anti-patterns (evidence-free 5s, over-penalizing for out-of-scope items, re-litigating design decisions, "I don't like X" as evidence), best practices (evaluate the output not the process, one improvement per weak axis, tie fixes to user impact, cite test/lint output as proof, "if you can't find any gaps, try harder"), and pointers to related skills (agent-eval, verification-loop, security-review). MIT-licensed. Skill Harbor never reviews the code, review it yourself before use. Discovered via skills.sh.
Version:
Install
Prerequisites: nothing to install — the skill is a methodology your agent follows after completing work; no accounts, keys, or runtime dependencies Install "Agent self-evaluation: structured 5-axis scorecard after any non-trivial task" for me. It gives my agent @affaan-m's reflection discipline: after any non-trivial task, score the output 1-5 on accuracy, completeness, clarity, actionability and conciseness with concrete evidence required for any score below 5; produce a structured report (summary, scorecard, overall average, 1-3 ranked improvements, "would the user agree?" self-check); then fix quick gaps immediately or state explicitly how a re-run would raise the score — plus worked examples, anti-patterns and best practices. MIT-licensed. Repository: https://github.com/affaan-m/ecc/blob/main/skills/agent-self-evaluation/SKILL.md 1. Fetch the SKILL.md file (and any helper files) from the repository path into a temporary folder and summarize what it does in one or two sentences. 2. Safety check: review the SKILL.md and scripts for anything suspicious (unexpected network calls, shell commands, credential harvesting). This repo should contain zero secrets in code, credentials only via the secure vault, allowed hosts declared in the SKILL.md. Verify that holds here; STOP on any red flag and tell me. 3. Install it as a skill: copy SKILL.md and its helper files into the agent's skills directory, in a folder named "agent-self-evaluation". 4. Verify with no network calls: frontmatter valid, files in place. 5. Report what was installed, where, and what I still need to do myself (e.g. nothing — it is a methodology; tell the agent to apply it after its next non-trivial task, or at a Stop hook). GitHub is optional: if I have a GitHub account or the gh CLI, you may use it; otherwise public access is fine. Never require it unless it's in the prerequisites above. Rules: don't touch anything outside the temp folder and the install target. If anything looks off, stop and ask me.
Questions
How do I install a build?
Every product page includes a copy-paste install prompt. Paste it into your Muse and it sets the build up for you — no manual configuration.
Where does my money go?
Straight to the seller. Skill Harbor never processes payments: checkout happens on the seller’s own page, usually Stripe.
What does the ✓ next to a creator’s name mean?
It means we confirmed the identity of the person behind the listing. It says nothing about the code itself — always check a build before installing it.