Regex vs LLM for Structured Text
Cheap document parsing for Muse: a regex-first pipeline that handles 95%+ of structured text, with a confidence scorer and LLM validation only for edge cases. ~95% cost savings.
- What
- Cheap document parsing for Muse: a regex-first pipeline that handles 95%+ of structured text, with a confidence scorer and LLM validation only for edge cases. ~95% cost savings.
- Cost
- Free
- Needs
- Use "Regex vs LLM for Structured Text" with your Muse.
- Install
- Copy the installer prompt below into your Muse — your agent does the rest.
Curated by Skill Harbor: a decision framework and hybrid pipeline for parsing structured text that keeps Muse from burning money on LLM calls regex could have handled. The key insight: regex handles 95-98% of cases cheaply and deterministically, reserve expensive LLM calls for the remaining edge cases. Architecture: a regex parser extracts structure first, a text cleaner removes noise (markers, page numbers, artifacts), a confidence scorer flags low-confidence extractions (few choices, missing answer, short text), and an LLM validator fixes only the flagged items, returning corrected JSON or CORRECT. Includes a complete Python implementation (frozen dataclasses, never mutate parsed items), real-world metrics from a production quiz pipeline (410 items: 98.0% regex success, 8 low-confidence items, about 5 LLM calls, ~95% cost savings versus all-LLM, 93% test coverage), a decision tree (consistent repeating format goes regex-first, free-form variable text goes LLM directly), best practices (start with regex even imperfect, cheapest Haiku-class model for validation, TDD for parsers, log pipeline metrics), and anti-patterns to avoid (sending everything to an LLM, regex for free-form text, skipping confidence scoring, mutating parsed objects). Use when parsing quizzes, forms, invoices, receipts or tables, choosing between regex and LLM for extraction, or optimizing extraction cost and accuracy. By @affaan-m, listed here with credit to its creator. From the affaan-m/ECC repository (MIT). Honest caveats: metrics come from one production pipeline, your mileage may vary with messier inputs ; LLM validation needs an API key and a client ; regex still needs tests for edge cases and encoding issues. Skill Harbor never reviews the code, review it yourself before use.
Version:
Install
Copy the install package below, then paste it into MuseThe install prompt below already includes the vetting steps: your agent follows the community checklist before installing anything with executable code. Want more?
Use "Regex vs LLM for Structured Text" with your Muse. Prerequisites: none to install. Pure guidance; an LLM API key helps if you want the validator stage (cheapest Haiku-class model is enough). 1. Open the skill: https://github.com/affaan-m/ECC/blob/main/skills/regex-vs-llm-structured-text/SKILL.md and copy the full SKILL.md text. 2. Paste it into a chat with Muse and add: "Design a hybrid regex-first extraction pipeline for: [describe your documents]." 3. Ask it to start with the regex parser and confidence scorer, add LLM validation only for flagged items, and log regex success rate versus LLM call count. Tip: run the decision tree first ("is this format consistent and repeating?") before writing any parser code. Safety: a skill is plain-text instructions; it runs nothing by itself. Review generated code before merging.
Saved to your recent installs. Find it anytime on /connect.
Questions
How do I install a build?
Every product page includes a copy-paste install prompt. Paste it into your Muse and it sets the build up for you — no manual configuration.
Where does my money go?
Straight to the seller. Skill Harbor never processes payments: checkout happens on the seller’s own page, usually Stripe.
What does the ✓ next to a creator’s name mean?
It means we confirmed the identity of the person behind the listing. It says nothing about the code itself — always check a build before installing it.