PluginBench
Skill
Pass
Audit score 90

regex-vs-llm-structured-text

affaan-m/ecc

Choose regex for structured text (95%+ accuracy), add LLM only for low-confidence edge cases to cut costs by ~95%.

What is regex-vs-llm-structured-text?

A decision framework for parsing structured text like quizzes, forms, and invoices. Start with regex to handle the majority of cases cheaply and deterministically, then use confidence scoring to identify edge cases that benefit from LLM validation. This hybrid approach balances accuracy and cost.

  • Provides a decision tree for choosing between regex and LLM based on text format consistency
  • Implements confidence scoring to flag low-confidence extractions programmatically
  • Demonstrates a hybrid pipeline that combines regex parsing with LLM validation for edge cases only
  • Includes production metrics showing 98% regex success rate with ~95% cost savings vs all-LLM approach
  • Offers best practices and anti-patterns for structured text parsing

How to install regex-vs-llm-structured-text

npx skills add null --skill regex-vs-llm-structured-text
Claude Code
Cursor
Windsurf
Cline

How to use regex-vs-llm-structured-text

  1. 1.Write regex patterns to match your structured text format (e.g., question ID, text, choices, answer)
  2. 2.Implement a confidence scorer that flags items with missing fields, unusual structure, or low match quality
  3. 3.Set a confidence threshold (e.g., 0.95) to identify items needing LLM review
  4. 4.Build a hybrid pipeline that runs regex first, scores confidence, then sends only low-confidence items to an LLM validator
  5. 5.Monitor metrics (regex success rate, LLM call count) to track pipeline health and adjust patterns as needed

Use cases

Good for
  • Parsing quiz or exam questions with multiple-choice answers from PDFs or documents
  • Extracting form data from structured templates where most entries follow a consistent pattern
  • Processing invoices or receipts to identify line items, amounts, and metadata
  • Validating document structure (headers, sections, tables) where some fields may be malformed
  • Building cost-efficient pipelines that handle 95%+ of cases with regex and delegate only ambiguous cases to LLM
Who it's for
  • Backend engineers building document processing pipelines
  • Data engineers optimizing cost and latency in text extraction workflows
  • Developers working with structured text formats (quizzes, forms, invoices)
  • Teams balancing accuracy requirements with API call budgets

regex-vs-llm-structured-text FAQ

Why start with regex instead of LLM for all text?

Regex handles 95-98% of structured text cases deterministically and costs ~95% less than LLM calls. Reserve expensive LLM validation for the remaining edge cases.

How do I know if my text is structured enough for regex?

If >90% of your text follows a repeating pattern (consistent field order, delimiters, formatting), regex is a good fit. Free-form or highly variable text should use LLM directly.

What confidence threshold should I use?

A threshold of 0.95 is a reasonable starting point. Adjust based on your accuracy requirements and cost tolerance—lower thresholds send more items to LLM but improve accuracy.

Which LLM model should I use for validation?

Use the cheapest available model (e.g., Claude Haiku) for validation tasks. Validation is a simpler task than initial parsing, so smaller models are sufficient.

How do I handle encoding issues or malformed input?

Add test cases for known edge cases (missing fields, unusual formatting, encoding artifacts) and refine your regex patterns or confidence scorer based on failures.

Full instructions (SKILL.md)

Source of truth, from affaan-m/ecc.


name: regex-vs-llm-structured-text description: Decision framework for choosing between regex and LLM when parsing structured text — start with regex, add LLM only for low-confidence edge cases. metadata: origin: ECC

Regex vs LLM for Structured Text Parsing

A practical decision framework for parsing structured text (quizzes, forms, invoices, documents). The key insight: regex handles 95-98% of cases cheaply and deterministically. Reserve expensive LLM calls for the remaining edge cases.

When to Activate

  • Parsing structured text with repeating patterns (questions, forms, tables)
  • Deciding between regex and LLM for text extraction
  • Building hybrid pipelines that combine both approaches
  • Optimizing cost/accuracy tradeoffs in text processing

Decision Framework

Is the text format consistent and repeating?
├── Yes (>90% follows a pattern) → Start with Regex
│   ├── Regex handles 95%+ → Done, no LLM needed
│   └── Regex handles <95% → Add LLM for edge cases only
└── No (free-form, highly variable) → Use LLM directly

Architecture Pattern

Source Text
    │
    ▼
[Regex Parser] ─── Extracts structure (95-98% accuracy)
    │
    ▼
[Text Cleaner] ─── Removes noise (markers, page numbers, artifacts)
    │
    ▼
[Confidence Scorer] ─── Flags low-confidence extractions
    │
    ├── High confidence (≥0.95) → Direct output
    │
    └── Low confidence (<0.95) → [LLM Validator] → Output

Implementation

1. Regex Parser (Handles the Majority)

import re
from dataclasses import dataclass

@dataclass(frozen=True)
class ParsedItem:
    id: str
    text: str
    choices: tuple[str, ...]
    answer: str
    confidence: float = 1.0

def parse_structured_text(content: str) -> list[ParsedItem]:
    """Parse structured text using regex patterns."""
    pattern = re.compile(
        r"(?P<id>\d+)\.\s*(?P<text>.+?)\n"
        r"(?P<choices>(?:[A-D]\..+?\n)+)"
        r"Answer:\s*(?P<answer>[A-D])",
        re.MULTILINE | re.DOTALL,
    )
    items = []
    for match in pattern.finditer(content):
        choices = tuple(
            c.strip() for c in re.findall(r"[A-D]\.\s*(.+)", match.group("choices"))
        )
        items.append(ParsedItem(
            id=match.group("id"),
            text=match.group("text").strip(),
            choices=choices,
            answer=match.group("answer"),
        ))
    return items

2. Confidence Scoring

Flag items that may need LLM review:

@dataclass(frozen=True)
class ConfidenceFlag:
    item_id: str
    score: float
    reasons: tuple[str, ...]

def score_confidence(item: ParsedItem) -> ConfidenceFlag:
    """Score extraction confidence and flag issues."""
    reasons = []
    score = 1.0

    if len(item.choices) < 3:
        reasons.append("few_choices")
        score -= 0.3

    if not item.answer:
        reasons.append("missing_answer")
        score -= 0.5

    if len(item.text) < 10:
        reasons.append("short_text")
        score -= 0.2

    return ConfidenceFlag(
        item_id=item.id,
        score=max(0.0, score),
        reasons=tuple(reasons),
    )

def identify_low_confidence(
    items: list[ParsedItem],
    threshold: float = 0.95,
) -> list[ConfidenceFlag]:
    """Return items below confidence threshold."""
    flags = [score_confidence(item) for item in items]
    return [f for f in flags if f.score < threshold]

3. LLM Validator (Edge Cases Only)

def validate_with_llm(
    item: ParsedItem,
    original_text: str,
    client,
) -> ParsedItem:
    """Use LLM to fix low-confidence extractions."""
    response = client.messages.create(
        model="claude-haiku-4-5-20251001",  # Cheapest model for validation
        max_tokens=500,
        messages=[{
            "role": "user",
            "content": (
                f"Extract the question, choices, and answer from this text.\n\n"
                f"Text: {original_text}\n\n"
                f"Current extraction: {item}\n\n"
                f"Return corrected JSON if needed, or 'CORRECT' if accurate."
            ),
        }],
    )
    # Parse LLM response and return corrected item...
    return corrected_item

4. Hybrid Pipeline

def process_document(
    content: str,
    *,
    llm_client=None,
    confidence_threshold: float = 0.95,
) -> list[ParsedItem]:
    """Full pipeline: regex -> confidence check -> LLM for edge cases."""
    # Step 1: Regex extraction (handles 95-98%)
    items = parse_structured_text(content)

    # Step 2: Confidence scoring
    low_confidence = identify_low_confidence(items, confidence_threshold)

    if not low_confidence or llm_client is None:
        return items

    # Step 3: LLM validation (only for flagged items)
    low_conf_ids = {f.item_id for f in low_confidence}
    result = []
    for item in items:
        if item.id in low_conf_ids:
            result.append(validate_with_llm(item, content, llm_client))
        else:
            result.append(item)

    return result

Real-World Metrics

From a production quiz parsing pipeline (410 items):

MetricValue
Regex success rate98.0%
Low confidence items8 (2.0%)
LLM calls needed~5
Cost savings vs all-LLM~95%
Test coverage93%

Best Practices

  • Start with regex — even imperfect regex gives you a baseline to improve
  • Use confidence scoring to programmatically identify what needs LLM help
  • Use the cheapest LLM for validation (Haiku-class models are sufficient)
  • Never mutate parsed items — return new instances from cleaning/validation steps
  • TDD works well for parsers — write tests for known patterns first, then edge cases
  • Log metrics (regex success rate, LLM call count) to track pipeline health

Anti-Patterns to Avoid

  • Sending all text to an LLM when regex handles 95%+ of cases (expensive and slow)
  • Using regex for free-form, highly variable text (LLM is better here)
  • Skipping confidence scoring and hoping regex "just works"
  • Mutating parsed objects during cleaning/validation steps
  • Not testing edge cases (malformed input, missing fields, encoding issues)

When to Use

  • Quiz/exam question parsing
  • Form data extraction
  • Invoice/receipt processing
  • Document structure parsing (headers, sections, tables)
  • Any structured text with repeating patterns where cost matters