seo-programmatic
agricidaniel/claude-seo
Plan and audit SEO pages generated at scale from structured data sources.
What is seo-programmatic?
Programmatic SEO skill for building and auditing pages created automatically from data sources like CSVs, APIs, or databases. Use when generating pages at scale to enforce quality gates, prevent thin content penalties, and avoid index bloat.
- Assess data source quality (CSV/JSON/API/database) for uniqueness and freshness
- Design templates with dynamic content blocks that produce distinct, valuable pages
- Plan URL patterns (tools, locations, integrations, glossary, templates) with uniqueness enforcement
- Automate internal linking via hub/spoke model, related items, breadcrumbs, and cross-linking
- Enforce thin content safeguards with quality gates (≥40% unique content, ≥300 words, human review)
- Generate sitemaps dynamically and prevent index bloat via noindex/canonical strategy
How to install seo-programmatic
npx skills add https://github.com/agricidaniel/claude-seo --skill seo-programmaticHow to use seo-programmatic
- 1.Assess your data source: count rows, check field uniqueness, verify freshness, flag duplicates (>80% overlap)
- 2.Design your template with variable injection points (title, H1, meta description, schema) and conditional logic
- 3.Choose a URL pattern matching your content type (e.g., /tools/[name], /[city]/[service])
- 4.Plan internal linking: define hub pages, related-item rules, breadcrumb schema, and anchor text strategy
- 5.Run uniqueness audit: ensure ≥40% of content differs between pages (exclude shared headers/footers)
- 6.Set quality gates: flag pages <300 words or <40% unique content; require human review for 5-10% sample
- 7.Generate sitemap with lastmod timestamps from data updates; split at 50k URLs or 50MB per file
- 8.Publish in batches (50-100 pages), monitor indexing and rankings for 2-4 weeks before scaling further
Use cases
- Building a tool directory with 500+ product pages from a database of tool attributes
- Creating location + service pages (e.g., '/denver/plumbing') from city and service data
- Generating integration landing pages with setup docs and API details from platform metadata
- Publishing glossary definitions at scale with unique examples and related terms per entry
- Scaling template or downloadable resource pages with distinct use cases and instructions
- SEO specialists managing large-scale content generation
- Product managers scaling content from structured data sources
- Developers building template-based page generators
- Content strategists auditing programmatic content for quality and uniqueness
seo-programmatic FAQ
Words unique to each page divided by total page words (excluding shared headers, footers, navigation). Template boilerplate text IS included. Metadata (title, description) is scored separately via string comparison heuristics.
Noindex pages that fail quality gates (<40% unique content, <300 words), true duplicates, or low-value filtered/paginated views. Every programmatic page must have a self-referencing canonical tag.
Ensure ≥30-40% genuine uniqueness between pages (not just keyword/city swaps), conduct 5-10% human review before publishing, publish in batches of 50-100 pages with 2-4 week monitoring, and verify each page has standalone value independent of similar pages.
CSV/JSON files, API endpoints, and database queries with sufficient unique attributes per record. Flag data with >80% field overlap as duplicates, and verify freshness; stale data produces stale pages.
Integration pages with real setup docs, template/tool pages with downloadable content, glossary pages with 200+ word definitions, product pages with unique specs, and data-driven pages with unique statistics. Avoid scaling location pages with only city-name swaps or AI-generated pages without human review.
Full instructions (SKILL.md)
Source of truth, from agricidaniel/claude-seo.
name: seo-programmatic description: > Programmatic SEO planning and analysis for pages generated at scale from data sources. Covers template engines, URL patterns, internal linking automation, thin content safeguards, and index bloat prevention. Use when user says "programmatic SEO", "pages at scale", "dynamic pages", "template pages", "generated pages", or "data-driven SEO". user-invocable: true argument-hint: "[url or plan]" license: MIT metadata: author: AgriciDaniel version: "2.4.1" category: seo
Programmatic SEO Analysis & Planning
Build and audit SEO pages generated at scale from structured data sources. Enforces quality gates to prevent thin content penalties and index bloat.
Data Source Assessment
Evaluate the data powering programmatic pages:
- CSV/JSON files: Row count, column uniqueness, missing values
- API endpoints: Response structure, data freshness, rate limits
- Database queries: Record count, field completeness, update frequency
- Data quality checks:
- Each record must have enough unique attributes to generate distinct content
- Flag duplicate or near-duplicate records (>80% field overlap)
- Verify data freshness; stale data produces stale pages
Template Engine Planning
Design templates that produce unique, valuable pages:
- Variable injection points: Title, H1, body sections, meta description, schema
- Content blocks: Static (shared across pages) vs dynamic (unique per page)
- Conditional logic: Show/hide sections based on data availability
- Supplementary content: Related items, contextual tips, user-generated content
- Template review checklist:
- Each page must read as a standalone, valuable resource
- No "mad-libs" patterns (just swapping city/product names in identical text)
- Dynamic sections must add genuine information, not just keyword variations
URL Pattern Strategy
Common Patterns
/tools/[tool-name]: Tool/product directory pages/[city]/[service]: Location + service pages/integrations/[platform]: Integration landing pages/glossary/[term]: Definition/reference pages/templates/[template-name]: Downloadable template pages
URL Rules
- Lowercase, hyphenated slugs derived from data
- Logical hierarchy reflecting site architecture
- No duplicate slugs; enforce uniqueness at generation time
- Keep URLs under 100 characters
- No query parameters for primary content URLs
- Consistent trailing slash usage (match existing site pattern)
Internal Linking Automation
- Hub/spoke model: Category hub pages linking to individual programmatic pages
- Related items: Auto-link to 3-5 related pages based on data attributes
- Breadcrumbs: Generate BreadcrumbList schema from URL hierarchy
- Cross-linking: Link between programmatic pages sharing attributes (same category, same city, same feature)
- Anchor text: Use descriptive, varied anchor text. Avoid exact-match keyword repetition
- Link density: 3-5 internal links per 1000 words (match seo-content guidelines)
Thin Content Safeguards
Quality Gates
| Metric | Threshold | Action |
|---|---|---|
| Pages without content review | 100+ | ⚠️ WARNING: require content audit before publishing |
| Pages without justification | 500+ | 🛑 HARD STOP: require explicit user approval and thin content audit |
| Unique content per page | <40% | ❌ Flag as thin content (likely penalty risk) |
| Word count per page | <300 | ⚠️ Flag for review (may lack sufficient value) |
Scaled Content Abuse: Enforcement Context (2025-2026)
Google's Scaled Content Abuse policy (introduced March 2024) saw major enforcement escalation in 2025:
- June 2025: Third-party reports described a wave of manual actions against sites publishing AI-generated content at scale (no Google announcement)
- August 2025: Third-party/SEO-community reporting described stronger SpamBrain detection for AI-generated link schemes and content farms
- Result: Google reported 45% reduction in low-quality, unoriginal content in search results post-March 2024 enforcement
Enhanced quality gates for programmatic pages:
- Content differentiation: ≥30-40% of content must be genuinely unique between any two programmatic pages (not just city/keyword string replacement)
- Human review: Minimum 5-10% sample review of generated pages before publishing
- Progressive rollout: Publish in batches of 50-100 pages. Monitor indexing and rankings for 2-4 weeks before expanding. Never publish 500+ programmatic pages simultaneously without explicit quality review.
- Standalone value test: Each page should pass: "Would this page be worth publishing even if no other similar pages existed?"
- Site reputation abuse: Google clarified site reputation abuse language on 2024-11-19; treat third-party/hosted programmatic content as a policy risk. Since 2026-08-30 (announced 2026-08-28) enforcement depends on the searcher: manual actions apply outside the EEA, while for EEA users the third-party section may be categorized separately from the main domain. Report the risk for both audiences.
Recommendation: The WARNING gate at
<40% unique contentremains appropriate. Consider a HARD STOP at<30%unique content to prevent scaled content abuse risk.
Safe Programmatic Pages (OK at scale)
✅ Integration pages (with real setup docs, API details, screenshots) ✅ Template/tool pages (with downloadable content, usage instructions) ✅ Glossary pages (200+ word definitions with examples, related terms) ✅ Product pages (unique specs, reviews, comparison data) ✅ Data-driven pages (unique statistics, charts, analysis per record)
Penalty Risk (avoid at scale)
❌ Location pages with only city name swapped in identical text ❌ "Best [tool] for [industry]" without industry-specific value ❌ "[Competitor] alternative" without real comparison data ❌ AI-generated pages without human review and unique value-add ❌ Pages where >60% of content is shared template boilerplate
Uniqueness Calculation
Unique content % = (words unique to this page) / (total words on page) × 100
Measure against all other pages in the programmatic set. Shared headers, footers, and navigation are excluded from the calculation. Template boilerplate text IS included.
Metadata is scored separately. This calculation covers body copy only, so a set that passes it can still carry one generated title/description shape on every URL. Run "${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run metadata_template.py --pairs-file <file> --json (heuristic, deterministic string comparison) over the whole set and treat a site_risk of high as a gate failure regardless of body uniqueness.
Canonical Strategy
- Every programmatic page must have a self-referencing canonical tag
- Parameter variations (sort, filter) canonical to the base URL when duplicate or low-value
- Paginated series: use self-canonical paginated pages when content differs; keep crawlable links
- If programmatic pages overlap with manual pages, the manual page is canonical
- No canonical to a different domain unless intentional cross-domain setup
Sitemap Integration
- Auto-generate sitemap entries for all programmatic pages
- Split at 50,000 URLs or 50MB uncompressed per sitemap file, whichever comes first (protocol limit)
- Use sitemap index if multiple sitemap files needed
<lastmod>reflects actual data update timestamp (not generation time)- Exclude noindexed programmatic pages from sitemap
- Register sitemap in robots.txt
- Update sitemap dynamically as new records are added to data source
Index Bloat Prevention
- Noindex low-value pages: Pages that don't meet quality gates
- Pagination: Reserve noindex/canonical consolidation for true duplicates or low-value filtered views
- Faceted navigation: Reserve noindex/canonical to base category for true duplicates or low-value filtered views
- Crawl budget: For sites with >10k programmatic pages, monitor crawl stats in Search Console
- Thin page consolidation: Merge records with insufficient data into aggregated pages
- Regular audits: Monthly review of indexed page count vs intended count
Output
Programmatic SEO Score: XX/100
Assessment Summary
| Category | Status | Score |
|---|---|---|
| Data Quality | ✅/⚠️/❌ | XX/100 |
| Template Uniqueness | ✅/⚠️/❌ | XX/100 |
| URL Structure | ✅/⚠️/❌ | XX/100 |
| Internal Linking | ✅/⚠️/❌ | XX/100 |
| Thin Content Risk | ✅/⚠️/❌ | XX/100 |
| Index Management | ✅/⚠️/❌ | XX/100 |
Critical Issues (fix immediately)
High Priority (fix within 1 week)
Medium Priority (fix within 1 month)
Low Priority (backlog)
Recommendations
- Data source improvements
- Template modifications
- URL pattern adjustments
- Quality gate compliance actions
Error Handling
| Scenario | Action |
|---|---|
| URL unreachable | Report connection error with status code. Suggest verifying URL accessibility and checking for authentication requirements. |
| No programmatic pages detected | Inform user that no template-generated or data-driven page patterns were found. Suggest checking if pages use client-side rendering or if the URL points to the correct section. |
| Thin content threshold exceeded | Trigger quality gate warning. Report the unique content percentage and flag pages below 40% uniqueness. Require user acknowledgment before proceeding. |
| Quality gate violation | Halt analysis at the HARD STOP threshold (500+ pages without justification or <30% unique content). Present findings and require explicit user approval to continue. |
Related skills
More from agricidaniel/claude-seo and the wider catalog.

seo-schema
Detect, validate, and generate Schema.org JSON-LD markup for SEO rich results.

seo-sitemap
Analyze existing XML sitemaps or generate new ones with validation and industry templates.

seo-sxo
Diagnose ranking problems by comparing your page type against what Google actually rewards in search results.

seo-technical
Audit technical SEO: crawlability, indexability, security, mobile, Core Web Vitals, and structured data.

day1-onboarding
AI Native Camp Day 1 onboarding: learn Claude Code through guided conversation and hands-on exercises.

day2-create-context-sync-skill
Build your own context sync skill to collect information from multiple external tools into a unified document.