PluginBench
Skill
Review
Audit score 70

defuddle

kepano/obsidian-skills

Extract clean Markdown from HTML pages, removing ads and clutter to reduce token usage.

What is defuddle?

Defuddle CLI parses web pages and extracts readable Markdown content, stripping navigation, ads, and other noise. Use it when you need clean text from standard web pages with minimal processing overhead.

  • Parse URLs and extract clean Markdown output with `--md` flag
  • Remove navigation, ads, and page clutter automatically
  • Save extracted content directly to Markdown files with `-o` option
  • Extract specific metadata properties (title, description, domain) from pages
  • Output in multiple formats: Markdown, JSON, or raw HTML

How to install defuddle

npx skills add https://github.com/kepano/obsidian-skills --skill defuddle
Prerequisites
  • Node.js and npm installed
  • Install globally with: `npm install -g defuddle`
Claude Code
Cursor
Windsurf
Cline

How to use defuddle

  1. 1.Install Defuddle globally: `npm install -g defuddle`
  2. 2.Parse a URL to Markdown: `defuddle parse <url> --md`
  3. 3.Save output to file: `defuddle parse <url> --md -o content.md`
  4. 4.Extract specific metadata: `defuddle parse <url> -p title` (or description, domain)
  5. 5.Choose output format: use `--md` for Markdown, `--json` for structured data, or omit flag for HTML

Use cases

Good for
  • Convert web articles to clean Markdown for note-taking or documentation
  • Extract article content while filtering out sidebars and ads before processing
  • Batch download and convert multiple web pages to Markdown files
  • Pull specific metadata like titles and descriptions from web pages
  • Reduce token usage by cleaning HTML before feeding to language models
Who it's for
  • Knowledge workers managing web research and notes
  • Developers building content pipelines or web scrapers
  • AI agents processing web content for analysis or summarization
  • Technical writers converting online resources to documentation

defuddle FAQ

When should I use Defuddle instead of WebFetch?

Use Defuddle for standard web pages where you want clean, readable content with ads and navigation removed. It's more efficient for typical web scraping and reduces token usage.

What output formats does Defuddle support?

Defuddle supports Markdown (--md), JSON (--json with both HTML and markdown), raw HTML (no flag), and specific metadata properties (-p flag).

Can I save the extracted content to a file?

Yes, use the `-o` flag followed by a filename: `defuddle parse <url> --md -o content.md`

How do I extract just the title or description from a page?

Use the `-p` flag with the property name: `defuddle parse <url> -p title` or `defuddle parse <url> -p description`

Full instructions (SKILL.md)

Source of truth, from kepano/obsidian-skills.


name: defuddle description: Extract clean Markdown from HTML pages with Defuddle CLI.

Defuddle

Use Defuddle CLI to extract clean readable content from web pages. Prefer over WebFetch for standard web pages — it removes navigation, ads, and clutter, reducing token usage.

If not installed: npm install -g defuddle

Usage

Always use --md for markdown output:

defuddle parse <url> --md

Save to file:

defuddle parse <url> --md -o content.md

Extract specific metadata:

defuddle parse <url> -p title
defuddle parse <url> -p description
defuddle parse <url> -p domain

Output formats

FlagFormat
--mdMarkdown (default choice)
--jsonJSON with both HTML and markdown
(none)HTML
-p <name>Specific metadata property