PluginBench
Skill
Review
Audit score 70

markdown-converter

intellectronica/agent-skills

Convert PDFs, Word, Excel, PowerPoint, images, audio, and more to Markdown for LLM processing.

What is markdown-converter?

Markdown Converter transforms documents and media files into Markdown format using markitdown. Use it when you need to extract content from PDFs, Office documents, images, audio, or web formats for analysis, processing by language models, or text-based workflows.

  • Convert PDFs, Word (.docx), PowerPoint (.pptx), and Excel (.xlsx, .xls) documents to Markdown
  • Extract text from images with OCR and EXIF metadata, and transcribe audio files
  • Process HTML, CSV, JSON, XML, ZIP archives, YouTube URLs, and EPub files
  • Preserve document structure including headings, tables, lists, and links in output
  • Support stdin/stdout piping for integration into command pipelines
  • Optionally use Azure Document Intelligence for improved PDF extraction quality

How to install markdown-converter

npx skills add https://github.com/intellectronica/agent-skills --skill markdown-converter
Prerequisites
  • uvx (comes with Python 3.9+)
  • For Azure Document Intelligence: valid endpoint and credentials (optional, for enhanced PDF extraction)
Claude Code
Cursor
Windsurf
Cline

How to use markdown-converter

  1. 1.Run `uvx markitdown <input-file>` to convert a file and print to stdout
  2. 2.Use `-o <output-file>` flag to save directly to a Markdown file, or pipe output with `> output.md`
  3. 3.For stdin input, use `-x <extension>` or `-m <mime-type>` to hint the file type
  4. 4.For complex PDFs, add `-d -e <endpoint>` flags to enable Azure Document Intelligence extraction
  5. 5.Check `--list-plugins` to see available plugins, or use `--use-plugins` to enable third-party extensions

Use cases

Good for
  • Convert a PDF report to Markdown for feeding into an LLM for analysis or summarization
  • Extract data from an Excel spreadsheet into a structured Markdown table format
  • Transcribe and convert audio files to Markdown text for processing
  • Batch convert multiple Office documents to Markdown for knowledge base indexing
  • Extract text and metadata from images using OCR for document digitization
Who it's for
  • Data analysts processing documents for LLM workflows
  • Developers building document processing pipelines
  • Researchers converting academic papers and reports to text format
  • Content teams digitizing and archiving documents
  • Anyone needing to extract structured text from mixed media formats

markdown-converter FAQ

Do I need to install markitdown separately?

No. The skill uses `uvx markitdown`, which downloads and runs markitdown without requiring installation. The first run caches dependencies for faster subsequent conversions.

What happens to document formatting like tables and lists?

Markdown Converter preserves document structure, converting tables to Markdown tables, lists to Markdown lists, and maintaining heading hierarchy.

Can I convert from stdin or pipe data in?

Yes. Use `cat file | uvx markitdown` or pipe from another command. Use `-x <extension>` or `-m <mime-type>` to hint the file type when using stdin.

When should I use Azure Document Intelligence?

Use the `-d` flag with `-e <endpoint>` for complex PDFs with poor text extraction, scanned documents, or when you need higher accuracy. It requires Azure credentials.

What image and audio formats are supported?

Images are processed with OCR and EXIF metadata extraction. Audio files are transcribed. Specific format support depends on markitdown's underlying libraries.

Full instructions (SKILL.md)

Source of truth, from intellectronica/agent-skills.


name: markdown-converter description: Convert documents and files to Markdown using markitdown. Use when converting PDF, Word (.docx), PowerPoint (.pptx), Excel (.xlsx, .xls), HTML, CSV, JSON, XML, images (with EXIF/OCR), audio (with transcription), ZIP archives, YouTube URLs, or EPubs to Markdown format for LLM processing or text analysis.

Markdown Converter

Convert files to Markdown using uvx markitdown — no installation required.

Basic Usage

# Convert to stdout
uvx markitdown input.pdf

# Save to file
uvx markitdown input.pdf -o output.md
uvx markitdown input.docx > output.md

# From stdin
cat input.pdf | uvx markitdown

Supported Formats

  • Documents: PDF, Word (.docx), PowerPoint (.pptx), Excel (.xlsx, .xls)
  • Web/Data: HTML, CSV, JSON, XML
  • Media: Images (EXIF + OCR), Audio (EXIF + transcription)
  • Other: ZIP (iterates contents), YouTube URLs, EPub

Options

-o OUTPUT      # Output file
-x EXTENSION   # Hint file extension (for stdin)
-m MIME_TYPE   # Hint MIME type
-c CHARSET     # Hint charset (e.g., UTF-8)
-d             # Use Azure Document Intelligence
-e ENDPOINT    # Document Intelligence endpoint
--use-plugins  # Enable 3rd-party plugins
--list-plugins # Show installed plugins

Examples

# Convert Word document
uvx markitdown report.docx -o report.md

# Convert Excel spreadsheet
uvx markitdown data.xlsx > data.md

# Convert PowerPoint presentation
uvx markitdown slides.pptx -o slides.md

# Convert with file type hint (for stdin)
cat document | uvx markitdown -x .pdf > output.md

# Use Azure Document Intelligence for better PDF extraction
uvx markitdown scan.pdf -d -e "https://your-resource.cognitiveservices.azure.com/"

Notes

  • Output preserves document structure: headings, tables, lists, links
  • First run caches dependencies; subsequent runs are faster
  • For complex PDFs with poor extraction, use -d with Azure Document Intelligence