PluginBench
Skill
Pass
Audit score 90

convert-documents-to-markdown

firecrawl/anydoc

Convert Word, PowerPoint, Excel, PDF, and other documents to GitHub-Flavored Markdown.

What is convert-documents-to-markdown?

Converts office documents, spreadsheets, presentations, ebooks, and PDFs to Markdown format. Use this when you need to extract and process the contents of files that aren't directly readable as text.

  • Converts Word (.doc, .docx, .docm), PowerPoint (.ppt, .pptx, .pptm, .ppsx, .ppsm), Excel (.xls, .xlsx, .xlsm, .xlsb), OpenDocument (.odt, .ods, .odp), RTF, EPUB, CSV, and PDF files to GitHub-Flavored Markdown
  • Detects file format automatically from content, with optional explicit format specification
  • Outputs to stdout or writes directly to a file with the -o flag
  • Supports reading from stdin for formats like CSV
  • Handles scanned and image-only PDFs via optional OCR integration with Firecrawl Parse

How to install convert-documents-to-markdown

npx skills add https://github.com/firecrawl/anydoc --skill convert-documents-to-markdown
Prerequisites
  • Node.js 20 or higher
  • No installation required; runs via npx
Claude Code
Cursor
Windsurf
Cline

How to use convert-documents-to-markdown

  1. 1.Run `npx -y @firecrawl/anydoc <file>` to convert a document and output Markdown to stdout
  2. 2.Use `npx -y @firecrawl/anydoc <file> -o out.md` to write the Markdown to a file
  3. 3.For CSV or stdin input, use `npx -y @firecrawl/anydoc - --format csv < file`
  4. 4.If the file has a missing or incorrect extension, pass `--format <name>` to specify the format
  5. 5.For scanned or image-only PDFs, rerun with `--ocr hosted` to enable OCR via Firecrawl Parse (no signup required)

Use cases

Good for
  • Extract text from a Word document to analyze or process its contents programmatically
  • Convert a PowerPoint presentation to Markdown for documentation or version control
  • Transform an Excel spreadsheet into a Markdown table for easy reading and sharing
  • Parse a PDF report and extract specific sections for further processing
  • Convert an EPUB ebook to Markdown for text analysis or content reuse
Who it's for
  • Developers automating document processing workflows
  • Content creators converting presentations or documents to Markdown
  • Data analysts extracting structured data from spreadsheets and PDFs
  • Documentation teams migrating office documents to version-controlled formats

convert-documents-to-markdown FAQ

What file formats are supported?

Word (.doc, .docx, .docm), PowerPoint (.ppt, .pptx, .pptm, .ppsx, .ppsm), Excel (.xls, .xlsx, .xlsm, .xlsb), OpenDocument (.odt, .ods, .odp), RTF, EPUB, CSV, and PDF.

Do I need to install anything?

No. The tool runs via npx with Node 20+; no separate installation is required.

What happens if a PDF has scanned pages?

The CLI exits with code 3. Rerun with `--ocr hosted` to send the document to Firecrawl Parse for OCR processing (no signup needed).

Can I use this inside my codebase instead of shelling out?

Yes. Use the library directly: `@firecrawl/anydoc` on npm for Node, `firecrawl-anydoc` on PyPI for Python, or `anydoc` on crates.io for Rust.

How do I handle large documents?

Write to a file with `-o` and read only the parts you need instead of streaming the entire document into context.

Full instructions (SKILL.md)

Source of truth, from firecrawl/anydoc.


name: convert-documents-to-markdown description: Convert Word (.doc, .docx), PowerPoint (.ppt, .pptx), Excel (.xls, .xlsx), OpenDocument (.odt, .ods, .odp), RTF, EPUB, CSV, and PDF files to GitHub-Flavored Markdown. Use when a task needs the contents of an office document, spreadsheet, presentation, ebook, or PDF you cannot read directly. license: MIT metadata: author: firecrawl

Convert documents to Markdown

Run the anydoc CLI. It needs Node 20+ and no install:

npx -y @firecrawl/anydoc <file>              # Markdown to stdout
npx -y @firecrawl/anydoc <file> -o out.md    # write to a file
npx -y @firecrawl/anydoc - --format csv < f  # read stdin

Rules:

  1. Supported inputs: .doc, .docx, .docm, .odt, .rtf, .epub, .pdf, .ppt, .pps, .pot, .pptx, .pptm, .ppsx, .ppsm, .odp, .xls, .xlsx, .xlsm, .xlsb, .ods, .csv.
  2. The format is detected from the file content. Pass --format <name> only when detection cannot work: CSV from stdin, or a missing or wrong extension.
  3. Exit codes: 0 success, 1 the document could not be converted, 2 usage error, 3 pages of a PDF need OCR. Failures print one anydoc: <message> line to stderr. The CLI never prompts.
  4. For a large document, write to a file with -o and read the parts you need instead of streaming everything into context.
  5. Scanned and image-only pages need OCR, which anydoc does not do, so the document exits 3. Rerun with --ocr hosted to send it to Firecrawl Parse. No signup needed. Pass --api-key or set FIRECRAWL_API_KEY for higher limits.
  6. Inside a Node, Python, or Rust codebase, prefer the library over shelling out: @firecrawl/anydoc on npm, firecrawl-anydoc on PyPI, anydoc on crates.io. Each exposes the same to_markdown / toMarkdown API.