convert-documents-to-markdown
firecrawl/anydoc
Convert Word, PowerPoint, Excel, PDF, and other documents to GitHub-Flavored Markdown.
What is convert-documents-to-markdown?
Converts office documents, spreadsheets, presentations, ebooks, and PDFs to Markdown format. Use this when you need to extract and process the contents of files that aren't directly readable as text.
- Converts Word (.doc, .docx, .docm), PowerPoint (.ppt, .pptx, .pptm, .ppsx, .ppsm), Excel (.xls, .xlsx, .xlsm, .xlsb), OpenDocument (.odt, .ods, .odp), RTF, EPUB, CSV, and PDF files to GitHub-Flavored Markdown
- Detects file format automatically from content, with optional explicit format specification
- Outputs to stdout or writes directly to a file with the -o flag
- Supports reading from stdin for formats like CSV
- Handles scanned and image-only PDFs via optional OCR integration with Firecrawl Parse
How to install convert-documents-to-markdown
npx skills add https://github.com/firecrawl/anydoc --skill convert-documents-to-markdown- Node.js 20 or higher
- No installation required; runs via npx
How to use convert-documents-to-markdown
- 1.Run `npx -y @firecrawl/anydoc <file>` to convert a document and output Markdown to stdout
- 2.Use `npx -y @firecrawl/anydoc <file> -o out.md` to write the Markdown to a file
- 3.For CSV or stdin input, use `npx -y @firecrawl/anydoc - --format csv < file`
- 4.If the file has a missing or incorrect extension, pass `--format <name>` to specify the format
- 5.For scanned or image-only PDFs, rerun with `--ocr hosted` to enable OCR via Firecrawl Parse (no signup required)
Use cases
- Extract text from a Word document to analyze or process its contents programmatically
- Convert a PowerPoint presentation to Markdown for documentation or version control
- Transform an Excel spreadsheet into a Markdown table for easy reading and sharing
- Parse a PDF report and extract specific sections for further processing
- Convert an EPUB ebook to Markdown for text analysis or content reuse
- Developers automating document processing workflows
- Content creators converting presentations or documents to Markdown
- Data analysts extracting structured data from spreadsheets and PDFs
- Documentation teams migrating office documents to version-controlled formats
convert-documents-to-markdown FAQ
Word (.doc, .docx, .docm), PowerPoint (.ppt, .pptx, .pptm, .ppsx, .ppsm), Excel (.xls, .xlsx, .xlsm, .xlsb), OpenDocument (.odt, .ods, .odp), RTF, EPUB, CSV, and PDF.
No. The tool runs via npx with Node 20+; no separate installation is required.
The CLI exits with code 3. Rerun with `--ocr hosted` to send the document to Firecrawl Parse for OCR processing (no signup needed).
Yes. Use the library directly: `@firecrawl/anydoc` on npm for Node, `firecrawl-anydoc` on PyPI for Python, or `anydoc` on crates.io for Rust.
Write to a file with `-o` and read only the parts you need instead of streaming the entire document into context.
Full instructions (SKILL.md)
Source of truth, from firecrawl/anydoc.
name: convert-documents-to-markdown description: Convert Word (.doc, .docx), PowerPoint (.ppt, .pptx), Excel (.xls, .xlsx), OpenDocument (.odt, .ods, .odp), RTF, EPUB, CSV, and PDF files to GitHub-Flavored Markdown. Use when a task needs the contents of an office document, spreadsheet, presentation, ebook, or PDF you cannot read directly. license: MIT metadata: author: firecrawl
Convert documents to Markdown
Run the anydoc CLI. It needs Node 20+ and no install:
npx -y @firecrawl/anydoc <file> # Markdown to stdout
npx -y @firecrawl/anydoc <file> -o out.md # write to a file
npx -y @firecrawl/anydoc - --format csv < f # read stdin
Rules:
- Supported inputs:
.doc,.docx,.docm,.odt,.rtf,.epub,.pdf,.ppt,.pps,.pot,.pptx,.pptm,.ppsx,.ppsm,.odp,.xls,.xlsx,.xlsm,.xlsb,.ods,.csv. - The format is detected from the file content. Pass
--format <name>only when detection cannot work: CSV from stdin, or a missing or wrong extension. - Exit codes: 0 success, 1 the document could not be converted, 2 usage error, 3 pages of a PDF need OCR. Failures print one
anydoc: <message>line to stderr. The CLI never prompts. - For a large document, write to a file with
-oand read the parts you need instead of streaming everything into context. - Scanned and image-only pages need OCR, which anydoc does not do, so the document exits 3. Rerun with
--ocr hostedto send it to Firecrawl Parse. No signup needed. Pass--api-keyor setFIRECRAWL_API_KEYfor higher limits. - Inside a Node, Python, or Rust codebase, prefer the library over shelling out:
@firecrawl/anydocon npm,firecrawl-anydocon PyPI,anydocon crates.io. Each exposes the sameto_markdown/toMarkdownAPI.
Related skills
More from firecrawl/anydoc and the wider catalog.
firecrawl
Search, scrape, and interact with the web via Firecrawl CLI—real-time content extraction and monitoring.
firecrawl-agent
AI-powered autonomous data extraction from complex websites, returning structured JSON.
firecrawl-browser
Interact with live webpages: click, fill forms, navigate, and extract data after scraping.
firecrawl-crawl
Bulk extract content from entire websites or site sections with depth and path filtering.
firecrawl-download
Download entire websites as local markdown, screenshots, or multiple formats organized in directories.
firecrawl-instruct
Interact with live browser sessions to click, fill forms, navigate, and extract data from JavaScript-heavy pages.