Documents

anydoc

Try it

Convert Word (.doc, .docx), PowerPoint (.ppt, .pptx), Excel (.xls, .xlsx), OpenDocument (.odt, .ods, .odp), RTF, EPUB, CSV, and PDF files to GitHub-Flavored Markdown. Use when a task needs the contents of an office document, spreadsheet, presentation, ebook, or PDF you cannot read directly.

What it does

Convert Word (.doc, .docx), PowerPoint (.ppt, .pptx), Excel (.xls, .xlsx), OpenDocument (.odt, .ods, .odp), RTF, EPUB, CSV, and PDF files to GitHub-Flavored Markdown. Use when a task needs the contents of an office document, spreadsheet, presentation, ebook, or PDF you cannot read directly.

The skill document

Convert documents to Markdown

Run the anydoc CLI and write Markdown to stdout or a file:

anydoc               # Markdown to stdout
anydoc  -o out.md    # write to a file
anydoc - --format csv < f  # read stdin

First-time install

If command -v anydoc fails, install the PATH-first wrapper used on this workstation:

mkdir -p ~/.local/bin
cat > ~/.local/bin/anydoc <<'EOF'
#!/usr/bin/env bash
set -euo pipefail

exec bunx @firecrawl/anydoc@latest "$@"
EOF
chmod +x ~/.local/bin/anydoc

Requirements and checks:

  1. Keep ~/.local/bin on PATH ahead of later package-manager bin dirs.
  2. Requires Bun so bunx can fetch @firecrawl/anydoc@latest.
  3. Verify with command -v anydoc and anydoc -V.
  4. Prefer this wrapper over a permanent Bun/npm global install of @firecrawl/anydoc.

Rules

  1. Supported inputs: .doc, .docx, .docm, .odt, .rtf, .epub, .pdf, .ppt, .pps, .pot, .pptx, .pptm, .ppsx, .ppsm, .odp, .xls, .xlsx, .xlsm, .xlsb, .ods, .csv.
  2. The format is detected from the file content. Pass --format only when detection cannot work: CSV from stdin, or a missing or wrong extension.
  3. Exit codes: 0 success, 1 the document could not be converted, 2 usage error. Failures print one anydoc: line to stderr. The CLI never prompts.
  4. For a large document, write to a file with -o and read the parts you need instead of streaming everything into context.
  5. Scanned and image-only PDFs need OCR, which anydoc does not do; they fail as unsupported. The hosted Firecrawl Parse API handles those.
  6. Inside a Node, Python, or Rust codebase, prefer the library over shelling out: @firecrawl/anydoc on npm, firecrawl-anydoc on PyPI, anydoc on crates.io. Each exposes the same to_markdown / toMarkdown API.

Related skills

Convert Markdown to DOCX, PDF, PPTX, XLSX, JSON and more from the command line.

87 installs3 stars

Converts PDF, DOCX, XLSX, PPTX, HTML, CSV, and other files to Markdown using the @covoyage/file2md CLI (@covoyage/file2md). Use when the user needs office do...

1 installs

Convert supported documents into Markdown for agent-side reading, with local output and optional image-preserving packages.

28 installs7 stars

Convert PDF and image files into 10 document and data formats, with OCR and layout controls.

by ComPDF32 installs95 stars

See: https://github.com/anyforge/anyparse Use the AnyParse API to extract content from various documents. Supports PDF, Word, Excel, CSV, TSV, images, PPT, HTML, Markdown, Epub, ipynb, RST, EML, and many other formats. Supports document orientation classification, layout analysis, and layout preservation.

Convert a Markdown file or raw Markdown string into a polished Word DOCX document. Supports custom Word template files, includes built-in DOCX templates, and...

23 installs1 stars