Design & media

Doc OCR

Try it

OCR (Optical Character Recognition) for Word documents (.docx) containing scanned pages or image-embedded content. Uses MinerU to extract text from Word file...

What it does

Use OCR to extract text from Word (.docx) files that contain scanned pages or image-embedded content, using MinerU.

The skill document

Doc OCR

Use OCR to extract text from Word (.docx) files that contain scanned pages or image-embedded content, using MinerU.

Install

npm install -g mineru-open-api
# or via Go (macOS/Linux):
go install github.com/opendatalab/MinerU-Ecosystem/cli/mineru-open-api@latest

Quick Start

# OCR extraction from .docx (requires token)
mineru-open-api extract report.docx --ocr -o ./out/

# With VLM model for better accuracy on complex image layouts
mineru-open-api extract report.docx --ocr --model vlm -o ./out/

Authentication

Token required:

mineru-open-api auth             # Interactive token setup
export MINERU_TOKEN="your-token" # Or via environment variable

Create token at: https://mineru.net/apiManage/token

Capabilities

  • Supported input: .docx (local file or URL)
  • OCR is only available via extract (requires token)
  • Use --ocr flag to enable OCR on image-embedded content
  • Use --model vlm for complex or mixed-content documents
  • Language hint with --language (default: ch, use en for English)

Notes

  • OCR is NOT available in flash-extract — use extract with --ocr
  • If the .docx has a normal text layer, OCR is not needed — use doc-extract instead
  • Output goes to stdout by default; use -o to save to a file or directory
  • All progress/status messages go to stderr; document content goes to stdout
  • MinerU is open-source by OpenDataLab (Shanghai AI Lab): https://github.com/opendatalab/MinerU

Related skills

Join a video meeting as an AI bot with voice, avatar, and screenshare across four operating modes.

by johnpatternai21 installs8 stars

pdf

Official

Handle PDF tasks in Python: merge, split, rotate, extract text/tables/images, OCR scans, and create new PDFs.

by Anthropic177.5k stars

chatgpt-apps

Official

Scaffold ChatGPT Apps SDK projects with docs-aligned MCP servers, widgets, and tool plans.

by OpenAI27.5k stars

Reuse the target Figma file's published design system to build or update full-page screens from code or description.

by OpenAI27.5k stars

hatch-pet

Official

Generate Codex-compatible animated pets and 9-state atlases from a concept, brand cue, or reference images.

by OpenAI27.5k stars

More from mzlzyca

Browse all skills

Convert PDF documents to Word (.docx) format using MinerU. Transforms PDF files into editable Word documents preserving layout, text, tables, and formatting....

by mzlzyca30 installs

Parse and extract structured content from Word documents (.doc, .docx) into well-organized Markdown using MinerU. Preserves the full document hierarchy: head...

by mzlzyca32 installs

Parse academic papers and research documents from PDF using MinerU. Extracts structured content including title, abstract, sections, figures, tables, formula...

by mzlzyca20 installs2 stars

OCR for photos and images using MinerU. Extract text from photographs, screenshots, camera captures, and image files with high accuracy. Features: image OCR...

by mzlzyca20 installs

Professional-grade OCR for PDFs and images using MinerU. Advanced text recognition with VLM (Vision Language Model) support for complex layouts, mixed conten...

by mzlzyca32 installs

Convert Word documents (.doc, .docx) to clean, well-structured Markdown using MinerU's document processing engine. Ideal for turning Microsoft Word files int...

by mzlzyca26 installs