Site audit, content writing, and competitor analysis for organic search rankings.
Design & media
PDF Analysis
Try itAnalyze the structure, layout, and content of PDF documents using MinerU. Returns structured output preserving headings, tables, images, formulas, and docume...
What it does
Analyze and extract structured content from PDF files using MinerU. Returns Markdown with layout, headings, and structure preserved.
The skill document
PDF Analysis
Analyze and extract structured content from PDF files using MinerU. Returns Markdown with layout, headings, and structure preserved.
Install
npm install -g mineru-open-api
# or via Go (macOS/Linux):
go install github.com/opendatalab/MinerU-Ecosystem/cli/mineru-open-api@latest
Quick Start
# Quick analysis, no token required (max 10 MB / 20 pages)
mineru-open-api flash-extract report.pdf
# Save to directory
mineru-open-api flash-extract report.pdf -o ./out/
# From URL
mineru-open-api flash-extract https://example.com/report.pdf
# With language hint
mineru-open-api flash-extract report.pdf --language en
# Full analysis with tables and formulas (requires token)
mineru-open-api extract report.pdf -o ./out/
Authentication
No token needed for flash-extract. Token required for extract:
mineru-open-api auth # Interactive token setup
export MINERU_TOKEN="your-token" # Or via environment variable
Create token at: https://mineru.net/apiManage/token
Capabilities
- Supported input: .pdf (local file or URL)
flash-extract: quick, no token, max 10 MB / 20 pages, Markdown output onlyextract: token required, full features (tables, formulas, OCR, multi-format output)- Language hint with
--language(default:ch, useenfor English) - Page range with
--pages(e.g.1-10)
Notes
- Use
flash-extractfor quick reads; useextractfor tables, formulas, or files over 10 MB - Output goes to stdout by default; use
-oto save to a file or directory - All progress/status messages go to stderr; document content goes to stdout
- MinerU is open-source by OpenDataLab (Shanghai AI Lab): https://github.com/opendatalab/MinerU
Related skills
Join a video meeting as an AI bot with voice, avatar, and screenshare across four operating modes.
Read, merge, split, extract from, generate, rotate, watermark, encrypt, and OCR PDFs using pypdf, pdfplumber, reportlab, and CLI tools.
chatgpt-apps
OfficialScaffold ChatGPT Apps SDK projects with a docs-grounded MCP server and widget, classified by archetype before any code.
figma-generate-design
OfficialBuild full-page Figma screens by reusing a published design system instead of drawing primitives with hardcoded values.
hatch-pet
OfficialGenerate Codex-compatible animated pets and 9-state atlases from a concept, brand cue, or reference images.
More from mzlzyca
Browse all skillsConvert PDF documents to Word (.docx) format using MinerU. Transforms PDF files into editable Word documents preserving layout, text, tables, and formatting....
Parse and extract structured content from Word documents (.doc, .docx) into well-organized Markdown using MinerU. Preserves the full document hierarchy: head...
Parse academic papers and research documents from PDF using MinerU. Extracts structured content including title, abstract, sections, figures, tables, formula...
OCR for photos and images using MinerU. Extract text from photographs, screenshots, camera captures, and image files with high accuracy. Features: image OCR...
Professional-grade OCR for PDFs and images using MinerU. Advanced text recognition with VLM (Vision Language Model) support for complex layouts, mixed conten...
Convert Word documents (.doc, .docx) to clean, well-structured Markdown using MinerU's document processing engine. Ideal for turning Microsoft Word files int...