Generate and edit Draw.io, Mermaid, and Excalidraw diagrams from natural language using a structured JSON spec.
Documents
HTML to Text
Try itConvert HTML to plain readable text using MinerU. Strips HTML markup and extracts clean text content from web pages and HTML files. Features: HTML to text co...
What it does
Extract plain readable text from HTML files or web pages using MinerU. MinerU outputs Markdown as the closest format to plain text.
The skill document
HTML to Text
Extract plain readable text from HTML files or web pages using MinerU. MinerU outputs Markdown as the closest format to plain text.
Install
npm install -g mineru-open-api
# or via Go (macOS/Linux):
go install github.com/opendatalab/MinerU-Ecosystem/cli/mineru-open-api@latest
Quick Start
# Extract text from a local HTML file (requires token)
mineru-open-api extract page.html -o ./out/
# Extract text from a web page (requires token)
mineru-open-api crawl https://example.com/article
# JSON output contains text fields (requires token)
mineru-open-api extract page.html -f json -o ./out/
Authentication
Token required:
mineru-open-api auth # Interactive token setup
export MINERU_TOKEN="your-token" # Or via environment variable
Create token at: https://mineru.net/apiManage/token
Capabilities
- Supported input: local .html file or web page URL
- HTML requires
extractorcrawl(token required) — not supported byflash-extract - MinerU does not have a
-f textoption; Markdown is the closest plain-text output - For truly plain text: use
extract -f jsonand read the text fields from JSON output - Language hint with
--language(default:ch, useenfor English)
Notes
- MinerU has no
-f textformat; use Markdown output or-f jsonfor text fields - HTML is NOT supported by
flash-extract - Output goes to stdout by default; use
-oto save to a file or directory - All progress/status messages go to stderr; document content goes to stdout
- MinerU is open-source by OpenDataLab (Shanghai AI Lab): https://github.com/opendatalab/MinerU
Related skills
Join a video meeting as an AI bot with voice, avatar, and screenshare across four operating modes.
Stores durable facts in a categorized, plain-markdown vault on disk, alongside your agent's built-in memory.
Write, debug, and tune Playwright specs with locator strategy, trace diagnosis, and CI-aware timeouts.
Save, search, and manage personal notes and knowledge bases in Get笔记 on explicit request.
Fetch raw ad creative, app, ranking, and revenue data from AdMapix as structured JSON.
More from mzlzyca
Browse all skillsConvert PDF documents to Word (.docx) format using MinerU. Transforms PDF files into editable Word documents preserving layout, text, tables, and formatting....
Parse and extract structured content from Word documents (.doc, .docx) into well-organized Markdown using MinerU. Preserves the full document hierarchy: head...
Parse academic papers and research documents from PDF using MinerU. Extracts structured content including title, abstract, sections, figures, tables, formula...
OCR for photos and images using MinerU. Extract text from photographs, screenshots, camera captures, and image files with high accuracy. Features: image OCR...
Professional-grade OCR for PDFs and images using MinerU. Advanced text recognition with VLM (Vision Language Model) support for complex layouts, mixed conten...
Convert Word documents (.doc, .docx) to clean, well-structured Markdown using MinerU's document processing engine. Ideal for turning Microsoft Word files int...