Extract text, metadata, and pages from PDF files using pypdf. Use for tasks such as reading PDF content, extracting specific pages, splitting or merging PDFs...
设计与多媒体
pdf-extraction
试用Extract text, tables, and metadata from PDFs. Auto-detects native text vs scanned image pages and routes to pdfplumber or Tesseract OCR.
它能做什么
Extract text, tables, and metadata from PDFs. Auto-detects native text vs scanned image pages and routes to pdfplumber or Tesseract OCR.
技能文档
PDF Extraction (auto text / OCR)
Extract content from PDFs without deciding whether each page is selectable text or a scan.
| Page type | Tool |
|---|---|
| Native text | pdfplumber (text + tables) |
| Scanned / image | PyMuPDF + Tesseract OCR |
Install
# System OCR engine (required for scanned pages)
# Ubuntu/Debian:
sudo apt install tesseract-ocr tesseract-ocr-eng
# Optional Traditional Chinese:
# sudo apt install tesseract-ocr-chi-tra
# CLI — from GitHub (recommended)
pip install "git+https://github.com/alex-ht/pdf-extraction.git@v1.0.0"
# Or install from this skill folder after ClawHub install:
# pip install -e .
Source: https://github.com/alex-ht/pdf-extraction
Quick start
# Auto-detect text vs OCR per page → print text
pdf-extract document.pdf
# Write to file
pdf-extract document.pdf -o out.txt
# See which mode each page will use
pdf-extract document.pdf --analyze-only
# Tables + Markdown
pdf-extract document.pdf --tables --format markdown -o out.md
# JSON (includes per-page mode)
pdf-extract document.pdf --format json -o out.json
# Force mode
pdf-extract scan.pdf --mode ocr --ocr-lang eng
pdf-extract text.pdf --mode text
# Page range
pdf-extract doc.pdf --pages 1-3,5
# Also: python -m pdf_extract document.pdf
Auto mode rules
For each page (--mode auto, default):
- Try native text via pdfplumber; count chars and embedded images
- Enough text → text
- Sparse text (default < 40 non-whitespace chars) or image-heavy → ocr
Tune with --min-text-chars. Override with --mode text or --mode ocr.
Agent usage
When the user provides a PDF and wants content extracted:
- Prefer
pdf-extract(auto mode). Do not ask whether it is text or scanned. - Use
--analyze-onlyif you only need routing diagnostics. - Use
--tableswhen tables matter (works best on native-text pages). - For Chinese scans, set
--ocr-lang eng+chi_traifchi_trais installed. - Surface stderr summary lines (which pages used text vs ocr) when useful.
CLI reference
| Flag | Meaning |
|---|---|
-o, --output | Write to file (default stdout) |
-f, --format | text | json | markdown |
--mode | auto | text | ocr |
--pages | e.g. 1-3,5 |
--tables | Extract tables on text pages |
--layout | Preserve native text layout |
--meta | Include PDF metadata |
--ocr-lang | Tesseract langs (default eng) |
--ocr-dpi | OCR render DPI (default 200) |
--analyze-only | Classification JSON only |
-q, --quiet | Suppress progress on stderr |
Dependencies
- Python 3.10+:
pdfplumber,pymupdf,Pillow(seepyproject.toml) - System:
tesseract(and language packs as needed)
相关技能
Extract text from PDF, DOCX, and TXT with encoding detection and page/paragraph structure
通过 ComPDF Cloud API 执行 50 多种 PDF 与文档处理操作。
Render PDF pages to images, extract embedded images, annotate PDFs, and perform advanced PDF inspection using pymupdf (fitz). Use for tasks such as exporting...
Recognize text from scanned PDFs and images.
Extract PDF content into structured JSON.