为单篇或一批 ArXiv 论文下载源文件与 PDF,再通读全文并按指定语言生成结构化的 summary.md。
文档
defuddle-web-cleaner
Extract and clean readable article content, metadata, and markdown from URLs or HTML for research, note taking, and web scraping.
它能做什么
Extract and clean readable article content, metadata, and markdown from URLs or HTML for research, note taking, and web scraping.
技能文档
name: defuddle-web-cleaner description: extract clean article content from web pages using defuddle. use when a user provides a url or html and wants the readable article text, markdown version, or structured metadata. helpful for web scraping, research workflows, note taking, obsidian clipping, and converting web pages to markdown.
Defuddle Web Cleaner
Extract the main readable content from a web page.
This skill removes unnecessary elements such as:
- navigation bars
- sidebars
- ads
- comments
- footers
- social buttons
The result is clean article content.
Supported Inputs
- URL
- Raw HTML
- Web page text
Output Format
Default output:
Title
Author
Site
Published date
Markdown article content
Alternative output (JSON):
{ title, author, site, description, published, content, contentMarkdown }
Processing Steps
- Detect input type
- Load page HTML
- Run Defuddle parser
- Extract metadata
- Convert to Markdown if requested
- Return clean content
Example
Input:
Output:
Title: AI is Changing Everything
Author: Jane Smith
Site: Example Blog
Markdown:
AI is Changing Everything
Artificial intelligence is transforming industries...
Tips
Use this skill when:
- saving articles to Obsidian
- building research datasets
- cleaning webpages for LLM processing
- summarizing articles
相关技能
压缩任意来源,保留每一条论断、对冲、数值与归属。
把 YouTube 视频整理成带章节、时间戳和要点的 Markdown 摘要
由模型规划查询并判断相关性,将多轮 arXiv 结果合并去重,最终产出可直接使用的主题论文集。
通过 md2wechat CLI 把 Markdown 转为微信公众号 HTML、封面、信息图和图文帖。
汇总过去 30 天 Reddit、X、YouTube 和网页上关于某个话题的真实讨论。