Browser

Huo15 Js Scraper

Try it

JavaScript渲染网站抓取工具。当需要抓取JS渲染的页面(如企微文档、Vue/React SPA)、企查查企业数据获取)、绕过反爬、或者普通curl/wget/web_fetch无法获取内容的网站时使用此技能。支持Playwright和scrapling双引擎自动切换。

What it does

JavaScript渲染网站抓取技能,支持Playwright和scrapling双引擎。

The skill document

huo15-js-scraper

JavaScript渲染网站抓取技能,支持Playwright和scrapling双引擎。

快速使用

# 基本用法(自动选择引擎)
python3 ~/.openclaw/workspace/skills/huo15-js-scraper/scripts/scrape.py 

# 指定选择器
python3 ~/.openclaw/workspace/skills/huo15-js-scraper/scripts/scrape.py  --selector ".content"

# 输出JSON
python3 ~/.openclaw/workspace/skills/huo15-js-scraper/scripts/scrape.py  --output json

# 强制使用scrapling引擎
python3 ~/.openclaw/workspace/skills/huo15-js-scraper/scripts/scrape.py  --engine scrapling

引擎选择策略

场景推荐引擎
企微文档 / 微信相关Playwright
Cloudflare保护站scrapling (stealth)
Vue/React SPAPlaywright
简单静态页scrapling (basic)
未知站Playwright(更稳定)

Python API

from huo15_js_scraper import scrape

# 方式1:自动选择(推荐)
result = scrape('https://example.com')
print(result['content'])

# 方式2:强制Playwright
result = scrape('https://developer.work.weixin.qq.com/document/path/91756', engine='playwright')

企业微信文档知识库

已构建完整的企微官方文档知识库,位于: ~/workspace/knowledge-base/企业微信文档/

知识库结构

企业微信文档/
├── README.md (索引)
├── 01-快速入门/      - 开发前必读
├── 02-服务端API/     - 通讯录、消息、客户联系、企业支付...
├── 03-客户端API/     - 小程序API、JS-SDK
├── 04-工具资源/       - WeUI、错误码、频率限制
└── 99-附录/          - FAQ、更新日志

更新企微文档知识库

# 列出所有可抓取文档
python3 ~/.openclaw/workspace/skills/huo15-js-scraper/scripts/wecom_docs_scraper.py --list

# 抓取单个文档
python3 ~/.openclaw/workspace/skills/huo15-js-scraper/scripts/wecom_docs_scraper.py --path-id 90556 --category "01-快速入门" --title "快速入门"

# 批量抓取(更新全部52个文档)
python3 ~/.openclaw/workspace/skills/huo15-js-scraper/scripts/wecom_docs_scraper.py --all

核心文档

文档路径ID说明
快速入门90556开发前必读
获取access_token91039API认证基础
发送应用消息90235消息推送核心
创建成员90195通讯录管理
客户联系概述92109客户管理基础
JS-SDK签名算法90506前端开发必备

企查查企业数据

企查查(qcc.com)企业信息查询,支持两种方式:

  1. ✅ 推荐:MCP方式(官方API,稳定可靠)
  2. 备用:直接抓取(需要账号登录,有反爬限制)

推荐方案:企查查MCP(官方API)

企查查提供官方MCP服务,支持 OpenClaw,已封装20+企业查询 SKILL。

数据规模:

  • 3.65亿+ 市场主体
  • 2.5亿+ 司法诉讼
  • 2.1亿+ 知识产权
  • 1.7亿+ 招投标

MCP Servers(4个):

Server别名主要能力
qcc-company企业基座工商登记、股权结构
qcc-risk风控大脑34项风险扫描工具
qcc-ipr知产引擎专利、商标、软著
qcc-operation经营罗盘招投标、资质、舆情

安装步骤:

# 1. 注册获取API Key
# 访问 https://agent.qcc.com 注册

# 2. 添加到OpenClaw配置
# 在OpenClaw插件配置中添加企查查MCP服务器

# MCP接入地址: https://agent.qcc.com/mcp
# 需要配置 API Key 认证

预置 SKILL(发送消息给AI即可加载):

请加载并使用这个 SKILL:https://github.com/duhu2000/financial-services-qcc

SKILL命令示例:

# KYB企业核验(~30秒)
/kyb-verification-qcc 华为技术有限公司

# IC Memo投资备忘录(~30秒)
/ic-memo-qcc 宁德时代 --round Series-B

# 企业画像速览(~3分钟)
/strip-profile-qcc 美团平台有限公司

# 知识产权尽调
/ip-due-diligence-qcc 企业名称 --peer 竞品

# 供应链风险评估
/supply-chain-risk-qcc 企业名称 --tier 1

# 关联方穿透
/related-party-qcc 企业名称 --depth 5

输出格式: 支持 .md / .docx / .pptx


备用方案:直接抓取

如无法使用MCP,可使用直接抓取方式(需要企查查账号)。

安装依赖

pip3 install playwright --break-system-packages
playwright install chromium

登录(首次使用)

# 生成二维码截图,扫码登录
python3 ~/.openclaw/workspace/skills/huo15-js-scraper/scripts/qichacha_scraper.py --login

登录后Cookie自动保存到 ~/.cache/huo15-js-scraper/qichacha_cookies.json

搜索企业

# 搜索企业
python3 ~/.openclaw/workspace/skills/huo15-js-scraper/scripts/qichacha_scraper.py --search "腾讯" --limit 10

# 输出JSON
python3 ~/.openclaw/workspace/skills/huo15-js-scraper/scripts/qichacha_scraper.py --search "腾讯" --output json

企业详情

# 获取企业详细信息(部分需要VIP)
python3 ~/.openclaw/workspace/skills/huo15-js-scraper/scripts/qichacha_scraper.py --company "https://www.qcc.com/firm/xxxxx.html"

返回信息示例

搜索结果(无需登录可查看基础信息):

  • 公司名称
  • 企业状态(开业/存续/吊销)
  • 行业分类
  • 注册资本
  • 法定代表人

详细信息(可能需要VIP):

  • 工商信息
  • 股东信息
  • 年报数据
  • 风险信息

注意事项

  • 企查查搜索功能需要登录才能访问
  • 详细信息(如年报、股东)需要VIP账号
  • Cookie有效期约7天,过期需重新登录
  • 建议设置 --wait 5 等待页面渲染

常见问题

Q: 企微文档怎么抓?

python3 ~/.openclaw/workspace/skills/huo15-js-scraper/scripts/scrape.py \
  "https://developer.work.weixin.qq.com/document/path/91756" \
  --wait 5

Q: 提示playwright未安装?

pip3 install playwright --break-system-packages
playwright install chromium

Q: scrapling安装?

pip3 install "scrapling[all]" --break-system-packages
scrapling install

Q: 内容为空或获取到跳转页面?

增加 --wait 时间,让JS有更多时间渲染:

python3 ...scrape.py  --wait 5

依赖安装

# Playwright(主引擎)
pip3 install playwright --break-system-packages
playwright install chromium

# scrapling(降级引擎)
pip3 install "scrapling[all]" --break-system-packages
scrapling install

工作原理

  1. 优先使用 Playwright(chromium headless)加载页面,等待networkidle
  2. 等待指定时间让JS渲染完成
  3. 通过CSS选择器提取内容
  4. 如果Playwright失败,自动降级到scrapling

Related skills

Generate and edit Draw.io, Mermaid, and Excalidraw diagrams from natural language using a structured JSON spec.

by nssa.io1.0k installs47 stars

Join a video meeting as an AI bot with voice, avatar, and screenshare across four operating modes.

by johnpatternai21 installs8 stars

Find why your productivity system keeps failing, then apply the smallest fix — capacity math, bottleneck routing, durable local notes.

by Iván854 installs69 stars

Stores durable facts in a categorized, plain-markdown vault on disk, alongside your agent's built-in memory.

by Iván555 installs18 stars

More from zhaobod1

Browse all skills

Generate enterprise Word and native PDF documents across 39 format presets, with 7 contract subtypes.

by zhaobod149 installs

Bootstrap a fresh OpenClaw workspace with a guided 4-step setup that writes five identity and preference files.

by zhaobod130 installs

Adds Claude Code-style agent experience to OpenClaw without modifying the host.

by zhaobod144 installs

麻省理工学院48小时学习法技能(青岛火一五信息科技有限公司)。完整还原 Ihtesham Ali 原始三问框架 + 反馈循环 + 完整 48h 三阶段时间线,叠加网上最佳实践(synthesis / contradictions / gaps / Feynman teach-back / weakness ana...

by zhaobod133 installs2 stars

通过火山方舟Ark API调用Seedance 2.0生成第一人称带货短视频,v2 新增剧本驱动的配音(edge-tts/火山TTS)、背景音乐自动混音、字幕烧录,内置 8 套人设模板(传统女、时尚主播、老中医、厨房主妇、美妆博主、健身教练、户外探店、数码博主)。触发词:生成视频、带货视频、产品视频、拍视频、剧本...

by zhaobod126 installs1 stars

规范 + 时尚的思维导图生成。输入 Markdown 大纲 / JSON / OPML / XMind,输出 XMind 2021+ (.xmind)、OPML、FreeMind (.mm)、Markdown、PNG、PDF、SVG;内置 13 种风格(modern / classic / dark / xiao...

by zhaobod136 installs