Browser

web-extract

Try it

Extract structured JSON from web pages, search engines, and entire sites in ONE call — {title, summary, sections, key_metrics, outgoing_links, author, date, page_type, ...} fields, no second LLM pass to parse HTML. Six endpoints: scrape (single URL), scrape-interactive (JS-rendered pages with click/scroll/type), search (Google SERP + deep-scrape), map (URL discovery), crawl + crawl-status (async recursive crawl). Markdown/raw HTML on request. USE when the user needs page DATA — product pricing/specs, article fields, link graphs, JS-heavy SPAs, Google results with content. Prefer over browser-act (automation/screenshots) and WebFetch (static, no JS, no structured fields). Not for citation-rich research (use deep-research). Trigger (EN): scrape this URL, extract data from page, crawl this site, deep-scrape search results, map a domain's URLs, render this JS page. 触发词:抓取/爬取/网页提取/结构化抽取/搜索带内容/全站爬取/JS 渲染抓取/点击后抓取. Requires ZOODATA_API_KEY (free key: https://zoodata.ai/en/api-keys).

What it does

Extract structured JSON from web pages, search engines, and entire sites in ONE call — {title, summary, sections, key_metrics, outgoing_links, author, date, page_type, ...} fields, no second LLM pass to parse HTML. Six endpoints: scrape (single URL), scrape-interactive (JS-rendered pages with click/scroll/type), search (Google SERP + deep-scrape), map (URL discovery), crawl + crawl-status (async recursive crawl). Markdown/raw HTML on request. USE when the user needs page DATA — product pricing/specs, article fields, link graphs, JS-heavy SPAs, Google results with content. Prefer over browser-act (automation/screenshots) and WebFetch (static, no JS, no structured fields). Not for citation-rich research (use deep-research). Trigger (EN): scrape this URL, extract data from page, crawl this site, deep-scrape search results, map a domain's URLs, render this JS page. 触发词:抓取/爬取/网页提取/结构化抽取/搜索带内容/全站爬取/JS 渲染抓取/点击后抓取. Requires ZOODATA_API_KEY (free key: https://zoodata.ai/en/api-keys).

The skill document

Web Extract — Structured Data from the Open Web

Backed by the ZooData WebTools API. Six HTTP endpoints. One API key. Structured JSON by default — no second LLM pass to parse fields.

Files

FilePurpose
{skill_base_dir}/scripts/webtools.pyThin CLI wrapper — one subcommand per endpoint. Has the Cloudflare-UA, crawl-warmup, and nested-data-shape quirks baked in. Run --help for params.
{skill_base_dir}/references/reference.mdFull request/response schemas, error codes, billing, edge cases. Load when you need exact field names.

Why pick this skill (vs the alternatives in this environment)

ToolWhat it gives youWhen to pick it
web-extract (this skill)Page → {title, summary, sections, key_metrics, outgoing_links, ...} JSON in one callYou need page DATA (price, specs, fields, link graphs) — downstream code or LLM can use the JSON directly without re-parsing
browser-actBrowser session: click, scroll, type, screenshotYou need to interact with a page (login flow, take a screenshot, fill a form) or visually verify rendering
WebFetch (built-in)Static URL → markdownYou need a single static page as prose, no JS rendering, no structured fields
deep-researchMulti-source research with citationsYou need a synthesized report drawing from many web sources, not raw data
monidGeneric tool-discovery layerYou're not sure which tool to use yet and want to browse options

Key advantage: every other tool above forces a second pass (re-LLM the markdown / re-parse the HTML / extract structure manually). The underlying API returns the structured fields directly — saves a round-trip and tokens.

Credential

Required: ZOODATA_API_KEY. Get a free key (1,000 credits) at zoodata.ai/en/api-keys.

export ZOODATA_API_KEY='hms_live_xxx'
# OR persist to disk (keep the file private — 0600; it holds a bearer credential):
mkdir -p ~/.zoodata && chmod 700 ~/.zoodata
(umask 077; echo '{"api_key":"hms_live_xxx"}' > ~/.zoodata/config.json)

Capabilities & Data Flow

  • Network: only https://api.zoodata.ai WebTools endpoints (Bearer ZOODATA_API_KEY). Target pages are fetched server-side by ZooData, not from this machine.
  • Execution: bundled CLI {skill_base_dir}/scripts/webtools.py (Python 3, stdlib-only) with exactly the documented subcommands (scrape, interactive, search, map, crawl, crawl-status, crawl-wait, check) — the full surface is the declared surface.
  • Local files: none; reads the optional credential store ~/.zoodata/config.json.
  • Sent to the API: the target URLs, search queries, and crawl options you request.
  • Credits: every call consumes credits (crawl bills per page). For large crawls or deep-scrapes, state the estimated cost and confirm before submitting.

Two ways to drive it

  1. webtools.py CLI (preferred — quirks baked in): UA header, warmup tolerance, retry/backoff, JSON-first defaults all handled. Just run python {skill_base_dir}/scripts/webtools.py .
  2. Raw curl / HTTP (when CLI isn't installed or for ad-hoc calls): every example below also shows the raw POST. Always set User-Agent: web-extract-skill/1.0 — the Cloudflare edge rejects the default Python-urllib UA with HTTP 403 (see Tips).

Endpoint base URL: https://api.zoodata.ai/openapi/v2/webtools/*. All POST with JSON body, except crawl/{id} (GET).

Default format: JSON

This skill defaults to formats: ["json"] for every scrape / scrape-interactive / search-with-scrapeOptions call. The API itself defaults to ["markdown"], but JSON gives the agent structured fields (title, summary, sections, key_metrics, outgoing_links, page_type, …) that compose better with downstream tools and are usually smaller to put back into context.

Switch to markdown only when the user explicitly asks for "the article text" / "clean prose" / "as markdown", or when the page is genuinely article-shaped (long-form blog post / news / docs) and structured fields wouldn't help. Switch to rawHtml only when the user needs DOM access (custom CSS selector extraction, table parsing the API didn't structure, embedded `` data, etc.). Request multiple formats in one call (["json","markdown"]) when you genuinely need both — but never silently fall back to markdown when the user didn't ask for it.

When to use

User intentEndpointCost
Fetch one URL as structured JSON (default)POST /scrape formats:["json"]1 credit
Fetch one URL as markdown (user explicitly asked)POST /scrape formats:["markdown"]1 credit
Fetch one URL as raw HTML (user needs DOM)POST /scrape formats:["rawHtml"]1 credit
Same as above, but page needs JS / login / click-through / scrollPOST /scrape-interactive1 credit
Google search → list of result URLsPOST /search (SERP-only)1 credit
Google search → URLs + structured JSON for eachPOST /search scrapeOptions:{format:"json"}1 credit per returned result
List all URLs reachable from a websitePOST /map1 credit
Recursively pull every page of a site (async)POST /crawl + GET /crawl/{id} poll loop1/submit + 1/poll

Shorthand: WT="python {skill_base_dir}/scripts/webtools.py"

# 1. Scrape one URL as structured JSON (default)
$WT scrape https://example.com
# → data.json = {title, summary, sections, key_metrics, outgoing_links, author, date, page_type, ...}

# 2. Scrape as markdown (user explicitly asked for prose)
$WT scrape https://example.com --format markdown

# 3. Need BOTH JSON and markdown in one call
$WT scrape https://example.com --format json --format markdown

# 4. JS-heavy page — wait, click "Load more", scrape
$WT interactive https://example.com --actions '[
  {"type":"wait","milliseconds":1500},
  {"type":"click","selector":"button.load-more"},
  {"type":"wait","milliseconds":1000}
]'

# 5. SERP-only search (cheap — 1 credit total, regardless of result count)
$WT search "best wireless earbuds 2026" --limit 10

# 6. Deep-scrape every search result as JSON (1 credit per result)
$WT search "best wireless earbuds 2026" --limit 5 --deep-scrape

# 7. Search restricted to specific domains, last 24h
$WT search "python async" --limit 20 --tbs qdr:d \
  --include-domains "python.org,realpython.com"

# 8. Discover every URL on a docs site under /api/
$WT map https://docs.example.com --include-paths "/api/.*" --limit 1000

# 9. Crawl + poll in one shot (handles warmup, paginates pages internally)
$WT crawl-wait https://blog.example.com --limit 200 --max-depth 3 \
  --poll-interval 10 --max-wait 1800

# 9b. Or submit + manual poll
$WT crawl https://blog.example.com --limit 200 --max-depth 3
$WT crawl-status job_xxx --skip 0 --limit 100

# 10. Verify auth + credits (1 credit — scrapes example.com)
$WT check

Run $WT --help for every flag on any subcommand.

Same calls via raw curl

When the CLI isn't available (no python3 / offline / restricted env), every endpoint also takes a plain JSON POST:

curl -sS -X POST https://api.zoodata.ai/openapi/v2/webtools/scrape \
  -H "Authorization: Bearer $ZOODATA_API_KEY" \
  -H "Content-Type: application/json" \
  -H "User-Agent: web-extract-skill/1.0"   `# REQUIRED — see Tips` \
  -d '{"url":"https://example.com","formats":["json"]}'

See references/reference.md for the full JSON body shape of every endpoint. Any non-curl HTTP client must set a custom User-Agent header — Cloudflare's edge blocks the default Python-urllib UA with HTTP 403 (error code 1010).

Key parameters at a glance

Full schemas are in references/reference.md. Load that when you need exact field names or are debugging.

/scrape: url (required), formats: ["json"|"markdown"|"rawHtml"]skill default ["json"] (API itself defaults to ["markdown"] if formats omitted; always pass it explicitly). Max 3 formats per call.

/scrape-interactive: url, formats, actions: [...] — 7 action types (wait/click/write/press/scroll/scrape/executeJavascript). Max 50 actions, cumulative wait ≤ 60s.

/search: query (supports inline Google operators), limit (1–20, default 10), sources (["web","news","images"]), tbs (qdr:d|w|m|y), includeDomains/excludeDomains (bare hostnames, max 20 each), scrapeOptions: {format:"markdown"|"json"} (presence triggers deep-scrape).

/map: url, limit (1–100000, default 5000), search (keyword rank filter), sitemapMode (include/only/skip), includeSubdomains (default true), includePaths/excludePaths (regex), ignoreQueryParameters (default true).

/crawl: url, limit (1–10000, default 100), maxDepth (1–10), includePaths/excludePaths, allowSubdomains, sitemapMode, ignoreQueryParameters (default false — opposite of /map), crawlEntireDomain.

GET /crawl/{job_id}: query params skip, limit (1–1000). data is an array of page-scrape blobs. meta.total is the running total; compare against data.length to detect completion.

Tips

  • 🚨 Cloudflare 1010 trap (Python / non-curl HTTP clients). The Cloudflare edge in front of api.zoodata.ai blocks default Python Python-urllib/X.Y User-Agent → HTTP 403. ALWAYS set an explicit User-Agent header (any non-default value works, e.g. User-Agent: web-extract-skill/1.0). curl is fine out of the box; Python / Go / Node fetch / requests all need it.
  • Crawl GET /crawl/{job_id} warmup gotcha: for ~5–10 seconds after POST /crawl returns the job_id, polling may return NOT_FOUND. This is the job still registering, not a real "not found." Tolerate it: retry with 5s sleep up to ~3 times before giving up.
  • crawl-status response shape: data is an object {id, status, completed, total, data: [pages...]} — the page array is nested at data.data. Status values: "queued", "running", "completed", "failed". (Failed success:false responses cost 0 credits.)
  • A success:true response can still contain an error page. Always check data.meta.statusCode (or per-result meta.statusCode for search) before trusting content.
  • Deep-scrape /search may return fewer than limit results — pages with no extractable content in the chosen format are dropped.
  • Domain filters are bare hostnames only. https://github.com or github.com/path → 422. Use github.com.
  • includePaths / excludePaths are regex, not glob. Use /blog/.* not /blog/**.
  • Crawl job IDs are tenant-scoped. Polling someone else's job returns 404 (indistinguishable from "job not found" — by design).
  • Refused requests (ACCESS_DENIED, RATE_LIMITED) are NOT billed. TIMEOUT and CONTENT_UNAVAILABLE are.
  • Polling costs 1 credit per call. For large crawls, batch your polls — fetch with limit=1000 and space polls ≥10s apart.
  • Format choice (skill default = JSON): json is the default — gives structured fields (title/summary/sections/key_metrics/outgoing_links/page_type/...) that downstream tools can consume directly. Use markdown only when the user asks for clean prose or the page is article-shaped (blog post / docs); use rawHtml only when DOM access is needed (custom selectors, table parsing the JSON layer doesn't expose).

On Missing Key (no credentials configured)

BEFORE calling any endpoint, verify a credential is configured. Reliable check: python {skill_base_dir}/scripts/webtools.py check — exits 2 if no key is found in env (ZOODATA_API_KEY) or the config file (~/.zoodata/config.json). A [ -z "$ZOODATA_API_KEY" ] test alone is NOT sufficient.

When no key is found through any mechanism:

  1. STOP. Do not call any endpoint.
  2. Do NOT fall back to "partial output from training data" / "describe the page from common knowledge" / "for reference only" preview. This skill's deliverable is structured extraction of THIS URL's actual content — without the data, there is no deliverable.
  3. Tell the user, in their language, all three of:
    • "ZOODATA_API_KEY is not set — I need this to fetch the page."
    • Get a free key (1,000 credits, no credit card): https://zoodata.ai/en/api-keys
    • Configure via one of:
      • export ZOODATA_API_KEY='hms_live_xxx' (session only)
      • mkdir -p ~/.zoodata && chmod 700 ~/.zoodata && (umask 077; echo '{"api_key":"hms_live_xxx"}' > ~/.zoodata/config.json) (persistent; keep the file private — 0600)

On 401 Invalid Key

When any endpoint returns HTTP 401:

  1. STOP further calls immediately. A rejected key won't be accepted on retry — every subsequent call will return 401 too.
  2. Tell the user: their ZOODATA_API_KEY was rejected (invalid, revoked, or expired). Direct them to zoodata.ai/en/api-keys.
  3. If you collected partial output before failure, show it and mark partial. Do not fabricate the rest — no "training-data fallback" / "page description from common knowledge" substitution.

On 402 Credit Exhausted

When any endpoint returns HTTP 402:

  1. STOP further calls.
  2. Report partial findings already gathered, plus the creditsRemaining number from the last successful call.
  3. Point the user at zoodata.ai/en/pricing to top up. Do not fabricate the rest — no "common-sense page summary" filler in place of a real fetch.

See also

  • zoodata — commerce endpoints (Amazon products / markets / reviews / brands). Pair with webtools when you need to validate Amazon findings against the open web (competitor sites, news, off-Amazon reviews).
  • amazon-market-entry-analyzer — uses the commerce side; webtools complements it for non-Amazon channel research.

Related skills

Web extraction for LLMs and agents. Scrape, crawl, map, search, extract, summarize, diff, monitor, and research any URL into clean Markdown, text, or JSON, i...

3 installs

Turn any web page into clean, readable content an agent can use — fetch a URL and get the main text/markdown back without ads, nav, or boilerplate. Use whene...

Fetch web pages and generate AI-powered summaries using Ollama LLM. Use when you need to (1) scrape and summarize web articles, (2) extract key points from w...

General-purpose web-intelligence utilities via the Crawlora API — scrape any URL to clean markdown/HTML, extract schema-conforming JSON from a page, fingerprint a site's tech stack, geocode addresses, compare cost of living between cities/countries (Numbeo), look up a company's import/export trade records (ImportYeti), check a domain's traffic (SimilarWeb), or resolve a brand's identity from its domain. Use for one-off utility lookups that don't fit a specific platform skill.

1 installs

Deep-crawl any website from start URLs, return per-page LLM-ready text/markdown/HTML plus metadata (title, description, author, language, canonical URL, OG) and in-scope outbound links. Use when user mentions deep crawl website, recursive crawl, crawl a whole site, scrape entire website, scrape docs

Extract structured data from HTML pages or URLs using CSS selectors, XPath, regex, or LLM natural...

1 installs