浏览器

Apify

试用

通过托管认证接入 Apify API,运行爬虫并管理 actors、数据集、键值存储与定时任务。

它能做什么

通过托管 API 网关连接 Apify 平台,令牌由网关自动注入。可以列出并运行 actors,管理 actor tasks,读取与写入数据集条目,操作键值存储与请求队列,并维护定时任务与 webhook。认证使用 Maton API key,多个连接可通过 connection 头指定,仅暴露 Apify v2 参考中列出的端点。默认只执行读与列表调用;actor 运行、定时任务修改、webhook 创建等写操作都需要用户确认,因为它们会消耗算力单位或产生会话结束后仍持续运行的后台资源。

什么时候用它

  • 运行已有的 Apify 爬虫 actor 处理一组 URL,并从默认数据集取回抓取结果
  • 列出已连接 Apify 账户中的数据集与键值存储,查看历史抓取任务的产物
  • 用 cron 表达式为 actor task 设置定时调度,用于周期性网页抓取
  • 创建 webhook,在 actor 运行成功或失败时接收回调通知

技能文档

Apify

Access the Apify API with managed authentication. Run web scrapers and actors, manage datasets, key-value stores, request queues, schedules, and webhooks.

Quick Start

# List your actors
python <<'EOF'
import urllib.request, os, json
req = urllib.request.Request('https://gateway.maton.ai/apify/v2/acts')
req.add_header('Authorization', f'Bearer {os.environ["MATON_API_KEY"]}')
print(json.dumps(json.load(urllib.request.urlopen(req)), indent=2))
EOF

Base URL

https://gateway.maton.ai/apify/v2/

All endpoints documented below are accessed under this base URL. The gateway proxies requests to api.apify.com and automatically injects your API token. Only the endpoints listed in the API Reference section below are supported.

Security & Permissions

  • The Maton API key acts as a credential — treat it as a secret and do not expose it in client-side code or public repositories
  • Read operations (GET) retrieve data from the connected Apify account
  • Write operations (POST, PUT, DELETE) create, modify, or delete resources — confirm with the user before destructive actions
  • Actor runs consume Apify compute units on the connected account — confirm before starting runs, especially on large inputs
  • Schedules and webhooks are persistent resources that continue operating after the session ends — confirm before creating, and review existing ones before modification
  • When multiple connections exist, always specify the Maton-Connection header to avoid acting on the wrong account

Authentication

All requests require the Maton API key in the Authorization header:

Authorization: Bearer $MATON_API_KEY

Environment Variable: Set your API key as MATON_API_KEY:

export MATON_API_KEY="YOUR_API_KEY"

Getting Your API Key

  1. Sign in or create an account at maton.ai
  2. Go to maton.ai/settings
  3. Copy your API key

Connection Management

Manage your Apify connections at https://ctrl.maton.ai.

List Connections

python <<'EOF'
import urllib.request, os, json
req = urllib.request.Request('https://ctrl.maton.ai/connections?app=apify&status=ACTIVE')
req.add_header('Authorization', f'Bearer {os.environ["MATON_API_KEY"]}')
print(json.dumps(json.load(urllib.request.urlopen(req)), indent=2))
EOF

Create Connection

python <<'EOF'
import urllib.request, os, json
data = json.dumps({'app': 'apify'}).encode()
req = urllib.request.Request('https://ctrl.maton.ai/connections', data=data, method='POST')
req.add_header('Authorization', f'Bearer {os.environ["MATON_API_KEY"]}')
req.add_header('Content-Type', 'application/json')
print(json.dumps(json.load(urllib.request.urlopen(req)), indent=2))
EOF

Get Connection

python <<'EOF'
import urllib.request, os, json
req = urllib.request.Request('https://ctrl.maton.ai/connections/{connection_id}')
req.add_header('Authorization', f'Bearer {os.environ["MATON_API_KEY"]}')
print(json.dumps(json.load(urllib.request.urlopen(req)), indent=2))
EOF

Response:

{
  "connection": {
    "connection_id": "{connection_id}",
    "status": "ACTIVE",
    "creation_time": "2026-04-07T21:20:16.974921Z",
    "last_updated_time": "2026-04-07T21:23:43.726795Z",
    "url": "https://connect.maton.ai/?session_token=...",
    "app": "apify",
    "metadata": {}
  }
}

Open the returned url in a browser to complete authentication setup.

Delete Connection

python <<'EOF'
import urllib.request, os, json
req = urllib.request.Request('https://ctrl.maton.ai/connections/{connection_id}', method='DELETE')
req.add_header('Authorization', f'Bearer {os.environ["MATON_API_KEY"]}')
print(json.dumps(json.load(urllib.request.urlopen(req)), indent=2))
EOF

Specifying Connection

If you have multiple Apify connections, specify which one to use with the Maton-Connection header:

python <<'EOF'
import urllib.request, os, json
req = urllib.request.Request('https://gateway.maton.ai/apify/v2/acts')
req.add_header('Authorization', f'Bearer {os.environ["MATON_API_KEY"]}')
req.add_header('Maton-Connection', '{connection_id}')
print(json.dumps(json.load(urllib.request.urlopen(req)), indent=2))
EOF

If you have multiple connections, always include this header to ensure requests go to the intended account.

API Reference

User

Get Current User

GET /apify/v2/users/me

Response:

{
  "data": {
    "id": "GgXk48GBlDInv62bA",
    "username": "my_username",
    "profile": {
      "name": "John Doe",
      "pictureUrl": "https://..."
    },
    "email": "john@example.com",
    "plan": {
      "id": "FREE",
      "description": "Free plan",
      "monthlyUsageCreditsUsd": 5
    },
    "createdAt": "2024-04-27T22:08:45.429Z"
  }
}

Actors

List Actors

GET /apify/v2/acts
GET /apify/v2/acts?limit=10&offset=0

Response:

{
  "data": {
    "total": 4,
    "count": 4,
    "offset": 0,
    "limit": 1000,
    "desc": false,
    "items": [
      {
        "id": "moJRLRc85AitArpNN",
        "name": "web-scraper",
        "username": "apify",
        "title": "Web Scraper",
        "createdAt": "2019-03-07T11:28:01.600Z",
        "modifiedAt": "2026-03-11T14:36:47.849Z",
        "stats": {
          "totalRuns": 2,
          "lastRunStartedAt": "2026-04-07T21:24:57.927Z"
        }
      }
    ]
  }
}

Get Actor

GET /apify/v2/acts/{actorId}

Run Actor

POST /apify/v2/acts/{actorId}/runs
Content-Type: application/json

{
  "startUrls": [{"url": "https://example.com"}],
  "maxRequestsPerCrawl": 10
}

Response:

{
  "data": {
    "id": "mxA2b6luHFdcxBZuG",
    "actId": "moJRLRc85AitArpNN",
    "status": "RUNNING",
    "startedAt": "2026-04-07T21:24:57.927Z",
    "defaultKeyValueStoreId": "qP9EdMQrEqNcC2PzZ",
    "defaultDatasetId": "E9O7dXhrNxgA06o5k",
    "defaultRequestQueueId": "N3xb1qGmNzoxPNaAW"
  }
}

Actor Runs

List Actor Runs

GET /apify/v2/actor-runs
GET /apify/v2/actor-runs?limit=10&desc=1

Response:

{
  "data": {
    "total": 1,
    "count": 1,
    "offset": 0,
    "limit": 1000,
    "items": [
      {
        "id": "mxA2b6luHFdcxBZuG",
        "actId": "moJRLRc85AitArpNN",
        "status": "SUCCEEDED",
        "startedAt": "2026-04-07T21:24:57.927Z",
        "finishedAt": "2026-04-07T21:25:08.086Z",
        "defaultDatasetId": "E9O7dXhrNxgA06o5k",
        "usageTotalUsd": 0.0037
      }
    ]
  }
}

Get Actor Run

GET /apify/v2/actor-runs/{runId}

Response:

{
  "data": {
    "id": "mxA2b6luHFdcxBZuG",
    "actId": "moJRLRc85AitArpNN",
    "status": "SUCCEEDED",
    "statusMessage": "Finished! Total 1 requests: 1 succeeded, 0 failed.",
    "startedAt": "2026-04-07T21:24:57.927Z",
    "finishedAt": "2026-04-07T21:25:08.086Z",
    "stats": {
      "durationMillis": 10009,
      "runTimeSecs": 10.009,
      "computeUnits": 0.011,
      "memAvgBytes": 254919122,
      "cpuAvgUsage": 14.67
    },
    "defaultKeyValueStoreId": "qP9EdMQrEqNcC2PzZ",
    "defaultDatasetId": "E9O7dXhrNxgA06o5k",
    "defaultRequestQueueId": "N3xb1qGmNzoxPNaAW"
  }
}

Abort Actor Run

POST /apify/v2/actor-runs/{runId}/abort

Resurrect Actor Run

POST /apify/v2/actor-runs/{runId}/resurrect

Actor Tasks

List Actor Tasks

GET /apify/v2/actor-tasks

Get Actor Task

GET /apify/v2/actor-tasks/{taskId}

Create Actor Task

POST /apify/v2/actor-tasks
Content-Type: application/json

{
  "actId": "moJRLRc85AitArpNN",
  "name": "my-scraping-task",
  "options": {
    "build": "latest",
    "memoryMbytes": 1024,
    "timeoutSecs": 300
  },
  "input": {
    "startUrls": [{"url": "https://example.com"}]
  }
}

Run Actor Task

POST /apify/v2/actor-tasks/{taskId}/runs

Update Actor Task

PUT /apify/v2/actor-tasks/{taskId}
Content-Type: application/json

{
  "name": "updated-task-name"
}

Delete Actor Task

DELETE /apify/v2/actor-tasks/{taskId}

Datasets

List Datasets

GET /apify/v2/datasets

Get Dataset

GET /apify/v2/datasets/{datasetId}

Create Dataset

POST /apify/v2/datasets
Content-Type: application/json

{
  "name": "my-dataset"
}

Get Dataset Items

GET /apify/v2/datasets/{datasetId}/items
GET /apify/v2/datasets/{datasetId}/items?format=json&limit=100

Response:

[
  {
    "title": "Example Domain",
    "url": "https://example.com",
    "#debug": {
      "requestId": "zYk68OuvhfdFudP",
      "statusCode": 200
    }
  }
]

Push Items to Dataset

POST /apify/v2/datasets/{datasetId}/items
Content-Type: application/json

[
  {"title": "Item 1", "url": "https://example1.com"},
  {"title": "Item 2", "url": "https://example2.com"}
]

Delete Dataset

DELETE /apify/v2/datasets/{datasetId}

Key-Value Stores

List Key-Value Stores

GET /apify/v2/key-value-stores

Get Key-Value Store

GET /apify/v2/key-value-stores/{storeId}

Response:

{
  "data": {
    "id": "qP9EdMQrEqNcC2PzZ",
    "name": null,
    "userId": "GgXk48GBlDInv62bA",
    "createdAt": "2026-04-07T21:24:57.930Z",
    "stats": {
      "readCount": 2,
      "writeCount": 6,
      "storageBytes": 2018
    }
  }
}

Create Key-Value Store

POST /apify/v2/key-value-stores
Content-Type: application/json

{
  "name": "my-store"
}

List Keys

GET /apify/v2/key-value-stores/{storeId}/keys

Get Record

GET /apify/v2/key-value-stores/{storeId}/records/{key}

Put Record

PUT /apify/v2/key-value-stores/{storeId}/records/{key}
Content-Type: application/json

{"data": "value"}

Delete Record

DELETE /apify/v2/key-value-stores/{storeId}/records/{key}

Delete Key-Value Store

DELETE /apify/v2/key-value-stores/{storeId}

Request Queues

List Request Queues

GET /apify/v2/request-queues

Get Request Queue

GET /apify/v2/request-queues/{queueId}

Create Request Queue

POST /apify/v2/request-queues
Content-Type: application/json

{
  "name": "my-queue"
}

Add Request to Queue

POST /apify/v2/request-queues/{queueId}/requests
Content-Type: application/json

{
  "url": "https://example.com",
  "uniqueKey": "example-key"
}

Delete Request Queue

DELETE /apify/v2/request-queues/{queueId}

Schedules

List Schedules

GET /apify/v2/schedules

Get Schedule

GET /apify/v2/schedules/{scheduleId}

Create Schedule

POST /apify/v2/schedules
Content-Type: application/json

{
  "name": "daily-scrape",
  "cronExpression": "0 0 * * *",
  "isEnabled": true,
  "actions": [
    {
      "type": "RUN_ACTOR_TASK",
      "actorTaskId": "task123"
    }
  ]
}

Update Schedule

PUT /apify/v2/schedules/{scheduleId}
Content-Type: application/json

{
  "isEnabled": false
}

Delete Schedule

DELETE /apify/v2/schedules/{scheduleId}

Webhooks

List Webhooks

GET /apify/v2/webhooks

Get Webhook

GET /apify/v2/webhooks/{webhookId}

Create Webhook

POST /apify/v2/webhooks
Content-Type: application/json

{
  "eventTypes": ["ACTOR.RUN.SUCCEEDED"],
  "requestUrl": "https://example.com/webhook",
  "condition": {
    "actorId": "moJRLRc85AitArpNN"
  }
}

Update Webhook

PUT /apify/v2/webhooks/{webhookId}
Content-Type: application/json

{
  "isAdHoc": false
}

Delete Webhook

DELETE /apify/v2/webhooks/{webhookId}

Pagination

Apify uses offset-based pagination:

GET /apify/v2/acts?limit=10&offset=20&desc=1

Parameters:

  • limit - Maximum items per response (default: 1000)
  • offset - Number of items to skip
  • desc - Set to 1 for descending order (newest first)

Response includes:

{
  "data": {
    "total": 100,
    "count": 10,
    "offset": 20,
    "limit": 10,
    "desc": true,
    "items": [...]
  }
}

For key-value stores, use key-based pagination:

GET /apify/v2/key-value-stores/{storeId}/keys?limit=100&exclusiveStartKey=lastKey

Code Examples

JavaScript

const response = await fetch(
  'https://gateway.maton.ai/apify/v2/acts',
  {
    headers: {
      'Authorization': `Bearer ${process.env.MATON_API_KEY}`
    }
  }
);
const data = await response.json();
console.log(data.data.items);

Python

import os
import requests

response = requests.get(
    'https://gateway.maton.ai/apify/v2/actor-runs',
    headers={'Authorization': f'Bearer {os.environ["MATON_API_KEY"]}'},
    params={'limit': 10, 'desc': 1}
)
runs = response.json()['data']['items']

Run Actor and Get Results

import os
import requests
import time

headers = {
    'Authorization': f'Bearer {os.environ["MATON_API_KEY"]}',
    'Content-Type': 'application/json'
}

# Run an actor
run_resp = requests.post(
    'https://gateway.maton.ai/apify/v2/acts/apify~web-scraper/runs',
    headers=headers,
    json={
        'startUrls': [{'url': 'https://example.com'}],
        'maxRequestsPerCrawl': 10
    }
)
run = run_resp.json()['data']
run_id = run['id']
dataset_id = run['defaultDatasetId']

# Wait for completion
while True:
    status_resp = requests.get(
        f'https://gateway.maton.ai/apify/v2/actor-runs/{run_id}',
        headers=headers
    )
    status = status_resp.json()['data']['status']
    if status in ['SUCCEEDED', 'FAILED', 'ABORTED']:
        break
    time.sleep(5)

# Get results
items_resp = requests.get(
    f'https://gateway.maton.ai/apify/v2/datasets/{dataset_id}/items',
    headers=headers
)
results = items_resp.json()
print(f'Scraped {len(results)} items')

Notes

  • Actor IDs can be specified as {username}~{actorName} (e.g., apify~web-scraper) or by ID
  • Run statuses: READY, RUNNING, SUCCEEDED, FAILED, ABORTING, ABORTED, TIMING-OUT, TIMED-OUT
  • Dataset items can be retrieved in various formats: json, jsonl, csv, xlsx, xml, rss
  • Key-value store records can store any content type
  • Schedule cron expressions follow standard cron format
  • IMPORTANT: When using curl commands, use curl -g when URLs contain brackets to disable glob parsing
  • IMPORTANT: When piping curl output to jq, environment variables may not expand correctly in some shells

Rate Limits

ScopeLimit
Global250,000 requests/minute
Default per-resource60 requests/second
Key-Value Store CRUD200 requests/second
Dataset & Queue operations400 requests/second

Error Handling

StatusMeaning
400Missing Apify connection or invalid request
401Invalid or missing Maton API key
404Resource not found
429Rate limited (use exponential backoff)
4xx/5xxPassthrough error from Apify API

Resources

常见问题

这个技能会自动处理认证吗?
会。所有请求都经过 Maton 网关,由网关自动注入你的 Apify API token。你只需要在 Authorization 头里带上 Maton API key,或通过 `maton` CLI 用 OAuth 登录。
如果我绑定了多个 Apify 账户怎么办?
可以在 ctrl.maton.ai 列出所有连接,并通过 Maton-Connection 头在每次请求时指定要操作的账户,避免误操作。
直接运行 actor 或修改定时任务是否安全?
Actor 运行会消耗 Apify 算力单位,定时任务与 webhook 属于会话结束后仍持续运行的资源。技能默认只做读和列表操作,遇到写操作或新建连接时会先向用户确认。

相关技能

通过 Maton 网关调用 Tavily API,完成网页搜索、内容提取、站点爬取与异步研究任务。

30 次安装

通过托管认证调用 Firecrawl API,提供网页抓取、整站爬取、URL 发现和带正文内容的搜索能力。

95 次安装1 星标

通过托管 OAuth 代理调用 Apollo.io API,完成人员与公司搜索、联系人补全和销售数据管理。

226 次安装5 星标

通过托管 OAuth 代理接入 Google Search Console,查询搜索分析数据、管理 sitemap 并查看站点表现。

265 次安装10 星标

通过托管密钥认证调用 Exa API,完成网页搜索、内容抓取、相似页查找与异步研究任务。

26 次安装

通过托管 OAuth,在 WordPress.com REST v1.1 API 上读写文章、页面、站点和内容。

158 次安装8 星标