数据分析

Kaggle

试用

通过托管 API 网关搜索、抓取与下载 Kaggle 数据集、模型、竞赛与 Notebook。

它能做什么

借助 Maton 托管网关调用 Kaggle 公开 RPC 接口,只需一个 Bearer 令牌(MATON_API_KEY)即可完成鉴权。所有请求均为 POST + JSON,路径遵循 `/v1/{ServiceName}/{MethodName}` 模式。覆盖数据集(列出、获取、列文件、元信息、下载)、模型(列出、获取)、竞赛(列出)以及 kernel(列出、获取)。连接在 ctrl.maton.ai 统一管理,Kaggle 凭据无需写入业务代码。

什么时候用它

  • 按关键词搜索 Kaggle 数据集
  • 列出某作者名下可用的模型
  • 下载 ZIP 前先获取数据集元信息
  • 按分类筛选当前可参与的竞赛

技能文档

Kaggle

Access Kaggle datasets, models, competitions, and notebooks via managed API authentication.

Quick Start

python <<'EOF'
import urllib.request, os, json
data = json.dumps({}).encode()
req = urllib.request.Request('https://gateway.maton.ai/kaggle/v1/datasets.DatasetApiService/ListDatasets', data=data, method='POST')
req.add_header('Authorization', f'Bearer {os.environ["MATON_API_KEY"]}')
req.add_header('Content-Type', 'application/json')
print(json.dumps(json.load(urllib.request.urlopen(req)), indent=2))
EOF

Base URL

https://gateway.maton.ai/kaggle/{native-api-path}

The gateway proxies requests to api.kaggle.com and automatically injects your credentials.

Authentication

All requests require the Maton API key:

Authorization: Bearer $MATON_API_KEY

Environment Variable: Set your API key as MATON_API_KEY:

export MATON_API_KEY="YOUR_API_KEY"

Getting Your API Key

  1. Sign in or create an account at maton.ai
  2. Go to maton.ai/settings
  3. Copy your API key

Connection Management

Manage your Kaggle connections at https://ctrl.maton.ai.

List Connections

python <<'EOF'
import urllib.request, os, json
req = urllib.request.Request('https://ctrl.maton.ai/connections?app=kaggle&status=ACTIVE')
req.add_header('Authorization', f'Bearer {os.environ["MATON_API_KEY"]}')
print(json.dumps(json.load(urllib.request.urlopen(req)), indent=2))
EOF

Create Connection

python <<'EOF'
import urllib.request, os, json
data = json.dumps({'app': 'kaggle'}).encode()
req = urllib.request.Request('https://ctrl.maton.ai/connections', data=data, method='POST')
req.add_header('Authorization', f'Bearer {os.environ["MATON_API_KEY"]}')
req.add_header('Content-Type', 'application/json')
print(json.dumps(json.load(urllib.request.urlopen(req)), indent=2))
EOF

Open the returned url in a browser to complete authentication. Kaggle uses API key authentication - you'll need to provide your Kaggle username and API key from kaggle.com/settings.

Delete Connection

python <<'EOF'
import urllib.request, os, json
req = urllib.request.Request('https://ctrl.maton.ai/connections/{connection_id}', method='DELETE')
req.add_header('Authorization', f'Bearer {os.environ["MATON_API_KEY"]}')
print(json.dumps(json.load(urllib.request.urlopen(req)), indent=2))
EOF

API Reference

Kaggle uses an RPC-style API. All requests are POST with JSON body.

POST /kaggle/v1/{ServiceName}/{MethodName}
Content-Type: application/json

Datasets

List Datasets

POST /kaggle/v1/datasets.DatasetApiService/ListDatasets
Content-Type: application/json

{}

Request Body Parameters:

  • search - Search term (optional)
  • user - Filter by username (optional)
  • pageSize - Results per page (optional)
  • pageToken - Pagination token (optional)

Example with search:

{
  "search": "covid"
}

Response:

{
  "datasets": [
    {
      "id": 9481458,
      "ref": "amar5693/screen-time-sleep-and-stress-analysis-dataset",
      "title": "Screen Time, Sleep & Stress Analysis Dataset",
      "subtitle": "ML-ready dataset analyzing smartphone usage and productivity.",
      "totalBytes": 787136,
      "downloadCount": 11659,
      "voteCount": 236,
      "usabilityRating": 1,
      "licenseName": "CC0: Public Domain",
      "ownerName": "Amar Tiwari",
      "tags": [...]
    }
  ]
}

Get Dataset

POST /kaggle/v1/datasets.DatasetApiService/GetDataset
Content-Type: application/json

{
  "ownerSlug": "amar5693",
  "datasetSlug": "screen-time-sleep-and-stress-analysis-dataset"
}

Response:

{
  "id": 9481458,
  "title": "Screen Time, Sleep & Stress Analysis Dataset",
  "subtitle": "ML-ready dataset analyzing smartphone usage and productivity.",
  "totalBytes": 787136,
  "downloadCount": 11659,
  "usabilityRating": 1
}

List Dataset Files

POST /kaggle/v1/datasets.DatasetApiService/ListDatasetFiles
Content-Type: application/json

{
  "ownerSlug": "amar5693",
  "datasetSlug": "screen-time-sleep-and-stress-analysis-dataset"
}

Response:

{
  "datasetFiles": [
    {
      "name": "Smartphone_Usage_Productivity_Dataset_50000.csv",
      "creationDate": "2026-02-13T06:56:19.803Z",
      "totalBytes": 2958561
    }
  ]
}

Get Dataset Metadata

POST /kaggle/v1/datasets.DatasetApiService/GetDatasetMetadata
Content-Type: application/json

{
  "ownerSlug": "amar5693",
  "datasetSlug": "screen-time-sleep-and-stress-analysis-dataset"
}

Response:

{
  "info": {
    "datasetId": 9481458,
    "datasetSlug": "screen-time-sleep-and-stress-analysis-dataset",
    "ownerUser": "amar5693",
    "title": "Screen Time, Sleep & Stress Analysis Dataset",
    "description": "...",
    "totalViews": 44291,
    "totalVotes": 236,
    "totalDownloads": 11661
  }
}

Download Dataset

POST /kaggle/v1/datasets.DatasetApiService/DownloadDataset
Content-Type: application/json

{
  "ownerSlug": "amar5693",
  "datasetSlug": "screen-time-sleep-and-stress-analysis-dataset"
}

Returns binary data (ZIP file). Response headers:

  • Content-Type: application/zip
  • Content-Length:

Models

List Models

POST /kaggle/v1/models.ModelApiService/ListModels
Content-Type: application/json

{}

Request Body Parameters:

  • owner - Filter by owner (optional)
  • search - Search term (optional)
  • pageSize - Results per page (optional)

Example:

{
  "owner": "google"
}

Response:

{
  "models": [
    {
      "id": 1,
      "owner": "google",
      "slug": "gemma",
      "title": "Gemma",
      "subtitle": "Gemma is a family of lightweight, state-of-the-art models",
      "instanceCount": 16,
      "framework": "transformers"
    }
  ]
}

Get Model

POST /kaggle/v1/models.ModelApiService/GetModel
Content-Type: application/json

{
  "ownerSlug": "google",
  "modelSlug": "gemma"
}

Response:

{
  "id": 1,
  "title": "Gemma",
  "slug": "gemma",
  "owner": "google",
  "subtitle": "Gemma is a family of lightweight, state-of-the-art models",
  "publishTime": "2024-02-21T16:00:00Z",
  "instanceCount": 16
}

Competitions

List Competitions

POST /kaggle/v1/competitions.CompetitionApiService/ListCompetitions
Content-Type: application/json

{}

Request Body Parameters:

  • search - Search term (optional)
  • category - Filter by category (optional)
  • pageSize - Results per page (optional)

Example:

{
  "search": "nlp"
}

Response:

{
  "competitions": [
    {
      "id": 118448,
      "ref": "https://www.kaggle.com/competitions/ai-mathematical-olympiad-progress-prize-3",
      "title": "AI Mathematical Olympiad - Progress Prize 3",
      "url": "https://www.kaggle.com/competitions/ai-mathematical-olympiad-progress-prize-3",
      "deadline": "2026-06-06T23:59:00Z",
      "category": "Featured",
      "reward": "$1,048,576",
      "teamCount": 1234,
      "userHasEntered": false
    }
  ]
}

Kernels (Notebooks)

List Kernels

POST /kaggle/v1/kernels.KernelsApiService/ListKernels
Content-Type: application/json

{}

Request Body Parameters:

  • search - Search term (optional)
  • user - Filter by username (optional)
  • language - Filter by language: python, r, etc. (optional)
  • pageSize - Results per page (optional)

Example:

{
  "search": "titanic"
}

Response:

{
  "kernels": [
    {
      "id": 5660537,
      "ref": "alexisbcook/titanic-tutorial",
      "title": "Titanic Tutorial",
      "author": "alexisbcook",
      "language": "Python",
      "totalVotes": 1234,
      "totalViews": 56789
    }
  ]
}

Get Kernel

POST /kaggle/v1/kernels.KernelsApiService/GetKernel
Content-Type: application/json

{
  "userName": "alexisbcook",
  "kernelSlug": "titanic-tutorial"
}

Response:

{
  "metadata": {
    "id": 5660537,
    "ref": "alexisbcook/titanic-tutorial",
    "title": "Titanic Tutorial",
    "author": "alexisbcook",
    "language": "Python"
  }
}

Code Examples

JavaScript

const response = await fetch(
  'https://gateway.maton.ai/kaggle/v1/datasets.DatasetApiService/ListDatasets',
  {
    method: 'POST',
    headers: {
      'Authorization': `Bearer ${process.env.MATON_API_KEY}`,
      'Content-Type': 'application/json'
    },
    body: JSON.stringify({ search: 'covid' })
  }
);
const data = await response.json();
console.log(data);

Python

import os
import requests

response = requests.post(
    'https://gateway.maton.ai/kaggle/v1/datasets.DatasetApiService/ListDatasets',
    headers={
        'Authorization': f'Bearer {os.environ["MATON_API_KEY"]}',
        'Content-Type': 'application/json'
    },
    json={'search': 'covid'}
)
print(response.json())

Notes

  • All API calls use POST method with JSON body
  • API follows RPC pattern: /v1/{ServiceName}/{MethodName}
  • Dataset refs use format: {owner}/{dataset-slug}
  • Model refs use format: {owner}/{model-slug}
  • Kernel refs use format: {user}/{kernel-slug}
  • Download endpoints return binary data (ZIP files)
  • Some operations require specific permissions (competition participation, kernel access)

Error Handling

StatusMeaning
200Success
400Invalid request parameters
401Invalid or missing authentication
403Permission denied
404Resource not found
429Rate limited

Resources

相关技能

Kaggle (kaggle.com). Use this skill for ANY Kaggle request — searching and reading data. Whenever a task involves Kaggle, use this skill instead of calling t...

通过托管 OAuth 代理接入 Google Search Console,查询搜索分析数据、管理 sitemap 并查看站点表现。

265 次安装10 星标

通过托管 OAuth 调用 Confluence Cloud API,管理页面、空间、博客、评论与附件。

43 次安装

通过托管 OAuth 连接 Google Tasks,统一 API 完成任务列表与任务的读写管理。

253 次安装10 星标

通过 Maton 网关调用 Tavily API,完成网页搜索、内容提取、站点爬取与异步研究任务。

30 次安装

通过托管认证接入 Apify API,运行爬虫并管理 actors、数据集、键值存储与定时任务。

26 次安装