编程

API Rate Limiting

试用

Rate limiting algorithms, implementation strategies, HTTP conventions, tiered limits, distributed patterns, and client-side handling. Use when protecting APIs from abuse, implementing usage tiers, or configuring gateway-level throttling.

它能做什么

| Algorithm | Accuracy | Burst Handling | Best For | |-----------|----------|----------------|----------| | **Token Bucket** | High | Allows controlled bursts | API rate limiting, traffic shaping | | **Leaky Bucket** | High | Smooths bursts entirely | Steady-rate processing, queues | | **Fixed Wind…

技能文档

Rate Limiting Patterns

Algorithms

AlgorithmAccuracyBurst HandlingBest For
Token BucketHighAllows controlled burstsAPI rate limiting, traffic shaping
Leaky BucketHighSmooths bursts entirelySteady-rate processing, queues
Fixed WindowLowAllows edge bursts (2x)Simple use cases, prototyping
Sliding Window LogVery HighPrecise controlStrict compliance, billing-critical
Sliding Window CounterHighGood approximationProduction APIs — best tradeoff

Fixed window problem: A user sends the full limit at 11:59 and again at 12:01, doubling the effective rate. Sliding window fixes this.

Token Bucket

Bucket holds tokens up to capacity. Tokens refill at a fixed rate. Each request consumes one.

class TokenBucket:
    def __init__(self, capacity: int, refill_rate: float):
        self.capacity = capacity
        self.tokens = capacity
        self.refill_rate = refill_rate  # tokens per second
        self.last_refill = time.monotonic()

    def allow(self) -> bool:
        now = time.monotonic()
        elapsed = now - self.last_refill
        self.tokens = min(self.capacity, self.tokens + elapsed * self.refill_rate)
        self.last_refill = now
        if self.tokens >= 1:
            self.tokens -= 1
            return True
        return False

Sliding Window Counter

Hybrid of fixed window and sliding window log — weights the previous window's count by overlap percentage:

def sliding_window_allow(key: str, limit: int, window_sec: int) -> bool:
    now = time.time()
    current_window = int(now // window_sec)
    position_in_window = (now % window_sec) / window_sec

    prev_count = get_count(key, current_window - 1)
    curr_count = get_count(key, current_window)

    estimated = prev_count * (1 - position_in_window) + curr_count
    if estimated >= limit:
        return False
    increment_count(key, current_window)
    return True

Implementation Options

ApproachScopeBest For
In-memorySingle serverZero latency, no dependencies
Redis (INCR + EXPIRE)DistributedMulti-instance deployments
API GatewayEdgeNo code, built-in dashboards
MiddlewarePer-serviceFine-grained per-user/endpoint control

Use gateway-level limiting as outer defense + application-level for fine-grained control.


HTTP Headers

Always return rate limit info, even on successful requests:

RateLimit-Limit: 1000
RateLimit-Remaining: 742
RateLimit-Reset: 1625097600
Retry-After: 30
HeaderWhen to Include
RateLimit-LimitEvery response
RateLimit-RemainingEvery response
RateLimit-ResetEvery response
Retry-After429 responses only

429 Response Body

{
  "error": {
    "code": "rate_limit_exceeded",
    "message": "Rate limit exceeded. Maximum 1000 requests per hour.",
    "retry_after": 30,
    "limit": 1000,
    "reset_at": "2025-07-01T12:00:00Z"
  }
}

Never return 500 or 503 for rate limiting — 429 is the correct status code.


Rate Limit Tiers

Apply limits at multiple granularities:

ScopeKeyExample LimitPurpose
Per-IPClient IP100 req/minAbuse prevention
Per-UserUser ID1000 req/hrFair usage
Per-API-KeyAPI key5000 req/hrService-to-service
Per-EndpointRoute + key60 req/min on /searchProtect expensive ops

Tiered pricing:

TierRate LimitBurstCost
Free100 req/hr10$0
Pro5,000 req/hr100$49/mo
Enterprise100,000 req/hr2,000Custom

Evaluate from most specific to least specific: per-endpoint > per-user > per-IP.


Distributed Rate Limiting

Redis-based pattern for consistent limiting across instances:

def redis_rate_limit(redis, key: str, limit: int, window: int) -> bool:
    pipe = redis.pipeline()
    now = time.time()
    window_key = f"rl:{key}:{int(now // window)}"
    pipe.incr(window_key)
    pipe.expire(window_key, window * 2)
    results = pipe.execute()
    return results[0] <= limit

Atomic Lua script (prevents race conditions):

local key = KEYS[1]
local limit = tonumber(ARGV[1])
local window = tonumber(ARGV[2])
local current = redis.call('INCR', key)
if current == 1 then
    redis.call('EXPIRE', key, window)
end
return current <= limit and 1 or 0

Never do separate GET then SET — the gap allows overcount.


API Gateway Configuration

NGINX:

http {
    limit_req_zone $binary_remote_addr zone=api:10m rate=10r/s;
    server {
        location /api/ {
            limit_req zone=api burst=20 nodelay;
            limit_req_status 429;
        }
    }
}

Kong:

plugins:
  - name: rate-limiting
    config:
      minute: 60
      hour: 1000
      policy: redis
      redis_host: redis.internal

Client-Side Handling

Clients must handle 429 gracefully:

async function fetchWithRetry(url: string, maxRetries = 3): Promise {
  for (let attempt = 0; attempt < maxRetries; attempt++) {
    const res = await fetch(url);
    if (res.status !== 429) return res;

    const retryAfter = res.headers.get('Retry-After');
    const delay = retryAfter
      ? parseInt(retryAfter, 10) * 1000
      : Math.min(1000 * 2 ** attempt, 30000);
    await new Promise(r => setTimeout(r, delay));
  }
  throw new Error('Rate limit exceeded after retries');
}
  • Always respect Retry-After when present
  • Use exponential backoff with jitter when absent
  • Implement request queuing for batch operations

Monitoring

Track these metrics:

  • Rate limit hit rate — % of requests returning 429 (alert if >5% sustained)
  • Near-limit warnings — requests where remaining < 10% of limit
  • Top offenders — keys/IPs hitting limits most frequently
  • Limit headroom — how close normal traffic is to the ceiling
  • False positives — legitimate users being rate limited

Anti-Patterns

Anti-PatternFix
Application-only limitingAlways combine with infrastructure-level limits
No retry guidanceAlways include Retry-After header on 429
Inconsistent limitsSame endpoint, same limits across services
No burst allowanceAllow controlled bursts for legitimate traffic
Silent droppingAlways return 429 so clients can distinguish from errors
Global single counterPer-endpoint counters to protect expensive operations
Hard-coded limitsUse configuration, not code constants

NEVER Do

  1. NEVER rate limit health check endpoints — monitoring systems will false-alarm
  2. NEVER use client-supplied identifiers as sole rate limit key — trivially spoofed
  3. NEVER return 200 OK when rate limiting — clients must know they were throttled
  4. NEVER set limits without measuring actual traffic first — you'll block legitimate users or set limits too high to matter
  5. NEVER share counters across unrelated tenants — noisy neighbor problem
  6. NEVER skip rate limiting on internal APIs — misbehaving internal services can take down shared infrastructure
  7. NEVER implement rate limiting without logging — you need visibility to tune limits and detect abuse

相关技能

诊断生产力系统反复失效的根因,给出最小干预——容量测算、瓶颈定位、可靠的本地记录。

作者 Iván854 次安装69 星标

按用户明确指令,在得到大脑(Get笔记)中保存、搜索并管理笔记与知识库。

作者 iswalle763 次安装66 星标

figma-use

官方

遵守 `use_figma` 脚本编写规范,避免在 Figma Plugin API 上踩常见陷阱导致静默失败。

作者 OpenAI27.6k 星标

按微软当前文档,选对 ASP.NET Core 应用模型,搭好主机、请求管道与实现方式。

作者 OpenAI27.6k 星标

围绕 SKILL.md 的编写、迭代与对照评测,逐步打磨技能质量。

作者 Anthropic177.6k 星标

为 Codex 搭一个可在任意目录下按命令名运行的长期 CLI,提供组合式子命令和稳定 JSON 输出。

作者 OpenAI27.6k 星标

wpank 的更多技能

浏览全部技能

Systematic code review patterns covering security, performance, maintainability, correctness, and testing — with severity levels, structured feedback guidance, review process, and anti-patterns to avoid. Use when reviewing PRs, establishing review standards, or improving review quality.

作者 wpank553 次安装20 星标

Pragmatic coding standards for writing clean, maintainable code — naming, functions, structure, anti-patterns, and pre-edit safety checks. Use when writing new code, refactoring existing code, reviewing code quality, or establishing coding standards.

作者 wpank198 次安装6 星标

Build reliable, fast E2E test suites with Playwright and Cypress. Critical user journey coverage, flaky test elimination, CI/CD integration.

作者 wpank336 次安装6 星标

Build scalable, themable Tailwind CSS component libraries using CVA for variants, compound components, design tokens, dark mode, and responsive grids.

作者 wpank217 次安装9 星标

Create software diagrams using Mermaid syntax. Use when users need to create, visualize, or document software through diagrams including class diagrams, sequence diagrams, flowcharts, ERDs, C4 architecture diagrams, state diagrams, git graphs, and other diagram types. Triggers include requests to diagram, visualize, model, map out, or show the flow of a system.

作者 wpank251 次安装5 星标

Provides backend architecture patterns (Clean Architecture, Hexagonal, DDD) for building maintainable, testable, and scalable systems with clear layering and...

作者 wpank152 次安装7 星标