Skip to main content
The limits exist to stop runaway clients, not to ration normal use. They sit far above what a working integration needs. If you meet one, your client is most likely looping or leaking connections.

Limits

Every key, evaluation or production, has the same limits.

How the request budget refills

A key’s budget refills evenly over each minute, and a full minute’s budget can be spent at once. A backfill can therefore burst, then settle to the steady rate. The IP limit sits above the key limit, so a fleet of your servers behind one outbound IP still gets its full key budget.

Approximate by design

Limits are counted per server, not across the whole fleet. Treat the numbers as ceilings to stay well under, not as exact quotas to run at.

Response headers

Once your key is accepted, X-RateLimit-Remaining reports your key’s budget. A response that ends before that point reports your IP address’s budget instead: a 401, a 429 from the IP limit, or a 503 when the key could not be checked. Pace your client on X-RateLimit-Remaining to slow down before you reach the limit.

Stream connections

A connection counts against your key from its handshake until it closes. Requests over the socket do not spend your request budget. When your key holds 100 open connections, the next handshake gets:
Waiting alone does not free a slot. Close a connection you no longer need. A single connection holds up to 100 subscriptions, so most integrations need only a few.

The OpenAPI document

GET /openapi.json needs no key, spends no budget, and carries no X-RateLimit-Remaining header.