How it works
One limiter (GCRA, atomic in a Valkey script) fronts /v1, /mcp and /oauth/*. Several
budgets are evaluated for every request and the tightest one wins the headers. Authenticated requests are keyed per
principal (the user plus the specific connection or API key) and per user; unauthenticated ones per source
IP (IPv6 addresses are bucketed per /64). Budgets refill continuously, so โ60/minโ means a burst of 60 and then one request
a second, not a hard reset on the minute.
When the limiter's store is unreachable, public paths fail closed: 503 rate_limiter_unavailable
(or {"error": "temporarily_unavailable"} on /oauth/*) with Retry-After: 5.
Budgets: authenticated paths
| Class | Applies to | Per principal | Per user | Notes |
|---|---|---|---|---|
| Reads | GET /v1/me, every /mcp request, and every future read endpoint | 60/min | 120/min | Live. Per principal = per (user, connection or key); per user across all of them. An alert-only ceiling of 3,000/min per client id never blocks. |
| Transcript requests | POST /v1/calls/{id}/transcript | 5/min | 5/min | Reserved: endpoint not live. |
| Quotes | POST /v1/quotes, voixcall_quote_call | 20/min | 20/min | Reserved: endpoint not live. |
| Call placement | POST /v1/calls, voixcall_place_call | 10/min | 10/min | Reserved: endpoint not live. Plus one concurrent callback session. |
| Message reads | GET /v1/messages, voixcall_read_messages | 60/min | 120/min | Reserved: endpoint not live. Plus 10/min per number. |
| Key management | /v1/api-keys (dashboard session) | โ | 30/min | Live. |
Budgets: unauthenticated paths and OAuth
| Class | Applies to | Limit | Notes |
|---|---|---|---|
| Metadata | /.well-known/*, /v1/openapi.json | 300/min per IP | Live. |
| Authorize and register | GET /oauth/authorize, POST /oauth/register | 30/min per IP | Live. |
| Register | POST /oauth/register | 20/hour per IP, on top of the row above | Live. Also a global cap of 5,000 unexpired self-registered clients; beyond it registration answers 503 with Retry-After: 3600. |
| Token flood guard | POST /oauth/token, POST /oauth/revoke | burst 600, then 120/min sustained per IP | Live. The burst is deliberately large so one platform egress address carrying many users is not the sharp edge; the RateLimit-Policy header advertises the burst (600;w=60). |
| Failed bearer checks | /v1, /mcp | 60/min per IP | Live. Consumed only when a presented credential fails the lookup. Requests with no credential (the discovery probe) and valid tokens never touch it. Once exhausted, the failure is answered 429 rate_limited instead of 401. |
OAuth lockouts
Penalties attach to the source or the credential, never to a public client id, so one misbehaving installation cannot lock every user of a shared client (Claude, ChatGPT) out of refreshes. Only guessing signals count: presenting a known-but-revoked token after a disconnect does not.
| Dimension | Limit | Effect |
|---|---|---|
| Failed exchanges per source IP | 10/min | 60 s back-off for that IP: 429 temporarily_unavailable before any grant work |
| Failed exchanges per presented credential | 3 per hour | the code or refresh token and its whole grant family are revoked |
| Failed exchanges per client id | 1,000/min | alert only, never blocks |
| Successful token requests per client id | 600/min | alert only, never blocks (a public client id is shared by every installation) |
| Successful token requests per grant family | 10/min | 429 temporarily_unavailable for that family (a looping refresh) |
Headers
Every limited response carries both the IETF pair (RateLimit and RateLimit-Policy, draft-ietf-httpapi-ratelimit-headers-07)
and the X-RateLimit-* trio most SDKs already parse. They describe the tightest budget that applied to that request.
RateLimit: limit=60, remaining=59, reset=1
RateLimit-Policy: 60;w=60
X-RateLimit-Limit: 60
X-RateLimit-Remaining: 59
X-RateLimit-Reset: 1790611021 RateLimit: limit=<n>, remaining=<n>, reset=<seconds>resetis seconds until the budget has room for one more request.RateLimit-Policy: <burst>;w=<window seconds>- The policy in force. For most classes burst equals the per-window rate (
60;w=60); the token flood guard advertises its burst,600;w=60, while refilling at 120 a minute. X-RateLimit-Limit,X-RateLimit-Remaining- Same numbers as
limitandremaining. X-RateLimit-Reset- Unix timestamp (seconds) of the same moment
resetdescribes. Retry-After: <seconds>- Only on 429 and 503. Integer seconds, never below 1. Sleep at least this long before retrying.
All of these are exposed to browsers through Access-Control-Expose-Headers, together with Request-Id,
WWW-Authenticate, Idempotent-Replayed, Deprecation and Sunset.
HTTP/1.1 429 Too Many Requests
Retry-After: 7
RateLimit: limit=60, remaining=0, reset=7
RateLimit-Policy: 60;w=60
X-RateLimit-Limit: 60
X-RateLimit-Remaining: 0
X-RateLimit-Reset: 1790611027
{"error":{"type":"rate_limit_error","code":"rate_limited","message":"Too many requests. Retry after 7 seconds.",
"param":null,"doc_url":"https://voixcall.com/developers/errors#rate_limited","request_id":"req_โฆ"}} Spend limits are separate
Request budgets bound traffic. Spend is bounded separately once paid operations ship: a per-user daily cap, per-key and
per-client daily caps, and velocity checks on destinations. Those return insufficient_credits_error or
permission_error, not 429, and are documented with the endpoints that spend.
What to do in a client
- Honour
Retry-After; add jitter if you run several workers. - Cache
/v1/meand the metadata documents; nothing there changes within a session. - Refresh tokens once per expiry, not on every request: more than 10 refreshes a minute on one grant is treated as a loop.
- Do not retry 401s in a loop: after 60 failed checks a minute your address is answered 429 for the rest of the window.