VoixCall

VoixCall for developers

Rate limits

Generous for people and assistants, tight for guessing. Every response tells you where you stand.

How it works

One limiter (GCRA, atomic in a Valkey script) fronts /v1, /mcp and /oauth/*. Several budgets are evaluated for every request and the tightest one wins the headers. Authenticated requests are keyed per principal (the user plus the specific connection or API key) and per user; unauthenticated ones per source IP (IPv6 addresses are bucketed per /64). Budgets refill continuously, so โ€œ60/minโ€ means a burst of 60 and then one request a second, not a hard reset on the minute.

When the limiter's store is unreachable, public paths fail closed: 503 rate_limiter_unavailable (or {"error": "temporarily_unavailable"} on /oauth/*) with Retry-After: 5.

Budgets: authenticated paths

ClassApplies toPer principalPer userNotes
ReadsGET /v1/me, every /mcp request, and every future read endpoint60/min120/minLive. Per principal = per (user, connection or key); per user across all of them. An alert-only ceiling of 3,000/min per client id never blocks.
Transcript requestsPOST /v1/calls/{id}/transcript5/min5/minReserved: endpoint not live.
QuotesPOST /v1/quotes, voixcall_quote_call20/min20/minReserved: endpoint not live.
Call placementPOST /v1/calls, voixcall_place_call10/min10/minReserved: endpoint not live. Plus one concurrent callback session.
Message readsGET /v1/messages, voixcall_read_messages60/min120/minReserved: endpoint not live. Plus 10/min per number.
Key management/v1/api-keys (dashboard session)โ€”30/minLive.

Budgets: unauthenticated paths and OAuth

ClassApplies toLimitNotes
Metadata/.well-known/*, /v1/openapi.json300/min per IPLive.
Authorize and registerGET /oauth/authorize, POST /oauth/register30/min per IPLive.
RegisterPOST /oauth/register20/hour per IP, on top of the row aboveLive. Also a global cap of 5,000 unexpired self-registered clients; beyond it registration answers 503 with Retry-After: 3600.
Token flood guardPOST /oauth/token, POST /oauth/revokeburst 600, then 120/min sustained per IPLive. The burst is deliberately large so one platform egress address carrying many users is not the sharp edge; the RateLimit-Policy header advertises the burst (600;w=60).
Failed bearer checks/v1, /mcp60/min per IPLive. Consumed only when a presented credential fails the lookup. Requests with no credential (the discovery probe) and valid tokens never touch it. Once exhausted, the failure is answered 429 rate_limited instead of 401.

OAuth lockouts

Penalties attach to the source or the credential, never to a public client id, so one misbehaving installation cannot lock every user of a shared client (Claude, ChatGPT) out of refreshes. Only guessing signals count: presenting a known-but-revoked token after a disconnect does not.

DimensionLimitEffect
Failed exchanges per source IP10/min60 s back-off for that IP: 429 temporarily_unavailable before any grant work
Failed exchanges per presented credential3 per hourthe code or refresh token and its whole grant family are revoked
Failed exchanges per client id1,000/minalert only, never blocks
Successful token requests per client id600/minalert only, never blocks (a public client id is shared by every installation)
Successful token requests per grant family10/min429 temporarily_unavailable for that family (a looping refresh)

Headers

Every limited response carries both the IETF pair (RateLimit and RateLimit-Policy, draft-ietf-httpapi-ratelimit-headers-07) and the X-RateLimit-* trio most SDKs already parse. They describe the tightest budget that applied to that request.

Response headers on an allowed request (reads budget)
RateLimit: limit=60, remaining=59, reset=1
RateLimit-Policy: 60;w=60
X-RateLimit-Limit: 60
X-RateLimit-Remaining: 59
X-RateLimit-Reset: 1790611021
RateLimit: limit=<n>, remaining=<n>, reset=<seconds>
reset is seconds until the budget has room for one more request.
RateLimit-Policy: <burst>;w=<window seconds>
The policy in force. For most classes burst equals the per-window rate (60;w=60); the token flood guard advertises its burst, 600;w=60, while refilling at 120 a minute.
X-RateLimit-Limit, X-RateLimit-Remaining
Same numbers as limit and remaining.
X-RateLimit-Reset
Unix timestamp (seconds) of the same moment reset describes.
Retry-After: <seconds>
Only on 429 and 503. Integer seconds, never below 1. Sleep at least this long before retrying.

All of these are exposed to browsers through Access-Control-Expose-Headers, together with Request-Id, WWW-Authenticate, Idempotent-Replayed, Deprecation and Sunset.

A denied request
HTTP/1.1 429 Too Many Requests
Retry-After: 7
RateLimit: limit=60, remaining=0, reset=7
RateLimit-Policy: 60;w=60
X-RateLimit-Limit: 60
X-RateLimit-Remaining: 0
X-RateLimit-Reset: 1790611027

{"error":{"type":"rate_limit_error","code":"rate_limited","message":"Too many requests. Retry after 7 seconds.",
          "param":null,"doc_url":"https://voixcall.com/developers/errors#rate_limited","request_id":"req_โ€ฆ"}}

Spend limits are separate

Request budgets bound traffic. Spend is bounded separately once paid operations ship: a per-user daily cap, per-key and per-client daily caps, and velocity checks on destinations. Those return insufficient_credits_error or permission_error, not 429, and are documented with the endpoints that spend.

What to do in a client

  • Honour Retry-After; add jitter if you run several workers.
  • Cache /v1/me and the metadata documents; nothing there changes within a session.
  • Refresh tokens once per expiry, not on every request: more than 10 refreshes a minute on one grant is treated as a loop.
  • Do not retry 401s in a loop: after 60 failed checks a minute your address is answered 429 for the rest of the window.