Rate limits
The API uses token-bucket rate
limiting. Each applicable limit allows a short burst followed by a sustained
request rate. Requests beyond an applicable limit receive
429 Too Many Requests.
How Token Buckets Work
Each limit is a bucket with two numbers:
- Capacity — the maximum burst. The bucket starts full; each admitted operation consumes one token from each applicable operation bucket.
- Refill / sec — the sustained rate at which tokens are replenished.
If at least one token is available, the API accepts the request and consumes a
token. If the bucket is empty, the API rejects the request with a 429. The
Retry-After response header tells you how many seconds to wait. An idle bucket
refills back to full, so you regain your full burst allowance after a quiet
period.
Example: a bucket with capacity 8 and refill 4/s lets you send 8 requests back-to-back, then accepts roughly 4 more per second sustained. If you drain it, waiting one second restores 4 tokens.
Limit Levels
Every admitted operation is checked against each applicable operation bucket. Authentication failures do not consume a token. An idempotency replay returns the stored response without re-executing or consuming a token.
Buckets are normally per command type. At the per-command and per-room levels, command types normally have their own bucket. The exception is the shared seek group: playback seek and Recording mask/unmask controls draw from one 20 / 20 bucket. A burst of one of those commands can drain the budget available to the others; other command types use independent buckets.
The 429 Response
When a hard limit is exceeded, the API returns 429 Too Many Requests. Read the
integer delay from the Retry-After response header. The JSON body keeps the
uniform code and message error shape:
Branch on code (the stable enum, always "rate_limited" for a 429), not
the human-readable message string, which may change. Wait at least the number
of seconds in Retry-After, then retry. Do not look for retry timing in the
JSON body. Include the traceparent response header when contacting support
about a request.
Per-Command Limits
The table below lists each command’s per-call bucket as
capacity / refill-per-second. The three commands in the shared seek group
draw from the same bucket rather than receiving a bucket apiece.
Room-member volume handlers do not use per-command-in-call buckets. Room-member add and delete endpoints use only the per-room limits below.
Per-Room Limits
Room commands are additionally checked against a per-room bucket that aggregates across all members of the room.
Per-App Limits (App-Global Commands)
Creating a room and listing rooms or sessions are limited per application, aggregated across all requests. Each limit is shown as capacity / refill-per-second.
Resource Limits (Not a Rate Limit)
This is a count cap, not a rate cap. Everything above bounds how fast you may call; this bounds how many rooms exist at once. The two are independent: staying under the creation rate limit does not exempt you from the room-count cap, and vice versa.
Each application has a cap on the number of rooms it may have concurrently live at one time. The cap counts every room your app currently owns, including rooms that are still being torn down — a room frees its slot only once it is deleted and its teardown has finalized.
The default room cap is 100 concurrently live rooms. Contact support for
higher limit.
When creating a room would exceed the cap, POST /api/v1/rooms is rejected with
403 Forbidden — not a 429. Because this is not a rate limit, the
response does not include a Retry-After header. The condition clears when you
delete a room and free a slot, not after a fixed delay.
Room member cap
A second count cap bounds how many members may be in a single room at once. It is independent of the room-count cap above and of every rate limit. The default is 250 members per room. Contact support for higher limit.
When a join (POST /api/v1/rooms/{room_id}/members) would exceed the cap, it is
rejected with 409 Conflict and code: room_full — not a 429 and not
the room-count 403. The member is not seated. As with the room-count cap, this
is not a rate limit, so the response does not include a Retry-After header.
The condition clears when a member leaves and frees a slot, not after a fixed
delay.
The cap applies to members added through the join endpoint.
Best Practices
- Honor
Retry-After. On a429, wait at least the number of seconds in theRetry-Afterresponse header before retrying. The JSON body does not contain retry timing. After the burst capacity is exhausted, requests above the refill rate receive429responses. - Wait for webhooks. For asynchronous operations, use the lifecycle webhook to detect completion instead of polling.
- Batch playback files. Send multiple URLs in one
playback/playcommand withtype: files. - Back off on repeated 429s. If a retry receives another
429, honor its newRetry-Afterresponse header and increase your backoff before retrying again.