Rate limits
Rate limits protect the platform from a single integration monopolising it, and protect you from a runaway loop burning your quota. They are per API key, so two keys never share a budget.
Why several budgets and not one
A single limit forces everything to compete: a nightly import of 5,000 documents would starve the live chat traffic that customers are waiting on. Separate budgets mean the expensive, bursty work and the latency-sensitive work cannot exhaust each other.
Reading that table:
- Per minute budgets refill continuously. Spending them evenly is better than sending everything in the first second.
- Per second is a burst brake. It stops a fast loop before the per-minute budget notices.
- The write budget applies to mutations plus a couple of expensive reads (semantic search runs a real retrieval pass, so it is metered like a write).
- The chat budget is separate because a chat turn costs a message credit as well as a request.
Backfills and imports on their own key keeps interactive traffic on its own budget. It also makes the request log readable: you can see which workload caused a spike.
What a 429 looks like
{
"message": "Too many requests.",
"error": {
"code": "RATE_LIMITED",
"message": "Too many requests.",
"details": []
},
"retry_after": 12
}The Retry-After header carries the same number of seconds.
Handling it
- Read Retry-After (seconds). It is not a suggestion — retrying sooner keeps you limited for longer.
- Sleep at least that long, plus a little random jitter so a fleet does not retry in lockstep.
- Retry the same request, with the same Idempotency-Key on any mutation.
- If you are hitting 429s steadily rather than in bursts, the fix is fewer requests, not faster retries.
# --retry handles the sleeping for you; --retry-all-errors covers 429.
curl --retry 5 --retry-delay 2 --retry-all-errors \
https://api.agentency.com/v1/chatbots \
-H "Authorization: Bearer <YOUR_KEY>"The official SDKs do all of this already — see SDKs.
Staying under the limit
- Page with a sensible
limit. One request for 100 rows beats a hundred requests for one. See Pagination. - Subscribe instead of polling. A webhook costs you zero requests; a poll loop checking training status every second costs 3,600 an hour and mostly learns nothing. See Webhooks.
- Cache what does not change. Chatbot settings and scope catalogues do not need refetching per operation.
- Batch where the API allows it. File uploads accept
files[].
Immediate retries without backoff turn a momentary limit into a sustained one, because each retry consumes budget that would otherwise have refilled. Always sleep, and always add jitter.
Higher limits
If a legitimate workload does not fit, get in touch through Support with your key id, the endpoints involved, and the throughput you need. It is usually a fast conversation.
What's next
- Errors — which failures are worth retrying.
- Idempotency — how to retry safely.
- Going live — the production checklist.