Going live
Everything up to now worked with a test key, where mistakes are cheap. This is what to check before a live key starts spending real credits on real customers.
None of it is exotic. It is the set of things that, in practice, are what breaks first.
Credentials
- Mint a separate live key. Do not promote the test key you have been pasting into terminals.
- Grant only the scopes the integration uses. If it never writes knowledge, do not give it `knowledge:write`.
- Store the secret in a secret manager, not in the repository, not in an env file that gets committed.
- Give staging its own key so revoking one does not take down the other.
Anything shipped to a browser is readable by anyone who opens devtools. Call the API from your server. For in-page chat, use the widget, which is designed to be public and carries no account credentials.
Rotation is built in: POST to the rotate endpoint in the dashboard issues a
new secret and keeps the old one working for 24 hours, so you can deploy
without a flag day. GET /v1/me tells a running process which key it is using.
If a key leaks, revoke it — from the dashboard, or DELETE /v1/api_keys/{id}
for the calling key itself. Revocation takes effect immediately.
Retries and idempotency
The one that costs money if you get it wrong.
- Send an
Idempotency-Keyon every write, and on every chat turn. A timeout is not a failure; it is an unknown. Retrying without a key can double-charge and double-create. - Use a new key for a genuinely new request, and the same key for a retry of the same one.
- Retry on
429,500,502,503and504. Do not retry4xxother than429— a422will fail the same way forever. - Back off exponentially with jitter. A fleet retrying in lockstep is how a brief blip becomes an outage.
A `429` tells you how long to wait. Ignoring it and retrying immediately keeps you rate-limited for longer than waiting would have.
Rate limits
There are separate budgets for general requests, writes, and chat, all per key. That means one runaway backfill cannot starve your live chat traffic — as long as they use different keys.
Backfills, imports and nightly syncs on their own key keeps interactive traffic on its own budget, and makes the request log readable when something goes wrong.
See Rate limits.
Webhooks
If you are polling in production, you are probably doing something the wrong way round.
- Register an endpoint over HTTPS and subscribe to the events you actually handle.
- Verify the signature on every delivery, including the timestamp.
- Deduplicate on the event id — a retry re-sends the same id.
- Return 2xx quickly and do the work asynchronously.
- Watch your delivery failures. Endpoints that fail repeatedly get disabled.
The events worth having before launch:
| Event | Why |
|---|---|
knowledge.dataset.trained / .failed | Know when content is usable, and when it is not |
crawl.completed / .failed | Same, for whole sites |
conversation.handoff.requested | A visitor is waiting for a human |
action.run.failed | Your own systems are failing mid-conversation |
billing.credits.low / .exhausted | Top up before chat stops working |
Monitoring
GET /v1/usage— credits used, credits left, storage. Alert before it reaches zero, not when chat starts refusing.GET /v1/request_logs— your own recent calls with status and error code. Auth failures are recorded too, which is what answers "why did my key stop working".GET /v1/status— unauthenticated health. Safe for an uptime monitor.- Delivery history on each webhook endpoint, for anything you expected and did not receive.
`billing.credits.exhausted` means chat is now refusing every visitor. Wiring `billing.credits.low` first is the difference between a top-up and an outage.
Data and privacy
- Conversations contain whatever your visitors typed. Treat them as personal data.
DELETE /v1/conversations/{id}and the customer anonymize endpoints exist for data-subject requests. Know which you need before someone asks.- Do not put secrets in a chat message or a knowledge source. Anything indexed can come back in an answer.
Versioning
The API is versioned by date, and the version is pinned per key, so a change on our side does not reshape a response under a running integration. Responses carry the version they were built with. Breaking changes ship as a new version with notice; additive fields can appear at any time, so parse defensively and ignore what you do not recognise.
See Versioning.
The short version
- Separate live key, least scopes, in a secret manager.
- Idempotency-Key on every write and every chat turn.
- Exponential backoff with jitter; honour Retry-After.
- A second key for bulk work.
- Webhooks for training, handoff, action failures and credits.
- Alerts on credit balance and delivery failures.
- Signature verification that checks the timestamp and accepts either secret during rotation.