Chat
Two endpoints: one returns the whole answer, one streams it as it is generated. Both run the same turn, cost the same, and produce the same conversation record.
POST /v1/chat
POST /v1/chat/stream
Both need chat:send.
Blocking or streaming
Use /v1/chat for anything that is not a live UI — a background job, a
Slack bot, an evaluation harness. One request, one JSON answer, no plumbing.
Use /v1/chat/stream when a person is watching a cursor blink. Tokens
arrive as they are produced, which makes a four-second answer feel immediate.
Send an `Idempotency-Key` on both endpoints. Without one, a timeout followed by a retry is two billed turns and two conversation entries.
A single turn
curl -X POST https://api.agentency.com/v1/chat \
-H "Authorization: Bearer <YOUR_KEY>" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: turn-8f21c0" \
-d '{
"chatbot_id": 1,
"message": "Do you ship to Ireland?"
}'{
"object": "message",
"id": 42,
"conversation_id": 7,
"session_id": "api_9f2c7ab1…",
"response": "Yes — Ireland is covered by our standard EU shipping.",
"status": "answered",
"confidence_score": 88.1,
"sources": [{ "title": "Shipping FAQ", "dataset_id": 3 }],
"actions": [],
"usage": { "input_tokens": 41, "output_tokens": 88, "total_tokens": 129 }
}Worth knowing about each field:
session_id— send it back to continue this conversation. Omit it and the next message starts a new one.sources— what the answer was grounded in. Empty means nothing relevant was retrieved, which is your signal that the answer is not from your content.confidence_score— how sure retrieval was. Low with non-empty sources usually means the content is nearly but not quite on topic.actions— anything the bot triggered this turn, such as a handoff request or a CTA. See Actions.usage— tokens for this turn, so you can attribute spend per customer.
Keeping a conversation going
Pass the session_id you got back. History is kept server-side; you do not
need to resend previous messages.
curl -X POST https://api.agentency.com/v1/chat \
-H "Authorization: Bearer <YOUR_KEY>" \
-H "Content-Type: application/json" \
-d '{
"chatbot_id": 1,
"session_id": "api_9f2c7ab1…",
"message": "How much is it?"
}'"How much is it?" only makes sense with the previous turn in view — that is what the session buys you.
You can only append to sessions the API started (`api_…`). Conversations that began in the widget or on a messaging channel are readable through `GET /v1/conversations`, but not writable from here: a visitor conversation belongs to that channel.
Streaming
POST /v1/chat/stream answers with text/event-stream. Frames arrive in this
order:
- `data:` frames carrying answer fragments as they are generated.
- An optional `event: status` frame while longer work happens.
- One `event: metadata` frame with query_id, session_id, sources, actions and usage.
- `data: [DONE]`.
curl -N -X POST https://api.agentency.com/v1/chat/stream \
-H "Authorization: Bearer <YOUR_KEY>" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: stream-8f21c0" \
-d '{"chatbot_id": 1, "message": "Do you ship to Ireland?"}'The trailing metadata frame is the one to keep — it carries the same
query_id, session_id, sources and usage the blocking endpoint returns
in its body.
Retrying a stream
An SSE body cannot be replayed frame by frame, so a retry with the same
Idempotency-Key answers with a JSON summary of the original turn and
Idempotency-Replayed: true instead of a second stream. You still recover
query_id and session_id, and you are not charged twice.
If persistence fails
If the answer was generated but could not be recorded, the stream ends with a
metadata frame carrying error: "turn_persistence_failed". The text you
received is still valid; there is no query_id to attach feedback to. Treat it
as a warning, not a failed request.
Reading it back
Everything a chat turn produces is readable afterwards:
GET /v1/conversations/{id}/messages— the whole thread.GET /v1/messages/{id}— one message, including itssources.
See Conversations.
Common mistakes
- No
Idempotency-Key. The single most expensive omission on the API. - Resending history yourself. The session already has it; extra context costs tokens and can confuse the answer.
- Ignoring empty
sources. That is the difference between "answered from your content" and "answered anyway". - Buffering the stream. A proxy that buffers turns streaming back into a slow blocking call. Disable buffering for this path.
What's next
- Conversations — read, review and export threads.
- Actions — let a turn call your systems.
- Rate limits — chat has its own budget.