Train a bot
Training is how a chatbot stops being a generic model and starts answering from your content. You attach sources, the platform extracts and indexes the text, and from then on answers are grounded in what you gave it.
This page covers every source type, what happens after you submit one, and how to tell the difference between "still working" and "broken".
What a knowledge source is
Each source is a dataset: a name, some text, and the chatbot it belongs to. Where the text came from — pasted, uploaded, or crawled — only matters at creation. After extraction they are all the same thing.
Datasets can be grouped into collections, which is useful once you have more than a handful. A collection is organisational; it does not change how retrieval works.
POST/v1/knowledge
Two fields decide the type:
| Field | What it means |
|---|---|
type | The content format: Text, PDF, Word, Excel, URL, Video, Audio, Image |
source_type | Where it comes from: text, upload, url, integration, api |
Raw text
The simplest source, and the best one for content you generate yourself — product data, FAQs assembled from a database, policy text.
curl -X POST https://api.agentency.com/v1/knowledge \
-H "Authorization: Bearer <YOUR_KEY>" \
-H "Content-Type: application/json" \
-d '{
"chatbot_id": 1,
"name": "Shipping FAQ",
"type": "Text",
"source_type": "text",
"content": "We ship within 2 business days. Express delivery arrives next day.",
"auto_train": true
}'There is a per-request size limit on inline text. For anything long, upload it as a file instead — you get the same result without fighting a payload cap.
Files
Upload first, then reference the returned id. Splitting it in two means a failed upload never leaves you with a half-created dataset, and one upload can back several datasets.
- POST /v1/files with the file as multipart form data.
- Take `id` from the response.
- POST /v1/knowledge with `source_type: "upload"` and `file_id`.
# 1. Upload
curl -X POST https://api.agentency.com/v1/files \
-H "Authorization: Bearer <YOUR_KEY>" \
-F "file=@handbook.pdf"
# 2. Attach
curl -X POST https://api.agentency.com/v1/knowledge \
-H "Authorization: Bearer <YOUR_KEY>" \
-H "Content-Type: application/json" \
-d '{
"chatbot_id": 1,
"name": "Handbook",
"type": "PDF",
"source_type": "upload",
"file_id": 1,
"auto_train": true
}'To import a folder, send them together and let one Idempotency-Key cover the
batch. The response is a list rather than a single object:
curl -X POST https://api.agentency.com/v1/files \
-H "Authorization: Bearer <YOUR_KEY>" \
-H "Idempotency-Key: import-2026-01-14" \
-F "files[]=@handbook.pdf" \
-F "files[]=@pricing.xlsx"Scanned PDFs and images go through OCR automatically. That takes longer than text extraction, which is worth knowing before you decide your polling interval.
A single web page
Give it a URL and the page is fetched, cleaned of navigation and boilerplate, and indexed.
curl -X POST https://api.agentency.com/v1/knowledge \
-H "Authorization: Bearer <YOUR_KEY>" \
-H "Content-Type: application/json" \
-d '{
"chatbot_id": 1,
"name": "Pricing page",
"type": "URL",
"source_type": "url",
"source_url": "https://example.com/pricing",
"auto_train": true
}'Use POST /v1/knowledge/{id}/recrawl to refetch it later — that is the right
call when a page changes, rather than creating a second dataset for the same
URL.
A whole site
For more than a page or two, use a crawler. Two kinds:
- Website crawler — follows links from a starting URL, within limits you set.
- Sitemap crawler — reads
sitemap.xmland fetches what it lists. More predictable, and the better choice when a sitemap exists.
POST/v1/sitemap_crawlers
Check the sitemap before committing to it. This costs nothing and mutates nothing, so it only needs read access:
curl -X POST https://api.agentency.com/v1/sitemap_crawlers/validate_url \
-H "Authorization: Bearer <YOUR_KEY>" \
-H "Content-Type: application/json" \
-d '{"sitemap_url": "https://example.com/sitemap.xml"}'{
"object": "sitemap_validation",
"valid": true,
"type": "urlset",
"url_count": 143,
"has_lastmod": true,
"errors": [],
"warnings": []
}Then create the crawler and start it with
POST /v1/sitemap_crawlers/{id}/scan. Crawls run in the background and can
take a while on a large site — subscribe to crawl.completed and
crawl.failed rather than polling for minutes.
URLs are validated before they are fetched, and internal or private addresses are refused. Crawl your own content, or content you have permission to use.
Cloud imports
POST /v1/knowledge/import/cloud pulls documents straight from a connected
provider. Your side supplies short-lived references; the platform downloads
each file and feeds it through the same pipeline as an upload. One
Idempotency-Key covers the whole batch.
Knowing when it is ready
Every dataset carries embedding_status:
| Status | Meaning |
|---|---|
not_started | Created, no training requested yet |
pending | Queued |
processing | Extracting and indexing |
completed | Usable in answers |
failed | Something went wrong; see processing_error |
Two endpoints read it:
GET/v1/knowledge/{id}/status
GET/v1/chatbots/{id}/readiness
The first is about one source. The second is about the bot: can it answer at all yet, and how much trained material does it have.
`knowledge.dataset.trained` and `knowledge.dataset.failed` fire the moment a source reaches a terminal state, from every path — upload, crawl, re-embed and retry alike. A poll loop that never sees `failed` will wait forever.
Changing and removing content
- Edit —
PATCH /v1/knowledge/{id}to change the name or metadata. - Replace the file —
POST /v1/knowledge/{id}/replace_filekeeps the dataset and swaps its contents, so anything referencing it stays valid. - Retry —
POST /v1/knowledge/{id}/retryafter a transient failure. - Preview a removal —
POST /v1/knowledge/removal_impacttells you what a delete would affect. Read-only, so a read-scoped key can call it. - Remove —
POST /v1/knowledge/remove_from_chatbotdetaches from one bot;POST /v1/knowledge/remove_from_systemdeletes it everywhere.
Removal is asynchronous too: the rows go quickly, the index catches up.
Common mistakes
- Chatting before training finishes. The call succeeds and the answer is
ungrounded —
sourcescomes back empty. Check readiness first. - Re-uploading instead of recrawling. You end up with two datasets saying slightly different things, and retrieval has to pick.
- Polling a failed dataset forever.
failedis terminal. Readprocessing_error, fix the cause, then retry. - Ignoring the quota. Storage is metered per plan.
GET /v1/usageshows where you are before an ingest starts failing.