Skip to content
Knowledge base7 min read

Dataset status and training

What each badge on the Knowledge table means, what the failure messages are telling you, and when to just wait.

Browse topics

What it is

Every source on the Knowledge table carries a badge telling you where it is in the pipeline between "you added it" and "the chatbot can answer from it". Reading that badge correctly is the difference between waiting patiently and chasing a problem that does not exist.

The chatbot only answers from sources that have finished training. A source that is still pending is invisible to it — not broken, just not ready.

The badges, in order

Not trained. The text has been extracted and stored, but no training has been run. You get this when you added the source with Start training turned off, or when a crawl collected pages without training them. Nothing is wrong; the source is simply waiting for you to say go.

Pending. Training is queued. A worker will pick it up shortly. This is the normal state for the first seconds or minutes after you add something.

Reading Document. A step that only appears for images and scanned PDFs: the document is being transcribed into text before it can be trained. The progress bar deliberately has no percentage. See Scanned documents and images.

Training. Actively being processed right now. Larger sources show a percentage as they go.

Trained. Finished, indexed, and answerable. This is the only state in which the chatbot can use the source.

Updating / re-syncing. You edited a source that was already trained. The important detail: the previous version keeps answering while the new one is processed. Editing never leaves the chatbot temporarily ignorant.

Failed. Something went wrong. Open the row to see a plain-language reason.

Removing. Deletion is in progress. The row stays visible until cleanup finishes, so you can see that it is genuinely going rather than wondering whether the click registered.

When you would use it

Read this page when Test Chatbot ignores something you just added, when the banner above the table has been spinning for a while, or when a row is showing a badge you have not seen before.

Where to find it

Open Dashboard → Chatbots → your chatbot → Knowledge.

Steps

  1. Open Knowledge and find the row.
  2. Read the badge and match it to the list above.
  3. Pending or training? Wait. Refreshing does not make it faster. Large PDFs and site crawls legitimately take minutes.
  4. Failed? Open the row and read the reason. Retry once from the row — transient failures are common and usually clear on a second attempt.
  5. Not trained? Use the train action on the row, or the bulk Train now action in the banner if several sources are waiting.
  6. Watch the banner above the table for whole-batch progress.
  7. When the badge reads trained, test with a question that uses the source's own wording.

What you will see

Badges sit on each table row. Above the table, a banner summarises anything currently in flight and clears when the work is done.

Two things that look like status but are not:

  • The usage meter measures size, not readiness. You can be far under your cap and still be training, and you can be at your cap with everything trained.
  • A scheduled sitemap crawl waiting for its next scan is idle, not working. It does not keep the banner running. See Crawl from a sitemap.

What the failure messages mean

The reasons shown on a failed row are deliberately plain. There is no server stack trace and no provider error text — those tell you nothing useful and can leak details that should not be on screen.

"No trainable content was found in this item." Extraction produced nothing. An empty document, an image-only PDF, a scan too poor to read, or a crawled page that was pure navigation. Fix the source, not the setting. For scans, see Scanned documents and images.

"The training service was temporarily unavailable. Try again." A transient outage in the processing service. Retry the row; this usually succeeds immediately.

"Training stalled and was auto-recovered. Try again." A worker died mid-run and the system noticed and cleaned up after itself. Retry.

"Removing this item failed. Try again." Cleanup did not complete. Retry the removal — see Remove knowledge.

"Training failed due to an unexpected error." Rare. Retry once; if it recurs on the same source and not on others, that document is the problem — try replacing it or re-creating the content as text.

A failure that mentions your plan limit is a different thing entirely: the source was fine, but it did not fit. See Knowledge limits.

Why training is asynchronous

It would be simpler if uploading returned a finished, trained source. It does not, and it should not.

Extracting and indexing a large document takes far longer than a browser request may reasonably last. Doing it in the background means an upload never times out, a slow document never blocks the rest of your work, and a failure can be retried automatically rather than making you start over.

The trade-off is the wait, and the confusion of a source that exists but cannot yet answer. That is what the badges are for.

Self-healing

Sources can get stranded — a worker is restarted, a dispatch is lost, a crawl finishes but one page never gets queued. Agentency checks for this on a schedule and re-runs anything that is genuinely stuck.

Practically: a row that has been pending for far longer than its size justifies will usually resolve itself without you doing anything. The recovery passes are deliberately patient (measured in tens of minutes) so they never re-run a job that is merely waiting its turn behind something big.

This is why "wait a bit longer" is a genuinely good first response to a stuck row, and why deleting and re-adding in a loop is the worst one — each cycle puts you at the back of the queue again.

Limits and plan notes

  • Training speed is not a plan feature. A bigger plan raises your capacity, not your queue position.
  • Message credits are not consumed by training. Credits are spent when the chatbot answers a visitor, not when you add knowledge.
  • The usual real blockers are the plan knowledge cap and whether the document contains extractable text at all.

Common problems

Stuck on training for a long time.

Large files and whole-site crawls take minutes to tens of minutes. Leave it. If it ends in failed, read the reason and retry once.

The badge says trained but answers miss the source.

Ask using the source's own words. Then confirm the table is not filtered to a different collection, and check the row's size — a source showing a tiny size did not really extract. If the whole chatbot has never completed a first training run, see The hosted page says the chatbot is not ready.

Everything says Not trained.

You added sources with Start training off. Use the bulk train action in the banner, or train each row.

The row failed, I retried, and it failed again the same way.

The document is the problem, not the platform. Replace the file, or re-create the content as text knowledge.

I edited a source and I am worried the chatbot is offline meanwhile.

It is not. The previous version keeps answering until the update finishes.

Common questions

What does the Reading Document badge mean?

A scanned or image-based document is being transcribed into text before it can be trained. The progress bar has no percentage because there is no honest halfway figure.

Does editing a trained source take the chatbot offline?

No. The previous version keeps answering while the updated one is processed, so there is never a window where the chatbot has forgotten the topic.

A row has been pending for ages. Should I delete and re-add it?

No — that puts you at the back of the queue again. Stranded rows are picked up automatically by a recovery pass, so waiting is the better move.

Why do all my rows say Not trained?

They were added with Start training switched off. Use the train action on each row, or the bulk train action in the banner above the table.

Was this article helpful?

Ready to try it on your own content?

Create a free workspace, add a document, and ask the questions your team is tired of answering.

Dataset status and training | Agentency Help