Skip to content
Website crawling8 min read

Crawl status and history

Every run is recorded with counters and per-page results. This is where you find out why a crawl stopped.

Browse topics

What it is

Every crawl you start is recorded. The history keeps the run, its status, its counters, and — for website crawls — the result for each individual URL.

This is where you go to answer "is it still working, did it work, and if not, why not?" It is far more useful than staring at the Knowledge table, because it distinguishes a page that failed from a page that was deliberately skipped.

When you would use it

  • You started a crawl and want to know whether to wait.
  • A crawl finished with fewer pages than expected.
  • You need to cancel a run.
  • You want to check whether a scheduled sitemap rescan is actually running.

Where to find it

Open Dashboard → Chatbots → your chatbot → Knowledge. The active run appears on the queue card. For everything else, open the crawl history drawer (?drawer=crawlers-history).

The statuses

Each run carries a status badge. Seven of them matter:

BadgeMeaning
PreparingBuilding the list of links to fetch. Normal, and can take a moment on a large site
CollectingActively fetching pages right now
ScanningA sitemap run working through its current batch
PendingQueued and waiting for a worker
ScheduledA recurring sitemap resting between scans. Shown with its Next scan time
PausedHalted part-way by you, and resumable
CompletedFinished
FailedEnded with an error. The row shows a plain-language reason

Scheduled is the one worth learning. A recurring sitemap crawl is idle between scans, and the product says so plainly: those runs sit in their own Scheduled rescans section with a Next scan time beside each. A crawl there is waiting, not stuck, and it does not keep the Knowledge training banner spinning.

Completed does not mean "collected everything." A crawl that stopped at its page limit, or at your knowledge cap, still finishes as Completed. To find out how much it actually got, read the counters.

Runs that ended badly are grouped under Needs attention, with an action on each row — Retry to run it again, or Upgrade when the run stopped because your knowledge cap was full.

The counters

Each row carries a running tally, and reading it tells you the story of the run. The labels on screen are:

CounterMeaning
collectedPages fetched and added to your knowledge
up to datePages identical to the last scan, so nothing was rewritten and no capacity was spent
skippedDeliberately not collected — a robots rule, an unsupported file type, or too little text. Not in your knowledge base
failed to fetchThe page could not be downloaded: broken link, server error, timeout, or too large
remainingDiscovered but not yet processed — still queued for a later batch

There is also a links figure in the form 120 / 400 links. It counts every processed page — collected, up to date, skipped, and failed alike — so it is normally higher than the collected count. That is not a discrepancy.

A healthy daily rescan on a stable site shows a large up to date number and very little else. That is success, not inactivity.

The counter to watch is failed to fetch. A handful is normal on any real site. A large number clustered together usually means the site started refusing the crawler part-way through. Note that this is a fetching failure and is different from a source that failed training — those are counted on the Knowledge status bar instead.

Steps

  1. Stay on Knowledge after starting a crawl. The queue card shows the active run.
  2. Open the crawl history drawer to see past runs, website and sitemap alike.
  3. Read the status, the progress, and the counters.
  4. Open the row's pages list to see the per-URL result — which were collected, which were skipped, which failed.
  5. Read the reason on a failure. It is a short plain-language message, not a server dump.
  6. If pages failed to fetch, use Retry failed pages on that row rather than crawling the whole site again.
  7. To stop an in-flight website crawl or switch off a scheduled sitemap, use delete/cancel on that history row.
  8. Do not start several crawls on the same site at once. Wait for the active one.

The pages list

Every run — website and sitemap alike — can open a Crawled pages view listing each URL with its own status: Pending, Crawled, Processed, Updated, Skipped, Up to date, or Failed.

Filter it to answer a specific question rather than scrolling:

  • All pages — everything discovered, with current status.
  • Remaining — discovered but not yet fetched.
  • Failed — the ones to investigate, with Retry failed pages offered above them.
  • Skipped — what was deliberately left out, and why.
  • Up to date — re-checked this scan and unchanged.

Each entry can be opened in a new tab, so you can see for yourself whether a page that was skipped for "too little text" really does have anything worth reading on it. Pages already in your knowledge are marked as such.

A page whose text matched another page already collected carries a Duplicate pill instead of a plain Skipped one — only one copy is kept so the same text is not counted twice. Where the crawl recorded which page it kept, the row also offers View similar pages, which opens the kept page together with every other page folded into it, so you can check that the right copy survived.

Retry failed pages re-queues just the failed URLs. It is available once a website run has finished, and at any time on a sitemap run. This is almost always the right response to a cluster of fetch failures — re-running the whole crawl re-fetches everything that already worked and duplicates it.

What you will see

The history is a paginated drawer titled Content collection history, listing every website and sitemap session with its URL, its progress, and its status badge.

Finished smart crawls also offer View training report — a separate drawer that shows how pages were classified, whether the Site information source was created, and the listing-pages / Site information switches you can still change. See Smart content extraction. History counters answer "did we fetch it?"; the report answers "what did we train?"

The delete action on a history row does two different things depending on the type: it cancels an in-flight website crawl, and it deactivates a sitemap session so it stops rescanning.

What it does not do — and this catches people every time — is remove the knowledge the crawl already produced. The confirmation dialog says so outright: the collected knowledge stays, and only the session and its history are removed. Those pages are ordinary sources on the Knowledge table until you delete them there. See Remove knowledge.

Reading a run that went wrong

Work through it in this order:

  1. Is the status failed, or completed with few pages? Failed means the run itself broke. Completed-but-small means a limit was reached.
  2. Is the knowledge meter full? If yes, that ended it. Free space or upgrade — see Knowledge limits.
  3. Are most URLs skipped? Robots rules or a content filter. See Robots.txt and blocked pages.
  4. Are most URLs failed? The site is blocking or timing out. Try a single page manually first.
  5. Did few pages get found at all? Depth was too shallow, or the site is JavaScript-rendered and there were no links to follow in the HTML.

Limits and plan notes

  • History is per chatbot. A crawl started on a different chatbot does not appear here.
  • Error codes are deliberately short and safe. They do not include the raw server response, which would leak details of the target host.
  • A crawl left running by a crashed worker is detected and recovered automatically — a website crawl is failed cleanly, a sitemap crawl with work remaining is resumed. You may see a run change state on its own; that is the recovery working.
  • A scheduled sitemap between scans does not keep the Knowledge training banner spinning. If the banner is running, something is genuinely being processed.
  • Crawl history does not consume message credits or knowledge capacity.

Common problems

History is empty.

You have not crawled on this chatbot, or you are looking at a different chatbot. Check which one is in the URL.

Status says failed.

Open the error code. The common ones are a blocked or unreachable URL, robots restrictions, and the knowledge cap. Fix the cause, then crawl again — see Recrawl or stop a crawl.

The crawl says completed but I only got 12 pages.

Something bounded it. Check the meter, then the skipped and failed counters, then your depth and page settings.

A sitemap crawl has been sitting there for hours.

Look at the badge. Scheduled with a Next scan time is the healthy resting state between scans — it is waiting, not working, and deleting it would only switch the schedule off. Pending with no scan time means genuinely queued; if it stays that way it is picked up by the automatic recovery pass.

I deleted the crawl and the pages are still there.

Expected. Deleting removes the crawl session; the pages it created are knowledge sources you remove from the Knowledge table.

Two runs of the same site are both listed.

You started it twice. Cancel the second and remove the duplicate sources it created.

Common questions

What is the difference between skipped and failed?

Skipped means Agentency chose not to collect the page — a robots rule or too little text. Failed means the fetch was attempted and errored. Different causes, different fixes.

My rescan reports mostly 'up to date'. Is it broken?

That is success. Those pages are identical to the last scan, so nothing was rewritten and no knowledge capacity was spent. A quiet site should look like that.

My sitemap crawl says Scheduled and nothing is happening. Is it stuck?

No. Scheduled is the resting state between scans, shown with the next scan time. It is waiting, not working — deleting it would only switch the schedule off.

Some pages failed to fetch. Do I have to crawl the whole site again?

No. Use Retry failed pages on that row, which re-queues only the failed URLs. Re-running the whole crawl would re-fetch everything that already worked and duplicate it.

Does deleting a run from history remove the pages it collected?

No. It cancels an in-flight website crawl or deactivates a sitemap schedule. The pages are knowledge sources and stay until you remove them.

Was this article helpful?

Ready to try it on your own content?

Create a free workspace, add a document, and ask the questions your team is tired of answering.