Skip to content
Website crawling5 min read

Recrawl or stop a crawl

Cancel a run in flight, switch off a sitemap schedule, or refresh content — without ending up with two of every page.

Browse topics

What it is

Two related jobs:

Recrawling means getting fresh content from a site you have already crawled. How you do that depends on the mode — a website crawl has to be started again by hand, while a sitemap crawl can do it on a schedule.

Stopping means ending something in progress: cancelling an in-flight website crawl, or switching off a sitemap schedule so it stops rescanning.

When you would use it

  • Your website was redesigned or rewritten and the chatbot is quoting the old copy.
  • You started a crawl on the wrong URL and want it stopped before it fills your knowledge cap.
  • A run appears stuck.
  • You no longer want a sitemap crawl rescanning every day.

Where to find it

Open Dashboard → Chatbots → your chatbot → Knowledge, then the crawl history drawer (?drawer=crawlers-history).

New crawls start from Knowledge → Add → Website.

Steps — stopping something in progress

  1. Open Knowledge → crawl history.
  2. Find the run. Running means it is fetching right now.
  3. Use delete/cancel on that row and confirm.
  4. For a website crawl, fetching stops. For a sitemap crawl, the session is deactivated and will not rescan again.
  5. Pages already collected stay. Remove them from the Knowledge table if you do not want them — see Remove knowledge.

Steps — refreshing content

  1. Decide whether the site changed enough to be worth it. If nothing has changed, a website recrawl just duplicates what you already have.
  2. For a website crawl: start a new one from Add → Website with the same URL and settings. Then remove the old sources, or you will have two copies of every page.
  3. For a sitemap crawl: if the frequency is anything other than Manual, do nothing — the next scan is already scheduled, and unchanged pages will be skipped automatically. See Crawl from a sitemap.
  4. For a handful of pages: do not recrawl the whole site. Crawl the changed pages individually in Single page mode, and remove the stale sources.
  5. Wait for the new pages to train before treating Test Chatbot as up to date.

The duplication trap

This is the single most common mistake on this page, so it is worth being blunt about it.

Recrawling a website does not replace the old pages. It adds new ones.

Run the same website crawl twice and you have two sources for every page. The chatbot now has two versions of your refund policy, and no way to know which is current — so it may answer from either, including the one you were trying to replace.

Two ways to avoid it:

  • Remove first, crawl second. Delete the old crawl's sources, then run the fresh crawl into the clean table.
  • Use sitemap mode. It compares content and skips pages that have not changed, updating the ones that have. It is the only mode that handles repeat runs properly, which is the main reason to prefer it for anything you intend to refresh.

Turning a schedule off without losing the knowledge

Deactivating a sitemap session from the history stops future scans. The pages it has already collected remain as knowledge and the chatbot keeps answering from them.

That is usually what you want: freeze the content as it is today, stop the fetching. If you want the content gone too, remove the sources separately.

What you will see

History rows expose a delete/cancel action and, for website runs, a list of the individual pages. Sitemap rows show their counters and their scan state.

A confirmation dialog appears before anything is cancelled, because there is no undo — a cancelled crawl cannot be resumed from where it stopped. You start a new one.

Rescan frequency is chosen when the sitemap crawl is created — Manual, Hourly, Daily, or Weekly.

Limits and plan notes

  • Cancelling does not roll back pages that already trained. Knowledge you no longer want has to be removed explicitly.
  • Deleting a crawl session and removing knowledge are two separate operations in two separate places, by design: you may well want to stop a schedule while keeping what it gathered.
  • Recrawling consumes knowledge capacity again unless you remove the old sources first. On Free's 400 KB, a careless second crawl can fill the cap on its own.
  • A crawl abandoned by a crashed worker is detected and cleaned up automatically — a website crawl is failed, a sitemap crawl with work left is resumed. Cancelling by hand is for crawls you no longer want, not for crawls you think are broken.
  • A scheduled sitemap sitting between scans is idle, not stuck. Do not delete it just because it says pending.

Common problems

I started two crawls and the table doubled.

Each successful page became a source, twice. Stop the extra run and remove the duplicates. Then switch to sitemap mode if this is content you refresh regularly.

Delete in history did not remove the Knowledge rows.

Correct. History deletes the crawl session. Trained pages are knowledge and are removed from the Knowledge table.

I cancelled but pages kept appearing for a minute.

Work already in flight finishes. It settles quickly.

My scheduled sitemap keeps re-adding pages I deleted.

That is what a schedule does — it puts back anything still listed in the sitemap. Sitemap mode has no exclude-pattern field, so the list itself is the only lever: remove those URLs from your sitemap, point the crawl at a narrower child sitemap, or switch the frequency to Manual so nothing is re-added behind your back.

The chatbot still quotes the old page after a recrawl.

The old source is still there. Recrawling adds; it does not replace. Find and remove the stale row.

Can I resume a cancelled crawl?

No. Start a new one. Narrow it with include patterns so you are not re-fetching what you already have.

Common questions

Does recrawling replace the old pages?

No — it adds new ones. Run the same website crawl twice and the chatbot has two versions of every page. Remove the old sources first, or use sitemap mode.

Can I resume a cancelled crawl?

No. Start a new one, and narrow it with include patterns so you are not re-fetching everything you already have.

If I switch off a sitemap schedule, do I lose the pages?

No. Deactivating stops future scans; the collected pages remain knowledge and the chatbot keeps answering from them.

My schedule keeps re-adding pages I deleted.

That is what a schedule does — it restores anything still listed in the sitemap. Sitemap mode has no exclude field, so remove those URLs, use a narrower child sitemap, or switch the frequency to Manual.

Was this article helpful?

Ready to try it on your own content?

Create a free workspace, add a document, and ask the questions your team is tired of answering.

Recrawl or stop a crawl | Agentency Help