Crawl from a sitemap
The only crawl mode that rescans on a schedule — hourly, daily, or weekly — and skips pages that have not changed.
Browse topics
What it is
A sitemap crawl reads your site's sitemap.xml — the machine-readable index most content systems publish automatically — and fetches every URL listed in it.
The reason to choose it over an ordinary crawl is not the reading of the file. It is that a sitemap crawl can rescan on a schedule. Set it to daily and your chatbot's knowledge follows your website as it changes, without anyone remembering to press a button.
An ordinary website crawl is a one-off. It captures the site as it was that afternoon, and it stays that way until you run it again by hand. If you only take one thing from this page: that is the difference.
When you would use it
- Your site publishes a sitemap and you want the chatbot to stay current as you publish.
- Your site is large and link-following would miss pages that are only reachable through search or filters.
- Your content changes on a predictable rhythm — a blog, a docs site, a product catalogue.
- A full crawl produced a strange subset and you want an explicit list of what to fetch.
Use an ordinary website crawl instead when there is no sitemap, when you only want a specific corner of a site, or when the content is genuinely static.
Where to find it
Open Dashboard → Chatbots → your chatbot → Knowledge → Add → Website, then choose Sitemap under "What should we crawl?".
Before you start
- Find your sitemap URL. It is usually
https://yoursite.com/sitemap.xml. WordPress, Shopify, and most site builders publish one by default. Open it in a browser to confirm it loads. - A sitemap index — a sitemap that lists other sitemaps — is fine; child sitemaps are expanded automatically.
- Check your remaining knowledge capacity. A large sitemap can consume it quickly, and a scheduled rescan will keep pulling as your site grows.
Steps
- Open Dashboard → Chatbots → your chatbot → Knowledge → Add → Website.
- Under "What should we crawl?" choose Sitemap.
- Paste the sitemap URL. The field expects the sitemap itself, not your homepage.
- Open Advanced options.
- Set Max total URLs if you want a ceiling. Leave it empty for no limit.
- Set Rescan frequency — the setting that makes this mode worth choosing.
- Submit. Watch the Knowledge table and the crawl history for progress.
Rescan frequency — the four options
| Option | What happens |
|---|---|
| Manual | Runs once. Never rescans. You start each later run yourself. |
| Hourly | Rescans about an hour after each run finishes. |
| Daily | Rescans about a day after each run finishes. This is the default. |
| Weekly | Rescans about a week after each run finishes. |
Two details about how the timing actually works:
The clock starts when a run finishes, not when it started. A daily crawl that takes two hours next runs roughly twenty-six hours after you began it, not twenty-four. Over time a slow crawl drifts later in the day. This is intentional — it guarantees a gap between runs rather than allowing them to pile up.
"Due" is not "instant". A background check looks for due crawls every few minutes, so a run may begin a little after its scheduled moment. That is normal.
Which to choose: Daily suits almost everyone. Hourly is worth it only for a genuinely fast-moving site — news, live stock — and it is a lot of fetching on your web server for very little gain otherwise. Weekly suits a stable marketing site. Manual suits a one-off import where you do not want anything happening behind your back.
You are not locked in. Manual now and daily later is a perfectly good plan once you have seen what one run produces.
Unchanged pages are skipped
On every rescan, each page is fetched and compared against what was collected last time. If the text is identical, nothing is written: no new source, no duplicate row, and no wasted knowledge capacity.
A rescan therefore costs you almost nothing in storage when your site has not changed. Only genuinely changed pages produce an update, and genuinely new URLs produce new sources.
The crawl history reflects this with separate counters for pages crawled, pages updated, pages unchanged, pages skipped, and pages failed. A healthy daily rescan on a quiet site shows a large "unchanged" number and very little else — that is the system working, not a broken run.
Max total URLs
Leave this empty and every URL in the sitemap is crawled. Set a number and the crawl stops at that many.
The form accepts up to 100,000. Do not treat that as a target: it is a safety ceiling, not a recommendation. Your plan's knowledge cap will stop you long before 100,000 pages, and a crawl that halts halfway because it ran out of knowledge space is a worse outcome than one you scoped deliberately.
Sensible practice for a first run is a few hundred, then look at what came back before opening it up.
What you will see
Sitemap is one of three modes next to Single page and Full website. The rescan frequency and total-URL cap only appear when Sitemap is selected, inside Advanced options.
The crawl history lists sitemap runs separately from website runs, with their own counters. Open the history drawer from the Knowledge tab. See Crawl status and history.
Between scans, a scheduled sitemap crawl sits in a scheduled state, not a running one. It does not keep the training banner spinning and it does not count as work in progress. A crawl showing "scheduled" is healthy and waiting, not stuck.
Limits and plan notes
- Collected pages count toward the chatbot's knowledge cap exactly like uploads: Free 400 KB through Agency 60 MB. A scheduled rescan can therefore fill your cap over time as your site grows — check the meter occasionally. See Knowledge limits.
robots.txtrules apply to sitemap crawls just as they do to ordinary crawls. See Robots.txt and blocked pages.- A sitemap index is expanded, with a bound on how many child sitemaps are followed and how many URLs are collected in total. A hostile or accidentally enormous index cannot fan out without limit.
- Very large sitemap documents are handled by the background crawler rather than checked in full when you submit the form, so a big sitemap may be accepted for crawling without an instant preview of its contents.
- Only pages Agentency can fetch anonymously are collected. Anything behind a login is out of reach whatever your sitemap says.
Common problems
The sitemap URL will not validate.
It must be a reachable sitemap file — typically sitemap.xml — or an index of sitemaps. Your homepage is not a sitemap. Open the URL in a signed-out browser window and confirm you see XML.
Only some URLs imported.
Your total-URL cap, robots.txt, or the knowledge budget stopped the rest. Check the history counters to see which.
Note that sitemap mode has no include or exclude pattern fields — those belong to the full website crawl. In sitemap mode the sitemap is the list, so you narrow a crawl by pointing it at a narrower document. Most sites publish several child sitemaps (post-sitemap.xml, page-sitemap.xml, product-sitemap.xml); naming the one you want instead of the top-level index is usually all it takes, and it also keeps video and image sitemaps out of the run.
Nothing has rescanned.
Check the frequency is not Manual — that is the one setting that never reschedules. Then confirm the crawl is still active in the history; a session that was deleted or deactivated does not come back.
A scheduled crawl looks stuck at "scheduled".
That is the resting state between scans. It is waiting, not working. Do not delete it unless you want the schedule gone.
My rescans keep adding duplicates.
They should not. Unchanged pages are skipped by content comparison, and cosmetic address variants are collapsed before that comparison even happens — a trailing slash, a #section anchor, and the usual advertising tracking parameters are all ignored, so those are not the cause.
Real duplicates mean the same content is genuinely published at two different addresses: a printer-friendly copy, a page that also exists under a category path, or a second sitemap listing the same articles. The fix is on your side — list each page once in your sitemap — then remove the duplicate sources from the Knowledge table.
My knowledge cap filled up over a few weeks.
A scheduled crawl on a growing site will do that. Set a total-URL cap, exclude the sections that do not answer questions, and prune the rows you do not need. See Remove knowledge.
Common questions
What are the rescan options and which is the default?
Manual, Hourly, Daily, or Weekly. Daily is the default. Manual runs once and never reschedules, which is the setting people forget they chose.
When exactly does the next scan happen?
The interval is measured from when the previous run finished, not when it started, so a slow crawl drifts later. A background check picks up due scans every few minutes.
Will rescanning duplicate all my pages?
No. Each page is compared with what was collected before. Identical pages are recorded as unchanged and nothing is written, so a rescan costs almost no capacity.
How many URLs can one sitemap crawl collect?
Leave Max total URLs empty for no limit, or set a ceiling up to 100,000. Treat that as a safety ceiling — your plan's knowledge cap will stop you far sooner.
My sitemap crawl says pending. Is it stuck?
Probably not. A scheduled crawl rests in a pending state between scans and does not keep the training banner spinning. That is the healthy waiting state.
Was this article helpful?
Related articles
Crawl a website
Fetch public pages and turn each one into knowledge. Start with a single page before you commit to a whole site.
Recrawl or stop a crawl
Cancel a run in flight, switch off a sitemap schedule, or refresh content — without ending up with two of every page.
Crawl status and history
Every run is recorded with counters and per-page results. This is where you find out why a crawl stopped.
Knowledge limits
Your cap counts extracted text, not file size. Free 400 KB through Agency 60 MB, per chatbot.
Ready to try it on your own content?
Create a free workspace, add a document, and ask the questions your team is tired of answering.