Skip to content
Website crawling7 min read

Crawl a website

Fetch public pages and turn each one into knowledge. Start with a single page before you commit to a whole site.

Browse topics

What it is

A website crawl visits public pages on a site, extracts the readable text from each one, and turns it into knowledge for one chatbot. Each page becomes its own source on the Knowledge table.

You choose how wide to go: a single page, or a full website that follows internal links to a depth and page limit you set.

When you would use it

Use a crawl when your content already lives on a public website and copying each page by hand would be absurd — a help centre, a documentation tree, a product catalogue, a policy section.

Do not use it for:

  • Pages behind a login. Agentency crawls anonymously; it cannot sign in as you.
  • Applications that render everything with JavaScript. There is no browser executing scripts, so a page whose text only appears after scripts run will come back nearly empty.
  • Content you do not have the right to copy.
  • Content that changes constantly, unless you use the sitemap mode, which is the only mode that can rescan on a schedule.

Where to find it

Open Dashboard → Chatbots → your chatbot → Knowledge → Add → Website.

Before you start

Three checks that save a wasted run:

  1. Open the page you intend to crawl in a private browsing window. That is roughly what Agentency sees — no session, no cookies. If it redirects you to a login, crawling will not work.
  2. Check whether the text is really in the page. Right-click and choose "view page source", then search for a sentence you can see on screen. If the sentence is not in the source, the page is rendered by JavaScript and will crawl poorly.
  3. Check your knowledge capacity. A crawl halts when the budget runs out, so start with enough room.

Steps

  1. Open Dashboard → Chatbots → your chatbot → Knowledge → Add → Website.
  2. Paste the starting URL. Use https://.
  3. Under "What should we crawl?", choose:
    • Single page — fetch only this exact URL. Perfect for a first test.
    • Full website — follow internal links up to the limits below.
    • Sitemap — read a sitemap.xml. See Crawl from a sitemap.
  4. Open Advanced options if you want to change depth, page count, delay, parallelism, or include and exclude patterns. Leave Respect robots.txt on unless you own the site and know a rule is too broad.
  5. Keep Start training on and submit.
  6. Stay on the Knowledge tab. Pages appear as sources and then train.
  7. Spot-check with Test Chatbot, using a phrase from a heading on one of the crawled pages.

Start with one page

The most useful habit on this whole page: crawl a single page first.

A one-page crawl finishes in seconds and tells you everything you need to know before you commit to hundreds. Did the text come through? Is it the article text, or is it navigation and cookie banners? Did the chatbot answer from it?

If a single page produces good knowledge, a hundred pages of the same site will too. If it produces 200 bytes of menu labels, no amount of raising the page limit will fix that — and you will have found out in ten seconds instead of twenty minutes.

The advanced options

OptionDefaultWhat it does
Max depth5How many clicks from the start URL to follow
Max pages100The real cap on how much is collected
Delay between requests1 secondHow gentle to be on the site
Pages in parallel1Fetch several at once to finish faster
Follow external linksoffWhether to leave the site
Respect robots.txtonHonour the site's crawler rules
Include URL patternsemptyOnly crawl URLs that match
Exclude URL patternsemptySkip URLs that match

Max pages is the setting that matters. Depth only limits how far link-following reaches; the page count is what actually bounds the crawl.

Include and exclude patterns are worth more than a big page limit. One pattern per line, with * as a wildcard. Crawling https://example.com/docs/* and excluding https://example.com/tag/* gets you the content and skips the archive noise. That spends your knowledge budget on pages that answer questions instead of on pagination.

Leave the delay alone unless you have a reason. One second per request is polite. Setting parallelism high and the delay to zero on someone else's server is a good way to get blocked.

Smart content extraction (on by default) is a separate group of switches in the same Advanced options: it strips repeated menus and footers from each page and keeps that footer contact text once in a Site information source. Listing pages stay off unless you need them. See Smart content extraction.

What you will see

Website is a URL field plus the mode picker. Advanced options are collapsed by default, with the conservative defaults above.

Once you submit, an active crawl card appears on the Knowledge tab and pages arrive as sources over the following minutes. Each is titled from the page's own title, so a well-titled site produces a readable Knowledge table and a badly-titled one produces a lot of rows called "Home".

The crawl history records the run, per-page results, and any failures. See Crawl status and history.

Limits and plan notes

  • Only public pages that Agentency is permitted to fetch. Login walls, private apps, and internal addresses are all out.
  • Collected pages count toward your knowledge cap: Free 400 KB through Agency 60 MB. A crawl stops when the cap is reached, which is deliberate — see Crawl limits.
  • There is no JavaScript rendering. Raising the page limit will not help a single-page application.
  • robots.txt is respected by default. See Robots.txt and blocked pages.
  • A website crawl runs once. Only sitemap mode reschedules itself.
  • Do not start several crawls on the same site at the same time. They compete, and you will end up with duplicate sources.

Common problems

The crawl finished with almost no pages.

Three usual causes: the site is JavaScript-rendered, robots.txt disallows the paths, or the start URL redirected somewhere Agentency will not follow. Test with a single static page to tell them apart.

Add rejects the URL.

Use a public https:// address. Internal hostnames, local addresses, and anything that looks unsafe to fetch are blocked outright.

I got hundreds of useless pages.

The crawl found your tag archives, pagination, and author pages. Remove them, then crawl again with an include pattern for the section you actually want.

Every page failed at the first URL.

The host refused the fetch, the certificate failed, or the address is not publicly reachable. Open it in a private window from a different network to confirm.

The pages came in but the answers are generic.

Look at what was extracted — a page whose source is mostly navigation gives you mostly navigation. Consider pasting the important content as text knowledge instead.

I crawled twice and now everything is duplicated.

Each successful page becomes a source, so a second run of a one-off crawl duplicates them. Remove the duplicates and use sitemap mode if you want repeatable updates — it skips unchanged pages.

The crawl is still running after an hour.

Large sites with a one-second delay genuinely take a long time. Check the history for progress. If it has truly stalled, it is recovered automatically, and you can cancel it from the history drawer — see Recrawl or stop a crawl.

Common questions

Why did my crawl return almost nothing?

Usually the site is rendered by JavaScript, robots.txt disallows those paths, or the start URL redirected. Crawl one static page to tell them apart in seconds.

Can Agentency crawl pages behind a login?

No. It crawls anonymously and cannot sign in as you. Upload the documents or paste the content as text instead.

Which setting actually limits how much is collected?

Max pages. Depth only bounds how far link-following reaches — the page count is the real cap, and your knowledge budget can stop it sooner.

Does a website crawl repeat itself?

No. It runs once and captures the site as it was. Only sitemap mode can rescan on a schedule and skip pages that have not changed.

How do I check whether a page will crawl well before I start?

Open it in a private window, view the page source, and search for a sentence you can see on screen. If it is not in the source, it will not extract.

Was this article helpful?

Ready to try it on your own content?

Create a free workspace, add a document, and ask the questions your team is tired of answering.