Skip to content
Website crawling5 min read

Crawl limits

Depth, pages, delay, and parallelism are settings you choose. Your knowledge cap is separate and can stop a crawl first.

Browse topics

What it is

Two completely different sets of numbers govern a crawl, and confusing them is the source of nearly every "why did it stop?" question.

Crawl limits are the settings on the form: how deep to follow links, how many pages, how fast, how many sitemap URLs. You choose these.

The knowledge cap is how much extracted text your chatbot may hold in total. Your plan sets this.

They are independent. A crawl can stop with 90 of its 100 pages unused because the knowledge cap filled, and it can stop with plenty of knowledge space left because it reached its page limit. Reading the right meter tells you which happened.

When you would use it

  • A crawl stopped early and you want to know which limit ended it.
  • You are planning a large crawl and want to scope it sensibly.
  • A field snapped your number back down and you want to know why.

Where to find it

The crawl settings are at Dashboard → Chatbots → your chatbot → Knowledge → Add → Website → Advanced options.

The knowledge meter is on the Knowledge tab and repeated inside the Add drawer.

The defaults and the ceilings

SettingDefaultMaximum
Max depth510
Max pages100100,000
Delay between requests1 second60 seconds
Pages in parallel15
Max total URLs (sitemap only)no limit100,000
Respect robots.txton—
Follow external linksoff—

The maximums are safety ceilings, not targets. Nothing good happens when you set 100,000 pages on a first run: you will exhaust your knowledge cap in the first few hundred, put a lot of load on the web server, and end up with a partial crawl full of pages you did not want.

Fields clamp themselves. Type 50 into depth and it becomes 10 — that is the form correcting you before the server rejects the request, not a bug.

The knowledge caps

PlanKnowledge per chatbot
Free400 KB
Starter10 MB
Standard20 MB
Pro40 MB
Agency60 MB

A crawl halts when this budget is exhausted, even if pages remain. That is the designed behaviour: a partial, trained crawl is more useful than a failed one. See Knowledge limits.

Some arithmetic to make it real. A substantial article extracts to perhaps 5-15 KB of text. On Free's 400 KB, that is roughly 30 to 80 pages before you are full — regardless of the 100-page default sitting in the form. On Starter's 10 MB, hundreds to low thousands. Set your page limit with your plan in mind, not with the form's ceiling in mind.

Steps

  1. Before you start, open Advanced options and read the current values.
  2. Check the knowledge meter and estimate how many pages will fit.
  3. Set Max pages to something realistic for your remaining capacity.
  4. Add exclude patterns for the sections that never answer questions: /tag/*, /author/*, /page/*, /search*.
  5. Or better, add an include pattern for the one section you want: https://example.com/docs/*.
  6. Start the crawl and watch both the page counter in the history and the knowledge meter.
  7. If it stops, compare the two. Meter full means the knowledge cap ended it; pages exhausted means your page limit did.

Patterns beat big numbers

The instinct when a crawl misses content is to raise the page limit. It is usually the wrong move.

A site-wide crawl at 1,000 pages will spend most of your knowledge budget on pagination, tag archives, author pages, and category listings — pages that contain no answers and actively dilute the good content. The same budget spent on 150 documentation pages produces a far better chatbot.

Include patterns are the highest-value setting on the form and almost nobody uses them.

Limits and plan notes

  • Rate limits apply to how many crawls one workspace may run at once. A busy workspace queues rather than running everything simultaneously.
  • Scheduled sitemap rescans across all customers are bounded too, so a due scan may start a little after its time. See Crawl from a sitemap.
  • There is no JavaScript rendering. This is not a tunable limit — raising page counts cannot extract text that is not in the HTML.
  • Pages that fail, are blocked by robots.txt, or return too little text are recorded but do not become knowledge. See Robots.txt and blocked pages.
  • Extremely large individual pages are bounded so a single enormous document cannot consume everything.
  • Sitemap crawls skip unchanged pages on a rescan, so a repeat scan costs almost no additional capacity.

Common problems

I set 10,000 pages and got a handful.

Robots rules, blocked hosts, or the knowledge cap. Check the history counters and the meter — between them they will tell you which.

The form snapped my number back down.

You exceeded a ceiling. Depth stops at 10, parallelism at 5, delay at 60 seconds.

The crawl says complete but half the site is missing.

Depth may have been too shallow for a deep tree, the missing pages may be unreachable by links (which is exactly what sitemap mode solves), or the budget ran out. Check whether the meter is full first.

Everything is under the limits and it still stopped.

Look for failures in the history. A site that starts rate-limiting or blocking mid-crawl will produce a run that ends early with a cluster of failures rather than a clean finish.

Can I raise the limits beyond the maximums?

No. Those ceilings are fixed for everyone. If you genuinely need more content than they allow, split it across several chatbots by subject.

How do I make a crawl finish faster?

Raise Pages in parallel (up to 5) and lower the delay — but only on a site you own. On someone else's server, being fast is being rude, and a blocked crawler is slower than a polite one.

Common questions

What are the defaults and the maximums?

Defaults are depth 5, 100 pages, 1 second delay, 1 page at a time. Maximums are depth 10, 100,000 pages, 60 seconds, and 5 in parallel.

How many pages actually fit in my plan?

A substantial article extracts to roughly 5-15 KB, so Free's 400 KB is about 30-80 pages. Set your page limit against your plan, not against the form's ceiling.

Why did the form change the number I typed?

The field clamps to the allowed ceiling. That is the form correcting you before the server rejects the request, not a bug.

Should I just raise the page limit when a crawl misses content?

Usually not. A bigger crawl spends your budget on tag and pagination pages. An include pattern for the section you want gives a far better chatbot.

Can I make a crawl finish faster?

Raise Pages in parallel to 5 and lower the delay — but only on a site you own. On someone else's server, a crawler that gets blocked is slower than a polite one.

Was this article helpful?

Ready to try it on your own content?

Create a free workspace, add a document, and ask the questions your team is tired of answering.

Crawl limits | Agentency Help