Robots.txt and blocked pages
Agentency obeys robots.txt by default, and refuses unsafe addresses always. Here is how to tell which stopped you.
Browse topics
What it is
robots.txt is a small file most websites publish at their root — https://example.com/robots.txt — telling automated visitors which paths they may and may not fetch. Agentency reads it and obeys it, because Respect robots.txt is on by default.
Separately, and regardless of that setting, Agentency refuses to fetch addresses it considers unsafe: internal hostnames, private network addresses, and anything that looks like it is trying to reach a machine that is not a public website.
Both mechanisms cause pages to be missed, and the crawl history distinguishes them: a page ruled out by robots is recorded as skipped, while a page that could not be fetched is recorded as failed.
When you would use it
- A crawl skipped exactly the pages you wanted.
- Everything failed at the first URL.
- You are wondering whether to turn Respect robots.txt off.
- You want to crawl one section of a site and nothing else.
Where to find it
Open Dashboard → Chatbots → your chatbot → Knowledge → Add → Website → Advanced options.
Steps
- Before starting, open Advanced options.
- Leave Respect robots.txt on unless you own the site and know a specific rule is too broad.
- Add include patterns if only one tree should be fetched — one pattern per line,
*as a wildcard. - Add exclude patterns for noise you never want:
/tag/*,/page/*,/search*. - Start the crawl.
- Open the crawl history and its pages list. Read the skipped entries separately from the failed ones — they have different causes and different fixes.
- For anything behind a login, stop trying to crawl it and upload the content or paste it as text instead.
Checking robots.txt yourself
Thirty seconds of investigation beats guessing. Open https://yoursite.com/robots.txt in a browser.
A block like this stops a crawl of your help centre:
User-agent: *
Disallow: /help/
And this stops everything:
User-agent: *
Disallow: /
The second one is more common than you would think — staging sites and newly-launched sites are often left that way by accident, and nobody notices until something tries to crawl them.
If you own the site and the rule is wrong, fix robots.txt. That is a better answer than turning the setting off in Agentency, because it fixes the problem for every crawler including search engines.
Should I turn Respect robots.txt off?
Only on a site you own or administer, and only when you know which rule is in the way.
Understand exactly what it buys you. Turning it off means Agentency will fetch paths the site asked crawlers to avoid. It does not:
- get you past a login wall;
- get you past a paywall;
- let Agentency reach internal or private addresses;
- change whether the page has readable text.
So if your pages need a session to view, this switch is not the answer and never will be. And if the site belongs to somebody else, respecting their rules is not optional politeness — you are responsible for what you crawl.
Skipped versus failed
Getting this distinction right saves a lot of wasted effort.
Skipped means Agentency chose not to collect the page: a robots rule, a file type that is not a web page, or a page with too little text to be worth storing. Fix by changing the rules or the patterns.
Failed means the fetch was attempted and did not work: the host refused, the certificate was invalid, the request timed out, the page was too large, or the address is not publicly reachable. Fix by looking at the site.
A page that returns almost no text is recorded as a retryable failure rather than a skip, because an empty response is usually a rate-limit page or a firewall challenge — a temporary condition — rather than a deliberate exclusion.
Include and exclude patterns
These are the most under-used settings in the crawl form and the highest-value.
They belong to Full website mode only. If you picked Sitemap, the two pattern boxes are not shown, because the sitemap already is an explicit list of URLs — you narrow that crawl by choosing a narrower sitemap document instead.
Include means only URLs matching a pattern are crawled. https://example.com/docs/* restricts a crawl to the documentation and nothing else.
Exclude means matching URLs are skipped. Excluding https://example.com/tag/* and https://example.com/page/* removes the archive noise from a blog crawl.
One pattern per line. Where both are set, include decides the candidate set and exclude removes from it.
The payoff is not just tidiness — it is your knowledge budget. Every archive page you skip is budget spent instead on a page that answers a question. See Crawl limits.
Limits and plan notes
- Respect robots.txt defaults to on for every crawl mode, including sitemap crawls.
- Turning it off never disables the safety checks on unreachable or non-public addresses. Those are not negotiable.
- Sites can also block by user agent or by rate limiting, independently of
robots.txt. A polite delay reduces the chance of tripping those. - You are responsible for having the right to copy what you crawl. Being technically able to fetch a page is not the same as being allowed to use it.
- Skipped and failed pages consume no knowledge capacity.
Common problems
Important pages under /help are skipped.
Check robots.txt in a browser. If a rule disallows that path, allow it on your site — that is the right fix, and it helps search engines too.
Everything failed at the first URL.
The host blocked the fetch, the certificate is invalid, or the address is not public. Test the exact URL in a private browsing window on a different network.
The crawl worked from my office and not for Agentency.
You have a session, or your office IP is allowed and others are not. Agentency crawls anonymously from elsewhere. Test in a private window and, if you can, from a different network.
Robots is off and pages are still skipped.
Then robots was not the cause. Look for pages with too little text, non-HTML file types, or your own exclude patterns.
My staging site refuses everything.
Staging sites are commonly published with Disallow: /. That is deliberate on your part or your host's. Crawl production, or upload the content instead.
I own the site — how do I let Agentency in specifically?
Allow the paths you want crawled in robots.txt for all crawlers, then leave the Agentency setting on. Crawler-specific allowances are fragile and easy to forget about later.
Common questions
Should I turn Respect robots.txt off?
Only on a site you own, and only when you know which rule is in the way. Fixing robots.txt on your own site is the better answer — it helps search engines too.
Will turning it off get me past a login wall?
No. It only ignores the site's crawler rules. Logins, paywalls, and non-public addresses are all still out of reach.
How do I crawl just one section of a site?
Use an include pattern such as https://example.com/docs/* — only matching URLs are fetched. It is the highest-value setting on the crawl form.
My staging site refuses everything. Why?
Staging sites are commonly published with a rule disallowing every path. Crawl production instead, or upload the content directly.
Was this article helpful?
Related articles
Crawl a website
Fetch public pages and turn each one into knowledge. Start with a single page before you commit to a whole site.
Crawl limits
Depth, pages, delay, and parallelism are settings you choose. Your knowledge cap is separate and can stop a crawl first.
Crawl status and history
Every run is recorded with counters and per-page results. This is where you find out why a crawl stopped.
Upload files
Add PDFs, Word files, spreadsheets, and images from your device. Ten megabytes per file, twenty files per batch.
Ready to try it on your own content?
Create a free workspace, add a document, and ask the questions your team is tired of answering.