Website crawling
Turn a public website into chatbot knowledge — one page, a whole site, or a sitemap that rescans on a schedule.
Browse topics
Crawl a website
Fetch public pages and turn each one into knowledge. Start with a single page before you commit to a whole site.
Crawl from a sitemap
The only crawl mode that rescans on a schedule — hourly, daily, or weekly — and skips pages that have not changed.
Crawl status and history
Every run is recorded with counters and per-page results. This is where you find out why a crawl stopped.
Robots.txt and blocked pages
Agentency obeys robots.txt by default, and refuses unsafe addresses always. Here is how to tell which stopped you.
Recrawl or stop a crawl
Cancel a run in flight, switch off a sitemap schedule, or refresh content — without ending up with two of every page.
Crawl limits
Depth, pages, delay, and parallelism are settings you choose. Your knowledge cap is separate and can stop a crawl first.
Smart content extraction
How Smart content extraction, listing pages, and the Site information dataset work, plus the training report after a crawl.
Ready to try it on your own content?
Create a free workspace, add a document, and ask the questions your team is tired of answering.