Skip to content
Website crawling9 min read

Smart content extraction

How Smart content extraction, listing pages, and the Site information dataset work, plus the training report after a crawl.

Browse topics

What it is

Smart content extraction is how a website or sitemap crawl decides what is this page's content and what is chrome that repeats across the site — menus, sidebars, cookie banners, and the footer.

With it on, each page is trained on the article or product text. The repeated prose that would otherwise be thrown away — the company address, opening hours, phone number, and similar footer contact text — is kept once in a source labelled Site information.

Three switches on the crawl form control this:

  • Smart content extraction — learn the layout and keep the page's own content.
  • Include listing pages — also train on category, tag, archive, and index pages.
  • Keep repeated site text as one Site information dataset — store that shared prose as one knowledge source.

After the crawl finishes, a Training report shows how pages were classified, which text was treated as repeated, and whether the Site information source was created. You can flip the last two switches there without starting a new crawl.

When you would use it

Leave the defaults for almost every public help centre, docs tree, or marketing site: Smart content extraction on, listing pages off, Site information on.

Turn Include listing pages on only when a category or archive page itself answers questions — a department page with a long introduction, a services index that is the only place the offering is described.

Turn Site information off only when you have already pasted hours and contact details as text knowledge and do not want a second copy competing.

Where to find it

Open Dashboard → Chatbots → your chatbot → Knowledge → Add → Website. The three switches sit inside Advanced options, for both Full website and Sitemap modes.

The Training report opens from the active collection card or the crawl history drawer (?drawer=crawlers-history) — the control is labelled View training report. The drawer title is Training report (?drawer=crawl-report).

On the Knowledge table the extra source is labelled Site information, not with an internal code.

Before you start

  • Crawl one page first if you have not used this site before, then run the full crawl with the same switches.
  • Check remaining knowledge capacity. The Site information source is small; listing pages are not.
  • If Advanced options has no Smart content extraction switch, this workspace is still using the standard extractor. Pages then lose their footers the old way, and no Site information source is created.

Steps

  1. Open Dashboard → Chatbots → your chatbot → Knowledge → Add → Website.
  2. Paste the start URL (or the sitemap URL) and choose the crawl mode.
  3. Open Advanced options.
  4. Leave Smart content extraction on unless you have a reason to use the older extractor.
  5. Leave Include listing pages off unless those listing pages themselves carry answers.
  6. Leave Keep repeated site text as one Site information dataset on so footer contact text is stored once.
  7. Submit, stay on Knowledge, and wait for the run to finish.
  8. Open View training report on that crawl. Confirm the Site information line, spot-check a content page and a listing, and only then flip a switch if the report disagrees with what you wanted.
  9. On the Knowledge table, find the Site information row, wait until it reads trained, and ask Test Chatbot for the phone number or opening hours that live in the footer.

The three switches

SwitchDefaultWhat it does
Smart content extractiononLearns each section's repeated layout and keeps the page's own content
Include listing pagesoffAlso trains category, tag, and archive pages. The articles they link to are collected either way
Keep repeated site text as one Site information datasetonStores address, hours, policies, and similar footer prose once, instead of dropping it

Smart content extraction is chosen when you start the crawl and cannot be changed afterwards. The other two can. Turning Smart content extraction off greys out the other two — they only apply to a smart crawl.

A listing page that has a real introduction (a paragraph above the cards) still keeps that introduction when listing pages are off. Only the link cards are dropped. Turn the switch on when the listing is the content.

What is the Site information dataset?

It is one extra knowledge source per crawl, labelled Site information on the Knowledge table.

When it appears. After a website or sitemap crawl that ran with Smart content extraction on and the Site information switch on. It is built when the crawl finishes, not page by page, so it shows up a short while after the last page. It then trains like any other source — the chatbot can use it only once the badge reads trained. See Dataset status and training.

What it contains. Prose that repeated across the site or a section and is worth answering from: the footer address, phone number, opening hours, a short about blurb, a returns policy in the footer. Navigation labels and lists of links are not kept. Each piece of text is stored once, so the footer contact block is not copied onto every page's source.

When it does not appear.

  • Smart content extraction was off — there is no inventory to build it from.
  • You switched Keep repeated site text as one Site information dataset off, at start or later from the report. Repeated text is still removed from each page; it is just not stored.
  • The crawl never finished, or it stopped so early that there was not enough repeated prose to keep.
  • You turned the switch on long after the crawl, and the report says the dataset will be built on the next crawl. Run the crawl again (sitemap mode if you want updates without duplicates — see Crawl from a sitemap).

Turning the Site information switch off from the Training report removes that source from the knowledge base. Turning it back on rebuilds it when the collected text is still available; otherwise wait for the next crawl.

The Site information row is not recrawled on its own. The next smart crawl of that site rebuilds it.

The training report

Open View training report once collection has finished (a run that has not started yet has no report). While a crawl is still collecting, the report is partial and labelled as such.

It shows:

  • Page outcomes — content pages, listing pages, media, too-short, duplicates, failed, excluded, and still pending.
  • Sections — each part of the site profiled on its own, so a blog's sidebar is not confused with the docs navigation.
  • Repeated site text — previews of what was removed from every page. Blocks kept in Site information are marked Kept in Site information.
  • Suggested exclude patterns — URL paths that were almost entirely listings. Copy one into exclude patterns on the next crawl if you want to skip that section entirely. See Robots.txt and blocked pages.
  • Sample pages — a few URLs per outcome so you can click through and check the classification.
  • Site information — whether the dataset exists, is off, or will be built on the next crawl.
  • Training options — Include listing pages and Keep Site information dataset. These are the same two switches as on the form, applied to this finished crawl.

Flipping Include listing pages on queues the refused listing pages for training (a sitemap crawl may refetch them on the next scan). Flipping it off removes those listing sources. The confirmation dialog says so before anything is deleted.

The Training report is separate from crawl history counters. History answers "did we fetch it?"; the report answers "what did we train?" See Crawl status and history.

What you will see

On the crawl form, the three switches are collapsed under Advanced options. Listing pages and Site information stay visible but disabled while Smart content extraction is off.

On Knowledge, a finished smart crawl adds the usual page sources plus, when the switch was on, one Site information row. Page sources are titled from the page title; the Site information row is not a page.

The Training report is a slide-over, not a new page. You keep your place on Knowledge.

Limits and plan notes

  • Smart content extraction applies to website and sitemap crawls. It does not change file uploads or pasted text.
  • Collected pages and the Site information source all count toward the knowledge cap. Listing pages spend the cap on indexes you often do not want — that is why the switch starts off. See Knowledge limits.
  • A crawl that hits the knowledge cap still finishes; it may halt before Site information is built. Free space and run again, or turn Site information on from the report if the report says it can be built now.
  • Removing the crawl session from history does not delete the Site information source. Remove it from the Knowledge table. See Remove knowledge.
  • Do not mention these switches to a visitor. They only affect how this chatbot is trained.

Common problems

The chatbot does not know our phone number or opening hours.

Those almost always live in the footer. With Smart content extraction on and Site information on, they belong in the Site information source — wait until that row is trained, then ask using the words on the footer. If the row is missing, the switch was off, or the crawl has not finished. If the switch is on and the report says the dataset will be built on the next crawl, run the crawl again.

Every page's source still starts with the same menu.

Smart content extraction was off for that crawl (it cannot be turned on afterwards). Start a new crawl with the switch on, and remove the old page sources so the two versions do not compete. Use sitemap mode if you want the next run to update in place.

I have hundreds of tag and category pages I did not want.

Include listing pages was on. Turn it off from the Training report to drop those sources, or remove them from Knowledge. Next time leave the switch off, and add an exclude pattern for /tag/* if the report suggested one.

I flipped Include listing pages on and nothing happened.

On a website crawl the refused listings are queued from text already stored; give training a few minutes. On a sitemap crawl they may wait for the next scan — the report says when a refetch is scheduled. A paused sitemap keeps the change for resume.

There is no View training report control.

The crawl never started, or this workspace does not offer Smart content extraction. History counters still work as in Crawl status and history.

I turned Site information off by mistake.

Turn Keep Site information dataset back on in the Training report. If the report says it will be built on the next crawl, run that crawl (sitemap mode to avoid duplicating pages).

Common questions

What is the Site information dataset?

One knowledge source per crawl, labelled Site information. It stores repeated site prose once — the footer address, phone number, opening hours, and similar contact text — instead of dropping that text as chrome from every page. It appears after a smart crawl finishes when Keep repeated site text as one Site information dataset is on.

What does Smart content extraction do?

It learns each section's repeated layout (menus, footers, sidebars) and trains each page on that page's own content. The switch is under Advanced options on the crawl form, on by default, and cannot be changed after the crawl has started.

Should I turn on Include listing pages?

Usually no. Category, tag, and archive pages are mostly links to articles the crawl collects anyway. Leave it off unless a listing page itself carries answers you need. You can turn it on later from the Training report.

Where is the training report?

On the Knowledge tab, from the active collection card or the crawl history drawer. Open View training report. The drawer shows page outcomes, repeated site text, the Site information status, and the two switches you can still change after the crawl.

Why doesn't the chatbot know our phone number after a crawl?

That text usually lives in the footer. With Smart content extraction and Site information on, it is stored once in the Site information source — wait until that row is trained, then ask using the words on the footer. If the row is missing, the switch was off or the crawl has not finished.

Was this article helpful?

Ready to try it on your own content?

Create a free workspace, add a document, and ask the questions your team is tired of answering.