Skip to content

Website crawl

Goal: Discover pages on your website, choose which to index, and add them to a knowledge base.

  1. Open your knowledge base → Website crawl.
  2. Enter the start URL (usually your homepage or docs root).
  3. Set crawl scope:
SettingGuidance
Same domain onlyRecommended — stays on your site
Max depthHow many link hops from the start URL
Max pagesCap discovery to stay within plan document limits
  1. Click Discover links.

Chatevo lists discovered URLs. Review and select the pages you want indexed:

  • Include help articles, FAQs, and product pages visitors ask about
  • Exclude login, checkout, account, and duplicate listing pages
  • Set classification per page or use the KB default

Click Index selected. Each page is fetched, snapshotted, and queued through the standard pipeline: extract → chunk → embed → Qdrant. See Processing and indexing.

StatusMeaning
QueuedWaiting to fetch
ProcessingExtracting and embedding
IndexedAvailable for retrieval
FailedBlocked, empty, or unreachable — retry or exclude

On Standard, Pro, and Enterprise plans, enable Auto-retrain to re-crawl and re-index when source pages change. Free and Starter plans re-crawl manually.

  • Start with a focused section (e.g. /help/) instead of the entire marketing site
  • Ensure robots.txt allows Chatevo’s crawler
  • After indexing, test with real visitor questions in Test chat