Skip to content

Supported file types

Chatevo extracts text from common business document formats and web content. All types follow the same pipeline: extract → chunk → embed → store in Qdrant.

FormatExtensionNotes
PDF.pdfBest for manuals, brochures, and scanned docs with selectable text
Word.docxHeadings and lists are preserved for chunking
Plain text.txtIdeal for FAQs and short reference content
HTML.html, .htmUseful for exported help-center pages
MethodWhat it importsGuide
Single URLOne page at a timeImport from URL
Website crawlDiscover and select multiple pages from a domainWebsite crawl

Crawled and imported URLs are stored and indexed like uploaded files. Each import creates a snapshot — an immutable copy of the page content at import time.

File size, document count per knowledge base, and total storage depend on your plan. Check Billing → Usage or Usage and quotas before large uploads.

PlanTypical limits
FreeSmallest storage and document counts
StarterHigher limits; tools enabled
Standard+Larger KBs, auto-retrain, full smart query understanding

Exact numbers are listed on Plans compared.

  • Password-protected PDFs
  • Image-only scans without OCR text
  • Binary formats (.xlsx, .pptx, .zip) — export to PDF or DOCX first
  • Pages blocked by robots.txt or login walls during crawl