Vectors
Advanced Settings

Advanced Vector Settings

The defaults work well for most content. Read this page if your answers feel incomplete or off-target.

Chunking: how your content is split

Before storing, every source is split into chunks — the units that get matched against a query. Three settings control this (in ⚙️ Advanced Parameters on the vector page):

Chunking Method

MethodBest for
Recursive Text Splitter (default)Almost everything — splits at paragraph, then sentence boundaries
Simple Text SplitterUniform text where structure doesn't matter
Markdown Text SplitterMarkdown files — respects headings and lists

Chunk Size

Default 700 characters. The effective range is 300–2000 — values below 300 are automatically raised to 300.

  • Smaller chunks (300–600) → more precise matches, but less context per chunk. Good for FAQs and short facts.
  • Larger chunks (1000–2000) → more context per match, but matching gets fuzzier. Good for long-form articles and policies.

Overlap Size

Default 105 characters. Overlap repeats the end of one chunk at the start of the next, so sentences aren't cut in half at a boundary. The effective overlap is capped at 15% of the chunk size — setting more has no effect.

Section-aware splitting (automatic)

If a document uses numbered headings (3, 3.2 Payment Terms …), Indite detects them and splits by section, attaching section and subsection labels to each chunk. This happens automatically and enables questions like:

  • "List the subsections of section 3"
  • "What does section 4.2 say?"

If you control the source documents, numbered headings noticeably improve structured retrieval.

Retrieval tips

  • Number of Documents (on the Vector Store block, default 2): raise it to 4–6 for broad questions that need multiple chunks; keep it low for short factual lookups. Retrieved chunks below a relevance threshold are dropped automatically, so a higher number never forces bad matches in.
  • Results are diversified automatically — the block avoids returning near-duplicate chunks.
  • Chunks from crawled pages carry a source URL; instruct your AI block to cite it (see Use Vectors).

Troubleshooting

SymptomTry
Answers miss obvious contentCheck the source's status isn't Failed; retrain after fixing
Answers cut off mid-thoughtIncrease chunk size, or raise Number of Documents
Wrong section retrievedUse numbered headings in the source, or split content into more focused sources
Website content missingThe crawler stays on the same domain and stops at 200 pages / 2 minutes — add deeper pages as Individual Links
Changes not showingEvery edit needs a Retrain — training fully rebuilds the index
Indite Documentation v1.7.1
PrivacyTermsSupport