Advanced Vector Settings
The defaults work well for most content. Read this page if your answers feel incomplete or off-target.
Chunking: how your content is split
Before storing, every source is split into chunks — the units that get matched against a query. Three settings control this (in ⚙️ Advanced Parameters on the vector page):
Chunking Method
| Method | Best for |
|---|---|
| Recursive Text Splitter (default) | Almost everything — splits at paragraph, then sentence boundaries |
| Simple Text Splitter | Uniform text where structure doesn't matter |
| Markdown Text Splitter | Markdown files — respects headings and lists |
Chunk Size
Default 700 characters. The effective range is 300–2000 — values below 300 are automatically raised to 300.
- Smaller chunks (300–600) → more precise matches, but less context per chunk. Good for FAQs and short facts.
- Larger chunks (1000–2000) → more context per match, but matching gets fuzzier. Good for long-form articles and policies.
Overlap Size
Default 105 characters. Overlap repeats the end of one chunk at the start of the next, so sentences aren't cut in half at a boundary. The effective overlap is capped at 15% of the chunk size — setting more has no effect.
Section-aware splitting (automatic)
If a document uses numbered headings (3, 3.2 Payment Terms …), Indite detects them and splits by section, attaching section and subsection labels to each chunk. This happens automatically and enables questions like:
- "List the subsections of section 3"
- "What does section 4.2 say?"
If you control the source documents, numbered headings noticeably improve structured retrieval.
Retrieval tips
- Number of Documents (on the Vector Store block, default 2): raise it to 4–6 for broad questions that need multiple chunks; keep it low for short factual lookups. Retrieved chunks below a relevance threshold are dropped automatically, so a higher number never forces bad matches in.
- Results are diversified automatically — the block avoids returning near-duplicate chunks.
- Chunks from crawled pages carry a source URL; instruct your AI block to cite it (see Use Vectors).
Troubleshooting
| Symptom | Try |
|---|---|
| Answers miss obvious content | Check the source's status isn't Failed; retrain after fixing |
| Answers cut off mid-thought | Increase chunk size, or raise Number of Documents |
| Wrong section retrieved | Use numbered headings in the source, or split content into more focused sources |
| Website content missing | The crawler stays on the same domain and stops at 200 pages / 2 minutes — add deeper pages as Individual Links |
| Changes not showing | Every edit needs a Retrain — training fully rebuilds the index |