_build_result applied the noise-label filter to every item in the range, including the ones the result matched on. A hit on a footnote or index entry returned its section with the matched text removed, the clip anchor could not find the evidence and fell back to a prefix window, and the un-merge path rebuilt through the same filter. Noise is now a set of positions computed once per group by _noise_positions: noise-labelled items minus the matched ones. _expand_outward and _build_result take that set instead of a flag, so the matched item is kept in content and counted toward the budget. Footnotes leave the noise set. They carry sources, cross-references and clarifications, and docling attaches table and figure footnotes to the table itself, so the filter was dropping part of the table. The noise set is page_header, page_footer and document_index. Refs #609
63 lines
4.4 KiB
Markdown
63 lines
4.4 KiB
Markdown
# Search and Question Answering
|
||
|
||
## Search Settings
|
||
|
||
Configure search behavior and context expansion:
|
||
|
||
```yaml
|
||
search:
|
||
limit: 5 # Default number of results to return
|
||
max_context_chars: 5000 # Maximum characters in expanded context
|
||
```
|
||
|
||
- **limit**: Default number of search results to return when no limit is specified. Used by CLI, MCP server, and QA. Default: 5
|
||
- **max_context_chars**: Hard limit on total characters in expanded content. Default: 5000.
|
||
|
||
Context expansion is automatic and section-aware. For structured documents (with section headers), expansion includes the entire section containing the match. For sections that exceed the budget or are too small (e.g., a title+authors area), expansion grows outward item-by-item from the match center, skipping noise labels (page headers, page footers, table of contents). This naturally crosses into adjacent sections until the budget is filled. Picture and table matches are exempt: they return their enclosing section as-is and never cross section boundaries. For unstructured documents, expansion grows outward item-by-item. Results without `doc_item_refs` (e.g., custom chunks passed to `import_document`) pass through unexpanded.
|
||
|
||
!!! note "Reranking behavior"
|
||
When a reranker is configured, search automatically retrieves 10x the requested limit, then reranks to return the final count. This improves result quality without requiring you to adjust `limit`.
|
||
|
||
## Question Answering Configuration
|
||
|
||
Configure the RAG capability (used by `client.ask` and `haiku-rag ask`):
|
||
|
||
```yaml
|
||
qa:
|
||
model:
|
||
provider: ollama
|
||
name: qwen3.8
|
||
enable_thinking: true
|
||
temperature: 0.3 # Default: 0.3
|
||
vision: true # Set false for text-only models
|
||
max_searches: 5 # Maximum search units per question
|
||
```
|
||
|
||
- **model**: LLM configuration (see [Providers](providers.md#model-settings))
|
||
- **model.vision**: Set to `true` for vision-capable models (`qwen2.5vl`, `qwen3.6`, `gpt-4o`, `claude-sonnet`, …). The capability's `search` tool only attaches picture bytes (`BinaryContent`) to its `ToolReturn` when this is `true`, otherwise picture bytes are withheld. See [Pictures × embedder × QA model](processing.md#pictures-embedder-qa-model-how-the-pieces-compose) for the full matrix.
|
||
- **max_searches**: Maximum number of search units a capability can spend per question (default: 5). Up to three searches emitted in the same model response share one unit, so a model that rephrases its query in one response spends one unit. A search in a later response starts a new unit, as does each further group of three within one response. Shared by the RAG and analysis capabilities. Searches in one response also deduplicate their returns: evidence a sibling search already showed collapses to a reference line, and each picture attaches once per response.
|
||
|
||
!!! note "Thinking on vLLM"
|
||
`enable_thinking` only applies to models with a pydantic-ai reasoning profile (o-series, gpt-5, gpt-oss). For other vLLM-served models such as Qwen3 or the Gemma family, the field is a silent no-op — set the chat template switch via [`extra_body`](providers.md#raw-provider-pass-through) instead.
|
||
|
||
## Analysis Configuration
|
||
|
||
Configure the analysis capability:
|
||
|
||
```yaml
|
||
analysis:
|
||
model:
|
||
provider: anthropic
|
||
name: claude-sonnet-4-20250514
|
||
temperature: 0.0 # Default: 0.0 (deterministic for code generation)
|
||
code_timeout: 60.0 # Per call: compute stops, no read or search starts past it
|
||
max_output_chars: 50000 # Truncate output after this many chars
|
||
max_executions: 15 # Max execute_code calls per question
|
||
```
|
||
|
||
- **model**: LLM configuration (see [Providers](providers.md#model-settings)). When unset, falls back to `qa.model`.
|
||
- **code_timeout**: Seconds a single `execute_code` call has (default: 60). Past it the sandbox starts no further host call, a document read or an in-code `search()` / `list_documents()`; one already running finishes. Code that computes without host calls is killed by the worker watchdog at the same limit. `code_timeout * max_executions` is the cumulative ceiling across all calls in one question.
|
||
- **max_output_chars**: Truncate code output after this many characters (default: 50000)
|
||
- **max_executions**: Maximum `execute_code` calls per question before the capability is told to answer from what it has (default: 15)
|
||
|
||
See [Analysis capability](../capabilities/analysis.md) for usage details.
|