Replaces gpt-oss on ModelConfig, qa.model and processing.title_model, and ministral-3 on the picture-description model. qa.model.vision follows the model and is now true. enable_thinking was gated on the gpt-oss name, so it did nothing for qwen3.8. With title_model's max_tokens of 100 the reasoning consumed the whole budget and title generation returned an empty string. The mapping now applies to any ollama model via reasoning_effort(): false sends "none", true sends "high". Measured on qwen3.8:27b-mlx, "low" does not disable thinking and "none" does; gpt-oss is the inverse, its template has no "none" level, so it keeps "low". Picture description bypasses get_model -- docling posts the request itself from a params dict -- so the flag was inert on that path too. vlm_api_params() carries reasoning_effort into both converters' request bodies. At max_tokens 200 the description survived either way, but the switch cut completion tokens from 141 to 45. test_search_tool_skips_binary_content_when_qa_model_is_text_only asserted the vision default rather than setting it; it now configures vision=False itself. docs/benchmarks.md keeps ministral-3: those are recorded measurements.
63 lines
4.3 KiB
Markdown
63 lines
4.3 KiB
Markdown
# Search and Question Answering
|
||
|
||
## Search Settings
|
||
|
||
Configure search behavior and context expansion:
|
||
|
||
```yaml
|
||
search:
|
||
limit: 5 # Default number of results to return
|
||
max_context_chars: 5000 # Maximum characters in expanded context
|
||
```
|
||
|
||
- **limit**: Default number of search results to return when no limit is specified. Used by CLI, MCP server, and QA. Default: 5
|
||
- **max_context_chars**: Hard limit on total characters in expanded content. Default: 5000.
|
||
|
||
Context expansion is automatic and section-aware. For structured documents (with section headers), expansion includes the entire section containing the match. For sections that exceed the budget or are too small (e.g., a title+authors area), expansion grows outward item-by-item from the match center, skipping noise labels (footnotes, page headers). This naturally crosses into adjacent sections until the budget is filled. Picture and table matches are exempt: they return their enclosing section as-is and never cross section boundaries. For unstructured documents, expansion grows outward item-by-item. Results without `doc_item_refs` (e.g., custom chunks passed to `import_document`) pass through unexpanded.
|
||
|
||
!!! note "Reranking behavior"
|
||
When a reranker is configured, search automatically retrieves 10x the requested limit, then reranks to return the final count. This improves result quality without requiring you to adjust `limit`.
|
||
|
||
## Question Answering Configuration
|
||
|
||
Configure the RAG capability (used by `client.ask`, `haiku-rag ask`, and the MCP `ask_question` tool):
|
||
|
||
```yaml
|
||
qa:
|
||
model:
|
||
provider: ollama
|
||
name: qwen3.8
|
||
enable_thinking: true
|
||
temperature: 0.3 # Default: 0.3
|
||
vision: true # Set false for text-only models
|
||
max_searches: 5 # Maximum search units per question
|
||
```
|
||
|
||
- **model**: LLM configuration (see [Providers](providers.md#model-settings))
|
||
- **model.vision**: Set to `true` for vision-capable models (`qwen2.5vl`, `qwen3.6`, `gpt-4o`, `claude-sonnet`, …). The capability's `search` tool only attaches picture bytes (`BinaryContent`) to its `ToolReturn` when this is `true`, otherwise picture bytes are withheld. See [Pictures × embedder × QA model](processing.md#pictures-embedder-qa-model-how-the-pieces-compose) for the full matrix.
|
||
- **max_searches**: Maximum number of search units a capability can spend per question (default: 5). Up to three searches emitted in the same model response share one unit, so a model that rephrases its query in one response spends one unit. A search in a later response starts a new unit, as does each further group of three within one response. Shared by the RAG and analysis capabilities. Searches in one response also deduplicate their returns: evidence a sibling search already showed collapses to a reference line, and each picture attaches once per response.
|
||
|
||
!!! note "Thinking on vLLM"
|
||
`enable_thinking` only applies to models with a pydantic-ai reasoning profile (o-series, gpt-5, gpt-oss). For other vLLM-served models such as Qwen3 or the Gemma family, the field is a silent no-op — set the chat template switch via [`extra_body`](providers.md#raw-provider-pass-through) instead.
|
||
|
||
## Analysis Configuration
|
||
|
||
Configure the analysis capability:
|
||
|
||
```yaml
|
||
analysis:
|
||
model:
|
||
provider: anthropic
|
||
name: claude-sonnet-4-20250514
|
||
temperature: 0.0 # Default: 0.0 (deterministic for code generation)
|
||
code_timeout: 60.0 # Max seconds a call may spend reading documents
|
||
max_output_chars: 50000 # Truncate output after this many chars
|
||
max_executions: 15 # Max execute_code calls per question
|
||
```
|
||
|
||
- **model**: LLM configuration (see [Providers](providers.md#model-settings)). When unset, falls back to `qa.model`.
|
||
- **code_timeout**: Seconds a single `execute_code` call may spend reading documents (default: 60). The sandbox refuses further reads past this point. Code that computes without reading is killed by the worker watchdog at the same limit. `code_timeout * max_executions` is the cumulative ceiling across all calls in one question.
|
||
- **max_output_chars**: Truncate code output after this many characters (default: 50000)
|
||
- **max_executions**: Maximum `execute_code` calls per question before the capability is told to answer from what it has (default: 15)
|
||
|
||
See [Analysis capability](../capabilities/analysis.md) for usage details.
|