Replaces gpt-oss on ModelConfig, qa.model and processing.title_model, and ministral-3 on the picture-description model. qa.model.vision follows the model and is now true. enable_thinking was gated on the gpt-oss name, so it did nothing for qwen3.8. With title_model's max_tokens of 100 the reasoning consumed the whole budget and title generation returned an empty string. The mapping now applies to any ollama model via reasoning_effort(): false sends "none", true sends "high". Measured on qwen3.8:27b-mlx, "low" does not disable thinking and "none" does; gpt-oss is the inverse, its template has no "none" level, so it keeps "low". Picture description bypasses get_model -- docling posts the request itself from a params dict -- so the flag was inert on that path too. vlm_api_params() carries reasoning_effort into both converters' request bodies. At max_tokens 200 the description survived either way, but the switch cut completion tokens from 141 to 45. test_search_tool_skips_binary_content_when_qa_model_is_text_only asserted the vision default rather than setting it; it now configures vision=False itself. docs/benchmarks.md keeps ministral-3: those are recorded measurements.
4.3 KiB
Search and Question Answering
Search Settings
Configure search behavior and context expansion:
search:
limit: 5 # Default number of results to return
max_context_chars: 5000 # Maximum characters in expanded context
- limit: Default number of search results to return when no limit is specified. Used by CLI, MCP server, and QA. Default: 5
- max_context_chars: Hard limit on total characters in expanded content. Default: 5000.
Context expansion is automatic and section-aware. For structured documents (with section headers), expansion includes the entire section containing the match. For sections that exceed the budget or are too small (e.g., a title+authors area), expansion grows outward item-by-item from the match center, skipping noise labels (footnotes, page headers). This naturally crosses into adjacent sections until the budget is filled. Picture and table matches are exempt: they return their enclosing section as-is and never cross section boundaries. For unstructured documents, expansion grows outward item-by-item. Results without doc_item_refs (e.g., custom chunks passed to import_document) pass through unexpanded.
!!! note "Reranking behavior"
When a reranker is configured, search automatically retrieves 10x the requested limit, then reranks to return the final count. This improves result quality without requiring you to adjust limit.
Question Answering Configuration
Configure the RAG capability (used by client.ask, haiku-rag ask, and the MCP ask_question tool):
qa:
model:
provider: ollama
name: qwen3.8
enable_thinking: true
temperature: 0.3 # Default: 0.3
vision: true # Set false for text-only models
max_searches: 5 # Maximum search units per question
- model: LLM configuration (see Providers)
- model.vision: Set to
truefor vision-capable models (qwen2.5vl,qwen3.6,gpt-4o,claude-sonnet, …). The capability'ssearchtool only attaches picture bytes (BinaryContent) to itsToolReturnwhen this istrue, otherwise picture bytes are withheld. See Pictures × embedder × QA model for the full matrix. - max_searches: Maximum number of search units a capability can spend per question (default: 5). Up to three searches emitted in the same model response share one unit, so a model that rephrases its query in one response spends one unit. A search in a later response starts a new unit, as does each further group of three within one response. Shared by the RAG and analysis capabilities. Searches in one response also deduplicate their returns: evidence a sibling search already showed collapses to a reference line, and each picture attaches once per response.
!!! note "Thinking on vLLM"
enable_thinking only applies to models with a pydantic-ai reasoning profile (o-series, gpt-5, gpt-oss). For other vLLM-served models such as Qwen3 or the Gemma family, the field is a silent no-op — set the chat template switch via extra_body instead.
Analysis Configuration
Configure the analysis capability:
analysis:
model:
provider: anthropic
name: claude-sonnet-4-20250514
temperature: 0.0 # Default: 0.0 (deterministic for code generation)
code_timeout: 60.0 # Max seconds a call may spend reading documents
max_output_chars: 50000 # Truncate output after this many chars
max_executions: 15 # Max execute_code calls per question
- model: LLM configuration (see Providers). When unset, falls back to
qa.model. - code_timeout: Seconds a single
execute_codecall may spend reading documents (default: 60). The sandbox refuses further reads past this point. Code that computes without reading is killed by the worker watchdog at the same limit.code_timeout * max_executionsis the cumulative ceiling across all calls in one question. - max_output_chars: Truncate code output after this many characters (default: 50000)
- max_executions: Maximum
execute_codecalls per question before the capability is told to answer from what it has (default: 15)
See Analysis capability for usage details.