haiku.rag/docs/configuration/index.md
Yiorgis Gozadinos 34180a0fd1
One reference and one placement for a database
DatabaseRef is a name and a location. The configuration places databases
through lancedb.databases alone; with none configured the default is the
entry haiku.rag under storage.data_dir, selectable like any other.
lancedb.uri is removed, and a config carrying it fails to load with the
replacement spelled out. A path passed from Python is valid where the
configuration places nothing and raises AmbiguousDatabaseError beside
lancedb.databases; haiku-rag --db and haiku-ingester --db construct the
scope directly, so a human's override keeps working. Every database
answers to a name, and a database given as a path keeps its own errors.
2026-09-03 15:12:08 +03:00

6.6 KiB

Configuration

Configuration is done through YAML configuration files.

!!! note haiku.rag enforces one hard rule on existing databases: the embedding vector_dim in your config must match the value stored in the db. A mismatch exits with ConfigMismatchError and you must rebuild to apply the change (see Rebuild Database).

Opening a database never writes to it, so the stored embedding identity is left untouched. Changing only `provider` or `name` (e.g. switching from Ollama to vLLM serving the same model) is treated as soft drift: read-only opens log a warning and continue, while writable opens exit with `ConfigMismatchError`. Reconcile the stored identity with your config by running `haiku-rag rebuild --set-embedder` (see [Rebuild Database](../cli.md#rebuild-database)). If the change was unintentional, revert your config instead.

Getting Started

Generate a configuration file with defaults:

haiku-rag init-config

This creates a haiku.rag.yaml file in your current directory with all available settings.

Configuration File Locations

haiku.rag searches for configuration files in this order:

  1. Path specified via --config flag: haiku-rag --config /path/to/config.yaml <command>
  2. ./haiku.rag.yaml (current directory)
  3. Platform-specific user directory:
    • Linux: ~/.local/share/haiku.rag/haiku.rag.yaml
    • macOS: ~/Library/Application Support/haiku.rag/haiku.rag.yaml
    • Windows: C:/Users/<USER>/AppData/Roaming/haiku.rag/haiku.rag.yaml

Environment Variables

Any string value can reference an environment variable, so secrets stay out of the file and one config can serve multiple deployments:

ingester:
  queue:
    dburi: postgresql+asyncpg://haiku:${POSTGRES_PASSWORD}@db:5432/haiku_rag
  • ${VAR} is replaced with the value of VAR. If VAR is unset, loading fails with an error naming the variable.
  • ${VAR:-default} uses default when VAR is unset or empty.
  • $$ produces a literal $.

Substitution happens after the YAML is parsed, so a value containing :, @, or # fills the string verbatim and never changes the document structure.

Minimal Configuration

A minimal configuration file with defaults:

# haiku.rag.yaml
environment: production

embeddings:
  model:
    provider: ollama
    name: qwen3-embedding:4b
    vector_dim: 2560

qa:
  model:
    provider: ollama
    name: gpt-oss
    enable_thinking: true

Complete Configuration Example

# haiku.rag.yaml
environment: production

storage:
  data_dir: ""  # Empty = use default platform location
  vacuum_retention_seconds: 86400

ingester:
  sources:
    - type: fs
      id: local-docs
      root: /path/to/documents
      ignore_patterns: []  # Gitignore-style patterns to exclude
      include_patterns: []  # Gitignore-style patterns to include
      delete_orphans: true

lancedb:
  databases: {}  # Name-to-location map; empty places haiku.rag under data_dir
  api_key: ""  # LanceDB Cloud (db://) credentials
  region: ""

embeddings:
  model:
    provider: ollama
    name: qwen3-embedding:4b
    vector_dim: 2560

reranking:
  # Omit this section, or set `model: null`, to disable reranking.
  model:
    provider: cross-encoder  # cross-encoder, cohere, zeroentropy, vllm, jina, jina-local
    name: cross-encoder/ms-marco-MiniLM-L-6-v2
  multimodal: false  # vllm only: send picture chunks to the reranker as images

qa:
  model:
    provider: ollama
    name: gpt-oss
    enable_thinking: true
    temperature: 0.3
  max_searches: 5

search:
  limit: 5                     # Default number of results to return
  max_context_chars: 5000     # Maximum characters in expanded context
  vector_index_metric: cosine  # cosine or l2
  vector_refine_factor: 30

doctor:
  duplicates:                    # Near-duplicate document detection (doctor command)
    similarity_threshold: 0.97   # cosine cutoff on document embedding centroids
    min_chunks: 3                # documents with fewer chunks are excluded

prompts:
  domain_preamble: ""  # Prepended to capability instructions

processing:
  converter: docling-local  # docling-local or docling-serve
  chunker: docling-local    # docling-local or docling-serve
  chunker_type: hybrid      # hybrid or hierarchical
  chunk_size: 256
  chunking_tokenizer: "Qwen/Qwen3-Embedding-0.6B"
  chunking_merge_peers: true
  chunking_use_markdown_tables: false
  auto_title: false              # Auto-generate titles on ingestion
  title_model:
    provider: ollama
    name: gpt-oss
    enable_thinking: false
    temperature: 0.3
    max_tokens: 100
  conversion_options:
    do_ocr: true
    force_ocr: false
    ocr_lang: []
    do_table_structure: true
    table_mode: accurate
    table_cell_matching: true
    images_scale: 2.0

providers:
  ollama:
    base_url: http://localhost:11434

  docling_serve:
    base_url: http://localhost:5001
    api_key: ""
    timeout: 300

Programmatic Configuration

When using haiku.rag as a Python library, you can pass configuration directly to the HaikuRAG client:

from haiku.rag.config import AppConfig
from haiku.rag.config.models import EmbeddingModelConfig, ModelConfig, QAConfig, EmbeddingsConfig
from haiku.rag.client import HaikuRAG

# Create custom configuration
custom_config = AppConfig(
    qa=QAConfig(
        model=ModelConfig(
            provider="openai",
            name="gpt-4o",
            temperature=0.3
        )
    ),
    embeddings=EmbeddingsConfig(
        model=EmbeddingModelConfig(
            provider="ollama",
            name="qwen3-embedding:4b",
            vector_dim=2560
        )
    ),
    processing={"chunk_size": 512}
)

# Pass configuration to the client
async with HaikuRAG(config=custom_config) as client:
    ...

If you don't pass a config, the client uses the global configuration loaded from your YAML file or defaults.

This is useful for:

  • Jupyter notebooks
  • Python scripts
  • Testing with different configurations
  • Applications that need multiple clients with different configurations

Configuration Topics

For detailed configuration of specific topics, see:

  • Providers - Model settings and provider-specific configuration (embeddings, reranking)
  • Search and Question Answering - Search settings and question answering
  • Document Processing - Document conversion and chunking
  • Ingester - Continuous ingestion from filesystem, HTTP, S3, and WebDAV sources
  • Storage - Database, remote storage, and vector indexing
  • Prompts - Customize agent prompts for your domain