haiku.rag/docs/configuration/storage.md
Yiorgis Gozadinos 483c0ec354
Give the docs an architecture page and one extras list
overview.md repeated the landing page: the same install-and-ask block and
five of six identical links. It was positioning prose, where the docs had no
page describing how the system works.

Rewrite it as Architecture, following the data through: source adapter,
converter, chunker, embedder, transaction; then storage and its versioning;
then retrieval, with the 10x rerank fetch and section-bounded expansion; then
the two capabilities; then laptop versus ingester. Retitled in the nav and on
the landing page, filename kept so existing links resolve.

Extras were listed in three places and none was complete.
docs/installation.md now carries a table of all fifteen slim extras, what each
provides, and which the full package already includes.
haiku_rag_slim/README.md names them and links there. The claim that other
providers need their own pydantic-ai extra was wrong: haiku.rag-slim defines
anthropic, google, groq, mistral, bedrock and vertexai itself.

configuration/storage.md opens with the four operational constraints, which
were either buried in an S3 section or undocumented: one writer per URI,
reader lag by read_consistency_interval_seconds, migrate after a
schema-changing upgrade, and the fixed embedding dimension with what
ConfigMismatchError means and which rebuild mode resolves it.

The one-writer rule is stated as a haiku.rag constraint, which is what it is:
the multi-table lock, version snapshot and rollback are process-local, so a
second writer can commit inside another's transaction and be reverted by its
rollback. storage.md and ingester.md both claimed it was a LanceDB property
that corrupts manifests. The S3 deployment section now links to the
constraint instead of restating it.

Get started reads index, Quickstart, Installation, Architecture. The landing
page's list was missing Installation.
2026-08-20 15:07:06 +03:00

11 KiB

Database and Storage

Operational constraints

Four things to know before deploying.

Run one writer per database. This is a haiku.rag constraint, not a LanceDB one. A write that spans several tables is serialized by an in-process lock and rolled back by restoring each table to the version it had when the write started. Both are process-local: a second writing process can commit between that snapshot and the mutation, and a rollback would then revert its work along with ours. Run a single writer, either the haiku-ingester service or your own application. Read-only consumers are unrestricted.

Readers lag by an interval. A connection always sees its own writes. It sees another process's writes after lancedb.read_consistency_interval_seconds (default 30).

Migrate after an upgrade that changes the schema. haiku-rag migrate applies pending migrations in place, and haiku-rag info lists what is pending. A release that needs it says so in the changelog.

The embedding dimension is fixed per database. Every chunk vector has the dimension the database was created with. Changing embeddings.model.vector_dim raises ConfigMismatchError on open, because stored vectors cannot be compared against new ones. Changing the provider or model name while keeping the dimension warns on a read-only open and raises on a writable one. haiku-rag rebuild --set-embedder adopts the new identity without re-embedding, and haiku-rag rebuild --embed-only re-embeds against the new model.

Local Storage

By default, haiku.rag uses a local LanceDB database:

storage:
  data_dir: /path/to/data  # Empty = use default platform location
  auto_vacuum: true  # Enable automatic vacuuming after operations
  vacuum_retention_seconds: 86400  # Cleanup threshold in seconds
  • data_dir: Directory for local database storage. When empty, uses platform-specific default locations
  • auto_vacuum: When enabled (default), automatically runs vacuum after document create/update/delete operations and database rebuilds. Background vacuums are throttled to at most one every 5 minutes, so sustained ingestion does not trigger continuous compaction, and a final vacuum runs when the client closes. Set to false to disable automatic vacuuming and rely on manual haiku-rag vacuum commands only. Disabling can help avoid potential crashes in high-concurrency scenarios
  • vacuum_retention_seconds: When vacuum runs, old table versions older than this threshold are removed. Default: 86400 seconds (1 day). Set to 0 for aggressive cleanup (removes all old versions immediately)

!!! warning "Vacuum Retention Threshold" The vacuum_retention_seconds value should be larger than the typical time it takes to process and write a document. If a concurrent operation is in progress while vacuum runs, setting this value too low can cause race conditions where vacuum removes table versions that an in-flight operation still needs. The default of 86400 seconds (1 day) is conservative and safe for most use cases.

Vacuum Memory Requirements

Vacuum compacts small data files into larger ones. LanceDB targets roughly one million rows per fragment, which a documents table holding multi-megabyte docling blobs never reaches, so each vacuum that follows new documents re-merges the whole existing fragment rather than only the new ones. Peak memory therefore scales with the total size of the documents table, not with how much was added.

Measured peak resident memory is about 5x the size of the documents table's data files. An 8.8 GB table peaked at 48.7 GB. Plan for 6x the size of documents/ on disk as available RAM, or the vacuum will be killed by the OOM killer partway through.

Check the current size with:

du -sh /path/to/database.lancedb/documents.lance

If that number times six exceeds available RAM, use one of:

  • Reduce images_scale (see Image Settings). Rendered page rasters dominate the size of documents, and their byte cost falls with the square of the scale factor.
  • Set generate_page_images: false if visual grounding through visualize_chunk() is not needed. This removes page rasters entirely.
  • Set auto_vacuum: false and run haiku-rag vacuum manually when the machine is otherwise idle, so the peak does not land alongside ingestion.

This is an upstream limitation rather than a haiku.rag setting. Compaction bounds itself by row count instead of bytes, and LanceDB's async API exposes no batch size or fragment target to override it. Tracked at lancedb/lancedb#2325. The requirement above will drop once compaction batches by bytes.

Database Creation

Databases must be explicitly created before use:

CLI:

# Create in default location (see Configuration File Locations below)
haiku-rag init

# Create at custom path
haiku-rag init --db /path/to/database.lancedb

Python:

# Create at custom path
async with HaikuRAG("/path/to/database.lancedb", create=True) as client:
    ...

# Create in default location
async with HaikuRAG(create=True) as client:
    ...

The default location is platform-specific (e.g., ~/Library/Application Support/haiku.rag/ on macOS).

Operations on non-existent databases raise FileNotFoundError. This prevents accidental database creation from typos or misconfigured paths.

Remote Storage

For remote storage, use the lancedb settings with various backends:

# LanceDB Cloud
lancedb:
  uri: db://your-database-name
  api_key: your-api-key
  region: us-west-2  # optional

# Amazon S3
lancedb:
  uri: s3://my-bucket/my-table
  storage_options:
    region: us-east-1

# Amazon S3 with explicit credentials
lancedb:
  uri: s3://my-bucket/my-table
  storage_options:
    aws_access_key_id: YOUR_ACCESS_KEY
    aws_secret_access_key: YOUR_SECRET_KEY
    region: us-east-1

# S3-compatible (SeaweedFS, Tigris, etc.)
lancedb:
  uri: s3://my-bucket/my-table
  storage_options:
    endpoint: http://localhost:8333
    aws_access_key_id: YOUR_ACCESS_KEY
    aws_secret_access_key: YOUR_SECRET_KEY
    region: us-east-1
    allow_http: "true"

# Azure Blob Storage
lancedb:
  uri: az://my-container/my-table

# Google Cloud Storage
lancedb:
  uri: gs://my-bucket/my-table

# HDFS
lancedb:
  uri: hdfs://namenode:port/path/to/table
  • LanceDB Cloud (db://): Requires api_key and region. Table optimization and indexing are managed server-side.
  • Object storage (s3://, gs://, az://, hdfs://): Uses storage_options for credentials and endpoint configuration. Authentication can also be provided via environment variables (AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, etc.) or cloud provider SDK defaults (AWS CLI, Azure CLI, gcloud).
  • S3-compatible stores (MinIO, Tigris, etc.): Set endpoint in storage_options. When using http:// endpoints, also set allow_http: "true".

The storage_options keys are case-insensitive and passed directly to the underlying object store library. Available keys depend on the backend. See the LanceDB storage docs for details.

Note: Table optimization is automatically handled by LanceDB Cloud (db:// URIs) and is disabled for better performance. For object storage backends (S3, Azure, GCS), optimization and vector indexing are still performed normally.

Caching and Read Consistency

lancedb:
  read_consistency_interval_seconds: 30   # null to never re-check
  index_cache_size_bytes: 536870912       # null for the LanceDB default
  metadata_cache_size_bytes: 268435456
  • read_consistency_interval_seconds: how often a connection checks for writes from another process. null never checks, so a long-lived reader never sees the ingester's writes. 0 checks on every read.
  • index_cache_size_bytes / metadata_cache_size_bytes: sizes for the caches held by the LanceDB session, which is shared across every connection in the process. The first vector query loads the index into it, so on object storage the cache is what stops the next connection refetching it. Size it for the total set of indexes a process keeps warm, against the memory available to it.

Deployment Pattern: One Writer, Many Readers

The one-writer constraint shapes the deployment: one writing process per database URI, any number of read-only consumers.

The recommended layout for production is "different buckets, same account, separate IAM roles per process":

  • Ingestion process — IAM role with s3:Get/List on the documents bucket and s3:Get/Put/Delete on the LanceDB bucket. Runs haiku-ingester serve (with ingester.sources[type=s3] pointing at the documents bucket). Exactly one such process per LanceDB URI.
  • Consumer processes (1..N) — IAM role with s3:Get/List on the LanceDB bucket only. Run haiku-rag --read-only mcp, the chat TUI, etc. They never see the documents bucket.

Each process picks up its own credentials from the AWS default chain (env vars, IAM instance role, AWS profile), so no credentials are hard-coded in the configuration files.

Vector Indexing

Configure vector search settings:

search:
  vector_index_metric: cosine  # cosine, l2, or dot
  vector_refine_factor: 30     # Re-ranking factor for accuracy

For search behavior settings (limit, max_context_chars), see Search and Question Answering.

  • vector_index_metric: Distance metric for vector similarity:
    • cosine: Cosine similarity (default, best for most embeddings)
    • l2: Euclidean distance
    • dot: Dot product similarity
  • vector_refine_factor: Improves accuracy when using a vector index by retrieving refine_factor * limit candidates (using approximate search) and re-ranking them with exact distances. Higher values increase accuracy but slow down queries. Default: 30
    • Only applies with a vector index - has no effect on brute-force search, which already returns exact results

!!! note Vector indexes are only necessary for large datasets with over 100,000 chunks. For smaller datasets, LanceDB's brute-force kNN search provides exact results with good performance. Only create an index if you notice search performance degradation on large datasets.

Index creation:

Vector indexes are not created automatically during document ingestion to avoid slowing down the process. After you've added documents (at least 256 chunks required), create the index manually:

haiku-rag create-index

This command:

  • Checks if you have enough data (minimum 256 chunks)
  • Creates an IVF_PQ index for fast approximate nearest neighbor (ANN) search
  • Uses LanceDB's automatic parameter calculation based on your dataset size and vector dimensions

Re-indexing:

Indexes are not automatically updated when you add new documents. After adding a significant amount of new data:

haiku-rag create-index  # Rebuilds the index with all data

Searches still work with stale indexes - LanceDB uses the index for old data (fast ANN) and brute-force kNN for new unindexed rows, then combines the results. However, performance degrades as more unindexed data accumulates.

For datasets with fewer than 256 chunks, searches use brute-force kNN scans (exact nearest neighbors, 100% recall) which work well for small datasets but don't scale beyond a few hundred thousand vectors.