haiku.rag/docs/installation.md
Yiorgis Gozadinos 72ef18e39d
Reject unknown and out-of-range configuration values
Every section inherited plain BaseModel, so unknown keys were dropped
silently: providers.docling_serve.timeout was documented for months while
being ignored, and a typo in any setting took the default. Sections now
derive from ConfigModel, which forbids extras, so a stale or misspelled key
fails with its path. This already found search.context_radius in a live app
config and providers.vllm in soliplex's example.

converter, chunker and chunker_type are Literals. Sizes, limits,
dimensions, token budgets, attempt counts and breaker thresholds must be
positive; retention, delays, intervals and cooldowns non-negative;
similarity_threshold within 0-1; port within 0-65535. port 0 keeps its
OS-assigned meaning and worker_count allows 0 for an API-and-reaper-only
process.

get_reranker caught ImportError and returned None, so a configured reranker
whose extra was missing silently disappeared. It now propagates.
raise_missing_extra names the install command and re-raises when the failure
came from inside an installed package, so a broken transitive import is not
reported as a missing one. zeroentropy imported bare and now guards like the
others.

The haiku.rag package declares the jina extra. jina-local already worked
there through cross-encoder's transitive transformers and torch; the
resolved package set is unchanged, but the support is now promised rather
than inherited.

Provider fields stay unconstrained: get_model ends in a pass-through to
pydantic-ai for any provider it supports, so a Literal there would reject
valid configurations.
2026-08-19 15:32:51 +03:00

3.7 KiB

Installation

Choose Your Package

haiku.rag is available in two packages:

uv pip install haiku.rag

The full package pulls the docling, voyageai, cohere, zeroentropy, cross-encoder, jina and tui extras:

  • Document processing (Docling) - PDF, DOCX, PPTX, images, and 40+ file formats
  • Embedding providers - VoyageAI and Cohere
  • Rerankers - local cross-encoders, local Jina, Cohere, Zero Entropy

It does not include the s3 or ingester extras:

uv pip install 'haiku.rag[ingester]'   # the haiku-ingester service
uv pip install 'haiku.rag[s3]'         # S3 and object storage

Slim Package (Minimal Dependencies)

# Minimal installation (no document processing)
uv pip install haiku.rag-slim

# With document processing
uv pip install haiku.rag-slim[docling]

# With specific providers
uv pip install haiku.rag-slim[docling,voyageai,cross-encoder]

The slim package has minimal dependencies and lets you install only what you need:

  • docling - PDF, DOCX, PPTX, images, and other document formats
  • voyageai - VoyageAI embeddings
  • cross-encoder - Local reranking via sentence-transformers
  • jina - Local Jina reranking (provider: jina-local). Needs transformers and torch, which cross-encoder also pulls
  • cohere - Cohere embeddings and reranking
  • zeroentropy - Zero Entropy reranking
  • s3 - S3 and object-storage access
  • ingester - The haiku-ingester service (also pulls s3)
  • tui - Terminal UI for chat and inspect commands

Built-in providers (no extras needed):

  • Ollama (default embedding provider)
  • OpenAI (GPT models for QA and embeddings)
  • vLLM and other OpenAI-compatible endpoints (embeddings, QA, reranking)
  • Jina reranking via provider: jina, which calls the Jina HTTP API

Other Pydantic AI providers need their own Pydantic AI extra. For Claude models, install pydantic-ai-slim[anthropic].

See Configuration for configuring providers including advanced options like vLLM.

Requirements

  • Python 3.12+
  • Ollama (for default embeddings and QA)

Pre-download Models (Optional)

You can prefetch all required runtime models before first use:

haiku-rag download-models

This will download:

  • Docling models for document processing
  • HuggingFace tokenizer models for chunking
  • Any Ollama models referenced by your current configuration

Remote Processing (Optional)

When using haiku.rag-slim, you can skip installing the docling extra and instead use docling-serve for remote document processing. This is useful for:

  • Keeping dependencies minimal
  • Offloading heavy document processing to a dedicated service
  • Production deployments with separate processing infrastructure

See Remote processing for setup instructions and Document Processing for configuration options.

Docker

Only the slim image is published. Build the full image yourself:

Slim Image (Minimal)

Pre-built slim image with minimal dependencies - use with external docling-serve for document processing:

docker pull ghcr.io/ggozad/haiku.rag-slim:latest

See examples/docker/docker-compose.yml for a complete setup with docling-serve.

Full Image (Self-contained)

Build locally to include all features and document processing without docling-serve:

docker build -f docker/Dockerfile -t haiku-rag .
docker run -p 8001:8001 \
  -v /path/to/haiku.rag.yaml:/app/haiku.rag.yaml \
  -v /path/to/data:/data \
  haiku-rag

See docker/README.md for complete build and configuration instructions, including how to run the ingester service for continuous document ingestion.