haiku.rag/docs/installation.md
Yiorgis Gozadinos 483c0ec354
Give the docs an architecture page and one extras list
overview.md repeated the landing page: the same install-and-ask block and
five of six identical links. It was positioning prose, where the docs had no
page describing how the system works.

Rewrite it as Architecture, following the data through: source adapter,
converter, chunker, embedder, transaction; then storage and its versioning;
then retrieval, with the 10x rerank fetch and section-bounded expansion; then
the two capabilities; then laptop versus ingester. Retitled in the nav and on
the landing page, filename kept so existing links resolve.

Extras were listed in three places and none was complete.
docs/installation.md now carries a table of all fifteen slim extras, what each
provides, and which the full package already includes.
haiku_rag_slim/README.md names them and links there. The claim that other
providers need their own pydantic-ai extra was wrong: haiku.rag-slim defines
anthropic, google, groq, mistral, bedrock and vertexai itself.

configuration/storage.md opens with the four operational constraints, which
were either buried in an S3 section or undocumented: one writer per URI,
reader lag by read_consistency_interval_seconds, migrate after a
schema-changing upgrade, and the fixed embedding dimension with what
ConfigMismatchError means and which rebuild mode resolves it.

The one-writer rule is stated as a haiku.rag constraint, which is what it is:
the multi-table lock, version snapshot and rollback are process-local, so a
second writer can commit inside another's transaction and be reverted by its
rollback. storage.md and ingester.md both claimed it was a LanceDB property
that corrupts manifests. The S3 deployment section now links to the
constraint instead of restating it.

Get started reads index, Quickstart, Installation, Architecture. The landing
page's list was missing Installation.
2026-08-20 15:07:06 +03:00

3.9 KiB

Installation

Choose Your Package

haiku.rag is available in two packages:

uv pip install haiku.rag

The full package pulls the docling, voyageai, cohere, zeroentropy, cross-encoder, jina and tui extras. It does not include s3 or ingester:

uv pip install 'haiku.rag[ingester]'   # the haiku-ingester service
uv pip install 'haiku.rag[s3]'         # S3 and object storage

Slim Package (Minimal Dependencies)

uv pip install haiku.rag-slim
uv pip install 'haiku.rag-slim[docling]'
uv pip install 'haiku.rag-slim[docling,voyageai,cross-encoder]'

Extras

Every extra haiku.rag-slim defines. The right-hand column marks the ones the full haiku.rag package already includes.

Extra Provides In haiku.rag
docling PDF, DOCX, PPTX, images and 40+ formats, converted locally yes
tui Terminal UI for chat and inspect yes
voyageai VoyageAI embeddings yes
cohere Cohere embeddings and reranking yes
zeroentropy Zero Entropy reranking yes
cross-encoder Local reranking via sentence-transformers yes
jina Local Jina reranking (provider: jina-local) yes
s3 S3 and object-storage access no
ingester The haiku-ingester service (also pulls s3) no
anthropic Anthropic Claude models no
google Google Gemini models no
groq Groq models no
mistral Mistral models no
bedrock AWS Bedrock models no
vertexai Google Vertex AI models no

Ollama and any OpenAI-compatible endpoint work with no extra at all.

Built-in providers (no extras needed):

  • Ollama (default embedding provider)
  • OpenAI (GPT models for QA and embeddings)
  • vLLM and other OpenAI-compatible endpoints (embeddings, QA, reranking)
  • Jina reranking via provider: jina, which calls the Jina HTTP API

Other providers come from the extras above, which pull the matching Pydantic AI extra. For Claude models, uv pip install 'haiku.rag-slim[anthropic]'.

See Configuration for configuring providers including advanced options like vLLM.

Requirements

  • Python 3.12+
  • Ollama (for default embeddings and QA)

Pre-download Models (Optional)

You can prefetch all required runtime models before first use:

haiku-rag download-models

This will download:

  • Docling models for document processing
  • HuggingFace tokenizer models for chunking
  • Any Ollama models referenced by your current configuration

Remote Processing (Optional)

When using haiku.rag-slim, you can skip installing the docling extra and instead use docling-serve for remote document processing. This is useful for:

  • Keeping dependencies minimal
  • Offloading heavy document processing to a dedicated service
  • Production deployments with separate processing infrastructure

See Remote processing for setup instructions and Document Processing for configuration options.

Docker

Only the slim image is published. Build the full image yourself:

Slim Image (Minimal)

Pre-built slim image with minimal dependencies - use with external docling-serve for document processing:

docker pull ghcr.io/ggozad/haiku.rag-slim:latest

See examples/docker/docker-compose.yml for a complete setup with docling-serve.

Full Image (Self-contained)

Build locally to include all features and document processing without docling-serve:

docker build -f docker/Dockerfile -t haiku-rag .
docker run -p 8001:8001 \
  -v /path/to/haiku.rag.yaml:/app/haiku.rag.yaml \
  -v /path/to/data:/data \
  haiku-rag

See docker/README.md for complete build and configuration instructions, including how to run the ingester service for continuous document ingestion.