No description
Find a file
Chris McDonough da8e7dc568 Fix run_batch hanging forever when all workers die
The drain loop in run_batch() polls counts_by_status() waiting for
queued and claimed counts to reach zero. If all worker tasks crash
(unhandled exception, OOM), claimed jobs stay claimed forever and
the loop never exits — the CLI command hangs.

Check live_workers during the drain loop. If claimed jobs exist but
no workers are alive to process them, log an error and break out.
The stranded jobs will be reaped on the next start.
2026-06-01 07:52:17 -04:00
.github Straight migration to zensical 2026-05-20 15:44:39 +03:00
app vb 2026-05-29 11:47:12 +03:00
docker Docker & docs updates 2026-05-26 11:44:45 +03:00
docs Results for analysis skill 2026-06-01 10:40:52 +03:00
evaluations Remove unecessary Mean Reciprocal Rank metric 2026-06-01 10:40:52 +03:00
examples Adapt docker compose example to use the published image 2026-05-27 12:59:32 +03:00
haiku_rag_slim Fix run_batch hanging forever when all workers die 2026-06-01 07:52:17 -04:00
overrides Landing page hero with TUI screen recording 2026-05-20 16:56:25 +03:00
scripts add SeaweedFS integration tests for S3Watcher 2026-05-11 11:28:40 +03:00
tests Fix run_batch hanging forever when all workers die 2026-06-01 07:52:17 -04:00
.dockerignore Update docker build 2025-11-04 18:43:41 +02:00
.gitignore fix README links and stale MkDocs reference 2026-05-20 17:00:40 +03:00
.pre-commit-config.yaml Fix precommit to use uv installed ruff 2026-03-12 12:05:25 +02:00
.python-version Use 3.13 for development 2025-10-09 10:18:10 +03:00
CHANGELOG.md Add haiku-ingester run-batch, remove run-once 2026-05-29 17:00:59 +03:00
LICENSE MIT license 2025-06-18 10:17:27 +02:00
pyproject.toml Remove repliqa from test setup 2026-06-01 10:40:52 +03:00
README.md Docs update 2026-05-27 13:22:35 +03:00
server.json Update mcp registry schema 2025-12-19 12:39:39 +02:00
uv.lock Remove repliqa from test setup 2026-06-01 10:40:52 +03:00
zensical.toml Docs update 2026-05-27 13:22:35 +03:00

Haiku RAG

Tests codecov

Agentic RAG built on LanceDB, Pydantic AI, and Docling.

New: vision and multimodal search. Picture-aware ingestion captures embedded figure bytes; vision-capable QA models receive them alongside text. Multimodal embedders put picture vectors in the same space as text, enabling text-as-query → figure hits and image-as-query retrieval.

Features

  • Hybrid search — Vector + full-text with Reciprocal Rank Fusion
  • Multimodal & cross-modal search — Multimodal embedders (vLLM) put picture vectors in the same space as text; supports text-as-query → figure hits and image-as-query
  • Question answering — RAG skill with citations (page numbers, section headings)
  • Vision QA — Vision-capable models receive figure bytes alongside chunk text
  • Reranking — MxBAI, Cohere, Zero Entropy, or vLLM
  • Analysis skill — Complex analytical tasks via sandboxed Python code execution (aggregation, computation, multi-document analysis)
  • Conversational RAG — Chat TUI and web application for multi-turn conversations with session memory
  • Document structure — Stores full DoclingDocument, enabling structure-aware context expansion
  • Multiple providers — Embeddings: Ollama, OpenAI, VoyageAI, LM Studio, vLLM (multimodal). QA: any model supported by Pydantic AI
  • Local-first — Embedded LanceDB, no servers required. Also supports S3, GCS, Azure, and LanceDB Cloud
  • CLI & Python API — Full functionality from command line or code
  • MCP server — Expose as tools for AI assistants (Claude Desktop, etc.)
  • Visual grounding — View chunks highlighted on original page images
  • Production ingester — Long-lived haiku-ingester service with persistent SQLite queue, async worker pool with retries and a dead-letter queue, FS / HTTP / S3 / WebDAV source adapters, FastAPI control plane, and a browser dashboard for operators. See docs/ingester.md.
  • Time travel — Query the database at any historical point with --before
  • Inspector — TUI for browsing documents, chunks, and search results

Installation

Python 3.12 or newer required

pip install haiku.rag

Includes all features: document processing, all embedding providers, and rerankers.

Using uv? uv pip install haiku.rag

Slim Package (Minimal Dependencies)

pip install haiku.rag-slim

Install only the extras you need. See the Installation documentation for available options.

Quick Start

Note

: Requires an embedding provider (Ollama, OpenAI, etc.). See the Tutorial for setup instructions.

# Index a PDF
haiku-rag add-src paper.pdf

# Search
haiku-rag search "attention mechanism"

# Ask questions with citations
haiku-rag ask "What datasets were used for evaluation?"

# Analyze — complex analytical tasks via code execution
haiku-rag analyze "How many documents mention transformers?"

# Interactive chat — multi-turn conversations with memory
haiku-rag chat

# Continuously ingest from configured sources (FS, HTTP, S3, WebDAV)
haiku-ingester serve

See Configuration for customization options.

Python API

from haiku.rag.client import HaikuRAG

async with HaikuRAG("knowledge.lancedb", create=True) as rag:
    # Index documents
    await rag.create_document_from_source("paper.pdf")
    await rag.create_document_from_source("https://arxiv.org/pdf/1706.03762")

    # Search — returns chunks with provenance
    results = await rag.search("self-attention")
    for result in results:
        print(f"{result.score:.2f} | p.{result.page_numbers} | {result.content[:100]}")

    # QA with citations
    answer, citations = await rag.ask("What is the complexity of self-attention?")
    print(answer)
    for cite in citations:
        print(f"  [{cite.chunk_id}] p.{cite.page_numbers}: {cite.content[:80]}")

For details on the skills the client wraps, see the Skills docs.

MCP Server

Use with AI assistants like Claude Desktop:

haiku-rag mcp --stdio

Add to your Claude Desktop configuration:

{
  "mcpServers": {
    "haiku-rag": {
      "command": "haiku-rag",
      "args": ["mcp", "--stdio"]
    }
  }
}

Provides tools for document management, search, QA, and analysis directly in your AI assistant.

Examples

See the examples directory for working examples:

  • Docker Setup - Complete Docker deployment with continuous ingestion (haiku-ingester) and MCP server
  • Web Application - Full-stack conversational RAG with CopilotKit frontend

Documentation

Full documentation at: https://ggozad.github.io/haiku.rag/

License

This project is licensed under the MIT License.

mcp-name: io.github.ggozad/haiku-rag