No description
Find a file
Yiorgis Gozadinos a6a60963c0
Add SSE heartbeat to prevent connection timeouts
Sends SSE comment every 15 seconds during long LLM operations
to keep the connection alive and prevent body timeout errors.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-12 12:36:30 +02:00
.github Add coverage and badges 2025-12-29 14:36:01 +02:00
app Add SSE heartbeat to prevent connection timeouts 2026-01-12 12:36:30 +02:00
docker Clarify folder mounts in docker docs, check monitor paths exist in FileWatcher 2025-12-19 11:58:08 +02:00
docs Clarify docs about using a VLM with docling-serve 2026-01-08 10:57:52 +02:00
evaluations vb 2026-01-12 12:08:28 +02:00
examples vb 2026-01-12 12:08:28 +02:00
haiku_rag_slim Remove obsolete migrations 2026-01-12 12:07:49 +02:00
scripts vb 2026-01-08 10:25:25 +02:00
tests Update client serialization for compressed docling documents 2026-01-11 09:23:56 +02:00
.dockerignore Update docker build 2025-11-04 18:43:41 +02:00
.gitignore Chunkers return chunks with metadata, i.e. (refs, labels, headings, page_numbers) 2025-12-08 15:54:31 +02:00
.pre-commit-config.yaml Update dependencies 2025-12-08 15:56:22 +02:00
.python-version Use 3.13 for development 2025-10-09 10:18:10 +03:00
CHANGELOG.md vb 2026-01-12 12:08:28 +02:00
LICENSE
mkdocs.yml Move cassettes to default location, add development.md to docs 2025-12-29 14:13:25 +02:00
pyproject.toml vb 2026-01-12 12:08:28 +02:00
README.md Add coverage and badges 2025-12-29 14:36:01 +02:00
server.json Update mcp registry schema 2025-12-19 12:39:39 +02:00
uv.lock vb 2026-01-12 12:08:28 +02:00

Haiku RAG

Tests codecov

Agentic RAG built on LanceDB, Pydantic AI, and Docling.

Features

  • Hybrid search — Vector + full-text with Reciprocal Rank Fusion
  • Reranking — MxBAI, Cohere, Zero Entropy, or vLLM
  • Question answering — QA agents with citations (page numbers, section headings)
  • Research agents — Multi-agent workflows via pydantic-graph: plan, search, evaluate, synthesize
  • Document structure — Stores full DoclingDocument, enabling structure-aware context expansion
  • Visual grounding — View chunks highlighted on original page images
  • Time travel — Query the database at any historical point with --before
  • Multiple providers — Embeddings: Ollama, OpenAI, VoyageAI, LM Studio, vLLM. QA/Research: any model supported by Pydantic AI
  • Local-first — Embedded LanceDB, no servers required. Also supports S3, GCS, Azure, and LanceDB Cloud
  • MCP server — Expose as tools for AI assistants (Claude Desktop, etc.)
  • File monitoring — Watch directories and auto-index on changes
  • Inspector — TUI for browsing documents, chunks, and search results
  • CLI & Python API — Full functionality from command line or code

Installation

Python 3.12 or newer required

uv pip install haiku.rag

Includes all features: document processing, all embedding providers, and rerankers.

Slim Package (Minimal Dependencies)

uv pip install haiku.rag-slim

Install only the extras you need. See the Installation documentation for available options

Quick Start

# Index a PDF
haiku-rag add-src paper.pdf

# Search
haiku-rag search "attention mechanism"

# Ask questions with citations
haiku-rag ask "What datasets were used for evaluation?" --cite

# Deep QA — decomposes complex questions into sub-queries
haiku-rag ask "How does the proposed method compare to the baseline on MMLU?" --deep

# Research mode — iterative planning and search
haiku-rag research "What are the limitations of the approach?" --verbose

# Interactive research — human-in-the-loop with decision points
haiku-rag research "Compare the approaches discussed" --interactive

# Watch a directory for changes
haiku-rag serve --monitor

See Configuration for customization options.

Python API

from haiku.rag.client import HaikuRAG

async with HaikuRAG("research.lancedb", create=True) as rag:
    # Index documents
    await rag.create_document_from_source("paper.pdf")
    await rag.create_document_from_source("https://arxiv.org/pdf/1706.03762")

    # Search — returns chunks with provenance
    results = await rag.search("self-attention")
    for result in results:
        print(f"{result.score:.2f} | p.{result.page_numbers} | {result.content[:100]}")

    # QA with citations
    answer, citations = await rag.ask("What is the complexity of self-attention?")
    print(answer)
    for cite in citations:
        print(f"  [{cite.chunk_id}] p.{cite.page_numbers}: {cite.content[:80]}")

For research agents and streaming with AG-UI, see the Agents docs.

MCP Server

Use with AI assistants like Claude Desktop:

haiku-rag serve --mcp --stdio

Add to your Claude Desktop configuration:

{
  "mcpServers": {
    "haiku-rag": {
      "command": "haiku-rag",
      "args": ["serve", "--mcp", "--stdio"]
    }
  }
}

Provides tools for document management, search, QA, and research directly in your AI assistant.

Examples

See the examples directory for working examples:

  • Interactive Research Assistant - Full-stack research assistant with Pydantic AI and AG-UI featuring human-in-the-loop approval and real-time state synchronization
  • Docker Setup - Complete Docker deployment with file monitoring and MCP server
  • A2A Server - Self-contained A2A protocol server package with conversational agent interface

Documentation

Full documentation at: https://ggozad.github.io/haiku.rag/

mcp-name: io.github.ggozad/haiku-rag