No description
Find a file
Yiorgis Gozadinos d97aa15af9
Keep footnotes and matched items in expanded context
_build_result applied the noise-label filter to every item in the range,
including the ones the result matched on. A hit on a footnote or index
entry returned its section with the matched text removed, the clip anchor
could not find the evidence and fell back to a prefix window, and the
un-merge path rebuilt through the same filter.

Noise is now a set of positions computed once per group by
_noise_positions: noise-labelled items minus the matched ones.
_expand_outward and _build_result take that set instead of a flag, so the
matched item is kept in content and counted toward the budget.

Footnotes leave the noise set. They carry sources, cross-references and
clarifications, and docling attaches table and figure footnotes to the
table itself, so the filter was dropping part of the table. The noise set
is page_header, page_footer and document_index.

Refs #609
2026-09-07 13:52:14 +03:00
.agents/plugins Package the plugin for Codex from the same bundle 2026-09-07 11:21:25 +03:00
.claude/skills Document the Logfire HTTP query path in the eval-debugging skill 2026-08-24 00:08:59 +03:00
.claude-plugin Package the plugin for Codex from the same bundle 2026-09-07 11:21:25 +03:00
.github Resolve file:// URIs to paths through url2pathname 2026-08-21 10:22:26 +03:00
app vb 2026-09-03 18:00:43 +03:00
docker Remove the MCP write tools; the server opens read-only 2026-09-04 12:59:48 +03:00
docs Keep footnotes and matched items in expanded context 2026-09-07 13:52:14 +03:00
evaluations Default to ollama:qwen3.8 2026-09-04 12:36:36 +03:00
examples Remove the MCP write tools; the server opens read-only 2026-09-04 12:59:48 +03:00
haiku_rag_slim Keep footnotes and matched items in expanded context 2026-09-07 13:52:14 +03:00
overrides Fix the MCP registry entry and fill in package and docs metadata 2026-08-18 14:37:05 +03:00
plugins/haiku-rag Align the MCP guidance with the analysis instructions and move to Monty 0.0.23 2026-09-07 12:02:33 +03:00
scripts Package the plugin for Codex from the same bundle 2026-09-07 11:21:25 +03:00
tests Keep footnotes and matched items in expanded context 2026-09-07 13:52:14 +03:00
.dockerignore Update docker build 2025-11-04 18:43:41 +02:00
.gitignore Add Logfire debugging skills and worker-breaker event 2026-07-10 13:23:17 +03:00
.pre-commit-config.yaml Fix precommit to use uv installed ruff 2026-03-12 12:05:25 +02:00
.python-version Use 3.13 for development 2025-10-09 10:18:10 +03:00
CHANGELOG.md Keep footnotes and matched items in expanded context 2026-09-07 13:52:14 +03:00
LICENSE MIT license 2025-06-18 10:17:27 +02:00
pyproject.toml vb 2026-09-03 18:00:43 +03:00
README.md Package the plugin for Codex from the same bundle 2026-09-07 11:21:25 +03:00
server.json Fix the MCP registry entry and fill in package and docs metadata 2026-08-18 14:37:05 +03:00
uv.lock Align the MCP guidance with the analysis instructions and move to Monty 0.0.23 2026-09-07 12:02:33 +03:00
zensical.toml Give the docs an architecture page and one extras list 2026-08-20 15:07:06 +03:00

haiku.rag

PyPI Python Downloads Docs Tests codecov

Agentic RAG that answers questions about your own documents with citations to page numbers and section headings. Runs locally on an embedded database, no server required.

Built on LanceDB, Pydantic AI, and Docling. Full documentation at ggozad.github.io/haiku.rag.

Features

  • Hybrid search — Vector + full-text with Reciprocal Rank Fusion
  • Multimodal & cross-modal search — Multimodal embedders (vLLM, VoyageAI, Cohere) put picture vectors in the same space as text; supports text-as-query → figure hits and image-as-query
  • Question answering — RAG capability with citations (page numbers, section headings)
  • Vision QA — Vision-capable models receive figure bytes alongside chunk text; attach your own images to questions in ask, analyze and the chat TUI
  • Reranking — local cross-encoders, Cohere, Zero Entropy, or vLLM
  • Analysis capability — Complex analytical tasks via sandboxed Python code execution (aggregation, computation, multi-document analysis)
  • Evidence compaction — Optional capability that replaces earlier questions' search results on the request with the evidence they cited, so long conversations stop resending everything they retrieved
  • Citation policy — Optional capability that requires every answer to declare what grounds it, including declaring that nothing does
  • Conversational RAG — Chat TUI and web application for multi-turn conversations with session memory
  • Document structure — Stores full DoclingDocument, enabling structure-aware context expansion
  • Multiple providers — Embeddings: Ollama, OpenAI, VoyageAI, Cohere, LM Studio, vLLM (multimodal via multimodal: true on vLLM/VoyageAI/Cohere). QA: any model supported by Pydantic AI
  • Multi-database search — Search, ask, analyze, or chat across named databases with source attribution on results and citations
  • Local-first — Embedded LanceDB, no servers required. Also supports S3, GCS, Azure, and LanceDB Cloud
  • CLI & Python API — Full functionality from command line or code
  • MCP server — Expose as tools for AI assistants (Claude Desktop, etc.)
  • Visual grounding — View chunks highlighted on original page images
  • Production ingester — Long-lived haiku-ingester service with persistent SQLite queue, async worker pool with retries and a dead-letter queue, FS / HTTP / S3 / WebDAV source adapters, FastAPI control plane, and a browser dashboard for operators. See docs/ingester.md.
  • Tags — Name database states with haiku-rag tag and roll back to them
  • Inspector — TUI for browsing documents, chunks, and search results

Installation

Python 3.12 or newer required

pip install haiku.rag

Includes all features: document processing, all embedding providers, and rerankers.

Using uv? uv pip install haiku.rag

Slim Package (Minimal Dependencies)

pip install haiku.rag-slim

Install only the extras you need. See the Installation documentation for available options.

Quick Start

Note

: Requires an embedding provider (Ollama, OpenAI, etc.). See the Tutorial for setup instructions.

# Index a PDF
haiku-rag add-src paper.pdf

# Search
haiku-rag search "attention mechanism"

# Ask questions with citations
haiku-rag ask "What datasets were used for evaluation?"

# Ask about an image (vision-capable model)
haiku-rag ask "Does this figure match the spec in the design doc?" --image figure.png

# Analyze — complex analytical tasks via code execution
haiku-rag analyze "How many documents mention transformers?"

# Interactive chat — multi-turn conversations with memory
haiku-rag chat

# Continuously ingest from configured sources (FS, HTTP, S3, WebDAV)
haiku-ingester serve

See Configuration for customization options.

Python API

from haiku.rag.client import HaikuRAG

async with HaikuRAG("knowledge.lancedb", create=True) as rag:
    # Index documents
    await rag.create_document_from_source("paper.pdf")
    await rag.create_document_from_source("https://arxiv.org/pdf/1706.03762")

    # Search — returns chunks with provenance
    results = await rag.search("self-attention")
    for result in results:
        print(f"{result.score:.2f} | p.{result.page_numbers} | {result.content[:100]}")

    # QA with citations
    answer, citations = await rag.ask("What is the complexity of self-attention?")
    print(answer)
    for cite in citations:
        print(f"  [{cite.chunk_id}] p.{cite.page_numbers}: {cite.content[:80]}")

For direct agent composition, see the capabilities documentation.

MCP Server

Use with AI assistants like Claude Code, Codex, and Claude Desktop:

haiku-rag mcp --stdio

In Claude Code, install the plugin, which registers the server and a skill:

claude plugin marketplace add ggozad/haiku.rag
claude plugin install haiku-rag

In Codex, install the same plugin from its marketplace:

codex plugin marketplace add ggozad/haiku.rag
codex plugin add haiku-rag@haiku-rag

Add to your Claude Desktop configuration:

{
  "mcpServers": {
    "haiku-rag": {
      "command": "haiku-rag",
      "args": ["mcp", "--stdio"]
    }
  }
}

Provides search, document reading, and analysis tools directly in your AI assistant.

Examples

See the examples directory for working examples:

  • Docker Setup - Complete Docker deployment with continuous ingestion (haiku-ingester) and MCP server
  • Web Application - Full-stack conversational RAG with CopilotKit frontend

Documentation

Full documentation at: https://ggozad.github.io/haiku.rag/

License

This project is licensed under the MIT License.

mcp-name: io.github.ggozad/haiku-rag