Sixty-three comments said what the next statement already said: # Connect to LanceDB above connect_lancedb, # Path object above isinstance(source, Path), # Get page numbers from provenance above the prov loop, # Clear and populate results above list_view.clear(). They cost a read and carry nothing. The line is whether a comment restates one statement or labels a phase. Phase labels stay: the migrations keep # Create staging table with new schema and # Copy from staging to final table in batches, each heading ten lines of a long procedure. So do comments carrying a fact the code cannot: the merge_insert update-only note on document_meta, why the poller builds sources eagerly, why create_document_from_source returns a list for directories, that indexes need training data, the field-group markers in the config models, and the file:// URL-encoding note in create_document_from_source. capabilities/ is untouched. Its docstrings sit next to prompt surface, and changing them needs an eval to back it. The cassette-recording docs were wrong three ways. They named tests/test_qa.py::test_qa_anthropic, which no longer exists; they targeted whole modules, so a rewrite would re-record cassettes for services the recorder is not running; and they used COHERE_API_KEY where the SDK reads CO_API_KEY. docs/development.md now names exact tests with -n0, and the keyed example is test_cohere_reranker, which owns the one cassette recording api.cohere.com. |
||
|---|---|---|
| .claude/skills | ||
| .github | ||
| app | ||
| docker | ||
| docs | ||
| evaluations | ||
| examples | ||
| haiku_rag_slim | ||
| overrides | ||
| scripts | ||
| tests | ||
| .dockerignore | ||
| .gitignore | ||
| .pre-commit-config.yaml | ||
| .python-version | ||
| CHANGELOG.md | ||
| LICENSE | ||
| pyproject.toml | ||
| README.md | ||
| server.json | ||
| uv.lock | ||
| zensical.toml | ||
haiku.rag
Agentic RAG that answers questions about your own documents with citations to page numbers and section headings. Runs locally on an embedded database, no server required.
Built on LanceDB, Pydantic AI, and Docling. Full documentation at ggozad.github.io/haiku.rag.
New: vision and multimodal search. Picture-aware ingestion captures embedded figure bytes; vision-capable QA models receive them alongside text. Multimodal embedders put picture vectors in the same space as text, enabling text-as-query → figure hits and image-as-query retrieval.
Features
- Hybrid search — Vector + full-text with Reciprocal Rank Fusion
- Multimodal & cross-modal search — Multimodal embedders (vLLM, VoyageAI, Cohere) put picture vectors in the same space as text; supports text-as-query → figure hits and image-as-query
- Question answering — RAG capability with citations (page numbers, section headings)
- Vision QA — Vision-capable models receive figure bytes alongside chunk text; attach your own images to questions in
ask,analyze, MCP, and the chat TUI - Reranking — local cross-encoders, Cohere, Zero Entropy, or vLLM
- Analysis capability — Complex analytical tasks via sandboxed Python code execution (aggregation, computation, multi-document analysis)
- Evidence compaction — Optional capability that replaces earlier questions' search results on the request with the evidence they cited, so long conversations stop resending everything they retrieved
- Citation policy — Optional capability that requires every answer to declare what grounds it, including declaring that nothing does
- Conversational RAG — Chat TUI and web application for multi-turn conversations with session memory
- Document structure — Stores full DoclingDocument, enabling structure-aware context expansion
- Multiple providers — Embeddings: Ollama, OpenAI, VoyageAI, Cohere, LM Studio, vLLM (multimodal via
multimodal: trueon vLLM/VoyageAI/Cohere). QA: any model supported by Pydantic AI - Local-first — Embedded LanceDB, no servers required. Also supports S3, GCS, Azure, and LanceDB Cloud
- CLI & Python API — Full functionality from command line or code
- MCP server — Expose as tools for AI assistants (Claude Desktop, etc.)
- Visual grounding — View chunks highlighted on original page images
- Production ingester — Long-lived
haiku-ingesterservice with persistent SQLite queue, async worker pool with retries and a dead-letter queue, FS / HTTP / S3 / WebDAV source adapters, FastAPI control plane, and a browser dashboard for operators. See docs/ingester.md. - Tags — Name database states with
haiku-rag tagand roll back to them - Inspector — TUI for browsing documents, chunks, and search results
Installation
Python 3.12 or newer required
Full Package (Recommended)
pip install haiku.rag
Includes all features: document processing, all embedding providers, and rerankers.
Using uv? uv pip install haiku.rag
Slim Package (Minimal Dependencies)
pip install haiku.rag-slim
Install only the extras you need. See the Installation documentation for available options.
Quick Start
Note
: Requires an embedding provider (Ollama, OpenAI, etc.). See the Tutorial for setup instructions.
# Index a PDF
haiku-rag add-src paper.pdf
# Search
haiku-rag search "attention mechanism"
# Ask questions with citations
haiku-rag ask "What datasets were used for evaluation?"
# Ask about an image (vision-capable model)
haiku-rag ask "Does this figure match the spec in the design doc?" --image figure.png
# Analyze — complex analytical tasks via code execution
haiku-rag analyze "How many documents mention transformers?"
# Interactive chat — multi-turn conversations with memory
haiku-rag chat
# Continuously ingest from configured sources (FS, HTTP, S3, WebDAV)
haiku-ingester serve
See Configuration for customization options.
Python API
from haiku.rag.client import HaikuRAG
async with HaikuRAG("knowledge.lancedb", create=True) as rag:
# Index documents
await rag.create_document_from_source("paper.pdf")
await rag.create_document_from_source("https://arxiv.org/pdf/1706.03762")
# Search — returns chunks with provenance
results = await rag.search("self-attention")
for result in results:
print(f"{result.score:.2f} | p.{result.page_numbers} | {result.content[:100]}")
# QA with citations
answer, citations = await rag.ask("What is the complexity of self-attention?")
print(answer)
for cite in citations:
print(f" [{cite.chunk_id}] p.{cite.page_numbers}: {cite.content[:80]}")
For direct agent composition, see the capabilities documentation.
MCP Server
Use with AI assistants like Claude Desktop:
haiku-rag mcp --stdio
Add to your Claude Desktop configuration:
{
"mcpServers": {
"haiku-rag": {
"command": "haiku-rag",
"args": ["mcp", "--stdio"]
}
}
}
Provides tools for document management, search, QA, and analysis directly in your AI assistant.
Examples
See the examples directory for working examples:
- Docker Setup - Complete Docker deployment with continuous ingestion (
haiku-ingester) and MCP server - Web Application - Full-stack conversational RAG with CopilotKit frontend
Documentation
Full documentation at: https://ggozad.github.io/haiku.rag/
- Quickstart - Provider setup and first ingestion
- Installation - Packages and extras
- Configuration - YAML reference
- CLI - Command reference
- Python API - Complete API docs
- Capabilities - Native Pydantic AI RAG and analysis capabilities
- Tuning - Retrieval and answer-quality tuning
- Ingester - Production ingester for continuous indexing from FS, HTTP, S3, and WebDAV
- MCP - Model Context Protocol integration
- Remote processing - Offload conversion to docling-serve
- Applications - Chat TUI, web app, and inspector
- Benchmarks - Performance benchmarks
- Changelog - Version history
License
This project is licensed under the MIT License.
mcp-name: io.github.ggozad/haiku-rag