haiku.rag/haiku_rag_slim
Yiorgis Gozadinos d97aa15af9
Keep footnotes and matched items in expanded context
_build_result applied the noise-label filter to every item in the range,
including the ones the result matched on. A hit on a footnote or index
entry returned its section with the matched text removed, the clip anchor
could not find the evidence and fell back to a prefix window, and the
un-merge path rebuilt through the same filter.

Noise is now a set of positions computed once per group by
_noise_positions: noise-labelled items minus the matched ones.
_expand_outward and _build_result take that set instead of a flag, so the
matched item is kept in content and counted toward the budget.

Footnotes leave the noise set. They carry sources, cross-references and
clarifications, and docling attaches table and figure footnotes to the
table itself, so the filter was dropping part of the table. The noise set
is page_header, page_footer and document_index.

Refs #609
2026-09-07 13:52:14 +03:00
..
haiku/rag Keep footnotes and matched items in expanded context 2026-09-07 13:52:14 +03:00
LICENSE
pyproject.toml Align the MCP guidance with the analysis instructions and move to Monty 0.0.23 2026-09-07 12:02:33 +03:00
README.md Give the docs an architecture page and one extras list 2026-08-20 15:07:06 +03:00

haiku.rag-slim

Opinionated agentic RAG powered by LanceDB, Pydantic AI, and Docling - Core package with minimal dependencies.

haiku.rag-slim is the core package for users who want to install only the dependencies they need. Document processing (docling), and reranker support are all optional extras.

For most users, we recommend installing haiku.rag instead, which includes all features out of the box.

Installation

Python 3.12 or newer required

Minimal Installation

uv pip install haiku.rag-slim

Core functionality with OpenAI/Ollama support, MCP server, and Logfire observability. Document processing (docling) is optional.

With Document Processing

uv pip install haiku.rag-slim[docling]

Adds support for 40+ file formats including PDF, DOCX, HTML, and more.

Available Extras

docling, tui, voyageai, cohere, zeroentropy, cross-encoder, jina, s3, ingester, and one per model provider: anthropic, google, groq, mistral, bedrock, vertexai. Ollama and any OpenAI-compatible endpoint need no extra.

What each provides, and which ones the full haiku.rag package already includes: Installation.

# Common combinations
uv pip install 'haiku.rag-slim[docling,anthropic,cross-encoder]'
uv pip install 'haiku.rag-slim[docling,groq]'

Usage

See the main haiku.rag repository for:

  • Quick start guide
  • CLI examples
  • Python API usage
  • MCP server setup

Documentation

Full documentation: https://ggozad.github.io/haiku.rag/