At 327 lines it interleaved reading the tables, deriving the lookups every check needs, and the bodies of ten checks. Six checks were already functions; the rest were inline, so none of them could be read or tested without the others around them. Each one is now a function taking exactly what it needs: _check_document_meta_parity, _check_orphaned_chunks, _check_orphaned_items, _check_documents_without_items, _check_dangling_item_refs, _check_vector_dimension, _check_unembedded_chunks, _check_picture_data, _check_settings_row and _check_pending_migrations. run_db_checks reads the tables, then appends results. _document_centroids takes the vector reduction. Passing the matrix as a parameter keeps it a local of run_db_checks, so the del before clustering still drops the last reference — measured at 6.2 MB allocated to reduce a 102 MB matrix, no second copy. There is no snapshot object: one holding vectors would keep the largest allocation alive past the del. The reduction also rebound doc_ids from the document-id set to the centroid id list halfway through the function. The centroid ids have their own name now. No test changes: the 80 doctor tests cover these through run_db_checks and pass unchanged. |
||
|---|---|---|
| .. | ||
| haiku/rag | ||
| LICENSE | ||
| pyproject.toml | ||
| README.md | ||
haiku.rag-slim
Opinionated agentic RAG powered by LanceDB, Pydantic AI, and Docling - Core package with minimal dependencies.
haiku.rag-slim is the core package for users who want to install only the dependencies they need. Document processing (docling), and reranker support are all optional extras.
For most users, we recommend installing haiku.rag instead, which includes all features out of the box.
Installation
Python 3.12 or newer required
Minimal Installation
uv pip install haiku.rag-slim
Core functionality with OpenAI/Ollama support, MCP server, and Logfire observability. Document processing (docling) is optional.
With Document Processing
uv pip install haiku.rag-slim[docling]
Adds support for 40+ file formats including PDF, DOCX, HTML, and more.
Available Extras
Document Processing:
docling- PDF, DOCX, HTML, and 40+ file formats
Embedding Providers:
voyageai- VoyageAI embeddings
Rerankers:
cross-encoder- Local reranking via sentence-transformerscohere- Coherezeroentropy- Zero Entropy
Model Providers:
- OpenAI/Ollama - included in core (OpenAI-compatible APIs)
anthropic- Anthropic Claudegroq- Groqgoogle- Google Geminimistral- Mistral AIbedrock- AWS Bedrockvertexai- Google Vertex AI
# Common combinations
uv pip install haiku.rag-slim[docling,anthropic,cross-encoder]
uv pip install haiku.rag-slim[docling,groq]
Usage
See the main haiku.rag repository for:
- Quick start guide
- CLI examples
- Python API usage
- MCP server setup
Documentation
Full documentation: https://ggozad.github.io/haiku.rag/
- Installation - Provider setup
- Configuration - YAML configuration
- CLI - Command reference
- Python API - Complete API docs