The VFS bridge suspends the Monty worker for the length of a read. Monty checks its duration budget between interpreter steps, so it cannot check while a read is in flight. Code that reads in a loop overran a 60s budget by minutes. A read takes about 20ms on a 2789-document corpus, so a full scan spends about 55s in reads alone. Check the deadline before each read. Raising from inside the callback answers the worker's suspension, which keeps the session usable. Monty also spends max_duration_secs across the session rather than per call, and the sandbox reuses the session so that variables persist. Budget it for code_timeout * max_executions. At the old per-call value the first slow call starved every later one. Do not wrap feed_run in asyncio.wait_for. Cancelling during pure compute is clean, but cancelling while a read waits for an answer wedges the session with a protocol RuntimeError that escapes execute(). A call that computes without reading stays bounded by the session budget alone. |
||
|---|---|---|
| .. | ||
| haiku/rag | ||
| LICENSE | ||
| pyproject.toml | ||
| README.md | ||
haiku.rag-slim
Opinionated agentic RAG powered by LanceDB, Pydantic AI, and Docling - Core package with minimal dependencies.
haiku.rag-slim is the core package for users who want to install only the dependencies they need. Document processing (docling), and reranker support are all optional extras.
For most users, we recommend installing haiku.rag instead, which includes all features out of the box.
Installation
Python 3.12 or newer required
Minimal Installation
uv pip install haiku.rag-slim
Core functionality with OpenAI/Ollama support, MCP server, and Logfire observability. Document processing (docling) is optional.
With Document Processing
uv pip install haiku.rag-slim[docling]
Adds support for 40+ file formats including PDF, DOCX, HTML, and more.
Available Extras
Document Processing:
docling- PDF, DOCX, HTML, and 40+ file formats
Embedding Providers:
voyageai- VoyageAI embeddings
Rerankers:
cross-encoder- Local reranking via sentence-transformerscohere- Coherezeroentropy- Zero Entropy
Model Providers:
- OpenAI/Ollama - included in core (OpenAI-compatible APIs)
anthropic- Anthropic Claudegroq- Groqgoogle- Google Geminimistral- Mistral AIbedrock- AWS Bedrockvertexai- Google Vertex AI
# Common combinations
uv pip install haiku.rag-slim[docling,anthropic,cross-encoder]
uv pip install haiku.rag-slim[docling,groq]
Usage
See the main haiku.rag repository for:
- Quick start guide
- CLI examples
- Python API usage
- MCP server setup
Documentation
Full documentation: https://ggozad.github.io/haiku.rag/
- Installation - Provider setup
- Configuration - YAML configuration
- CLI - Command reference
- Python API - Complete API docs