No description
| src/haiku/rag | ||
| tests | ||
| .gitignore | ||
| .pre-commit-config.yaml | ||
| .python-version | ||
| pyproject.toml | ||
| README.md | ||
| uv.lock | ||
Haiku SQLite RAG
A SQLite-based Retrieval-Augmented Generation (RAG) system built for efficient document storage, chunking, and hybrid search capabilities.
Features
- Document Management: Store and manage documents with automatic content parsing
- Smart Updates: Intelligent file/URL monitoring with MD5-based change detection
- Hybrid Search: Full-text search (FTS5) combined with vector embeddings
- Multi-format Support: Parse 40+ file formats including PDF, DOCX, HTML, Markdown, and more
- Web Content: Direct URL ingestion with automatic content type detection
- Vector Embeddings: Uses sqlite-vec for efficient similarity search
- Automatic Chunking: Intelligent document segmentation for better retrieval
Installation
uv pip install haiku.rag
or for development, checkout the repository and then,
# Install dependencies
uv sync
# Activate virtual environment
source .venv/bin/activate
Quick Start
from pathlib import Path
from haiku.rag.client import HaikuRAG
# Initialize client with database path
client = HaikuRAG("path/to/database.db")
# Or use in-memory database for testing
client = HaikuRAG(":memory:")
# Create document from text
doc = await client.create_document(
content="Your document content here",
uri="doc://example",
metadata={"source": "manual", "topic": "example"}
)
# Create document from file (auto-parses content)
doc = await client.create_document_from_source("path/to/document.pdf")
# Create document from URL
doc = await client.create_document_from_source("https://example.com/article.html")
# Retrieve documents
doc = await client.get_document_by_id(1)
doc = await client.get_document_by_uri("file:///path/to/document.pdf")
# List all documents with pagination
docs = await client.list_documents(limit=10, offset=0)
# Update document content
doc.content = "Updated content"
await client.update_document(doc)
# Delete document
await client.delete_document(doc.id)
# Clean up
client.close()
Smart Document Updates
The system automatically tracks file changes using MD5 hashes:
# First call - creates new document
doc1 = await client.create_document_from_source("document.txt")
# Second call - no changes, returns existing document (no processing)
doc2 = await client.create_document_from_source("document.txt")
assert doc1.id == doc2.id
# After file modification - automatically updates existing document
# File content changed...
doc3 = await client.create_document_from_source("document.txt")
assert doc1.id == doc3.id # Same document
assert doc3.content != doc1.content # Updated content
Supported File Formats
The system supports 40+ file formats through MarkItDown:
- Documents: PDF, DOCX, PPTX, XLSX
- Web: HTML, XML
- Text: TXT, MD, CSV, JSON, YAML
- Code: PY, JS, TS, C, CPP, JAVA, GO, RS, and more
- Media: MP3, WAV (transcription)
Document Metadata
Documents automatically include metadata:
doc = await client.create_document_from_source("example.pdf")
print(doc.metadata)
# {
# "contentType": "application/pdf",
# "md5": "abc123...",
# "custom_field": "value" # Your custom metadata
# }
Contributing
- Fork the repository
- Create a feature branch
- Add tests for new functionality
- Ensure all tests pass:
pytest - Run type checking:
pyright - Run linting:
ruff check - Submit a pull request