No description
Find a file
2025-06-17 12:48:26 +02:00
src/haiku/rag Add a context manager to the client 2025-06-17 12:48:26 +02:00
tests Add a context manager to the client 2025-06-17 12:48:26 +02:00
.gitignore test coverage 2025-06-16 17:32:00 +02:00
.pre-commit-config.yaml init repo 2025-06-15 13:21:17 +02:00
.python-version init repo 2025-06-15 13:21:17 +02:00
pyproject.toml Handle files/url with create_document_from_source in client 2025-06-17 11:26:02 +02:00
README.md Add a context manager to the client 2025-06-17 12:48:26 +02:00
uv.lock Handle files/url with create_document_from_source in client 2025-06-17 11:26:02 +02:00

Haiku SQLite RAG

A SQLite-based Retrieval-Augmented Generation (RAG) system built for efficient document storage, chunking, and hybrid search capabilities.

Features

  • Document Management: Store and manage documents with automatic content parsing
  • Smart Updates: Intelligent file/URL monitoring with MD5-based change detection
  • Hybrid Search: Full-text search (FTS5) combined with vector embeddings
  • Multi-format Support: Parse 40+ file formats including PDF, DOCX, HTML, Markdown, and more
  • Web Content: Direct URL ingestion with automatic content type detection
  • Vector Embeddings: Uses sqlite-vec for efficient similarity search
  • Automatic Chunking: Intelligent document segmentation for better retrieval

Installation

uv pip install haiku.rag

or for development, checkout the repository and then,

# Install dependencies
uv sync

# Activate virtual environment
source .venv/bin/activate

Quick Start

from pathlib import Path
from haiku.rag.client import HaikuRAG

# Use as async context manager (recommended)
async with HaikuRAG("path/to/database.db") as client:
    # Create document from text
    doc = await client.create_document(
        content="Your document content here",
        uri="doc://example",
        metadata={"source": "manual", "topic": "example"}
    )

    # Create document from file (auto-parses content)
    doc = await client.create_document_from_source("path/to/document.pdf")

    # Create document from URL
    doc = await client.create_document_from_source("https://example.com/article.html")

    # Retrieve documents
    doc = await client.get_document_by_id(1)
    doc = await client.get_document_by_uri("file:///path/to/document.pdf")

    # List all documents with pagination
    docs = await client.list_documents(limit=10, offset=0)

    # Update document content
    doc.content = "Updated content"
    await client.update_document(doc)

    # Delete document
    await client.delete_document(doc.id)

    # Search documents using hybrid search (vector + full-text)
    results = await client.search("machine learning algorithms", limit=5)
    for chunk, score in results:
        print(f"Score: {score:.3f}")
        print(f"Content: {chunk.content}")
        print(f"Document ID: {chunk.document_id}")
        print("---")


# Or use without the context manager.
client = HaikuRAG(":memory:")
try:
    # ... operations ...
finally:
    client.close()

Search Functionality

haiku.rag provides hybrid search combining vector similarity and full-text search:

  1. Vector Search: Uses embeddings to find semantically similar content
  2. Full-text Search: Uses SQLite FTS5 for exact keyword matching
  3. Hybrid Ranking: Combines both using Reciprocal Rank Fusion (RRF)
  4. Chunked Results: Returns relevant document chunks with scores
async with HaikuRAG("database.db") as client:
    # Basic search
    results = await client.search("your query here")

    # Search with custom parameters
    results = await client.search(
        query="machine learning",
        limit=10,  # Maximum results to return
        k=60       # RRF parameter for reciprocal rank fusion
    )

    # Process results
    for chunk, relevance_score in results:
        print(f"Relevance: {relevance_score:.3f}")
        print(f"Content: {chunk.content}")
        print(f"From document: {chunk.document_id}")

Smart Document Updates

The system automatically tracks file changes using MD5 hashes:

async with HaikuRAG("database.db") as client:
    # First call - creates new document
    doc1 = await client.create_document_from_source("document.txt")

    # Second call - no changes, returns existing document (no processing)
    doc2 = await client.create_document_from_source("document.txt")
    assert doc1.id == doc2.id

    # After file modification - automatically updates existing document
    # File content changed...
    doc3 = await client.create_document_from_source("document.txt")
    assert doc1.id == doc3.id  # Same document
    assert doc3.content != doc1.content  # Updated content

Supported File Formats

The system supports 40+ file formats through MarkItDown:

  • Documents: PDF, DOCX, PPTX, XLSX
  • Web: HTML, XML
  • Text: TXT, MD, CSV, JSON, YAML
  • Code: PY, JS, TS, C, CPP, JAVA, GO, RS, and more
  • Media: MP3, WAV (transcription)

Document Metadata

Documents automatically include metadata:

doc = await client.create_document_from_source("example.pdf")
print(doc.metadata)
# {
#   "contentType": "application/pdf",
#   "md5": "abc123...",
#   "custom_field": "value"  # Your custom metadata
# }

Contributing

  1. Fork the repository
  2. Create a feature branch
  3. Add tests for new functionality
  4. Ensure all tests pass: pytest
  5. Run type checking & linting with pyright & ruff check
  6. Submit a pull request