Both optional capabilities read what earlier questions retrieved and cited from the capability's state, so a host that carries only the message history hands every run an empty record. Compaction then replaced the earlier evidence with receipts and retained nothing, and the loss was invisible: the citations the host already displayed were still there. It now refuses when it finds evidence from an earlier question and no record of what that question cited. `state_carried` reaches the optional capabilities through discovery, so the refusal distinguishes a host that never carries state from a question that simply cited nothing. The documentation taught the pattern that breaks: the compose example is now stateful and the requirement is stated where each capability is introduced. The app's browser storage was doing exactly this, keeping only the fields the UI reads. It now persists the whole namespace map, so the citation policy's violations survive a reload as well as the evidence record. |
||
|---|---|---|
| .. | ||
| backend | ||
| frontend | ||
| .env.example | ||
| docker-compose.dev.yml | ||
| docker-compose.yml | ||
| haiku.rag.yaml.example | ||
| README.md | ||
haiku.rag Chat App
A conversational RAG interface built with CopilotKit and pydantic-ai's AG-UI protocol.
Note: An illustrative example meant as a starting point, with no authentication. The compose files bind the backend to
127.0.0.1; don't expose it to an untrusted network.
Prerequisites
- Docker and Docker Compose
- A haiku.rag database (created via the
haiku-ragCLI) - An LLM API key (Anthropic, OpenAI, or local Ollama)
Quick Start
-
Set up environment variables:
cp .env.example .env # Edit .env with your API keys and database path -
Configure the LLM and embedding models:
cp haiku.rag.yaml.example haiku.rag.yaml # Edit haiku.rag.yaml to configure your models -
Start the app:
docker compose up -d -
Open the chat interface: http://localhost:3000
Configuration
Environment Variables
| Variable | Description | Required |
|---|---|---|
DB_PATH |
Path to your haiku.rag LanceDB database | Yes |
ANTHROPIC_API_KEY |
Anthropic API key | One LLM key required |
OPENAI_API_KEY |
OpenAI API key | One LLM key required |
OLLAMA_BASE_URL |
Ollama server URL (default: http://host.docker.internal:11434) |
For local models |
LOGFIRE_TOKEN |
Pydantic Logfire token for debugging | No |
haiku.rag.yaml
Configure the LLM, embeddings, and search settings:
qa:
model:
provider: anthropic # or openai, ollama
name: claude-sonnet-4-20250514
embeddings:
model:
provider: ollama
name: nomic-embed-text
search:
limit: 10
See haiku.rag.yaml.example for all options.
Development
For local development with hot reloading:
docker compose -f docker-compose.dev.yml up -d --build
- Backend code changes reload automatically
- Frontend available at http://localhost:3000
- Backend API at http://localhost:8001
Architecture
┌─────────────────┐ ┌─────────────────┐ ┌─────────────────┐
│ Frontend │────▶│ Backend │────▶│ haiku.rag │
│ (CopilotKit) │ │ (pydantic-ai) │ │ (LanceDB) │
│ localhost:3000 │ │ localhost:8001 │ │ │
└─────────────────┘ └─────────────────┘ └─────────────────┘
Backend Endpoints
| Endpoint | Method | Description |
|---|---|---|
/v1/chat/stream |
POST | AG-UI chat streaming |
/api/documents |
GET | List documents in database |
/api/info |
GET | Database statistics |
/api/visualize/{chunk_id} |
GET | Visual grounding for chunks |
/health |
GET | Health check |
Chat Capabilities
The chat can:
- Search your documents with hybrid vector + full-text search
- Answer questions with citations from your knowledge base
- Filter by document when you ask about specific files
- Show visual grounding for PDF/image sources