In Claude Code the client is the model, so the server no longer runs one. execute_code runs a Python program per call in the analysis sandbox over the selected documents and returns what it printed; the sandbox is created and closed per call so Monty's cumulative budget and a frozen mount never outlive a program. --no-agents goes with the two tools, and format_citations in haiku.rag.utils goes with its only caller. The sandbox exposes chunk metadata to code: chunk_meta on search results, metadata on list_documents rows and in metadata.json, and chunks.jsonl per document. A host-side failure inside a program, a document read or an in-code search raising, reaches the program by exception type only and is logged with its traceback. recovery_hint moves to haiku.rag.sandbox. Closes #604.
6.7 KiB
Model Context Protocol (MCP)
The MCP server exposes haiku.rag as MCP tools for compatible MCP clients like Claude Desktop.
Starting MCP Server
The MCP server supports Streamable HTTP and stdio transports:
# Default streamable HTTP transport on 127.0.0.1:8001
haiku-rag mcp
# Custom port
haiku-rag mcp --port 9000
# Bind to all interfaces (e.g. inside a container)
haiku-rag mcp --host 0.0.0.0 --port 8001
# stdio transport (for Claude Desktop)
haiku-rag mcp --stdio
--host defaults to 127.0.0.1 (loopback only). Bind to 0.0.0.0 only
when you want the MCP server reachable from outside the local machine —
e.g. inside a Docker container with port mapping, or on a trusted LAN.
The server opens the database read-only. Ingestion goes through the CLI
(haiku-rag add, add-src, delete) or haiku-ingester.
Collections
With several databases in lancedb.databases, the server covers all of
them, as haiku-rag search does. Results, documents and citations name
theirs in source. sources on the search and question tools restricts a
call to a subset; source on get_document names the database holding the
document. A name the server does not cover is an error.
haiku-rag --db-name NAME mcp serves one. See
Multiple Databases.
Claude Code
The repository ships a plugin that registers the server and a skill telling Claude when and how to use it:
claude plugin marketplace add ggozad/haiku.rag
claude plugin install haiku-rag
The plugin runs haiku-rag mcp --stdio, so haiku-rag must be on the PATH
and the configuration decides the database. The skill pre-approves every tool
and is also invocable as /haiku-rag. To register the server without the
plugin:
claude mcp add haiku-rag -- haiku-rag mcp --stdio
The skill works with that registration too: copy claude-plugin/skills/haiku-rag
into ~/.claude/skills/ and change the tool prefix in its allowed-tools from
mcp__plugin_haiku-rag_haiku-rag__ to mcp__haiku-rag__.
Claude Desktop Integration
Add to your Claude Desktop configuration (claude_desktop_config.json):
{
"mcpServers": {
"haiku-rag": {
"command": "haiku-rag",
"args": ["mcp", "--stdio"]
}
}
}
With a custom database path:
{
"mcpServers": {
"haiku-rag": {
"command": "haiku-rag",
"args": ["mcp", "--stdio", "--db", "/path/to/database.lancedb"]
}
}
}
After restarting Claude Desktop, you can ask Claude to search your documents or answer questions using your knowledge base.
Tools
Every tool is read-only and says so in its annotations. Each parameter carries a description in the tool schema, so the listing below names them without repeating it.
| Tool | Registered | Parameters |
|---|---|---|
search_documents |
always | query, limit, include_images, filter, sources |
search_documents_by_image |
multimodal embedder only | image_base64, limit, include_images, filter, sources |
get_document |
always | document_id, source |
get_document_outline |
always | document_id, source |
get_document_section |
always | document_id, section_id, source |
list_documents |
always | limit, offset, filter |
execute_code |
always | code, filter, sources |
search_documents runs hybrid search, vector and full-text. Its text content
is the rendering the in-process agents read: results best first, each with its
rank, Document ID, Collection when the server covers several, the document
title, section headings, the matched chunk's metadata when it has any, and the
passage expanded to its section the way the agents get it
(search.max_context_chars caps it). Pictures in the results follow as
image blocks, one per distinct picture, each preceded by a line naming its
result; include_images: false leaves them out. Search results carry no
structured content, so every client shows the model the same text and
images. Scores are not comparable across
queries or search types, so rank is the signal. search_documents_by_image
embeds the query image and searches by vector similarity alone.
get_document returns a document whole, in reading order. For a long one,
get_document_outline returns the heading tree with page numbers and
get_document_section the text of one section, subsections included; a
node's id in the outline is the section_id. A document without headings
has an empty outline. list_documents returns titles, URIs and metadata,
which is how a client learns what a filter can match.
execute_code runs a Python program in the sandbox of the
analysis capability, over the documents filter
and sources select, and returns what it printed. The program reads
/documents/{document_id}/ (metadata.json, content.txt, items.jsonl,
chunks.jsonl, toc.json) and can await search() and
await list_documents(); the tool description spells out the fields and the
interpreter's limits. Each call is one program: nothing carries over between
calls, and the sandbox is created and closed per call. A failing program is a
tool error carrying the interpreter's message and any output printed before
it. analysis.code_timeout bounds a call and analysis.max_output_chars its
output; no model runs on the server. Claude Code moves a call still running
after about two minutes to a background task.
Filters
filter is a SQL WHERE clause over the document columns id, uri, title,
metadata, created_at, updated_at. metadata is a JSON string, so match
its keys with LIKE:
metadata LIKE '%"author": "Smith"%'
uri LIKE '%.pdf'
title = 'Q3 report'
Errors
A failure is an MCP error, never an empty result. Expected failures carry a
message: a document or section id that matches nothing, a collection the
server does not cover, a filter the query engine rejects (with its message),
invalid base64, and a program that fails in execute_code. A failure on the
server inside a program, a database read or an in-code search raising, reaches
the program and the client as its exception type only; the traceback goes to
the server log.
Anything else reaches the client as Error calling tool 'name' and its
traceback goes to the server log.
Instructions
The server publishes instructions describing the knowledge base: what it
holds, when to reach for it, the collection names when it covers several, and
prompts.domain_preamble when set. Claude Code shows them to the model. Claude
Desktop does not, so every tool description stands on its own.
Continuous ingestion
For continuous document ingestion (filesystem watch, S3 polling, HTTP
sources, a job queue with retries), run haiku-ingester
as a separate process against the same LanceDB.