# Model Context Protocol (MCP) The MCP server exposes `haiku.rag` as MCP tools for compatible MCP clients like Claude Desktop. ## Starting MCP Server The MCP server supports Streamable HTTP and stdio transports: ```bash # Default streamable HTTP transport on 127.0.0.1:8001 haiku-rag mcp # Custom port haiku-rag mcp --port 9000 # Bind to all interfaces (e.g. inside a container) haiku-rag mcp --host 0.0.0.0 --port 8001 # stdio transport (for Claude Desktop) haiku-rag mcp --stdio ``` `--host` defaults to `127.0.0.1` (loopback only). Bind to `0.0.0.0` only when you want the MCP server reachable from outside the local machine — e.g. inside a Docker container with port mapping, or on a trusted LAN. The server opens the database read-only. Ingestion goes through the CLI (`haiku-rag add`, `add-src`, `delete`) or [`haiku-ingester`](ingester.md). ## Collections With several databases in `lancedb.databases`, the server covers all of them, as `haiku-rag search` does. Results, documents and citations name theirs in `source`. `sources` on `search_documents`, `search_documents_by_image` and `execute_code` restricts a call to a subset; `source` on `get_document` names the database holding the document. A name the server does not cover is an error. `haiku-rag --db-name NAME mcp` serves one. See [Multiple Databases](configuration/storage.md#multiple-databases). ## Claude Code The repository ships a plugin that registers the server and a skill telling Claude when and how to use it: ```bash claude plugin marketplace add ggozad/haiku.rag claude plugin install haiku-rag ``` The plugin runs `haiku-rag mcp --stdio`, so `haiku-rag` must be on the PATH and the configuration decides the database. The skill pre-approves every tool and is also invocable as `/haiku-rag`. To register the server without the plugin: ```bash claude mcp add haiku-rag -- haiku-rag mcp --stdio ``` The skill works with that registration too: copy `plugins/haiku-rag/skills/haiku-rag` into `~/.claude/skills/` and change the tool prefix in its `allowed-tools` from `mcp__plugin_haiku-rag_haiku-rag__` to `mcp__haiku-rag__`. ## Codex The repository's Codex plugin registers the server and installs the same Agent Skill: ```bash codex plugin marketplace add ggozad/haiku.rag codex plugin add haiku-rag@haiku-rag ``` The plugin runs `haiku-rag mcp --stdio`, so `haiku-rag` must be on the PATH. Invoke the skill as `$haiku-rag`. Codex can also select it automatically from its description. To register the server without the plugin: ```bash codex mcp add haiku-rag -- haiku-rag mcp --stdio ``` The skill works with that registration too: copy `plugins/haiku-rag/skills/haiku-rag` into `~/.agents/skills/`. The `allowed-tools` field supplies Claude Code's tool pre-approval and may be ignored by other Agent Skills clients. Codex configures MCP tool approvals separately in `config.toml`. ## Claude Desktop Integration Add to your Claude Desktop configuration (`claude_desktop_config.json`): ```json { "mcpServers": { "haiku-rag": { "command": "haiku-rag", "args": ["mcp", "--stdio"] } } } ``` With a custom database path: ```json { "mcpServers": { "haiku-rag": { "command": "haiku-rag", "args": ["mcp", "--stdio", "--db", "/path/to/database.lancedb"] } } } ``` After restarting Claude Desktop, you can ask Claude to search your documents or answer questions using your knowledge base. ## Tools Every tool is read-only and says so in its annotations. Each parameter carries a description in the tool schema, so the listing below names them without repeating it. | Tool | Registered | Parameters | |---|---|---| | `search_documents` | always | `query`, `limit`, `include_images`, `filter`, `sources` | | `search_documents_by_image` | multimodal embedder only | `image_base64`, `limit`, `include_images`, `filter`, `sources` | | `get_document` | always | `document_id`, `source` | | `get_document_outline` | always | `document_id`, `source` | | `get_document_section` | always | `document_id`, `section_id`, `source` | | `list_documents` | always | `limit`, `offset`, `filter` | | `execute_code` | always | `code`, `filter`, `sources` | `search_documents` runs hybrid search, vector and full-text. Its text content is the rendering the in-process agents read: results best first, each with its rank, `Document ID`, `Collection` when the server covers several, the document title, section headings, the matched chunk's metadata when it has any, and the passage expanded to its section the way the agents get it (`search.max_context_chars` caps it). Pictures in the results follow as image blocks, one per distinct picture, each preceded by a line naming its result; `include_images: false` leaves them out. Search results carry no structured content, so every client shows the model the same text and images. Scores are not comparable across queries or search types, so rank is the signal. `search_documents_by_image` embeds the query image and searches by vector similarity alone. `get_document` returns a document whole, in reading order. For a long one, `get_document_outline` returns the heading tree with page numbers and `get_document_section` the text of one section, subsections included; a node's `id` in the outline is the `section_id`. A document without headings has an empty outline. `list_documents` returns titles, URIs and metadata, which is how a client learns what a filter can match. `execute_code` runs a Python program in the sandbox of the [analysis capability](capabilities/analysis.md), over the documents `filter` and `sources` select, and returns what it printed. The program reads `/documents/{document_id}/` (`metadata.json`, `content.txt`, `items.jsonl`, `chunks.jsonl`, `toc.json`) and can `await search()` and `await list_documents()`; the tool description spells out the fields and the interpreter's limits. Each call is one program: nothing carries over between calls, and the sandbox is created and closed per call. A failing program is a tool error carrying the interpreter's message and any output printed before it. `analysis.code_timeout` bounds a call and `analysis.max_output_chars` its output; no model runs on the server. Claude Code moves a call still running after about two minutes to a background task. ### Filters `filter` is a SQL WHERE clause over the document columns `id`, `uri`, `title`, `metadata`, `created_at`, `updated_at`. `metadata` is a JSON string, so match its keys with LIKE: ```sql metadata LIKE '%"author": "Smith"%' uri LIKE '%.pdf' title = 'Q3 report' ``` ### Errors A failure is an MCP error, never an empty result. Expected failures carry a message: a document or section id that matches nothing, a collection the server does not cover, a filter the query engine rejects (with its message), invalid base64, and a program that fails in `execute_code`. A failure on the server inside a program, a database read or an in-code search raising, reaches the program and the client as its exception type only; the traceback goes to the server log. Anything else reaches the client as `Error calling tool 'name'` and its traceback goes to the server log. ### Instructions The server publishes `instructions` describing the knowledge base: what it holds, when to reach for it, the collection names when it covers several, and `prompts.domain_preamble` when set. Claude Code and Codex show them to the model. Claude Desktop does not, so every tool description stands on its own. ## Continuous ingestion For continuous document ingestion (filesystem watch, S3 polling, HTTP sources, a job queue with retries), run [`haiku-ingester`](ingester.md) as a separate process against the same LanceDB.