diff --git a/CHANGELOG.md b/CHANGELOG.md index f7296518..de459832 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -3,6 +3,7 @@ ### Added +- **MCP image-query tool + CLI `--image PATH`.** New MCP tool `search_documents_by_image(image_base64, limit, include_images)` routes a base64-encoded image through `client.search()`. Registered only when the configured embedder supports images, so non-multimodal MCP servers don't expose a tool that would always fail. The `haiku-rag search` CLI gains an `--image PATH` flag that reads the file and runs the same image-as-query path. - **Image-as-query search.** `client.search()` now accepts `str | bytes | PIL.Image.Image`. Bytes/PIL queries embed via the multimodal embedder's `embed_image_query` and dispatch to vector-only chunk search (FTS doesn't apply to non-text queries; reranking is also skipped). Raises a clear error if the configured embedder is text-only. `ChunkRepository.search()` gains an optional `query_vector` parameter that bypasses `embed_query` and forces the vector-only path. - **`vision: bool` flag on `ModelConfig`.** Tracks whether a configured language model can interpret images. Default `False`. The agent's `search` tool only attaches picture bytes (as `BinaryContent`) to the `ToolReturn` when `qa.model.vision = True`. Without the gate, sending image content to a text-only model behaves inconsistently across providers — Ollama silently accepts and the model hallucinates a confident wrong answer; OpenAI returns 400; others vary. Capability detection from a probe or a model-name whitelist is unreliable, so `vision` is an explicit user-set capability declaration. Set it to `True` for vision-capable QA models (`qwen2.5vl`, `qwen3.6`, `gpt-4o`, `claude-sonnet`, ...). - **Synthetic picture chunks at ingest under multimodal embedders.** `build_picture_chunks` (in `client/processing.py`) walks a `DoclingDocument`'s `pictures` and emits one synthetic `Chunk` per `PictureItem` with available bytes. Bytes ride on a `Chunk._picture_data` PrivateAttr (not serialized) so `embed_chunks` can route them through `embed_images` while text chunks keep going through `embed_documents`. Wired into the three ingest paths (`create_document`, `_create_document_from_file`, `_create_or_update_document_from_url`, `_update_document_with_chunks`, and `_rebuild_rechunk`) — guarded by `embedder.supports_images` so text-only configurations are unaffected. Snapshot/merge with `existing_picture_data` keeps picture chunks alive across rebuild round-trips. Picture chunks land in the same `chunks` table with the same vector dim as text chunks, so cross-modal search reuses the existing hybrid+RRF pipeline. diff --git a/docs/cli.md b/docs/cli.md index 211ce3b9..953815c3 100644 --- a/docs/cli.md +++ b/docs/cli.md @@ -134,6 +134,13 @@ haiku-rag search "transformers" --filter "title = 'Deep Learning Guide'" haiku-rag search "AI" --filter "uri LIKE '%.pdf' AND title LIKE '%paper%'" ``` +Image-as-query (requires a multimodal embedder): +```bash +haiku-rag search --image path/to/figure.png --limit 5 +``` + +When `--image` is used, the positional query is omitted. Pass one or the other, not both. + ## Question Answering Ask questions about your documents: diff --git a/docs/configuration/qa-research.md b/docs/configuration/qa-research.md index 9911ae29..49bf4dc9 100644 --- a/docs/configuration/qa-research.md +++ b/docs/configuration/qa-research.md @@ -29,10 +29,12 @@ qa: name: gpt-oss enable_thinking: true temperature: 0.3 # Default: 0.3 + vision: false # Set true for vision-capable QA models max_searches: 3 # Maximum search tool calls per question ``` - **model**: LLM configuration (see [Providers](providers.md#model-settings)) +- **model.vision**: Set to `true` for vision-capable QA models (`qwen2.5vl`, `qwen3.6`, `gpt-4o`, `claude-sonnet`, …). The agent's `search` tool only attaches picture bytes (`BinaryContent`) to its `ToolReturn` when this is `true`; otherwise picture bytes are withheld. See [Pictures × embedder × QA model](processing.md#pictures--embedder--qa-model-how-the-pieces-compose) for the full matrix. - **max_searches**: Maximum number of search tool calls the QA agent can make per question (default: 3) ## Research Configuration diff --git a/docs/mcp.md b/docs/mcp.md index dc1328a0..6fdf6847 100644 --- a/docs/mcp.md +++ b/docs/mcp.md @@ -38,6 +38,12 @@ The MCP server exposes `haiku.rag` as MCP tools for compatible MCP clients like - **`search_documents`** - Search using hybrid search (vector + full-text) - `query` (required): Search query - `limit` (optional): Maximum results (uses config default if not specified) + - `include_images` (optional, default `true`): Attach base64-encoded picture bytes to picture-labeled results + +- **`search_documents_by_image`** - Search using an image as the query (registered only when the configured embedder supports images) + - `image_base64` (required): Base64-encoded image (PNG/JPEG bytes) + - `limit` (optional): Maximum results + - `include_images` (optional, default `true`) ### Question Answering diff --git a/docs/python.md b/docs/python.md index e0605c94..ca98b6df 100644 --- a/docs/python.md +++ b/docs/python.md @@ -357,6 +357,28 @@ results = await client.search( - `created_at`, `updated_at` - Timestamps - `metadata` - Document metadata (as string, use LIKE for pattern matching) +### Image queries + +`client.search()` accepts an image instead of a text query when the configured embedder is multimodal (e.g. `provider: vllm` against a vision-language embedding model). The image is embedded once and the chunks table is searched vector-only — full-text search and reranking don't apply without a text query. + +```python +from PIL import Image + +# Bytes +results = await client.search( + open("figure.png", "rb").read(), + limit=5, +) + +# PIL.Image works equivalently +results = await client.search( + Image.open("figure.png"), + limit=5, +) +``` + +Image queries surface picture chunks (synthetic per-figure chunks emitted at ingest under a multimodal embedder) and any text chunks whose vectors land near the image vector in the shared embedding space. Calling `client.search(bytes)` against a text-only embedder raises a `ValueError`. + ### Expanding Search Context Expand search results with surrounding content from the document: diff --git a/haiku_rag_slim/haiku/rag/app.py b/haiku_rag_slim/haiku/rag/app.py index 768a6e5c..7c83f557 100644 --- a/haiku_rag_slim/haiku/rag/app.py +++ b/haiku_rag_slim/haiku/rag/app.py @@ -357,15 +357,35 @@ class HaikuRAGApp: # pragma: no cover ) async def search( - self, query: str, limit: int | None = None, filter: str | None = None + self, + query: str | None = None, + limit: int | None = None, + filter: str | None = None, + image: Path | None = None, ): + if query is None and image is None: + self.console.print( + "[red]Provide either a query argument or --image PATH.[/red]" + ) + return + if query is not None and image is not None: + self.console.print("[red]Pass either a query or --image, not both.[/red]") + return + + search_input: str | bytes + if image is not None: + search_input = image.read_bytes() + else: + assert query is not None + search_input = query + async with HaikuRAG( db_path=self.db_path, config=self.config, read_only=self.read_only, before=self.before, ) as self.client: - results = await self.client.search(query, limit=limit, filter=filter) + results = await self.client.search(search_input, limit=limit, filter=filter) if not results: self.console.print("[yellow]No results found.[/yellow]") return diff --git a/haiku_rag_slim/haiku/rag/cli.py b/haiku_rag_slim/haiku/rag/cli.py index d103cda9..29bb874e 100644 --- a/haiku_rag_slim/haiku/rag/cli.py +++ b/haiku_rag_slim/haiku/rag/cli.py @@ -296,8 +296,9 @@ _cli.command("rm", help="Alias for delete: remove a document by its ID")( @_cli.command("search", help="Search for documents by a query") def search( # pragma: no cover - query: str = typer.Argument( - help="The search query to use", + query: str | None = typer.Argument( + None, + help="The search query (omit when using --image)", ), limit: int | None = typer.Option( None, @@ -311,6 +312,11 @@ def search( # pragma: no cover "-f", help="SQL WHERE clause to filter documents (e.g., \"uri LIKE '%arxiv%'\")", ), + image: Path | None = typer.Option( + None, + "--image", + help="Path to an image file to use as the query (requires a multimodal embedder)", + ), db: Path | None = typer.Option( None, "--db", @@ -318,7 +324,7 @@ def search( # pragma: no cover ), ): app = create_app(db) - asyncio.run(app.search(query=query, limit=limit, filter=filter)) + asyncio.run(app.search(query=query, limit=limit, filter=filter, image=image)) @_cli.command("visualize", help="Show visual grounding for a chunk") diff --git a/haiku_rag_slim/haiku/rag/mcp.py b/haiku_rag_slim/haiku/rag/mcp.py index 574bb711..5e84634f 100644 --- a/haiku_rag_slim/haiku/rag/mcp.py +++ b/haiku_rag_slim/haiku/rag/mcp.py @@ -108,6 +108,39 @@ def create_mcp_server( except Exception: return [] + # Image-as-query tool, only registered when the configured embedder + # supports image embeddings. + from haiku.rag.embeddings import get_embedder + + if get_embedder(config).supports_images: + + @mcp.tool() + async def search_documents_by_image( + image_base64: str, + limit: int | None = None, + include_images: bool = True, + ) -> list[SearchResult]: + """Search the RAG system using an image as the query. + + ``image_base64`` is a base64-encoded image (PNG/JPEG bytes). The + image is embedded via the configured multimodal embedder and the + chunks table is searched vector-only. ``include_images`` controls + whether picture bytes are attached to picture-labeled results. + """ + import base64 + + try: + raw = base64.b64decode(image_base64) + except Exception: + return [] + try: + async with HaikuRAG(db_path, config=config, read_only=read_only) as rag: + return await rag.search( + raw, limit=limit, include_images=include_images + ) + except Exception: + return [] + @mcp.tool() async def get_document(document_id: str) -> Document | None: """Get a document by its ID.""" diff --git a/tests/test_mcp.py b/tests/test_mcp.py index 08664590..c9220005 100644 --- a/tests/test_mcp.py +++ b/tests/test_mcp.py @@ -193,3 +193,50 @@ class TestMCPWriteTools: result = await delete_doc(document_id="nonexistent-id") assert result is False + + +class TestMCPImageQuery: + """search_documents_by_image is registered only when the embedder is multimodal.""" + + @pytest.mark.asyncio + async def test_image_query_tool_absent_for_text_only_embedder(self, mcp_db): + """Default text-only embedder must not expose the image-query tool.""" + mcp = create_mcp_server(mcp_db, read_only=True) + names = {t.name for t in await mcp.list_tools()} + assert "search_documents_by_image" not in names + + @pytest.mark.asyncio + async def test_image_query_tool_registered_for_multimodal_embedder( + self, mcp_db, monkeypatch + ): + """When the embedder reports supports_images=True, the tool exists + and routes a base64 image through ``client.search``.""" + from haiku.rag.embeddings import EmbedderWrapper + + class StubMultimodal(EmbedderWrapper): + supports_images = True + + def __init__(self): + super().__init__(embedder=None, vector_dim=2560) + + async def embed_image_query(self, image): + # Produce a deterministic-ish vector of the right dim. + return [0.0] * 2560 + + monkeypatch.setattr( + "haiku.rag.embeddings.get_embedder", + lambda *a, **kw: StubMultimodal(), + ) + + mcp = create_mcp_server(mcp_db, read_only=True) + names = {t.name for t in await mcp.list_tools()} + assert "search_documents_by_image" in names + + search_by_image = await _get_tool(mcp, "search_documents_by_image") + # Standalone PNG header (won't decode to a real image but our stub doesn't care). + import base64 + + png_b64 = base64.b64encode(b"\x89PNG\r\n\x1a\n").decode("ascii") + results = await search_by_image(image_base64=png_b64) + # Empty list is fine (the stub vector won't match the toy fixture). + assert isinstance(results, list)