diff --git a/CHANGELOG.md b/CHANGELOG.md index 8d8760c8..6d18f527 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ ### Changed -- `provider: vllm` is text-only unless `embeddings.model.multimodal: true` is set. Existing multimodal vLLM configs must add the flag. +- `provider: vllm` is text-only unless `embeddings.model.multimodal: true` is set. Existing multimodal vLLM configs must add the flag; `multimodal` is not part of the stored embedding identity, so changing it raises no drift error — re-ingest or `rebuild` to add or drop picture chunks. ## [0.60.0] - 2026-06-22 diff --git a/README.md b/README.md index f7d3aff3..5040cd2b 100644 --- a/README.md +++ b/README.md @@ -10,14 +10,14 @@ Agentic RAG built on [LanceDB](https://lancedb.com/), [Pydantic AI](https://ai.p ## Features - **Hybrid search** — Vector + full-text with Reciprocal Rank Fusion -- **Multimodal & cross-modal search** — Multimodal embedders (vLLM) put picture vectors in the same space as text; supports text-as-query → figure hits and image-as-query +- **Multimodal & cross-modal search** — Multimodal embedders (vLLM, VoyageAI, Cohere) put picture vectors in the same space as text; supports text-as-query → figure hits and image-as-query - **Question answering** — RAG skill with citations (page numbers, section headings) - **Vision QA** — Vision-capable models receive figure bytes alongside chunk text - **Reranking** — MxBAI, Cohere, Zero Entropy, or vLLM - **Analysis skill** — Complex analytical tasks via sandboxed Python code execution (aggregation, computation, multi-document analysis) - **Conversational RAG** — Chat TUI and web application for multi-turn conversations with session memory - **Document structure** — Stores full [DoclingDocument](https://docling-project.github.io/docling/concepts/docling_document/), enabling structure-aware context expansion -- **Multiple providers** — Embeddings: Ollama, OpenAI, VoyageAI, LM Studio, vLLM (multimodal). QA: any model supported by Pydantic AI +- **Multiple providers** — Embeddings: Ollama, OpenAI, VoyageAI, Cohere, LM Studio, vLLM (multimodal via `multimodal: true` on vLLM/VoyageAI/Cohere). QA: any model supported by Pydantic AI - **Local-first** — Embedded LanceDB, no servers required. Also supports S3, GCS, Azure, and LanceDB Cloud - **CLI & Python API** — Full functionality from command line or code - **MCP server** — Expose as tools for AI assistants (Claude Desktop, etc.) diff --git a/docs/configuration/processing.md b/docs/configuration/processing.md index f8ddf80e..1174caaa 100644 --- a/docs/configuration/processing.md +++ b/docs/configuration/processing.md @@ -283,9 +283,11 @@ Three independent settings drive ingest, retrieval, and QA: | Setting | Question it answers | Values | |---|---|---| | `processing.pictures` | Generate and/or describe pictures at ingest? | `none` / `description` / `image` (default) | -| `embeddings.model.provider` | Can the embedder index image content? | text-only (`ollama`, `openai`, `cohere`, `sentence-transformers`) vs `vllm` (multimodal) | +| `embeddings.model.multimodal` | Can the embedder index image content? | `false` (default, text-only) / `true` (supported on `vllm`, `voyageai`, `cohere`) | | `qa.model.vision` | Can the QA model interpret images? | `false` (default) / `true` | +The Embedder column below is driven by `embeddings.model.multimodal`, not the provider name — a vision-capable model under a text-only configuration still indexes no images, and an image-only document then produces zero chunks. See [Multimodal embedders](providers.md#multimodal-embedders). + **What gets stored** by `pictures` × embedder: | `pictures` | Embedder | Text chunks | Synthetic picture chunks | diff --git a/docs/configuration/providers.md b/docs/configuration/providers.md index 63aea039..e0a17b28 100644 --- a/docs/configuration/providers.md +++ b/docs/configuration/providers.md @@ -228,11 +228,15 @@ embeddings: base_url: http://localhost:1234/v1 ``` -**Note:** The `base_url` must include the `/v1` path for OpenAI-compatible endpoints. +**Note:** The `base_url` must include the `/v1` path for OpenAI-compatible endpoints. This path is text-only. For a vision-language model served by vLLM, use `provider: vllm` with `multimodal: true` (below), not `provider: openai`. -### vLLM (multimodal) +### Multimodal embedders -For cross-modal retrieval (text and pictures share a single vector space), use the dedicated `vllm` provider against a vLLM server hosting a multimodal embedding model: +For cross-modal retrieval (text and pictures share a single vector space), set `embeddings.model.multimodal: true`. Capability is decided by this flag, not the provider name: each provider passes images in its own wire format, so multimodal is supported only on `vllm`, `voyageai`, and `cohere`. Setting it on any other provider raises at startup. + +A model produces picture chunks at ingest only when its embedder is multimodal. Without the flag, an image-only document produces zero chunks and is not retrievable. Switching `multimodal` on or off does not change the stored embedding identity, so it raises no drift error; re-ingest or `rebuild` to add or drop picture chunks. + +**vLLM** — a vLLM server hosting a multimodal embedding model. Text inputs use the standard OpenAI `input` field; image inputs use vLLM's `messages`-with-`image_url` superset. Tested with `Qwen/Qwen3-VL-Embedding-8B` (4096-dim) and `jinaai/jina-embeddings-v4` (2048-dim). Run vLLM separately; haiku.rag adds no Python ML dependencies for this path. ```yaml embeddings: @@ -241,11 +245,34 @@ embeddings: name: Qwen/Qwen3-VL-Embedding-8B vector_dim: 4096 base_url: http://localhost:8000/v1 + multimodal: true ``` -Tested with `Qwen/Qwen3-VL-Embedding-8B` (4096-dim) and `jinaai/jina-embeddings-v4` (2048-dim). Run vLLM separately. haiku.rag adds no Python ML dependencies for this path. Text inputs use the standard OpenAI `input` field. Image inputs use vLLM's `messages`-with-`image_url` superset, transparently to the caller. +**VoyageAI** — `voyage-multimodal-3` (1024-dim) via the `voyageai` extra. Reads `VOYAGE_API_KEY` from the environment. -Picture chunks for retrieval are emitted at ingest under any embedder reporting `supports_images=True`. See [Picture Handling](processing.md#picture-handling). +```yaml +embeddings: + model: + provider: voyageai + name: voyage-multimodal-3 + vector_dim: 1024 + multimodal: true +``` + +**Cohere** — `embed-v4.0` (configurable `vector_dim`, e.g. 1536) via the `cohere` extra. Reads `CO_API_KEY` from the environment. + +```yaml +embeddings: + model: + provider: cohere + name: embed-v4.0 + vector_dim: 1536 + multimodal: true +``` + +A text-only model served by vLLM uses `provider: vllm` without the flag (or `provider: openai` with a `base_url`). + +Picture chunks for retrieval are emitted at ingest under any multimodal embedder. See [Picture Handling](processing.md#picture-handling). ## Question Answering Providers diff --git a/docs/python.md b/docs/python.md index cc028fb1..5d2d8dc9 100644 --- a/docs/python.md +++ b/docs/python.md @@ -260,7 +260,7 @@ results = await client.search( ### Image queries -`client.search()` accepts an image instead of a text query when the configured embedder is multimodal (e.g. `provider: vllm` against a vision-language embedding model). The image is embedded once and the chunks table is searched vector-only. Full-text search and reranking don't apply without a text query. +`client.search()` accepts an image instead of a text query when the configured embedder is multimodal (`embeddings.model.multimodal: true` on a vLLM, VoyageAI, or Cohere model). The image is embedded once and the chunks table is searched vector-only. Full-text search and reranking don't apply without a text query. ```python from PIL import Image