Merge pull request #364 from ggozad/fix/remote-image-ingestion

Fix HTML image ingestion and wire conversion options across all formats
This commit is contained in:
Yiorgis Gozadinos 2026-05-12 15:58:05 +03:00 committed by GitHub
commit bb39eb3809
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
14 changed files with 566 additions and 109 deletions

View file

@ -3,13 +3,22 @@
### Added ### Added
- **`processing.conversion_options.fetch_remote_images`** (default `true`). Controls whether docling fetches images referenced by URL in HTML and Markdown inputs. docling-local only — docling-serve cannot fetch external images via its API regardless of this flag.
- **`s3://` is a first-class document source.** `create_document_from_source`, the CLI `haiku-rag add-src`, and the MCP `add_document_from_url` tool all dispatch on the `s3` URL scheme. Two-stage change detection keeps `metadata["md5"]` semantically uniform across all sources: HEAD ETag matching the stored `metadata["etag"]` short-circuits without GET; if ETag differs but bytes hash to the same MD5 (multipart re-upload, server-side `CopyObject`, SSE mode change), only the etag refreshes — no re-chunk or re-embed. Closes #357. - **`s3://` is a first-class document source.** `create_document_from_source`, the CLI `haiku-rag add-src`, and the MCP `add_document_from_url` tool all dispatch on the `s3` URL scheme. Two-stage change detection keeps `metadata["md5"]` semantically uniform across all sources: HEAD ETag matching the stored `metadata["etag"]` short-circuits without GET; if ETag differs but bytes hash to the same MD5 (multipart re-upload, server-side `CopyObject`, SSE mode change), only the etag refreshes — no re-chunk or re-embed. Closes #357.
- **S3 / object-storage monitoring.** `monitor.s3: list[S3MonitorEntry]` adds a polling watcher per bucket prefix alongside the existing local-directory watcher. Each entry has its own `poll_interval`, `include_patterns`, `ignore_patterns`, `delete_orphans`, and `storage_options`. The same `serve --monitor` flag enables both. Orphan deletion is per-entry (scoped via `uri LIKE 's3://bucket/prefix/%'`); other buckets and prefixes are never touched. - **S3 / object-storage monitoring.** `monitor.s3: list[S3MonitorEntry]` adds a polling watcher per bucket prefix alongside the existing local-directory watcher. Each entry has its own `poll_interval`, `include_patterns`, `ignore_patterns`, `delete_orphans`, and `storage_options`. The same `serve --monitor` flag enables both. Orphan deletion is per-entry (scoped via `uri LIKE 's3://bucket/prefix/%'`); other buckets and prefixes are never touched.
- **`[s3]` optional extra** (`obstore>=0.9`). Required for `s3://` sources and the S3 watcher. Uses obstore — the Python binding to the same Rust `object_store` crate that LanceDB uses internally — so `monitor.s3[*].storage_options` accepts the same dict shape as `lancedb.storage_options`. Empty/missing options fall back to the AWS default credential chain. - **`[s3]` optional extra** (`obstore>=0.9`). Required for `s3://` sources and the S3 watcher. Uses obstore — the Python binding to the same Rust `object_store` crate that LanceDB uses internally — so `monitor.s3[*].storage_options` accepts the same dict shape as `lancedb.storage_options`. Empty/missing options fall back to the AWS default credential chain.
- **`scripts/run-integration-tests.sh`** — wraps `docker compose up --wait`, `pytest -m integration`, and tear-down so the SeaweedFS-backed integration suite is a one-liner. - **`scripts/run-integration-tests.sh`** — wraps `docker compose up --wait`, `pytest -m integration`, and tear-down so the SeaweedFS-backed integration suite is a one-liner.
### Fixed
- **Conversion options now apply to non-PDF formats.** `DoclingLocalConverter` previously wired its `PdfPipelineOptions` only to `InputFormat.PDF`, so user settings (OCR knobs, `picture_description.enabled`, `images_scale`, etc.) silently no-op'd for HTML, Markdown, DOCX, PPTX, and IMAGE inputs. The converter now shares a single `PdfPipelineOptions` instance across PDF, IMAGE, HTML, MD, DOCX, and PPTX `FormatOption`s. SimplePipeline-backed formats ignore the PDF-specific fields; `ConvertPipelineOptions`-level enrichments (picture description / classification / chart extraction) now run uniformly. HTML and Markdown additionally receive `HTMLBackendOptions` / `MarkdownBackendOptions` gated on `fetch_remote_images`.
- **HTML text ingest path picks up converter options.** `convert_text(format="html"/"md")` previously used a bare `DoclingDocConverter()` with zero format options — the wix corpus ingest path. It now uses the same shared `_build_format_options()` helper as the file path.
- **Relative `<img>` paths resolve during URL ingest.** `HaikuRAG.convert()` and the converter `convert_file` / `convert_text` methods now thread a `source_uri` through to `HTMLBackendOptions.source_uri` / `MarkdownBackendOptions.source_uri`. URL ingest uses the originating URL; file ingest uses `file://`; raw text accepts an optional override. docling-serve accepts the kwarg as a no-op (its API has no equivalent option).
### Documentation ### Documentation
- New "External image fetching" subsection in `docs/configuration/processing.md` documenting `fetch_remote_images`, the SSRF / size / timeout guards inherited from docling, and a per-format table of which conversion options actually apply (PDF, IMAGE, HTML, MD, DOCX/PPTX, others).
- New "HTML Image Fetching" section in `docs/remote-processing.md` calling out that docling-serve cannot fetch external `<img>` URLs and recommending docling-local for HTML ingest when picture bytes matter.
- New "S3 / Object Storage Monitoring" section in `docs/server.md` and `docs/configuration/processing.md` covering the `[s3]` extra, polling cadence, ETag semantics, credentials, and CLI usage. - New "S3 / Object Storage Monitoring" section in `docs/server.md` and `docs/configuration/processing.md` covering the `[s3]` extra, polling cadence, ETag semantics, credentials, and CLI usage.
- New "Deployment Pattern: One Writer, Many Readers" subsection in `docs/configuration/storage.md` documenting the recommended IAM split (one ingestion process + N read-only consumers). - New "Deployment Pattern: One Writer, Many Readers" subsection in `docs/configuration/storage.md` documenting the recommended IAM split (one ingestion process + N read-only consumers).

View file

@ -99,10 +99,36 @@ conversion_options:
conversion_options: conversion_options:
images_scale: 2.0 # Image resolution scale factor images_scale: 2.0 # Image resolution scale factor
generate_page_images: true # Include rendered page images generate_page_images: true # Include rendered page images
fetch_remote_images: true # Fetch external <img src> URLs in HTML/MD
``` ```
- **images_scale**: Scale factor for extracted images. Higher values = better quality but larger size. Typical range: 1.0-3.0. - **images_scale**: Scale factor for extracted images. Higher values = better quality but larger size. Typical range: 1.0-3.0.
- **generate_page_images**: When `true` (default), rendered images of each PDF page are included in the document. Required for `visualize_chunk()` to show visual grounding. When `false`, page images are excluded to reduce document size. - **generate_page_images**: When `true` (default), rendered images of each PDF page are included in the document. Required for `visualize_chunk()` to show visual grounding. When `false`, page images are excluded to reduce document size.
- **fetch_remote_images**: When `true` (default), HTML and Markdown inputs have their external `<img src="https://...">` URLs fetched and stored as picture bytes. Set `false` for air-gapped ingest. Applies only to docling-local; see [Remote processing](../remote-processing.md#html-image-fetching) for the docling-serve limitation.
#### External image fetching
For HTML and Markdown inputs, docling fetches images referenced by URL when `fetch_remote_images: true`. Pictures end up in `document_items.picture_data` alongside the ones extracted from PDF/DOCX/PPTX. Inherited from docling:
- **SSRF guard**: hostnames must resolve to a global IP. Loopback, private (RFC1918), link-local, reserved, multicast, and unspecified addresses are rejected.
- **Size cap**: 20 MB per image (sent as a `Range` header), enforced again when streaming the response body.
- **Timeouts**: 5 s connect, 30 s read.
- **SVGs are skipped** (PIL cannot rasterize them).
- **`data:` URIs** are decoded inline (no network).
- **`file://` URIs** are *not* fetched — `enable_local_fetch` stays off to keep the SSRF surface narrow for arbitrary HTML/MD content.
Per-image failures (404, timeout, oversized, unreadable) leave that picture as a placeholder with `picture_data=NULL` — the rest of the document still ingests.
**Scope of conversion options across formats:**
| Input | OCR / table options | `images_scale` / `generate_page_images` | `picture_description` | `fetch_remote_images` |
|---|---|---|---|---|
| `.pdf` | ✅ | ✅ | ✅ | n/a |
| `.png` / `.jpg` / `.jpeg` / `.bmp` / `.tiff` / `.webp` | ✅ | ✅ | ✅ | n/a |
| `.html` / `.xhtml` | n/a (markup-based) | n/a | ✅ on embedded pictures | ✅ |
| `.md` / `.qmd` / `.rmd` | n/a | n/a | ✅ on embedded pictures | ✅ (only `<img>` HTML blocks; native `![alt](url)` syntax is not fetched by docling) |
| `.docx` / `.pptx` | n/a | n/a | ✅ on embedded pictures | n/a |
| Other (`.csv`, `.xlsx`, `.adoc`, `.tex`, `.xml`) | n/a | n/a | n/a | n/a |
#### Picture Handling #### Picture Handling

View file

@ -134,6 +134,14 @@ processing:
- `false` (default): Tables as narrative text - `false` (default): Tables as narrative text
- `true`: Tables as markdown format - `true`: Tables as markdown format
## HTML Image Fetching
docling-serve does **not** fetch external `<img src="https://...">` URLs in HTML inputs. The `ConvertDocumentsOptions` API exposes no equivalent of docling-local's `HTMLBackendOptions.fetch_images` / `enable_remote_fetch`, and the server-side `DoclingConverterManager` registers `format_options` only for PDF and IMAGE — HTML falls through to docling's defaults (`fetch_images=False`).
Consequence: ingesting HTML with external image references through docling-serve produces picture items with `picture_data=NULL`. The same input through docling-local fetches the bytes (subject to the SSRF / size / timeout guards documented in [Configuration → External image fetching](configuration/processing.md#external-image-fetching)).
To preserve image bytes when ingesting HTML or Markdown that references remote images, use `converter: docling-local`. The `processing.conversion_options.fetch_remote_images` flag has no effect on docling-serve and the `source_uri` kwarg on `HaikuRAG.convert()` is accepted but ignored on this path.
## VLM Picture Description with docling-serve ## VLM Picture Description with docling-serve
When using VLM picture description with docling-serve, the VLM API calls are made by the docling-serve container, not by haiku.rag. This requires additional configuration. When using VLM picture description with docling-serve, the VLM API calls are made by the docling-serve container, not by haiku.rag. This requires additional configuration.

View file

@ -139,19 +139,25 @@ class HaikuRAG:
# ========================================================================= # =========================================================================
@overload @overload
async def convert(self, source: Path) -> "DoclingDocument": ... async def convert(
self, source: Path, *, source_uri: str | None = None
) -> "DoclingDocument": ...
@overload @overload
async def convert( async def convert(
self, source: str, *, format: str = "md" self, source: str, *, format: str = "md", source_uri: str | None = None
) -> "DoclingDocument": ... ) -> "DoclingDocument": ...
async def convert( async def convert(
self, source: Path | str, *, format: str = "md" self,
source: Path | str,
*,
format: str = "md",
source_uri: str | None = None,
) -> "DoclingDocument": ) -> "DoclingDocument":
from haiku.rag.client.processing import convert from haiku.rag.client.processing import convert
return await convert(self._config, source, format=format) return await convert(self._config, source, format=format, source_uri=source_uri)
async def chunk( async def chunk(
self, self,

View file

@ -411,7 +411,7 @@ async def _create_or_update_document_from_url(
temp_path = Path(temp_file.name) temp_path = Path(temp_file.name)
try: try:
docling_document = await client.convert(temp_path) docling_document = await client.convert(temp_path, source_uri=url)
chunks = await client.chunk(docling_document) chunks = await client.chunk(docling_document)
embedded_chunks = await embed_chunks(chunks, client._config) embedded_chunks = await embed_chunks(chunks, client._config)
finally: finally:
@ -547,7 +547,7 @@ async def _create_or_update_document_from_s3(
return existing_doc return existing_doc
try: try:
docling_document = await client.convert(temp_path) docling_document = await client.convert(temp_path, source_uri=url)
chunks = await client.chunk(docling_document) chunks = await client.chunk(docling_document)
embedded_chunks = await embed_chunks(chunks, client._config) embedded_chunks = await embed_chunks(chunks, client._config)
finally: finally:

View file

@ -52,7 +52,11 @@ def _warn_if_descriptions_missing(
async def convert( async def convert(
config: AppConfig, source: Path | str, *, format: str = "md" config: AppConfig,
source: Path | str,
*,
format: str = "md",
source_uri: str | None = None,
) -> "DoclingDocument": ) -> "DoclingDocument":
"""Convert a file, URL, or text to DoclingDocument. """Convert a file, URL, or text to DoclingDocument.
@ -66,6 +70,10 @@ async def convert(
Defaults to "md". Use "plain" for plain text without parsing. Defaults to "md". Use "plain" for plain text without parsing.
Only used when source is raw text (not a file path or URL). Only used when source is raw text (not a file path or URL).
Files and URLs determine format from extension/content-type. Files and URLs determine format from extension/content-type.
source_uri: Origin URI used by docling's HTML/Markdown backends to
resolve relative `<img src="/path">` references. When omitted,
defaults to the URL (URL ingest) or `file://` URI (file ingest);
raw text input has no origin so no default is derived.
Returns: Returns:
DoclingDocument from the converted source. DoclingDocument from the converted source.
@ -82,7 +90,8 @@ async def convert(
raise ValueError(f"File does not exist: {source}") raise ValueError(f"File does not exist: {source}")
if source.suffix.lower() not in converter.supported_extensions: if source.suffix.lower() not in converter.supported_extensions:
raise ValueError(f"Unsupported file extension: {source.suffix}") raise ValueError(f"Unsupported file extension: {source.suffix}")
doc = await converter.convert_file(source) effective_uri = source_uri or source.absolute().as_uri()
doc = await converter.convert_file(source, source_uri=effective_uri)
_warn_if_descriptions_missing(config, doc, str(source)) _warn_if_descriptions_missing(config, doc, str(source))
return doc return doc
@ -113,7 +122,8 @@ async def convert(
temp_path = Path(temp_file.name) temp_path = Path(temp_file.name)
try: try:
doc = await converter.convert_file(temp_path) effective_uri = source_uri or source
doc = await converter.convert_file(temp_path, source_uri=effective_uri)
_warn_if_descriptions_missing(config, doc, source) _warn_if_descriptions_missing(config, doc, source)
return doc return doc
finally: finally:
@ -126,14 +136,15 @@ async def convert(
raise ValueError(f"File does not exist: {file_path}") raise ValueError(f"File does not exist: {file_path}")
if file_path.suffix.lower() not in converter.supported_extensions: if file_path.suffix.lower() not in converter.supported_extensions:
raise ValueError(f"Unsupported file extension: {file_path.suffix}") raise ValueError(f"Unsupported file extension: {file_path.suffix}")
doc = await converter.convert_file(file_path) effective_uri = source_uri or file_path.absolute().as_uri()
doc = await converter.convert_file(file_path, source_uri=effective_uri)
_warn_if_descriptions_missing(config, doc, str(file_path)) _warn_if_descriptions_missing(config, doc, str(file_path))
return doc return doc
else: else:
# Raw text content — HTML and markdown can still embed pictures # Raw text content — HTML and markdown can still embed pictures
# via <img>/![](...) so the same description check applies. # via <img>/![](...) so the same description check applies.
doc = await converter.convert_text(source, format=format) doc = await converter.convert_text(source, format=format, source_uri=source_uri)
_warn_if_descriptions_missing(config, doc, "<text input>") _warn_if_descriptions_missing(config, doc, "<text input>")
return doc return doc

View file

@ -161,6 +161,10 @@ class ConversionOptions(BaseModel):
images_scale: float = 2.0 images_scale: float = 2.0
generate_page_images: bool = True generate_page_images: bool = True
# Fetch images referenced by URL in HTML and Markdown inputs.
# docling-local only — docling-serve cannot fetch external images.
fetch_remote_images: bool = True
picture_description: PictureDescriptionConfig = Field( picture_description: PictureDescriptionConfig = Field(
default_factory=PictureDescriptionConfig default_factory=PictureDescriptionConfig
) )

View file

@ -26,11 +26,18 @@ class DocumentConverter(ABC):
pass pass
@abstractmethod @abstractmethod
async def convert_file(self, path: Path) -> "DoclingDocument": async def convert_file(
self, path: Path, source_uri: str | None = None
) -> "DoclingDocument":
"""Convert a file to DoclingDocument format. """Convert a file to DoclingDocument format.
Args: Args:
path: Path to the file to convert. path: Path to the file to convert.
source_uri: Optional origin URI (e.g. the URL the file was
downloaded from) used by docling's HTML/Markdown backends to
resolve relative `<img src="/path">` references. Ignored by
converters that have no equivalent backend option (notably
docling-serve).
Returns: Returns:
DoclingDocument representation of the file. DoclingDocument representation of the file.
@ -44,7 +51,11 @@ class DocumentConverter(ABC):
@abstractmethod @abstractmethod
async def convert_text( async def convert_text(
self, text: str, name: str = "content.md", format: str = "md" self,
text: str,
name: str = "content.md",
format: str = "md",
source_uri: str | None = None,
) -> "DoclingDocument": ) -> "DoclingDocument":
"""Convert text content to DoclingDocument format. """Convert text content to DoclingDocument format.
@ -53,6 +64,8 @@ class DocumentConverter(ABC):
name: The name to use for the document (defaults to "content.md"). name: The name to use for the document (defaults to "content.md").
format: The format of the text content ("md", "html", or "plain"). format: The format of the text content ("md", "html", or "plain").
Defaults to "md". Use "plain" for plain text without parsing. Defaults to "md". Use "plain" for plain text without parsing.
source_uri: Optional origin URI used by docling's HTML/Markdown
backends to resolve relative image references.
Returns: Returns:
DoclingDocument representation of the text. DoclingDocument representation of the text.

View file

@ -2,13 +2,15 @@
import asyncio import asyncio
from pathlib import Path from pathlib import Path
from typing import TYPE_CHECKING, ClassVar, cast from typing import TYPE_CHECKING, ClassVar
from haiku.rag.config import AppConfig from haiku.rag.config import AppConfig
from haiku.rag.converters.base import DocumentConverter from haiku.rag.converters.base import DocumentConverter
from haiku.rag.converters.text_utils import TextFileHandler from haiku.rag.converters.text_utils import TextFileHandler
if TYPE_CHECKING: if TYPE_CHECKING:
from docling.datamodel.base_models import InputFormat
from docling.document_converter import FormatOption
from docling_core.types.doc.document import DoclingDocument from docling_core.types.doc.document import DoclingDocument
from haiku.rag.config.models import ConversionOptions, ModelConfig from haiku.rag.config.models import ConversionOptions, ModelConfig
@ -103,23 +105,17 @@ class DoclingLocalConverter(DocumentConverter):
case _: # "auto" or any other value case _: # "auto" or any other value
return OcrAutoOptions(force_full_page_ocr=force_ocr, lang=lang) return OcrAutoOptions(force_full_page_ocr=force_ocr, lang=lang)
def _sync_convert_docling_file(self, path: Path) -> "DoclingDocument": def _build_pipeline_options(self):
"""Synchronous conversion of docling-supported files.""" """Build the shared PdfPipelineOptions instance applied to every wired
from docling.backend.docling_parse_backend import DoclingParseDocumentBackend FormatOption. SimplePipeline-backed formats ignore the PDF-specific
from docling.datamodel.base_models import InputFormat fields; the ConvertPipelineOptions-level picture description /
classification / chart-extraction settings apply uniformly."""
from docling.datamodel.pipeline_options import ( from docling.datamodel.pipeline_options import (
PdfPipelineOptions, PdfPipelineOptions,
PictureDescriptionApiOptions, PictureDescriptionApiOptions,
TableFormerMode, TableFormerMode,
TableStructureOptions, TableStructureOptions,
) )
from docling.document_converter import (
DocumentConverter as DoclingDocConverter,
)
from docling.document_converter import (
FormatOption,
PdfFormatOption,
)
opts = self.config.processing.conversion_options opts = self.config.processing.conversion_options
pic_desc = opts.picture_description pic_desc = opts.picture_description
@ -146,8 +142,6 @@ class DoclingLocalConverter(DocumentConverter):
if runs_vlm: if runs_vlm:
from pydantic import AnyUrl from pydantic import AnyUrl
prompt = self.config.prompts.picture_description
pipeline_options.enable_remote_services = True pipeline_options.enable_remote_services = True
pipeline_options.picture_description_options = PictureDescriptionApiOptions( pipeline_options.picture_description_options = PictureDescriptionApiOptions(
url=AnyUrl(self._get_vlm_api_url(pic_desc.model)), url=AnyUrl(self._get_vlm_api_url(pic_desc.model)),
@ -155,29 +149,97 @@ class DoclingLocalConverter(DocumentConverter):
model=pic_desc.model.name, model=pic_desc.model.name,
max_completion_tokens=pic_desc.max_tokens, max_completion_tokens=pic_desc.max_tokens,
), ),
prompt=prompt, prompt=self.config.prompts.picture_description,
timeout=pic_desc.timeout, timeout=pic_desc.timeout,
) )
format_options = cast( return pipeline_options
dict[InputFormat, FormatOption],
{ def _build_format_options(
InputFormat.PDF: PdfFormatOption( self, source_uri: str | None = None
pipeline_options=pipeline_options, ) -> "dict[InputFormat, FormatOption]":
backend=DoclingParseDocumentBackend, """Per-format options shared between file and text conversion paths.
)
}, Every wired FormatOption gets the same `PdfPipelineOptions` instance so
picture-description / classification / chart settings apply uniformly
across PDF, IMAGE, HTML, MD, DOCX, PPTX. HTML and Markdown additionally
receive backend options gated on `fetch_remote_images`.
Args:
source_uri: Origin URI used by the HTML and Markdown backends to
resolve relative `<img src="/path">` references (e.g. when
ingesting a downloaded HTML page).
"""
from docling.backend.docling_parse_backend import DoclingParseDocumentBackend
from docling.datamodel.backend_options import (
HTMLBackendOptions,
MarkdownBackendOptions,
)
from docling.datamodel.base_models import InputFormat
from docling.document_converter import (
HTMLFormatOption,
ImageFormatOption,
MarkdownFormatOption,
PdfFormatOption,
PowerpointFormatOption,
WordFormatOption,
)
from pydantic import AnyUrl
opts = self.config.processing.conversion_options
pipeline_options = self._build_pipeline_options()
fetch = opts.fetch_remote_images
source_url = AnyUrl(source_uri) if source_uri else None
return {
InputFormat.PDF: PdfFormatOption(
pipeline_options=pipeline_options,
backend=DoclingParseDocumentBackend,
),
InputFormat.IMAGE: ImageFormatOption(pipeline_options=pipeline_options),
InputFormat.HTML: HTMLFormatOption(
pipeline_options=pipeline_options,
backend_options=HTMLBackendOptions(
fetch_images=fetch,
enable_remote_fetch=fetch,
source_uri=source_url,
),
),
InputFormat.MD: MarkdownFormatOption(
pipeline_options=pipeline_options,
backend_options=MarkdownBackendOptions(
fetch_images=fetch,
enable_remote_fetch=fetch,
source_uri=source_url,
),
),
InputFormat.DOCX: WordFormatOption(pipeline_options=pipeline_options),
InputFormat.PPTX: PowerpointFormatOption(pipeline_options=pipeline_options),
}
def _sync_convert_docling_file(
self, path: Path, source_uri: str | None = None
) -> "DoclingDocument":
"""Synchronous conversion of docling-supported files."""
from docling.document_converter import (
DocumentConverter as DoclingDocConverter,
) )
converter = DoclingDocConverter(format_options=format_options) converter = DoclingDocConverter(
format_options=self._build_format_options(source_uri=source_uri)
)
result = converter.convert(path) result = converter.convert(path)
return result.document return result.document
async def convert_file(self, path: Path) -> "DoclingDocument": async def convert_file(
self, path: Path, source_uri: str | None = None
) -> "DoclingDocument":
"""Convert a file to DoclingDocument using local docling. """Convert a file to DoclingDocument using local docling.
Args: Args:
path: Path to the file to convert. path: Path to the file to convert.
source_uri: Optional origin URI used by docling's HTML/Markdown
backends to resolve relative image references.
Returns: Returns:
DoclingDocument representation of the file. DoclingDocument representation of the file.
@ -189,21 +251,33 @@ class DoclingLocalConverter(DocumentConverter):
file_extension = path.suffix.lower() file_extension = path.suffix.lower()
if file_extension in self.docling_extensions: if file_extension in self.docling_extensions:
return await asyncio.to_thread(self._sync_convert_docling_file, path) return await asyncio.to_thread(
self._sync_convert_docling_file, path, source_uri
)
elif file_extension in TextFileHandler.text_extensions: elif file_extension in TextFileHandler.text_extensions:
content = await asyncio.to_thread(path.read_text, encoding="utf-8") content = await asyncio.to_thread(path.read_text, encoding="utf-8")
prepared_content = TextFileHandler.prepare_text_content( prepared_content = TextFileHandler.prepare_text_content(
content, file_extension content, file_extension
) )
return await self.convert_text(prepared_content, name=f"{path.stem}.md") return await self.convert_text(
prepared_content,
name=f"{path.stem}.md",
source_uri=source_uri,
)
else: else:
content = await asyncio.to_thread(path.read_text, encoding="utf-8") content = await asyncio.to_thread(path.read_text, encoding="utf-8")
return await self.convert_text(content, name=f"{path.stem}.md") return await self.convert_text(
content, name=f"{path.stem}.md", source_uri=source_uri
)
except Exception: except Exception:
raise ValueError(f"Failed to parse file: {path}") raise ValueError(f"Failed to parse file: {path}")
async def convert_text( async def convert_text(
self, text: str, name: str = "content.md", format: str = "md" self,
text: str,
name: str = "content.md",
format: str = "md",
source_uri: str | None = None,
) -> "DoclingDocument": ) -> "DoclingDocument":
"""Convert text content to DoclingDocument using local docling. """Convert text content to DoclingDocument using local docling.
@ -212,6 +286,8 @@ class DoclingLocalConverter(DocumentConverter):
name: The name to use for the document (defaults to "content.md"). name: The name to use for the document (defaults to "content.md").
format: The format of the text content ("md", "html", or "plain"). format: The format of the text content ("md", "html", or "plain").
Defaults to "md". Use "plain" for plain text without parsing. Defaults to "md". Use "plain" for plain text without parsing.
source_uri: Optional origin URI used by docling's HTML/Markdown
backends to resolve relative image references.
Returns: Returns:
DoclingDocument representation of the text. DoclingDocument representation of the text.
@ -219,4 +295,43 @@ class DoclingLocalConverter(DocumentConverter):
Raises: Raises:
ValueError: If the text cannot be converted or format is unsupported. ValueError: If the text cannot be converted or format is unsupported.
""" """
return await TextFileHandler.text_to_docling_document(text, name, format) if format not in TextFileHandler.SUPPORTED_FORMATS:
raise ValueError(
f"Unsupported format: {format}. "
f"Supported formats: {', '.join(TextFileHandler.SUPPORTED_FORMATS)}"
)
doc_name = f"content.{format}" if name == "content.md" else name
if format == "plain":
return TextFileHandler._create_simple_docling_document(text, doc_name)
try:
return await asyncio.to_thread(
self._sync_convert_docling_text, text, doc_name, source_uri
)
except Exception as e:
raise ValueError(f"Failed to convert text to DoclingDocument: {e}")
def _sync_convert_docling_text(
self, text: str, doc_name: str, source_uri: str | None = None
) -> "DoclingDocument":
"""Synchronous text-to-DoclingDocument using the shared format options."""
from io import BytesIO
from docling.document_converter import (
DocumentConverter as DoclingDocConverter,
)
from docling.exceptions import ConversionError
from docling_core.types.io import DocumentStream
bytes_io = BytesIO(text.encode("utf-8"))
doc_stream = DocumentStream(name=doc_name, stream=bytes_io)
converter = DoclingDocConverter(
format_options=self._build_format_options(source_uri=source_uri)
)
try:
result = converter.convert(doc_stream)
return result.document
except ConversionError:
return TextFileHandler._create_simple_docling_document(text, doc_name)

View file

@ -226,11 +226,16 @@ class DoclingServeConverter(DocumentConverter):
) )
return self._parse_zip_to_docling(zip_bytes, name) return self._parse_zip_to_docling(zip_bytes, name)
async def convert_file(self, path: Path) -> "DoclingDocument": async def convert_file(
self, path: Path, source_uri: str | None = None
) -> "DoclingDocument":
"""Convert a file to DoclingDocument using docling-serve. """Convert a file to DoclingDocument using docling-serve.
Args: Args:
path: Path to the file to convert. path: Path to the file to convert.
source_uri: Ignored. docling-serve has no API path for
``HTMLBackendOptions.source_uri``; the kwarg is accepted for
interface compatibility with docling-local.
Returns: Returns:
DoclingDocument representation of the file. DoclingDocument representation of the file.
@ -261,7 +266,11 @@ class DoclingServeConverter(DocumentConverter):
SUPPORTED_FORMATS = ("md", "html", "plain") SUPPORTED_FORMATS = ("md", "html", "plain")
async def convert_text( async def convert_text(
self, text: str, name: str = "content.md", format: str = "md" self,
text: str,
name: str = "content.md",
format: str = "md",
source_uri: str | None = None,
) -> "DoclingDocument": ) -> "DoclingDocument":
"""Convert text content to DoclingDocument via docling-serve. """Convert text content to DoclingDocument via docling-serve.
@ -272,6 +281,9 @@ class DoclingServeConverter(DocumentConverter):
name: The name to use for the document (defaults to "content.md"). name: The name to use for the document (defaults to "content.md").
format: The format of the text content ("md", "html", or "plain"). format: The format of the text content ("md", "html", or "plain").
Defaults to "md". Use "plain" for plain text without parsing. Defaults to "md". Use "plain" for plain text without parsing.
source_uri: Ignored. docling-serve has no API path for
``HTMLBackendOptions.source_uri``; the kwarg is accepted for
interface compatibility with docling-local.
Returns: Returns:
DoclingDocument representation of the text. DoclingDocument representation of the text.

View file

@ -1,7 +1,5 @@
"""Shared utilities for text file handling in converters.""" """Shared utilities for text file handling in converters."""
import asyncio
from io import BytesIO
from typing import TYPE_CHECKING, ClassVar from typing import TYPE_CHECKING, ClassVar
if TYPE_CHECKING: if TYPE_CHECKING:
@ -183,61 +181,3 @@ class TextFileHandler:
doc = DoclingDocument(name=doc_name) doc = DoclingDocument(name=doc_name)
doc.add_text(label=DocItemLabel.TEXT, text=text) doc.add_text(label=DocItemLabel.TEXT, text=text)
return doc return doc
@staticmethod
def _sync_text_to_docling_document(
text: str, name: str = "content.md", format: str = "md"
) -> "DoclingDocument":
"""Synchronous implementation of text to DoclingDocument conversion."""
from docling.document_converter import DocumentConverter as DoclingDocConverter
from docling.exceptions import ConversionError
from docling_core.types.io import DocumentStream
if format not in TextFileHandler.SUPPORTED_FORMATS:
raise ValueError(
f"Unsupported format: {format}. "
f"Supported formats: {', '.join(TextFileHandler.SUPPORTED_FORMATS)}"
)
# Derive document name from format to tell docling which parser to use
doc_name = f"content.{format}" if name == "content.md" else name
# Plain text doesn't need parsing - create document directly
if format == "plain":
return TextFileHandler._create_simple_docling_document(text, doc_name)
bytes_io = BytesIO(text.encode("utf-8"))
doc_stream = DocumentStream(name=doc_name, stream=bytes_io)
converter = DoclingDocConverter()
try:
result = converter.convert(doc_stream)
return result.document
except ConversionError:
# Docling's format detection fails for plain text without markdown syntax.
# Fall back to creating a simple document directly.
return TextFileHandler._create_simple_docling_document(text, doc_name)
@staticmethod
async def text_to_docling_document(
text: str, name: str = "content.md", format: str = "md"
) -> "DoclingDocument":
"""Convert text to DoclingDocument using docling's parser.
Args:
text: The text content to convert.
name: The name to use for the document.
format: The format of the text content ("md", "html", or "plain").
Defaults to "md". Use "plain" for plain text without parsing.
Returns:
DoclingDocument representation of the text.
Raises:
ValueError: If the conversion fails or format is unsupported.
"""
try:
return await asyncio.to_thread(
TextFileHandler._sync_text_to_docling_document, text, name, format
)
except Exception as e:
raise ValueError(f"Failed to convert text to DoclingDocument: {e}")

View file

@ -272,3 +272,31 @@ processing:
assert cfg.processing.conversion_options.picture_description.enabled is False assert cfg.processing.conversion_options.picture_description.enabled is False
assert "generate_picture_images" not in data["processing"]["conversion_options"] assert "generate_picture_images" not in data["processing"]["conversion_options"]
assert any("generate_picture_images" in r.getMessage() for r in handler.records) assert any("generate_picture_images" in r.getMessage() for r in handler.records)
def test_fetch_remote_images_default_true():
"""`fetch_remote_images` defaults to True and round-trips through YAML."""
from haiku.rag.config.models import ConversionOptions
assert ConversionOptions().fetch_remote_images is True
cfg = AppConfig()
assert cfg.processing.conversion_options.fetch_remote_images is True
data = generate_default_config()
assert data["processing"]["conversion_options"]["fetch_remote_images"] is True
def test_fetch_remote_images_override_via_yaml(tmp_path):
"""User can disable image fetching via YAML."""
config_file = _write(
tmp_path,
"""
processing:
conversion_options:
fetch_remote_images: false
""",
)
data = load_yaml_config(config_file)
cfg = AppConfig.model_validate(data)
assert cfg.processing.conversion_options.fetch_remote_images is False

View file

@ -349,6 +349,283 @@ class TestDoclingLocalConverter:
converter.config.processing.conversion_options.generate_page_images is False converter.config.processing.conversion_options.generate_page_images is False
) )
@pytest.mark.asyncio
async def test_convert_text_html_fetches_data_uri_image(self, config):
"""`fetch_remote_images=True` decodes inline `data:` URIs into picture
bytes via the HTML backend. Default behavior."""
png_b64 = (
"iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAQAAAC1HAwCAAAAC0lEQVR42mNk"
"+M9QDwADhgGAWjR9awAAAABJRU5ErkJggg=="
)
html = f'<html><body><p>before</p><img src="data:image/png;base64,{png_b64}" alt="dot"/><p>after</p></body></html>'
converter = DoclingLocalConverter(config)
doc = await converter.convert_text(html, format="html")
assert doc.pictures, "HTML with <img> should yield picture items"
pics_with_image = [p for p in doc.pictures if p.image is not None]
assert len(pics_with_image) == len(doc.pictures), (
"All <img> with valid data: URIs should have decoded bytes"
)
@pytest.mark.asyncio
async def test_convert_text_html_no_fetch_when_disabled(self, config):
"""`fetch_remote_images=False` produces placeholder pictures with no
bytes even for inline `data:` URIs (docling's `fetch_images` gates
all image decoding, not just remote fetches)."""
png_b64 = (
"iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAQAAAC1HAwCAAAAC0lEQVR42mNk"
"+M9QDwADhgGAWjR9awAAAABJRU5ErkJggg=="
)
html = f'<html><body><img src="data:image/png;base64,{png_b64}"/></body></html>'
config.processing.conversion_options.fetch_remote_images = False
converter = DoclingLocalConverter(config)
doc = await converter.convert_text(html, format="html")
assert doc.pictures, "Picture placeholders are still emitted"
for pic in doc.pictures:
assert pic.image is None, (
"fetch_remote_images=False must leave picture.image=None"
)
@pytest.mark.asyncio
async def test_convert_text_md_html_block_fetches_data_uri_image(self, config):
"""Markdown with an embedded `<img>` HTML block produces picture bytes
proves the MarkdownBackendOptions wiring delegates to the HTML
backend with our `fetch_images` / `enable_remote_fetch` settings.
Note: docling's md backend does NOT fetch images from native
`![alt](url)` syntax only from embedded HTML blocks. That's an
upstream limitation, not something this PR can address.
"""
png_b64 = (
"iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAQAAAC1HAwCAAAAC0lEQVR42mNk"
"+M9QDwADhgGAWjR9awAAAABJRU5ErkJggg=="
)
md = (
f"# Title\n\nIntro paragraph.\n\n"
f'<img src="data:image/png;base64,{png_b64}" alt="dot"/>\n\n'
f"Trailing paragraph.\n"
)
converter = DoclingLocalConverter(config)
doc = await converter.convert_text(md, format="md")
assert doc.pictures, "MD with <img> HTML block should yield picture items"
pics_with_image = [p for p in doc.pictures if p.image is not None]
assert len(pics_with_image) == len(doc.pictures)
def test_build_format_options_covers_pdf_image_html_md_docx_pptx(self, config):
"""`_build_format_options()` registers every format we care about and
shares the same `PdfPipelineOptions` instance across them so
picture-description / classification / chart settings apply uniformly."""
from docling.datamodel.base_models import InputFormat
converter = DoclingLocalConverter(config)
options = converter._build_format_options()
wired = {
InputFormat.PDF,
InputFormat.IMAGE,
InputFormat.HTML,
InputFormat.MD,
InputFormat.DOCX,
InputFormat.PPTX,
}
assert wired <= set(options.keys()), (
f"Missing format options: {wired - set(options.keys())}"
)
pdf_opts = options[InputFormat.PDF].pipeline_options
for fmt in wired:
assert options[fmt].pipeline_options is pdf_opts, (
f"{fmt} must share the PDF pipeline_options instance"
)
def test_build_format_options_propagates_fetch_remote_images(self, config):
"""HTML and Markdown FormatOptions reflect `fetch_remote_images`."""
from docling.datamodel.backend_options import (
HTMLBackendOptions,
MarkdownBackendOptions,
)
from docling.datamodel.base_models import InputFormat
config.processing.conversion_options.fetch_remote_images = True
opts = DoclingLocalConverter(config)._build_format_options()
html_bo = opts[InputFormat.HTML].backend_options
md_bo = opts[InputFormat.MD].backend_options
assert isinstance(html_bo, HTMLBackendOptions)
assert html_bo.fetch_images is True
assert html_bo.enable_remote_fetch is True
assert isinstance(md_bo, MarkdownBackendOptions)
assert md_bo.fetch_images is True
assert md_bo.enable_remote_fetch is True
config.processing.conversion_options.fetch_remote_images = False
opts = DoclingLocalConverter(config)._build_format_options()
html_bo = opts[InputFormat.HTML].backend_options
md_bo = opts[InputFormat.MD].backend_options
assert isinstance(html_bo, HTMLBackendOptions)
assert html_bo.fetch_images is False
assert html_bo.enable_remote_fetch is False
assert isinstance(md_bo, MarkdownBackendOptions)
assert md_bo.fetch_images is False
assert md_bo.enable_remote_fetch is False
def test_build_format_options_threads_source_uri(self, config):
"""`_build_format_options(source_uri=...)` plumbs the URI into the
HTML and Markdown backend options so docling can resolve relative
`<img src="/foo.jpg">` paths during URL ingest."""
from docling.datamodel.backend_options import (
HTMLBackendOptions,
MarkdownBackendOptions,
)
from docling.datamodel.base_models import InputFormat
converter = DoclingLocalConverter(config)
opts = converter._build_format_options(source_uri="https://example.com/article")
html_bo = opts[InputFormat.HTML].backend_options
md_bo = opts[InputFormat.MD].backend_options
assert isinstance(html_bo, HTMLBackendOptions)
assert str(html_bo.source_uri) == "https://example.com/article"
assert isinstance(md_bo, MarkdownBackendOptions)
assert str(md_bo.source_uri) == "https://example.com/article"
# No source_uri ⇒ both stay None
opts_no_uri = converter._build_format_options()
html_bo = opts_no_uri[InputFormat.HTML].backend_options
md_bo = opts_no_uri[InputFormat.MD].backend_options
assert isinstance(html_bo, HTMLBackendOptions)
assert html_bo.source_uri is None
assert isinstance(md_bo, MarkdownBackendOptions)
assert md_bo.source_uri is None
@pytest.mark.asyncio
async def test_convert_text_html_mixed_img_sources(self, config, monkeypatch):
"""End-to-end: HTML with a mix of remote http, data:, broken http, and
file:// `<img>` sources. Remote and data: URIs land as picture bytes;
broken URLs and file:// stay as placeholder pictures. Models the
wix-style ingest where most images are remote URLs with a handful of
broken or local-only references mixed in."""
import base64
from docling.backend import html_backend as html_backend_module
canned_png = base64.b64decode(
"iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAQAAAC1HAwCAAAAC0lEQVR42mNk"
"+M9QDwADhgGAWjR9awAAAABJRU5ErkJggg=="
)
good_url = "https://cdn.example.com/static/cat.png"
broken_url = "https://cdn.example.com/missing.png"
def fake_load_image_data(self, src_loc: str):
if src_loc == good_url:
return canned_png
if src_loc == broken_url:
# Simulate a 404: docling's _create_image_ref swallows HTTPError
# via its except clause and returns None for the picture.
import requests
resp = requests.Response()
resp.status_code = 404
raise requests.HTTPError(response=resp)
return None # data: and file:// fall back to docling's own path
# Wrap rather than replace so data: URIs still decode through the real
# `_load_image_data`. Only intercept when src is one of our test URLs.
original = html_backend_module.HTMLDocumentBackend._load_image_data
def wrapped(self, src_loc: str):
if src_loc in (good_url, broken_url):
return fake_load_image_data(self, src_loc)
return original(self, src_loc)
monkeypatch.setattr(
html_backend_module.HTMLDocumentBackend, "_load_image_data", wrapped
)
png_b64 = (
"iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAQAAAC1HAwCAAAAC0lEQVR42mNk"
"+M9QDwADhgGAWjR9awAAAABJRU5ErkJggg=="
)
html = (
"<html><body>"
f'<img src="{good_url}" alt="cat"/>'
f'<img src="data:image/png;base64,{png_b64}" alt="dot"/>'
f'<img src="{broken_url}" alt="missing"/>'
'<img src="file:///etc/passwd" alt="local"/>'
"</body></html>"
)
converter = DoclingLocalConverter(config)
doc = await converter.convert_text(html, format="html")
assert len(doc.pictures) == 4, f"expected 4 pictures, got {len(doc.pictures)}"
with_bytes = sum(1 for p in doc.pictures if p.image is not None)
without_bytes = sum(1 for p in doc.pictures if p.image is None)
assert with_bytes == 2, (
f"good remote + data: should produce bytes; got {with_bytes}"
)
assert without_bytes == 2, (
f"broken http + file:// should stay placeholder; got {without_bytes}"
)
def test_image_format_shares_pdf_pipeline_options(self, config):
"""IMAGE FormatOption shares the same PdfPipelineOptions instance as
PDF. Without this, `do_ocr` / `picture_description.enabled` / etc.
silently no-op when ingesting raw `.png` / `.jpg` files (which run
through StandardPdfPipeline). End-to-end image conversion is covered
by the PDF picture test both paths feed the same pipeline class."""
from docling.datamodel.base_models import InputFormat
from docling.datamodel.pipeline_options import PdfPipelineOptions
config.processing.conversion_options.do_ocr = False
config.processing.conversion_options.images_scale = 3.5
fmt_opts = DoclingLocalConverter(config)._build_format_options()
pdf_pipe = fmt_opts[InputFormat.PDF].pipeline_options
image_pipe = fmt_opts[InputFormat.IMAGE].pipeline_options
assert image_pipe is pdf_pipe
assert isinstance(pdf_pipe, PdfPipelineOptions)
assert pdf_pipe.do_ocr is False
assert pdf_pipe.images_scale == 3.5
@pytest.mark.asyncio
async def test_convert_text_html_source_uri_resolves_relative_img(
self, config, monkeypatch
):
"""`convert_text(..., source_uri=...)` lets docling resolve a relative
`<img src="/path">` against the source URL. We patch the docling HTML
backend's `_load_image_data` to capture the resolved absolute URL
instead of doing a real network fetch."""
captured: list[str] = []
def fake_load_image_data(self, src_loc: str):
captured.append(src_loc)
return None # docling treats as a fetch failure → placeholder
from docling.backend import html_backend as html_backend_module
monkeypatch.setattr(
html_backend_module.HTMLDocumentBackend,
"_load_image_data",
fake_load_image_data,
)
html = '<html><body><img src="/static/cat.jpg"/></body></html>'
converter = DoclingLocalConverter(config)
await converter.convert_text(
html, format="html", source_uri="https://example.com/article"
)
assert captured, "_load_image_data should have been invoked"
assert captured[0] == "https://example.com/static/cat.jpg", (
f"Expected absolute URL resolved via source_uri, got {captured[0]!r}"
)
@pytest.mark.asyncio @pytest.mark.asyncio
async def test_convert_pdf_with_picture_images(self, config): async def test_convert_pdf_with_picture_images(self, config):
"""Picture bytes are produced by the local converter for PDFs that """Picture bytes are produced by the local converter for PDFs that

View file

@ -143,11 +143,15 @@ async def test_convert_emits_warning_via_chokepoint(
def supported_extensions(self) -> list[str]: def supported_extensions(self) -> list[str]:
return [".pdf"] return [".pdf"]
async def convert_file(self, path: Path): async def convert_file(self, path: Path, source_uri: str | None = None):
return _doc_with_pictures(with_descriptions=False) return _doc_with_pictures(with_descriptions=False)
async def convert_text( async def convert_text(
self, text: str, name: str = "content.md", format: str = "md" self,
text: str,
name: str = "content.md",
format: str = "md",
source_uri: str | None = None,
): ):
return _doc_without_pictures() return _doc_without_pictures()
@ -180,11 +184,15 @@ async def test_convert_text_path_also_warns(monkeypatch, caplog_warnings):
def supported_extensions(self) -> list[str]: def supported_extensions(self) -> list[str]:
return [".html"] return [".html"]
async def convert_file(self, path: Path): async def convert_file(self, path: Path, source_uri: str | None = None):
return _doc_without_pictures() return _doc_without_pictures()
async def convert_text( async def convert_text(
self, text: str, name: str = "content.md", format: str = "md" self,
text: str,
name: str = "content.md",
format: str = "md",
source_uri: str | None = None,
): ):
return _doc_with_pictures(with_descriptions=False) return _doc_with_pictures(with_descriptions=False)