haiku.rag/evaluations
Yiorgis Gozadinos a75d89122e
Drop the unused per-dataset filter default
No `DatasetSpec` declared `search_filter`, so `resolve_search_filter`
and the `--filter ""` clearing rule reconciled the flag against a
default that never existed. The flag alone covers the case. An empty
clause reaches `ChunkRepository.search`, which already treats it as
unfiltered.

Rename to `document_filter` throughout, matching
`run_capability_question`'s parameter and the metadata key that lands
in Logfire.

`_stub_spec` merges its overrides, so a test can override a loader
instead of rebuilding the whole spec.
2026-08-17 10:29:28 +03:00
..
configs Remove the wix evaluation dataset 2026-08-14 14:46:51 +03:00
evaluations Drop the unused per-dataset filter default 2026-08-17 10:29:28 +03:00
scripts Add T²-RAGBench leaderboard submission exporter 2026-06-08 11:51:02 +03:00
tests Drop the unused per-dataset filter default 2026-08-17 10:29:28 +03:00
LICENSE Restructure into uv workspace to support minimal and full installations 2025-11-04 17:59:12 +02:00
pyproject.toml vb 2026-08-13 16:16:35 +03:00
README.md Remove the wix evaluation dataset 2026-08-14 14:46:51 +03:00

Haiku RAG - Evaluations

Internal benchmarking and evaluation scripts for haiku.rag.

This package is not published to PyPI and is only used for development and testing purposes.

Overview

Contains evaluation scripts for benchmarking RAG retrieval and QA performance. Available datasets:

  • HotpotQA (hotpotqa) — multi-hop QA over Wikipedia paragraphs (distractor validation split, 7,405 questions, two gold documents per question)
  • OpenRAG Bench, two variants:
    • orb_text — text embedder (qwen3-embedding:4b, 2560-dim) with VLM picture descriptions baked into chunk content at ingest. Use for text-only retrieval/QA against figure-rich corpora.
    • orb_multimodal — multimodal embedder (qwen3-vl-embedding-8b, 4096-dim) with picture vectors in the same space as text. Use for cross-modal retrieval (text-as-query → figure hits, image-as-query) and vision QA where the figure itself is the answer.

Usage

After installing the package, you can run evaluations using the evaluations command:

# Run retrieval + QA benchmarks
evaluations run hotpotqa
evaluations run orb_text

# Use a custom config file
evaluations run hotpotqa --config /path/to/haiku.rag.yaml

# Override the database path
evaluations run hotpotqa --db /path/to/custom.lancedb

# Skip database population and run only benchmarks
evaluations run hotpotqa --skip-db

# Skip specific benchmarks
evaluations run hotpotqa --skip-retrieval
evaluations run hotpotqa --skip-qa

# Limit the number of test cases
evaluations run hotpotqa --limit 100

Choosing the target

evaluations run benchmarks --target rag-capability by default. Use --target analysis-capability to benchmark the analysis capability against the same datasets and judge:

evaluations run hotpotqa --target rag-capability
evaluations run hotpotqa --target analysis-capability --capability-model ollama:gpt-oss

--capability-model "provider:name" overrides the capability model independently from the judge (defaults to qa.model, or analysis.model when set for the analysis-capability target). A citation retrieval metric (cited_map) is computed alongside QA accuracy from the URIs the capability registered via the cite tool.

Debugging runs in Logfire

With LOGFIRE_TOKEN set, runs ship spans under service_name = 'evals'. The debug-evals skill in .claude/skills/ turns these into ready-made Logfire queries (recent runs, per-case pass rate and cited_map, failing and slowest cases) for use from Claude Code.

Pre-built Databases

Download pre-built evaluation databases from HuggingFace:

evaluations download hotpotqa
evaluations download all
evaluations download hotpotqa --force

Upload databases (maintainer only):

evaluations upload hotpotqa
evaluations upload all

Database Storage

By default, evaluation databases are stored in the haiku.rag data directory:

  • Linux: ~/.local/share/haiku.rag/evaluations/dbs/
  • macOS: ~/Library/Application Support/haiku.rag/evaluations/dbs/
  • Windows: C:/Users/<USER>/AppData/Roaming/haiku.rag/evaluations/dbs/

You can override this with the --db option.