RAGState and AnalysisState each declared the same five fields, so the generic base could not name them: StateT was bound to BaseModel, and every access went through cast(Any, state), a getattr by string, or a loop clearing fields by name so it could skip the one only AnalysisState has. EvidenceState declares them once. RAGState adds nothing, AnalysisState adds executions and overrides begin_invocation to clear them. StateT binds to EvidenceState, which removes all ten casts and both state-shape getattrs; the three getattr(ctx.deps, "state") probes stay, since those check a host-supplied object rather than our own state. discover_evidence reached into capability.state for two fields. It now asks through evidence_record() and citation_index(), alongside the evidence_tool_names() and cite_available accessors it already used. The eval runner's _RagLikeState protocol and the chat app's getattr reads described this shape from outside and are gone. Compatibility is semantic JSON-object equivalence, not bytes: field names and nesting are unchanged, so a dict stored by 0.75.0 loads and re-dumps equal, but deriving from a shared base reorders AnalysisState's keys. Nothing serializes, hashes or string-compares this state — every carry point re-validates by key. |
||
|---|---|---|
| .. | ||
| configs | ||
| evaluations | ||
| scripts | ||
| tests | ||
| LICENSE | ||
| pyproject.toml | ||
| README.md | ||
haiku.rag - Evaluations
Internal benchmarking and evaluation scripts for haiku.rag.
This package is not published to PyPI and is only used for development and testing purposes.
Overview
Contains evaluation scripts for benchmarking RAG retrieval and QA performance. Available datasets:
- HotpotQA (
hotpotqa) — multi-hop QA over Wikipedia paragraphs (distractor validation split, 7,405 questions, two gold documents per question) - MTRAG ClapNQ (
mtrag_clapnq,mtrag_clapnq_rewrite) — IBM's multi-turn RAG benchmark, ClapNQ (Wikipedia) domain: 183,408 passages, 208 retrieval queries with binary qrels, 224 generation tasks. The base key retrieves with the raw last user turn; the_rewritevariant uses the human standalone rewrites (both share one database). Retrieval reports Recall@5/@10, nDCG@5/@10, and MAP against IBM's published setup. QA replays each task's reference conversation prefix as message history and answers the final turn; the judge sees the conversation as a transcript, citation MAP is scored only on turns with gold passages, and refusal precision/recall is reported against the answerability labels. Generation scores are internal (our judge and rubric), not comparable with IBM's published generation numbers. Themtrag_clapnq_livekey replays whole conversations (one case per conversation,--limitcounts conversations) through a single capability session, carrying the model's own answers and tool history across turns; it reports the same outcomes per turn plus micro (per-turn) and macro (per-conversation) aggregates. - OpenRAG Bench, two variants:
orb_text— text embedder (qwen3-embedding:4b, 2560-dim) with VLM picture descriptions baked into chunk content at ingest. Use for text-only retrieval/QA against figure-rich corpora.orb_multimodal— multimodal embedder (qwen3-vl-embedding-8b, 4096-dim) with picture vectors in the same space as text. Use for cross-modal retrieval (text-as-query → figure hits, image-as-query) and vision QA where the figure itself is the answer.
Usage
After installing the package, you can run evaluations using the evaluations command:
# Run retrieval + QA benchmarks
evaluations run hotpotqa
evaluations run orb_text
# Use a custom config file
evaluations run hotpotqa --config /path/to/haiku.rag.yaml
# Override the database path
evaluations run hotpotqa --db /path/to/custom.lancedb
# Skip database population and run only benchmarks
evaluations run hotpotqa --skip-db
# Skip specific benchmarks
evaluations run hotpotqa --skip-retrieval
evaluations run hotpotqa --skip-qa
# Limit the number of test cases
evaluations run hotpotqa --limit 100
Choosing the target
evaluations run benchmarks --target rag-capability by default. Use
--target analysis-capability to benchmark the analysis capability against the same
datasets and judge:
evaluations run hotpotqa --target rag-capability
evaluations run hotpotqa --target analysis-capability --capability-model ollama:gpt-oss
--capability-model "provider:name" overrides the capability model independently from
the judge (defaults to qa.model, or analysis.model when set for the
analysis-capability target). A citation retrieval metric (cited_map) is computed
alongside QA accuracy from the URIs the capability registered via the cite tool.
Debugging runs in Logfire
With LOGFIRE_TOKEN set, runs ship spans under service_name = 'evals'. The
debug-evals skill in .claude/skills/ turns these into ready-made Logfire
queries (recent runs, per-case pass rate and cited_map, failing and slowest
cases) for use from Claude Code.
Pre-built Databases
Download pre-built evaluation databases from HuggingFace:
evaluations download hotpotqa
evaluations download all
evaluations download hotpotqa --force
Upload databases (maintainer only):
evaluations upload hotpotqa
evaluations upload all
Database Storage
By default, evaluation databases are stored in the haiku.rag data directory:
- Linux:
~/.local/share/haiku.rag/evaluations/dbs/ - macOS:
~/Library/Application Support/haiku.rag/evaluations/dbs/ - Windows:
C:/Users/<USER>/AppData/Roaming/haiku.rag/evaluations/dbs/
You can override this with the --db option.