haiku.rag/docs/benchmarks.md
2025-09-30 11:31:52 +03:00

4.1 KiB

Benchmarks

We use the repliqa dataset for the evaluation of haiku.rag.

You can perform your own evaluations with the Typer CLI in evaluations/benchmark.py, for example python -m evaluations.benchmark repliqa. The evaluation flow is orchestrated with pydantic-evals, which we leverage for dataset management, scoring, and report generation.

Recall

In order to calculate recall, we load the News Stories from repliqa_3 (1035 documents) and index them. Subsequently, we run a search over the question field for each row of the dataset and check whether we match the document that answers the question. Questions for which the answer cannot be found in the documents are ignored.

The recall obtained is ~0.79 for matching in the top result, raising to ~0.91 for the top 3 results with the "bare" default settings (Ollama qwen3, mxbai-embed-large embeddings, no reranking).

Embedding Model Document in top 1 Document in top 3 Reranker
Ollama / qwen3-embedding 0.81 0.95 None
Ollama / qwen3-embedding 0.91 0.98 mxbai-rerank-base-v2
Ollama / mxbai-embed-large 0.79 0.91 None
Ollama / mxbai-embed-large 0.90 0.95 mxbai-rerank-base-v2
Ollama / nomic-embed-text-v1.5 0.74 0.90 None

Question/Answer evaluation

Again using the same dataset, we use a QA agent to answer the question. pydantic-evals runs each case and coordinates an LLM judge (Ollama qwen3) to determine whether the answer is correct. The obtained accuracy is as follows:

Embedding Model QA Model Accuracy Reranker
Ollama / qwen3-embedding. Ollama / gpt-oss 0.93 None
Ollama / mxbai-embed-large Ollama / qwen3 0.85 None
Ollama / mxbai-embed-large Ollama / qwen3 0.87 mxbai-rerank-base-v2
Ollama / mxbai-embed-large Ollama / qwen3:0.6b 0.28 None

Note the significant degradation when very small models are used such as qwen3:0.6b.

Wix dataset

We also track retrieval performance on WixQA, a dataset of real customer support questions paired with curated answers from Wix. The benchmark follows the evaluation protocol described in the WixQA paper and gives us a view into how the system handles conversational, product-specific support queries.

For recall, we index the reference answer passages shipped with the dataset and run retrieval against each user question. Each sample supplies one or more relevant passage URIs; we count how many of those URIs land inside the top k retrieved documents, divide by the number of relevant passages for that query, and average across all queries.

The results for recall using the WixQA dataset are as follows:

Embedding Model Document in top 1 Document in top 3 Reranker
qwen3-embedding 0.36 0.57 mxbai-rerank-base-v2

And for QA accuracy,

Embedding Model QA Model Accuracy Reranker
qwen3-embedding gpt-oss 0.75 mxbai-rerank-base-v2