haiku.rag/docs/benchmarks.md
2025-07-08 12:51:58 +03:00

889 B

haiku.rag benchmarks

We use repliqa for the evaluation of haiku.rag

  • Recall

We load the News Stories from repliqa_3 which is 1035 documents, using tests/generate_benchmark_db.py, using the mxbai-embed-large Ollama embeddings.

Subsequently, we run a search over the question for each row of the dataset and check whether we match the document that answers the question. The recall obtained is ~0.75 for matching in the top result, raising to ~0.75 for the top 3 results.

  • Question/Answer evaluation

We use the News Stories from repliqa_3 using the mxbai-embed-large Ollama embeddings, with a QA agent also using Ollama with the qwen3 model (8b). For each story we ask the question and use an LLM judge (also qwen3) to evaluate whether the answer is correct or not. Thus we obtain accuracy of ~0.54.