haiku.rag/evaluations/evaluations
2026-08-06 13:17:58 +03:00
..
datasets Restore hotpotqa evaluation dataset 2026-07-17 16:25:27 +03:00
evaluators replace haiku.skills with native Pydantic AI capabilities 2026-07-24 15:26:17 +03:00
__init__.py Restructure into uv workspace to support minimal and full installations 2025-11-04 17:59:12 +02:00
benchmark.py Pin the eval judge sampling and standardise on Qwen3-Reranker 2026-08-06 13:17:58 +03:00
capability_runner.py Scope the limit notice and spend the cite window on own turns only 2026-07-30 19:14:14 +03:00
config.py Add deterministic Number-Match QA scoring for T²-RAGBench 2026-06-06 14:52:04 +03:00
numbers.py Normalize unicode signs 2026-06-06 14:52:04 +03:00
submission.py replace haiku.skills with native Pydantic AI capabilities 2026-07-24 15:26:17 +03:00