haiku.rag/evaluations/tests
2026-08-14 14:46:51 +03:00
..
__init__.py Tests for gepa 2026-03-12 12:05:25 +02:00
test_benchmark.py Remove the wix evaluation dataset 2026-08-14 14:46:51 +03:00
test_capability_runner.py Scope the limit notice and spend the cite window on own turns only 2026-07-30 19:14:14 +03:00
test_citation_evaluators.py Remove unecessary Mean Reciprocal Rank metric 2026-06-01 10:40:52 +03:00
test_config.py remove dataset-specific system prompts 2026-04-28 14:33:25 +03:00
test_datasets.py Remove the wix evaluation dataset 2026-08-14 14:46:51 +03:00
test_evaluators.py Match the numeric scale convention in Number-Match 2026-06-06 14:52:05 +03:00
test_numbers.py Normalize unicode signs 2026-06-06 14:52:04 +03:00
test_reference_configs.py Pin the eval judge sampling and standardise on Qwen3-Reranker 2026-08-06 13:17:58 +03:00
test_submission.py Add T²-RAGBench leaderboard submission exporter 2026-06-08 11:51:02 +03:00