haiku.rag/evaluations/tests
2026-06-08 11:51:02 +03:00
..
__init__.py
test_benchmark.py Add --filter-ids to run QA on a case-id subset 2026-06-08 09:36:01 +03:00
test_citation_evaluators.py
test_config.py
test_datasets.py Add T²-RAGBench TAT-DQA subset; generalize subset layout 2026-06-06 14:52:05 +03:00
test_evaluators.py Match the numeric scale convention in Number-Match 2026-06-06 14:52:05 +03:00
test_numbers.py Normalize unicode signs 2026-06-06 14:52:04 +03:00
test_skill_runner.py
test_submission.py Add T²-RAGBench leaderboard submission exporter 2026-06-08 11:51:02 +03:00