haiku.rag/evaluations/tests
2026-08-17 11:19:59 +03:00
..
__init__.py
test_benchmark.py Exclude fully unjudged conversations from the macro pass rate 2026-08-17 11:19:59 +03:00
test_capability_runner.py Thread the document filter through the live QA runner 2026-08-17 11:03:52 +03:00
test_citation_evaluators.py Add MTRAG ClapNQ multi-turn evaluation 2026-08-17 10:53:16 +03:00
test_config.py Add MTRAG ClapNQ multi-turn evaluation 2026-08-17 10:53:16 +03:00
test_conversation_evaluator.py Add MTRAG ClapNQ multi-turn evaluation 2026-08-17 10:53:16 +03:00
test_datasets.py Remove the wix evaluation dataset 2026-08-14 14:46:51 +03:00
test_evaluators.py Add MTRAG ClapNQ multi-turn evaluation 2026-08-17 10:53:16 +03:00
test_mtrag.py Simplify the eval harness and share the embed-fill path. 2026-08-17 10:56:26 +03:00
test_numbers.py
test_reference_configs.py Pin the eval judge sampling and standardise on Qwen3-Reranker 2026-08-06 13:17:58 +03:00
test_submission.py