haiku.rag/evaluations/evaluations
2026-06-03 16:03:42 +03:00
..
datasets exclude corrupt mi_phone.pdf from MMLongBench-Doc 2026-06-03 16:03:42 +03:00
evaluators Remove unecessary Mean Reciprocal Rank metric 2026-06-01 10:40:52 +03:00
__init__.py Restructure into uv workspace to support minimal and full installations 2025-11-04 17:59:12 +02:00
benchmark.py constrain MMLongBench-Doc QA retrieval to the target document 2026-06-03 15:00:02 +03:00
config.py remove dataset-specific system prompts 2026-04-28 14:33:25 +03:00
skill_runner.py Lift analysis-skill citation rate via SKILL.md tightening 2026-06-01 18:58:52 +03:00