haiku.rag/evaluations/evaluations
2026-05-20 12:56:22 +03:00
..
datasets split open_rag_bench dataset into orb_text and orb_multimodal variants 2026-05-06 12:49:03 +03:00
evaluators Bump pydantic-ai, prepare for 2.* 2026-05-18 15:14:29 +03:00
__init__.py Restructure into uv workspace to support minimal and full installations 2025-11-04 17:59:12 +02:00
benchmark.py evaluations: judge model moves to config.evaluations.judge 2026-05-20 12:56:22 +03:00
config.py remove dataset-specific system prompts 2026-04-28 14:33:25 +03:00
skill_runner.py Drop the multi-agent research workflow 2026-05-20 12:46:48 +03:00