haiku.rag/evaluations/evaluations
2026-06-01 10:40:51 +03:00
..
datasets Refresh benchmarks doc and remove unused eval datasets 2026-06-01 10:40:51 +03:00
evaluators bump pydantic-ai-slim to 1.100 and haiku.skills to 0.17 2026-05-21 13:20:54 +03:00
__init__.py Restructure into uv workspace to support minimal and full installations 2025-11-04 17:59:12 +02:00
benchmark.py evaluations: judge model moves to config.evaluations.judge 2026-05-20 12:56:22 +03:00
config.py remove dataset-specific system prompts 2026-04-28 14:33:25 +03:00
skill_runner.py Drop the multi-agent research workflow 2026-05-20 12:46:48 +03:00