haiku.rag/evaluations/tests
2026-06-01 10:40:51 +03:00
..
__init__.py Tests for gepa 2026-03-12 12:05:25 +02:00
test_benchmark.py Refresh benchmarks doc and remove unused eval datasets 2026-06-01 10:40:51 +03:00
test_citation_evaluators.py bump pydantic-ai-slim to 1.100 and haiku.skills to 0.17 2026-05-21 13:20:54 +03:00
test_config.py remove dataset-specific system prompts 2026-04-28 14:33:25 +03:00
test_datasets.py Refresh benchmarks doc and remove unused eval datasets 2026-06-01 10:40:51 +03:00
test_evaluators.py Test evaluations 2026-03-12 12:05:50 +02:00
test_skill_runner.py Drop the multi-agent research workflow 2026-05-20 12:46:48 +03:00