haiku.rag/evaluations/evaluations
Yiorgis Gozadinos 503db3271c
Name the evaluations set check for what it answers
`covers_a_set` is true for a mapping of one, which is a configured database
like any other; `uses_configured_databases` says that. Its population guard
said "several" for the same reason. Document evaluating the configured set
with `--skip-db`, against population, which writes one database and needs a
path.
2026-08-26 13:04:06 +03:00
..
datasets Add stable ids to FRAMES question rows 2026-08-24 09:03:44 +03:00
evaluators Simplify the eval harness and share the embed-fill path. 2026-08-17 10:56:26 +03:00
__init__.py Restructure into uv workspace to support minimal and full installations 2025-11-04 17:59:12 +02:00
artifacts.py Split the evaluation benchmark by responsibility 2026-08-20 14:08:09 +03:00
benchmark.py Name the evaluations set check for what it answers 2026-08-26 13:04:06 +03:00
capability_runner.py Record the database each citation came from 2026-08-24 10:03:47 +03:00
config.py Name the evaluations set check for what it answers 2026-08-26 13:04:06 +03:00
experiment.py Split the evaluation benchmark by responsibility 2026-08-20 14:08:09 +03:00
numbers.py Normalize unicode signs 2026-06-06 14:52:04 +03:00
population.py Split the evaluation benchmark by responsibility 2026-08-20 14:08:09 +03:00
qa.py Name the evaluations set check for what it answers 2026-08-26 13:04:06 +03:00
retrieval.py Name the evaluations set check for what it answers 2026-08-26 13:04:06 +03:00
submission.py replace haiku.skills with native Pydantic AI capabilities 2026-07-24 15:26:17 +03:00