haiku.rag/evaluations/evaluations
Yiorgis Gozadinos 09a7076b7e
State what the code does, not what it replaced
Comments and docstrings across the branch narrated rejected
alternatives, consequences and history; each now states the current
contract. Renames test_a_legacy_uri_client_keeps_its_error to
test_an_unnamed_database_keeps_its_error. Documents the Sandbox
connection paths, the citation header's database segment, both
AmbiguousDatabaseError conditions on create_app, and run_inspector's
scope parameter. Doc paragraphs added by the branch in python.md,
storage.md and cli.md are one physical line each.
2026-08-28 15:13:52 +03:00
..
datasets Add stable ids to FRAMES question rows 2026-08-24 09:03:44 +03:00
evaluators Simplify the eval harness and share the embed-fill path. 2026-08-17 10:56:26 +03:00
__init__.py Restructure into uv workspace to support minimal and full installations 2025-11-04 17:59:12 +02:00
artifacts.py Split the evaluation benchmark by responsibility 2026-08-20 14:08:09 +03:00
benchmark.py Name the evaluations set check for what it answers 2026-08-26 13:04:06 +03:00
capability_runner.py State what the code does, not what it replaced 2026-08-28 15:13:52 +03:00
config.py Pin the over-fetch rule, and assert the type a lookup raises 2026-08-28 14:03:13 +03:00
experiment.py Split the evaluation benchmark by responsibility 2026-08-20 14:08:09 +03:00
numbers.py Normalize unicode signs 2026-06-06 14:52:04 +03:00
population.py Split the evaluation benchmark by responsibility 2026-08-20 14:08:09 +03:00
qa.py State what the code does, not what it replaced 2026-08-28 15:13:52 +03:00
retrieval.py Name the evaluations set check for what it answers 2026-08-26 13:04:06 +03:00
submission.py replace haiku.skills with native Pydantic AI capabilities 2026-07-24 15:26:17 +03:00