The two hard gates are pass/fail, not rates: any scope leak or attribution error is a bug. The report names the offending cases and prints a greppable GATES: PASSED / FAILED line, since this dataset is an acceptance gate rather than a number to watch drift on. The attribution gate covers only families where exactly one database is correct. B5 is answered by two, so citing both is right there and policing it would report a bug that is not one. Surface coverage comes from the model's own Python, now carried on the run result: counting executions said how often code ran but never which surface it reached. Coverage is measured, not required, so a surface nothing reaches is a finding about whether it earns its place in the VFS. answer_correct is recorded as a score rather than a boolean, since booleans become assertions and the run reads its accuracy headline from scores. The two gates stay boolean assertions, which is what they are. DatasetSpec grows a report_hook so this lives with the dataset instead of becoming a key check in the shared QA reporting. The reference-config invariant said a dataset with a deterministic evaluator must declare no judge. That is not true here: RefusalJudge runs on any case carrying an answerability label, which B6 and B7 do, so its sampling has to be pinned. The assertion now requires a pinned block wherever a judge runs and allows its absence only as the claim that none does.
68 lines
2.1 KiB
YAML
68 lines
2.1 KiB
YAML
# Reference config for the `multidb_surfaces` analysis families (sandbox surfaces).
|
|
# lancedb.databases. Build the corpus first, then run with --skip-db:
|
|
# uv run python -m evaluations.datasets.multidb --config configs/multidb_surfaces.yaml
|
|
# evaluations run multidb_surfaces --config configs/multidb_surfaces.yaml --skip-db --skip-retrieval
|
|
# For the no-reranker arm, copy this and drop the reranking block; a reference
|
|
# config's filename has to name a dataset, so that arm lives outside configs/.
|
|
|
|
environment: development
|
|
|
|
storage:
|
|
auto_vacuum: false
|
|
|
|
lancedb:
|
|
# Order matters: RRF ties resolve to the order listed here, so the first
|
|
# database wins ties. B2 rotates the order it passes per case.
|
|
databases:
|
|
northern: ${HOME}/.local/share/haiku.rag/evaluations/dbs/multidb_northern.lancedb
|
|
southern: ${HOME}/.local/share/haiku.rag/evaluations/dbs/multidb_southern.lancedb
|
|
equipment: ${HOME}/.local/share/haiku.rag/evaluations/dbs/multidb_equipment.lancedb
|
|
|
|
processing:
|
|
# Pinned: the builder asserts no single chunk holds all twelve monthly
|
|
# readings, so S3 cannot be answered from one search hit.
|
|
chunk_size: 256
|
|
|
|
search:
|
|
# Three databases at limit 5 means RRF fuses 15 candidates to 5, so the
|
|
# truncation bites and the no-reranker arm can lose evidence.
|
|
limit: 5
|
|
|
|
embeddings:
|
|
model:
|
|
provider: openai
|
|
name: qwen3-embedding-4b
|
|
vector_dim: 2560
|
|
base_url: http://vllm:11431/v1
|
|
|
|
reranking:
|
|
model:
|
|
provider: vllm
|
|
name: Qwen/Qwen3-Reranker-4B
|
|
base_url: http://vllm:11455
|
|
|
|
qa:
|
|
model:
|
|
provider: openai
|
|
name: Inferact/Qwen3.8-27B-NVFP4
|
|
base_url: http://vllm:11439/v1
|
|
max_tokens: 16384
|
|
extra_body:
|
|
chat_template_kwargs:
|
|
reasoning_effort: low
|
|
|
|
evaluations:
|
|
# Scoring is deterministic, but RefusalJudge runs on the B6 and B7 cases,
|
|
# which carry answerability labels.
|
|
judge:
|
|
provider: openai
|
|
name: Inferact/Qwen3.8-27B-NVFP4
|
|
base_url: http://vllm:11439/v1
|
|
temperature: 0.6
|
|
max_tokens: 16384
|
|
extra_body:
|
|
top_p: 0.95
|
|
top_k: 20
|
|
min_p: 0
|
|
chat_template_kwargs:
|
|
reasoning_effort: low
|