The acceptance dataset for lancedb.databases needs a corpus where the expected database is known per question, which no existing dataset gives. Three databases: northern and southern hold station reports on one schema, equipment holds spec sheets that share the vocabulary and answer nothing. Names are invented throughout. A real station lets the model answer from priors, which would measure memorisation rather than retrieval. Three near-name pairs and one shared entity carry the attribution cases. Pair members are identical apart from the name and the numbers, since a difference in instrument or technician would hand the model a free discriminator. Elevations and years are unique across the corpus so a number identifies one station, and document counts differ per database so a count cannot be right by luck while attribution is wrong. Questions and gold answers both derive from STATIONS and INSTRUMENTS, so they cannot drift apart. Readings are derived with SHA-256 over database and name: hash() is salted per process, and keying on the name alone gave the two Station Auk reports identical tables. evaluations run populates one database and refuses a configured set, so the builder opens each database itself and the run is --skip-db. It asserts no single chunk holds all twelve monthly readings, because without that S3 collapses into a search question and the guarantee has to survive a chunker change. |
||
|---|---|---|
| .. | ||
| __init__.py | ||
| test_benchmark.py | ||
| test_capability_runner.py | ||
| test_citation_evaluators.py | ||
| test_config.py | ||
| test_conversation_evaluator.py | ||
| test_datasets.py | ||
| test_evaluators.py | ||
| test_mtrag.py | ||
| test_multidb_corpus.py | ||
| test_numbers.py | ||
| test_reference_configs.py | ||
| test_submission.py | ||