A run over several databases could report which documents were cited but not which database grounded the answer: `_result_from_run` walked the citation index for `document_uri` and dropped `Citation.source`. The distribution is not recoverable from the report afterwards, so a sharded run would have measured everything except attribution. `cited_sources` is one entry per cited chunk, in citation order, empty where the database is unnamed. |
||
|---|---|---|
| .. | ||
| datasets | ||
| evaluators | ||
| __init__.py | ||
| artifacts.py | ||
| benchmark.py | ||
| capability_runner.py | ||
| config.py | ||
| experiment.py | ||
| numbers.py | ||
| population.py | ||
| qa.py | ||
| retrieval.py | ||
| submission.py | ||