Retrieval and QA only read from the database, but the benchmark opened it
writable, where an embedder identity differing from the stored one aborts
instead of warning. Running a pre-built database against a different
serving stack then needed a `rebuild --set-embedder` first.
Correct the debug-evals skill alongside it: the pydantic-ai span names are
`execute_tool {tool_name}` and `invoke_agent agent`, targets are
`{rag,analysis}-capability`, and no `skill_model` metadata key exists.
|
||
|---|---|---|
| .. | ||
| datasets | ||
| evaluators | ||
| __init__.py | ||
| benchmark.py | ||
| capability_runner.py | ||
| config.py | ||
| numbers.py | ||
| submission.py | ||