Searches a model emits in one response share a budget unit, up to FREE_SIBLINGS_PER_ROUND (3) per unit; sequential searches pay one unit each, as before. Grouping keys on RunContext.run_step, which pydantic-ai increments once per model request. A budget-rejected round fails all its remaining siblings, and tracking resets per run. Glimmer opens most questions with a burst of ~3 rephrasings in a single response (95.8% of its three-search ORB cases are one-response bursts), spending 3 of 5 searches before reading anything. Pass rate at 3 calls equals 1 call, so a burst is priced as one probe. |
||
|---|---|---|
| .. | ||
| __init__.py | ||
| conftest.py | ||
| helpers.py | ||
| test_capabilities.py | ||
| test_citations.py | ||
| test_documents.py | ||
| test_expansion.py | ||
| test_lifecycle.py | ||
| test_scope.py | ||
| test_search.py | ||