litellm was doing nothing this project needs. Nine call sites, all the same
shape — model, messages, a temperature, an api_base pointing at the proxy — and
no streaming, tools, response_format, fallbacks, retries, Router or cost
tracking anywhere. `_proxy_model()` prefixed every model with `openai/`
specifically to stop litellm routing by provider, which is to say the SDK was
configured to behave like the OpenAI client it now is. Embeddings, model
discovery, speech and transcription already went over plain httpx.
The client is built in one place instead of thirteen assembled kwargs dicts,
and three things about it are deliberate: the base URL normalises to end in
`/v1`, because the SDK appends to whatever root it gets and litellm happened to
tolerate the bare host; `max_retries=0`, because the SDK retries twice by
default and would have turned the hand-written three attempts in
`extract_questions` into nine; and a placeholder key when none is configured,
so an unconfigured deployment fails at the request with the 502 every call site
expects rather than inside the constructor with a 500.
Verified against the live proxy rather than only against mocks: sync client,
async client and `_call_model` each returned from llm.danvics.com, and the
service boots clean. Nine distributions dropped, 156 to 147.
This also unblocked requirements.txt, which could not be edited at all while
litellm==1.28.13 — withdrawn from PyPI — was pinned in it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN