pdf-quiz-generator/docs/reranking.md
Daniel 031de53034 feat: rerank what a learner is shown, with Cohere through the proxy
Retrieval fused a bi-encoder and BM25 by reciprocal rank. A bi-encoder embeds a
document long before the question exists, so the two never meet: it is good at
"same topic" and mediocre at "answers this". A cross-encoder reads the pair.

The proxy already serves three — `cohere-rerank-v4.0-pro` is the default and
measurably better than the fast variant. Query text goes exactly where the
embeddings already go, and nothing new was signed up for.

It found a defect nobody was looking for. In AI Mode each finder scored
`1/(1+rank)` *within its own corpus*, so the best article, section, question and
card all scored 1.0 and the shortlist was a meaningless round-robin. A
cross-encoder is the first thing in this system that can compare a question
with a section. Candidates per kind widened so it can select rather than merely
reorder.

Measured against labels neither ranker produced. Questions, 60 disease tags:
precision@3 0.394 → 0.483. Sections, 60 article titles: 0.772 → 0.833.
"Management of bronchiolitis" led with influenza transmission and a pregnancy
question; "when do you image a first febrile seizure" returned the definition
rather than the sentence saying imaging is unnecessary.

And the honest negative, in docs/reranking.md: board vignettes are written
*not* to name their diagnosis, so on "what causes croup" it prefers a question
that says the word in passing over the barking-cough vignette that never says
it. Some of the bi-encoder's strength is traded away.

Not on the typeahead. A page of results is a choice being made and worth a
third of a second; a typeahead is a word being finished, runs on every
keystroke, and has nothing to judge yet.

The three-state thresholds stay on cosine, argued at the constant: a reranker
only ever sees a shortlist and structurally cannot answer the corpus-wide
question those numbers ask, and whether an answer claims to come from the
library is a promise that must not depend on a network hop.

Every failure returns None and leaves the order alone — unconfigured, no proxy,
connect error, bare 502, timeout, non-JSON, a duplicate or out-of-range index,
a non-numeric score, a list the wrong length. Verified against the running site
with a bogus model name: same results, fused order, no error to the reader.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
2026-09-12 18:23:06 +02:00

181 lines
10 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Reranking: the second opinion on order
Retrieval fuses two rankers that never see the query and the document together.
A bi-encoder embedded every row long before anybody typed anything, so cosine
distance answers "is this the same topic"; `ts_rank_cd` answers "do these words
coincide". Neither answers "does this document answer this question", and that
is the question a learner is actually asking.
A cross-encoder answers it, at the cost of one forward pass per pair, per query,
which is why it can only ever be given a shortlist the cheap rankers produced.
It reorders. It never decides what exists.
## What is served, and where the query goes
The proxy at `LITELLM_API_BASE` already serves three rerankers — `/model/info`
reports `"mode": "rerank"` for each:
| model | notes |
|---|---|
| `cohere-rerank-v4.0-pro` | **in use** — the best of the three here, see the numbers below |
| `cohere-rerank-v4.0-fast` | roughly 150 ms quicker, and measurably worse on this corpus |
| `jina-reranker-v2-base-multilingual` | open-weight, but rate-limited to the point of unusability on this proxy: a 60-query sweep got 429s on nearly every call |
Both `POST /rerank` and `POST /v1/rerank` answer, with the body Cohere defined —
`{model, query, documents, top_n}` in, `{results: [{index, relevance_score}]}`
out. Jina, BGE and LiteLLM's own rerank route all speak it, so replacing the
model behind `LITELLM_RERANK_MODEL` with a self-hosted open-weight one is a
configuration change and nothing else.
## Where it is applied
| place | reranked | why |
|---|---|---|
| AI Mode shortlist (`ai_mode_service.retrieve`) | yes | The shortlist *is* the evidence the answer may use, and the turn is already waiting on a chat model. It is also the only place where four corpora have to be compared with each other — see below. |
| Question search (`/questions?q=`, `/search`) | yes | A page of results is a choice being made. |
| Article search (`/articles?q=`, `/search`) | yes, through its sections | An article row is a title, a summary and a topical vector; the prose is in its sections. Handing a cross-encoder the article row is handing it a stub and asking about a document it cannot see. Its best section already decides its place in the fusion, so reranking the sections reranks the articles. |
| Test builder from a description | yes | A sentence about what somebody wants to study is exactly the query shape a cross-encoder is trained on, and the top of that list becomes the test rather than a page to scroll past. |
| Test builder from an uploaded file | **no** | The "query" is a whole handout. Clipped to 2000 characters it judges every candidate against whichever part that happened to be — confidently, and about the wrong thing. |
| Flashcard and media search | **no** | Short fields (front/back, title/caption) where the lexical ranker is already doing the whole job, and both are small personal collections. Latency for nothing. |
| Typeahead (`/search/suggest`) | **no** | It runs on every keystroke and has half a word to judge. It is prefix matching on titles, which is what finishing a word wants. Adding a 400 ms round trip per character would make the fastest thing on the site the slowest. |
## Latency
Measured from the backend container against `cohere-rerank-v4.0-pro`, median of
three, real question stems as documents:
| candidates | round trip |
|---|---|
| 10 | 250 ms |
| 30 | 380 ms |
| 50 | 430 ms |
| 100 | 580 ms |
Mostly fixed cost, then roughly 2 ms a pair. `MAX_CANDIDATES = 50`: the fused
pool is 60 wide and a page shows ten, so past fifty the model is being paid to
confirm that rows neither cheap ranker wanted are still not wanted.
In place, including reading the candidate text out of Postgres, over 120 live
queries never asked before: **median 480 ms, p90 550 ms**. AI Mode's whole
retrieval step goes from about 85 ms to about 500 ms — one call over roughly
forty candidates across all four corpora.
Two caches sit in front of that, and both matter more than the raw figure. Ours
is Redis, keyed on model, query and the document text, for a day: **35 ms** on
a hit, which is most repeat queries and every reload of a results page. The
proxy keeps its own cache of identical rerank requests and answers those in
about 30 ms, so a query one learner has already asked is cheap for the next even
after our key expires. A full four-corpus search page measured 94 ms warm,
236 ms with our cache cleared but the proxy's still warm, and about a second on
a genuinely new query, which is two corpora reranked one after the other.
Running the per-corpus finders concurrently would recover half of that and has
not been done.
## Does it help
Two measurements, both against labels neither ranker produced.
**Questions**, query = a disease tag's name, relevant = the questions carrying
that tag, 60 tags sampled at random from those with 340 questions:
| | precision@3 | MRR@10 |
|---|---|---|
| fused | 0.394 | 0.595 |
| + `cohere-rerank-v4.0-pro` | **0.483** | **0.641** |
| + `cohere-rerank-v4.0-fast` | 0.439 | 0.642 |
Top result changed on 38 of 60; precision@3 improved on 21 and fell on 9.
**Article sections**, query = an article's title, relevant = that article's
sections, 60 articles with three or more sections:
| | precision@3 | MRR@10 |
|---|---|---|
| fused | 0.772 | 0.847 |
| + `cohere-rerank-v4.0-pro` | **0.833** | **0.908** |
| + `cohere-rerank-v4.0-fast` | 0.811 | 0.917 |
Top result changed on 41 of 60; improved on 12, fell on 4.
### Queries where the old top result was wrong
* *"management of bronchiolitis in an infant"* — fused first three sections were
influenza transmission, foreign body ingestion, and the management of a
pregnancy. Reranked: three supportive-management sections, the third of them
bronchiolitis's own and the first two transient tachypnoea of the newborn.
Better, and not right: on this corpus the reranker will take a *Management*
section from a neighbouring topic over a differently-titled one from the right
topic.
* *"when do you image a first febrile seizure"* — fused returned the definition
of a febrile seizure. Reranked returned "routine laboratory studies and acute
neuroimaging are unnecessary…", which is the answer.
* *"teenager with knee pain worse after sport"* — fused led with a 6-year-old
girl's case and, in the articles, with bone pain that wakes a child at night.
Reranked led with a 15-year-old's sports knee pain, and with patellofemoral
pain syndrome.
* *"delayed passage of meconium"* — fused returned functional constipation.
Reranked returned "obstruction presents with failure to pass meconium within
2448 hours of birth".
* *"child limping with a fever and refusing to bear weight"* — fused led with a
urinary tract infection's presentation and acute gastroenteritis. Reranked led
with osteomyelitis, septic arthritis, and the differential for a limping
toddler.
* *"why does my patient with milk in the bottle have anemia"* (AI Mode) — fused
led with the general anemia article; reranked led with iron deficiency anemia.
### Where it does not help, and one place it is worse
The gain is uneven across corpora, and the reason is worth keeping in mind: a
bank question is a vignette written so as *not* to name its diagnosis. Asked
"what causes croup", the reranker prefers a question that says the word croup
in passing (a post-influenza bacterial tracheitis case) over the vignette with
the barking cough that never says it. That is defensible — someone searching
"croup" is not badly served by either — but it is the bi-encoder's strength being
partly traded away, and it is why the question-corpus gain (+0.09 precision@3)
is smaller than the prose-corpus one relative to how wrong the old order was.
The failure to watch for is a query the library does not cover. "How is croup
treated at home", against a corpus with no croup treatment section, is answered
with *Heat-Related Illness Treatment*, *Tetanus Treatment*, *Patellofemoral
Pain Syndrome Treatment* — the cross-encoder matching the shape of the
question when it cannot match the subject. Fusion's answer was no better, but it
was wrong in a way that looked wrong. Two things contain this: AI Mode's
closeness gate is measured on cosine over the whole corpus and still says
"nothing here covers this" regardless of what the reranker thought (see
retrieval-thresholds.md), and search results carry their own snippet. A minimum
relevance score would be the third, and it is deliberately not implemented: it
would need the same measured calibration the cosine thresholds have, on a scale
that changes with the model, and getting it wrong empties a page.
## When it is not there
Unset `LITELLM_RERANK_MODEL` (or the `settings:rerank_model` override in Redis,
which wins) and every list is in fusion order, with no error anywhere. The same
is true of a proxy that is down, slow, rate-limiting, or answering with
something that does not line up with the request: `rerank()` returns `None` and
every caller reads that as "keep what you had". The 429 storm from the Jina
model above is what that looks like in practice — a warning line per call and
results in their previous order.
`POST /admin/rerank/test` checks the configured model end to end, including
whether it puts an obviously relevant document above an obviously irrelevant
one. A reranker that answers 200 and ranks the decoy first is worse than one
that is switched off, and nothing else on the site would ever tell you.
## Re-measuring
```
docker compose exec -T backend python - <<'EOF'
from app.database import SessionLocal
from app.services import search_service as ss
db = SessionLocal()
for q in ["management of bronchiolitis in an infant", "delayed passage of meconium"]:
ranked, _ = ss.hybrid_ids(db, q, "article_section", limit=200)
print(q, "\n before", ranked[:3], "\n after ", ss.rerank_ids(db, q, "article_section", ranked)[:3])
EOF
```
The precision figures above come from labels that already exist in the database
— disease tags on questions, and an article owning its sections. Re-run that
comparison rather than judging by eye after a corpus change; ten queries read by
hand will agree with whichever ordering was looked at second.