pdf-quiz-generator/backend/scripts/reindex_article_search.py
Daniel 8362d706ac feat: search that looks at what an article says
Both halves of hybrid retrieval were reading the same 331 titles and summaries.
The lexical half was fixed earlier; this is the semantic one. `content` is NULL
for 323 articles because the generator writes into `sections`, so the vector for
98% of the library described the heading and nothing under it.

Depth is carried by the section index, where the longest section in the corpus
is under the embedding clamp — so every sentence of every body is embedded whole
somewhere, and nothing is truncated at that level at all. The article vector is
a topical signal instead: title, summary, the full outline, and an even slice of
every section's opening, budgeted so the clamp never silently fires. Round-robin
rather than head-and-tail, because truncating the head of a twelve-section
article stops in the pathophysiology and drops treatment and management — which
is where the words somebody actually searches for live.

`article_section_index` is populated and stays populated. The rebuild was a
private helper in one router, so the three other writers that save sections —
the generation task, the pipeline script and the seeds — silently skipped it.
That is how 323 articles came to have no rows at all. The generator itself is
one line poorer for it now.

A retrieval bug found on the way: the section-to-article rollup concatenated
rather than fused, so a section matching at rank 1 landed behind every weak
whole-article match and never reached the page. And `/articles/?q=` had no
rollup at all.

3,833 vectors in 332 seconds, batched 32 to a request — a normal article save
is now one round trip rather than fourteen. Proved against the vectors restored
from backup: "surgery for infant stridor that fails to improve" found
Laryngomalacia at rank 159, below the floor and invisible; it is rank 1 now, and
the section corpus answers it at rank 1 having previously been unable to answer
it at all.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
2026-09-12 15:57:42 +02:00

139 lines
6.3 KiB
Python

"""Make the semantic half of article search look at what an article says.
Two holes, one cause. Everything the generator writes goes into the `sections`
JSON column and `articles.content` stays NULL, so:
* the article's vector was built from title and summary alone for 323 of the
331 articles — enough to tell "Asthma" from "Migraine" and nothing finer;
* `article_section_index`, the projection that carries section-level retrieval,
held 24 rows covering the 8 hand-seeded samples. The generated articles had
no rows in it at all, so a term appearing once in one section — a drug, a
procedure, an eponym — was unreachable by meaning and reachable lexically
only after the `search_vector` fix in migration d6e7f8091a2b.
This re-embeds articles under the composition rule in
`embedding_service.article_embedding_text` and projects every section into the
index, embedding each one whole.
Resumable and idempotent by construction. `embed_records` stamps `embedded_at`,
so an article finished by a previous run is skipped by the same freshness test
that skips one nobody has edited; sections carry their own text, so an
unchanged section is compared and left alone rather than re-embedded. A run
that dies at article 200 costs nothing but the articles it had not reached.
docker compose exec backend python -m scripts.reindex_article_search
docker compose exec backend python -m scripts.reindex_article_search --apply
# after a change to how the article vector is composed, not just to the text:
docker compose exec backend python -m scripts.reindex_article_search --apply --all
Back up first; the writes touch every article row.
docker compose exec -T postgres sh -lc \\
'pg_dump -U "$POSTGRES_USER" "$POSTGRES_DB" -t articles -t article_section_index --data-only' \\
> backups/articles-$(date +%Y%m%d%H%M).sql
"""
import sys
import time
from datetime import datetime
from sqlalchemy import text as sa_text
from app.database import SessionLocal
from app.models.article import Article, ArticleSectionIndex
from app.services import article_service, embedding_service
# Commit boundary. Small enough that a crash loses seconds of work, large
# enough that the commit is not what the run spends its time on.
CHUNK = 20
def _needs_article_vector(article, active_model: str, started: datetime, force: bool) -> bool:
"""Whether this article's own vector is worth paying for again.
`embedded_at < updated_at` catches an edit. `--all` catches a change to the
composition rule, which no column records — bounded by the run's start time
so that resuming does not begin again from the top.
"""
if article.embedding is None or article.embedding_model != active_model:
return True
if article.embedded_at is None or article.embedded_at < (article.updated_at or datetime.min):
return True
return force and article.embedded_at < started
def _settle(db, articles) -> None:
"""Record that the stored vector matches the row as it now stands.
`Article.updated_at` carries `onupdate`, so writing the vector bumps it a
few milliseconds past the `embedded_at` the same write set. An exact
`embedded_at < updated_at` test therefore reports every article stale
forever, and a re-run pays for the whole corpus again. Raw SQL because a
Core update would fire `onupdate` once more and lose the race a second time.
"""
if not articles:
return
db.execute(sa_text("UPDATE articles SET embedded_at = GREATEST(embedded_at, updated_at) "
"WHERE id = ANY(:ids)"), {"ids": [a.id for a in articles]})
db.commit()
def main() -> int:
apply_changes = "--apply" in sys.argv
force = "--all" in sys.argv
started = datetime.utcnow()
db = SessionLocal()
try:
active = embedding_service._get_embedding_model()
articles = db.query(Article).order_by(Article.id).all()
indexed = {row[0] for row in db.query(ArticleSectionIndex.article_id).distinct()}
sections = sum(len(a.sections or []) for a in articles)
todo = [a for a in articles if _needs_article_vector(a, active, started, force)]
print(f" model : {active}")
print(f" articles : {len(articles)}")
print(f" sections in JSON : {sections}")
print(f" articles indexed : {len(indexed)} ({len(articles) - len(indexed)} with no section rows)")
print(f" article vectors : {len(todo)} to (re)generate")
if not apply_changes:
print("\n Re-run with --apply to write. Add --all to recompose every article vector.")
return 0
began = time.monotonic()
article_vectors, section_vectors, failed = 0, 0, []
for start in range(0, len(articles), CHUNK):
chunk = articles[start:start + CHUNK]
try:
wanted = [a for a in chunk if _needs_article_vector(a, active, started, force)]
article_vectors += embedding_service.embed_records(wanted, "article")
for article in chunk:
section_vectors += article_service.rebuild_section_index(db, article)
db.commit()
_settle(db, wanted)
except Exception as exc:
# One bad article must not cost the run; it is reported and the
# next chunk carries on, because a sweep that has to be restarted
# from zero is a sweep nobody runs.
db.rollback()
failed.extend(a.id for a in chunk)
print(f" chunk at {start} failed: {exc}", flush=True)
done = min(start + CHUNK, len(articles))
print(f"{done}/{len(articles)} "
f"{article_vectors} article + {section_vectors} section vectors", flush=True)
elapsed = time.monotonic() - began
rows = db.query(ArticleSectionIndex).count()
print(f"\n article vectors : {article_vectors}")
print(f" section vectors : {section_vectors}")
print(f" index rows now : {rows}")
print(f" elapsed : {elapsed:.0f}s")
if failed:
print(f" failed articles : {sorted(set(failed))}")
print(" Re-run to pick them up; finished work is skipped.")
finally:
db.close()
return 0
if __name__ == "__main__":
sys.exit(main())