Commit graph

3 commits

Author SHA1 Message Date
Daniel
e1b580386c feat: short / long / clinical, and stop broad topics retrieving index lines
Renames the middle view to Short and puts it first: it is the quickest way to
tell whether this is the article you wanted, and the full text is one click
away. Existing generated articles were migrated in place.

The prompt now asks for bullets that each carry a fact, because "X is important
to recognise" is a bullet that survives revision and teaches nothing.

The retrieval bug that made the last run mostly skips
The prose filter — drop chunks under 200 characters, since they are headings and
index lines — ran *after* taking the top fourteen hits. A broad query like
"Immunodeficiency" or a specialty name matches chapter titles first, so all
fourteen were index lines and the filter left nothing: the topic was skipped as
having no source material when the library holds plenty. Retrieval now asks for
five times what it needs and keeps the first passages that are actually prose.
Immunodeficiency went from 0 passages to 14, Pediatric Cardiology 0 to 14.

That is the same mistake the folder filter has a comment warning about — filter
inside the ranking, not after it — made two functions later.

Two things I got wrong and corrected rather than worked around: a `LIKE
'%key_points%'` check reported the migration had failed, when `_` is a
single-character wildcard and it was matching the title "Key points"; and a
variant count showing no Short sections was taken against the old image, where
Short was not yet a known variant and was being coerced to Long.

203 backend, 234 frontend green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
2026-09-10 18:21:21 +02:00
Daniel
af59fdb960 fix: MinIO name collision, and refuse to write an article from fragments
MinIO was resolving to the wrong container
Putting the backend on danvics_milvus to reach the clinical index gave it a
second service called `minio`, and Docker resolved that one first. Every object
read failed with InvalidAccessKeyId while the bucket simply looked empty — all
435 stem images unservable, and nothing in the logs saying why. The quiz MinIO
now answers to `quiz-minio`, which nothing else on this host claims.

A topic named after a shelf retrieved headings, not prose
"Pediatric Pulmonology" returned ten chunks whose top hit was 29 characters —
`**270** Pediatric Pulmonology`, an index line. Chapter titles rank well against
a query that looks like a chapter title. The model was handed a prompt with
citations and no content and said so, which was the correct response and read as
a JSON failure.

Two gates, both stated in the code. A chunk under 200 characters is a heading or
a running header rather than something to write from. A topic whose passages
total under 3,000 characters is skipped with the count in the reason, rather than
asking a model to write a medical article out of fragments — it will either
refuse or invent, and only one of those is visible.

The 71 generated drafts are deleted at the user's request. Nothing linked to
them and generation is resumable, so the cost was model calls rather than work.

Question bank corrections, from the agent that ran alongside:
262 questions had OCR-mangled units repaired — `inEq/L`, `mrnol/L`, flattened
`10⁹` superscripts and the rest — each with a version snapshot written first, so
every edit is reversible from the existing question editor. 94 stem images that
belonged to the explanation were removed; PREP's own `Item Q37A` / `Item C37B`
labels turned out to be a far better signal than word cues, taking the confident
split from 69/58/308 to 300/81/54. 13 uncertain images are listed for a person.

203 backend, 223 frontend green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
2026-09-10 17:31:53 +02:00
Daniel
025e5bb4ac feat: article CMS, three reading views, and articles written from the library
Standardises cross-references the way we agreed, and puts a CMS around articles
so hundreds of generated drafts are reviewable rather than merely present.

Links, made rename-proof
`[[7|Febrile seizures]]` resolves by id and displays the text — the id is the
part that must not change, the text is what keeps prose readable while you write
it. `[[old-slug]]` still resolves and is rewritten to the id form on save, not in
a migration: an article nobody has touched is not broken, and rewriting prose no
one asked to change is how an editor stops trusting the editor. Every slug an
article has ever had is kept, so a rename redirects instead of 404ing, and a save
reports markers pointing at nothing — at the moment the person who wrote the link
is still looking at it.

Three views of one topic
The full article to study from, the key points to revise from, the clinical view
to act from, with doses. They are views of one article rather than three
articles, so the numbers cannot drift apart and a question linked to the topic
still means one thing. Each section carries its variant; articles written before
this are the long view, unchanged.

CMS
draft → in review → published, with an author able to submit and only a
moderator able to publish. Every save snapshots what was there, restorable, and
restoring is itself snapshotted or the way back from a mistaken restore is gone.
The editorial queue is work rather than inventory: waiting for review, generated
and unread, published without sources, published with nothing to practise,
barely written. An empty bucket is drawn as good news, not as an alert.

Articles from the clinical library
The library index is 1.8M chunks of reference texts embedded with bge-m3 — the
same model PedsHub already uses, so our query vectors are directly comparable and
nothing had to be re-indexed. Retrieval supplies the facts and the provenance;
the model supplies the prose. References are built from the metadata of the
passages actually retrieved, never from the model, so a reference cannot be
invented — the same property that makes an AI Mode citation trustworthy. A topic
with fewer than three grounding passages is skipped rather than written from
memory. Everything lands as a draft.

Two things worth naming. The generated text is original writing grounded in those
books, not extracts from them: their facts are usable, their sentences are their
publishers'. And there are two Milvus servers on this host — the collection with
the data is the one reached as `milvus`, not the similarly named one on the other
stack, which I wired up first and which silently refused.

Also fixed along the way: `litellm==1.28.13` has been withdrawn from PyPI, so
requirements.txt could no longer be resolved from scratch and the image only
built because of a cached layer. Later additions go in their own layer until the
pins are refreshed.

182 backend, 223 frontend green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
2026-09-10 17:13:07 +02:00