Four things that share a spine, so they arrive together.
**Folders.** A hand-picked set of questions, and the fourth thing a grant can
name beside exam, discipline and category. Deliberately not `user_collections`
with a sharing flag: a library is a consequence of access — you save what you
can already see — while a folder is a source of it, and one table holding
thousands of private lists beside a handful that confer permission is one
mistake away from a leak. Built from the question manager, granted on /access.
Membership stays with the owner and moderators so a grantee cannot widen their
own reach, and deleting a folder takes its grants with it.
Two live constraints had to be rewritten to accept it: `ck_grant_has_a_dimension`
and `uq_grant_dimensions` both predate `folder_id`, so a folder-only grant
failed the check and two folder grants collided on the unique index.
**Per-question feedback.** The learner's half already existed. What was wrong
was who could read it: any grant at all let an educator list and delete reports
about the whole bank. Reports are now scoped by `question_scope_predicate`, the
same predicate that decides which questions that educator can see, and a reply
thread makes the report a conversation the learner can follow rather than a
form that swallows what they said.
**Per-section notes and article feedback.** Two tables on purpose:
`article_section_notes` is private to whoever wrote it, `article_feedback` goes
to whoever maintains the article. Both point at the section id inside
`articles.sections` rather than at `article_section_index`, whose rows are
dropped on unpublish — a cascade from there would delete a learner's writing
because an educator took an article down for an afternoon. A rename keeps a
note attached; a deleted section leaves it marked orphaned under the heading it
was written on, for its writer alone to remove.
The header's feedback badge covers both, because questions and reading are the
same job to whoever is doing it.
Migration i9f0a1b2c3d4. 556 backend and 572 frontend tests pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
Retrieval fused a bi-encoder and BM25 by reciprocal rank. A bi-encoder embeds a
document long before the question exists, so the two never meet: it is good at
"same topic" and mediocre at "answers this". A cross-encoder reads the pair.
The proxy already serves three — `cohere-rerank-v4.0-pro` is the default and
measurably better than the fast variant. Query text goes exactly where the
embeddings already go, and nothing new was signed up for.
It found a defect nobody was looking for. In AI Mode each finder scored
`1/(1+rank)` *within its own corpus*, so the best article, section, question and
card all scored 1.0 and the shortlist was a meaningless round-robin. A
cross-encoder is the first thing in this system that can compare a question
with a section. Candidates per kind widened so it can select rather than merely
reorder.
Measured against labels neither ranker produced. Questions, 60 disease tags:
precision@3 0.394 → 0.483. Sections, 60 article titles: 0.772 → 0.833.
"Management of bronchiolitis" led with influenza transmission and a pregnancy
question; "when do you image a first febrile seizure" returned the definition
rather than the sentence saying imaging is unnecessary.
And the honest negative, in docs/reranking.md: board vignettes are written
*not* to name their diagnosis, so on "what causes croup" it prefers a question
that says the word in passing over the barking-cough vignette that never says
it. Some of the bi-encoder's strength is traded away.
Not on the typeahead. A page of results is a choice being made and worth a
third of a second; a typeahead is a word being finished, runs on every
keystroke, and has nothing to judge yet.
The three-state thresholds stay on cosine, argued at the constant: a reranker
only ever sees a shortlist and structurally cannot answer the corpus-wide
question those numbers ask, and whether an answer claims to come from the
library is a promise that must not depend on a network hop.
Every failure returns None and leaves the order alone — unconfigured, no proxy,
connect error, bare 502, timeout, non-JSON, a duplicate or out-of-range index,
a non-numeric score, a list the wrong length. Verified against the running site
with a bogus model name: same results, fused order, no error to the reader.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
Single sign-on wrote a random string nobody would ever know. That reads as
"has a password" to everything that asks — so Settings demanded a current
password before it would let those accounts set their first, and the only way
through was to click "forgot password" for a password they never had. The same
trap was waiting for anybody who only ever signs in with a code.
Null says the true thing. Signing in refuses an account with no password the
way it refuses a wrong one, because which accounts have one is not a question
that endpoint answers. Setting a first password asks for no current one;
changing an existing password still does. `/auth/me` reports whether there is
one at all and nothing about it, because Settings has to choose between "Set a
password" and "Change password" and cannot tell from the outside.
The random strings already written are left alone. They are unguessable, so
nothing can sign in with them, and clearing them would mean deciding from
outside which accounts were meant to have one.
Identity is the email address throughout, so the three ways in are three ways
into the same account: single sign-on, a code, or a password — and a person may
acquire or drop the third at any point without losing the other two.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
A password is a thing to remember and a thing to lose. Somebody who can read
their own mail can now sign in without one: ask, receive six characters, type
them into the page that is already open.
A code rather than a link, and the difference is not cosmetic. The token in a
link was 256 bits, unguessable however long it lived, so its length, its expiry
and its rate limit were three independent decisions. Six characters is 2^30,
and the three stop being independent — so they are argued together:
* six characters of the invite alphabet, imported rather than copied, because
there should be one answer to which characters a person may be asked to
retype and that one already drops O/0 and I/1;
* a code answers five guesses and is then retired, not slowed — whoever is
typing has lost the mail or does not own it, and both are one click from a
new one;
* one code live per person, since several would mean one guess tested against
all of them;
* ten verify attempts per address per fifteen minutes, so nobody buys five
fresh guesses at a time by asking again.
Tens of guesses an hour against a billion, and the victim gets a mail for every
code burned. Eight characters would buy a thousandfold against an attack the
guess budget has already ended, and cost every person two more characters.
The attempt count lives in the row, not the cache. The Redis limiter fails open
when Redis is down, which is right for what it usually guards and wrong for the
only thing standing between a patient stranger and six characters.
Verifying is scoped to the address. A short code looked up on its own would be
tried against every code live on the site at once — the short code's one real
weakness, closed by knowing whose code it should be before comparing.
Fifteen minutes, because a first mail between strangers is routinely greylisted
five to ten and a code that expires before it arrives is not a sign-in method.
Shortening it buys nothing: one code is live and it answers five guesses
however long it sits there.
Nothing distinguishes an address with an account from one without — same
message, same status, same duration, and both rate limits counted before the
account is looked up, so a 429 cannot become the tell. Redis keys are
fingerprints, and the table holds a fingerprint rather than the code.
SSO stays first where it is configured, and a password is still one click away
for anybody who has one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
`/auth/forgot-password` and `/auth/resend-verification` both take care to say
"if that email exists" and both then answered the question anyway.
The reset limiter returned early for an unknown address, so it counted nothing
for one and counted for the other: ask four times and a registered address
gets 429 while an unknown one gets 200 for ever. It counts either way now — in
Redis for an address with no rows to count, keyed by a fingerprint, because a
list of addresses somebody tried is itself worth not keeping.
Resend answered "Email already verified." for a known verified address and "if
that email exists" for everything else, which is not a hint but an answer. One
sentence for every outcome now.
And Editorial has a way to write something. Drafting was only reachable from
the library — a page about reading, behind a button an educator arriving to
work has no reason to look for — so the two panels now open from a link.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
Three things landed together; the message names all of them, because a commit
that mentions one is a commit nobody finds the other two in.
**Figures.** Thirty-four JPEG 2000 files — 21 on questions, the rest unattached
in the media library — are WebP now, with `questions.image_path`,
`questions.explanation_image_path` and `media_assets.path` repointed together.
Serving already converted them on the way out, so nothing was broken; this
removes the step and makes what is stored the same thing that is served. The
originals stay: they are the only copy of what came out of the PDF, they cost a
few megabytes between them, and a conversion nobody can undo is not one to run
against a live bank. Paths are found by what the columns say rather than by
listing a bucket, because three tables record them and updating two would be
worse than none.
**The openai SDK is gone.** Ten call sites — one more than the map said, the
Celery article drafter — every one of them a POST with a JSON body, and not one
reading usage, cost, tool calls or logprobs. Every other call to the same proxy
was already plain httpx: embeddings, the ChromaDB embedding function, speech
both ways, model discovery, the vision probe. So this deletes an abstraction
rather than swapping one for another, and leaves one HTTP client instead of
two. `chat()` and `achat()` return the message content; a `ProxyError` carries
the status and the first 500 characters of the body, which is where the proxy
explains itself.
Behaviour is preserved deliberately, including a 600-second fallback timeout
for the four call sites that were running on the SDK's ten-minute default.
Lowering that is a real change and belongs in its own commit.
Proved against the live proxy on both services rather than only against mocks:
a completion, an async completion, a real 400 the vision probe still classifies
as a refusal, 407 models read from the catalogue, and a word read off an image.
**Voice.** A chosen voice is honoured whatever serves it. The prefix check only
accepted a locally served one, so a site adding a hosted voice would offer it
in Settings, save the learner's choice, and then quietly read every question in
the default voice. The list has always come from the database — adding a voice
is a row in Settings → AI models, never a code change.
And the sign-in page stops offering a locked door: `signup-policy` reports
whether registration is open at all, and the Sign up link goes when it is not.
The switch existed and the only way to discover it was to fill the form in.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
The voice picker was a dropdown in the quiz player, beside the question — the
one control on that screen with nothing to do with answering it, and one a
learner sets once and never touches. It is a setting now, on the user rather
than in a Redis blob, with a play button beside each voice because a voice is
worth hearing before it is chosen. Choosing nothing stays a real choice: it
means whatever an administrator marked default, so a site that changes its
default reaches everybody without a row being edited.
The tutor reads figures from `question_media` rather than the two legacy path
columns. Those agree exactly today, so nothing was being lost — the first
question given a second figure in the editor would have been the one that
broke it, silently and only for the tutor. The legacy columns remain as a
fallback for anything not projected into that table yet.
And the retrieval thresholds are written down in docs/retrieval-thresholds.md:
the three answers, the sixteen queries they were measured against, why they are
deliberately not the retrieval floor, and how to re-measure when the corpus
grows. Worth keeping the headline in mind — "discuss love" scores 0.491,
alongside "tell me a joke". A number in the 0.4s is noise, not a weak signal.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
Twenty-one stem figures are JPEG 2000. Chrome dropped it in 2015, Firefox and
Edge never had it, and the slim base image ships no MIME table — so
`guess_type` returned nothing, the fallback was `application/octet-stream`, and
`nosniff` finished the job. Those figures rendered nowhere but Safari.
The bytes were never the problem: Pillow decodes JP2 here perfectly well. Only
the delivery had to change, so it changes the way everything else already does
— through the thumbnail machinery, as a cached WebP derivative, stored beside
the original. A format no browser draws now asks for conversion whatever size
it was requested at, decided by the file's own magic rather than by the query
string. The 41 KB original comes back as an 83 KB full-size WebP or a 5 KB
thumbnail, and the stored file is untouched.
`.jp2`, `.jpx`, `.jpf` and `.webp` are registered at import, because a
container with no `/etc/mime.types` is a container that mislabels every one of
them. `.webp` had no figures behind it yet and would have failed the same way.
Two calls could hang for ten minutes. The SDK reads for that long by default
and this client retries nothing, so a stalled connection is a stalled request —
three of them in extraction, which does its own retrying. Both now pass an
explicit two-minute timeout.
Also removed: `EMBEDDING_PROVIDER`, which looks like a switch between a local
encoder and a remote one and is read nowhere, with a comment claiming
embeddings run locally when they have always gone over the network to the
proxy; and a `.replace("openai/", "")` that existed only to undo a prefix
nothing adds any more. The JPEG 2000 comment named the wrong mechanism — the
filename is no guide because there is no MIME table, not because it lies.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
Retrieval could not say "nothing". `hybrid_ids` fuses two rankers by reciprocal
rank and throws the distances away, and it returns the union — so the shortlist
was never empty, the "nothing matches" branch never fired, and a question about
photosynthesis came back with six paediatric sources and an instruction to
answer only from them.
So the fix is not more scenarios in the prompt. It is one calibrated number,
and three short prompts chosen by it in code. Asking a model to work out which
situation it is in is the part that does not work, and it is also the part that
makes prompts long.
Measured against this corpus with the bodies now embedded — eight clearly
on-topic questions and eight clearly off-topic:
off-topic 0.339 – 0.499 the French revolution … photosynthesis
on-topic 0.586 – 0.740 what causes croup … posterior urethral valves
The thresholds sit in the gap. They are deliberately not the retrieval floor:
that one decides what is worth putting in a list, where a weak hit costs a
reader a glance. These decide whether an answer claims to come from the
library, and a wrong claim costs them their trust in every other answer.
Above 0.55 the answer is sourced and cited, as before. Between 0.50 and 0.55 it
says nothing covers this directly, names what the closest material is, and
marks which parts came from where. Below, it says so in one line and then helps
anyway from general knowledge, citing nothing — refusing outright reads as a
broken assistant rather than a careful one, and the shortlist is not handed to
a model that has just been told the library does not cover the question.
An unmeasurable closeness is not a low one. No vector database or a downed
encoder returns None, and retrieval still found its rows by other means, so
those are still cited; dropping every citation because the ruler is missing
would be the worse failure.
Also: only published articles are indexed now. A draft is unfinished by
definition and has no business in a search result or in that shortlist. The
index follows publication both ways, and the fifteen-minute sweeper drops rows
whose article has been deleted or unpublished — an article that is never edited
again would otherwise keep its rows for good.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
Both halves of hybrid retrieval were reading the same 331 titles and summaries.
The lexical half was fixed earlier; this is the semantic one. `content` is NULL
for 323 articles because the generator writes into `sections`, so the vector for
98% of the library described the heading and nothing under it.
Depth is carried by the section index, where the longest section in the corpus
is under the embedding clamp — so every sentence of every body is embedded whole
somewhere, and nothing is truncated at that level at all. The article vector is
a topical signal instead: title, summary, the full outline, and an even slice of
every section's opening, budgeted so the clamp never silently fires. Round-robin
rather than head-and-tail, because truncating the head of a twelve-section
article stops in the pathophysiology and drops treatment and management — which
is where the words somebody actually searches for live.
`article_section_index` is populated and stays populated. The rebuild was a
private helper in one router, so the three other writers that save sections —
the generation task, the pipeline script and the seeds — silently skipped it.
That is how 323 articles came to have no rows at all. The generator itself is
one line poorer for it now.
A retrieval bug found on the way: the section-to-article rollup concatenated
rather than fused, so a section matching at rank 1 landed behind every weak
whole-article match and never reached the page. And `/articles/?q=` had no
rollup at all.
3,833 vectors in 332 seconds, batched 32 to a request — a normal article save
is now one round trip rather than fourteen. Proved against the vectors restored
from backup: "surgery for infant stridor that fails to improve" found
Laryngomalacia at rank 159, below the floor and invisible; it is rank 1 now, and
the section corpus answers it at rank 1 having previously been unable to answer
it at all.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
The colon pattern found 36. A verb-presence sweep found 75 more, and it was
wrong in both directions: it spared 21 genuine discipline overviews whose verbs
were simply not on the list, and it passed catalogues whose nouns are spelled
like verbs — "Mechanism, staging, and management of hypoxic-ischemic
encephalopathy, the leading cause of neonatal brain injury" satisfies a test
for "cause" and contains no verb at all.
A whitelist cannot tell those apart, so the first sentence of all 241 remaining
summaries was read rather than filtered, which found 41 more. 131 of 331 are
now claims instead of contents lists, in the shape of the one that worked:
what the condition is and who gets it, then what changes management.
The eight seeded demo articles all carried the same "Starter article for
demonstration" line as their summary. Each now has a real one written from its
own body — see the note below, because that line was doing a second job.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
**Prepared sessions.** Most of this existed: unanswered first, weakest topic
next, wrong-before-right after that, all scaled by what share of the real paper
each topic carries. What it could not do was change with time, say anything
about itself, or be reached without filling in a form.
Evidence now decays on a thirty-day half-life. Exponential rather than a fixed
window because memory has a slope, not a cliff — under a window, 29 days counts
fully and 31 counts for nothing — and because it is memoryless, so an answer's
weight does not shift when unrelated questions are answered, which is what lets
the preview stay a valid forecast. Spring is worth an eighth of last week. Two
things decay: a question's recall probability, drifting towards even rather
than past it, so an old right answer becomes eligible rather than wrong; and a
topic's accuracy, against a prior of two "no idea" answers, which fixes "right
once, known forever".
Strict unanswered-first meant that on a bank of 2,900 nothing was ever
recycled — spaced repetition existed and was unreachable. Review now takes up
to two fifths of a session. And the damping that spread the picks across topics
was applied only to seen material, so a learner with no history was handed the
heaviest domain entire instead of a spread; that was live.
The plan is the product. It is computed, shown, and then the session is built
from that plan's own ids and the plan returned with it, so the two cannot
differ; every figure in it is a tally over the chosen questions rather than a
forecast. No model touches the ranking — a learner asking "why these twenty"
has to get the same answer twice.
**Vision.** The proxy's own `/model/info` says which models can see, so nothing
is hard-coded: 77 report yes, 11 no, and 328 say nothing at all, which means
absent rather than incapable — so those are asked once with an 8px PNG and the
refusal cached. The deployment's main model turns out not to see, and questions
carry figures the learner is looking at, so the tutor was answering about an
image it had never been shown. It routes to a configured tool model now, folds
the description back in as text saying plainly where it came from, and caches
on the bytes because the same figure is re-sent every turn.
Also fixed on the way: `article` was missing from the admin's task list, so
article drafting always ran on the fallback model whatever an administrator
chose; and `.jpx` stem images were sent as JPEG because `mimetypes` guesses
that from the name, so the provider rejected them two hops later.
An administrator must pick a tool model in Settings → AI models. Until then the
tutor says a figure exists that nothing could read, rather than describing one
it cannot see.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
`search_vector` weighted title, summary and `content` — but `content` is NULL
for 323 of 331 articles, because everything the generator writes goes into the
`sections` JSON and only the eight hand-seeded samples ever used the column. For
98% of the library the body contributed nothing to full-text search, so a term
that appears only in a section — a drug name, a diagnostic criterion, an
eponym — returned nothing, and did so silently.
A generated column cannot contain a subquery, so the extraction is an IMMUTABLE
function it can call, and `content` stays in the expression for the eight that
use it. Proved rather than assumed: "supraglottoplasty" appears in no title or
summary in the corpus and now finds Laryngomalacia; before this it found
nothing.
Uploads are capped at 2 MB rather than 10. A document here is a query, never
content — read once to find matching questions in the bank and then
discarded — so the cap is about how much text is worth reading, and past two
megabytes somebody is uploading a textbook.
The previous commit's message covers only the litellm removal; it also carried
the 36 rewritten article summaries and the prompt rule behind them, which were
finished in the same window.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
**Linking.** A question could be tied to an article only from the article, by
typing the question's number into a box — so opening a question you had just
linked showed no sign of the link, and there was no control to add one. Both
ends now search: find the article by title from the question, find the question
by stem from the article, pick which section of the article the link lands on,
and see what is already linked. One shared finder, so the two ends of one
relationship cannot describe it differently. `GET /questions/{id}/articles`
mirrors the endpoint that already existed the other way, and `GET
/articles/linked` is retired — it answered this question by shipping the whole
prose of every linked article to the quiz player for a list of titles.
"Practise this topic" is a reader's control and no longer appears on an editing
screen.
**The player.** The rail was a bordered card floating in the page with a
scrollbar of its own, so a session had two scrollbars side by side and a
collapse handle tucked inside the card's padding. It is a column now: flush,
full height, its own background rather than its own border, the handle on the
boundary it moves, and a progress bar under the count. The bar at the foot is
the bottom edge of the window — three flush segments, no gaps, no pills —
because Exit as a small grey pill beside a large blue Next made leaving look
like the accident.
Study mode no longer asks whether you are sure. Leaving suspends: every answer
is saved, nothing is graded, and it is waiting where you left it — so the
dialog asked permission for something reversible, under a name for something
that does not happen. An exam still asks once, because a block has a clock, and
it now says what it is: "Leave this block?", not "End Session".
Options are lettered. The explanations already are — a stem extracted from a
board PDF says "Preferred Response: E" — so numbering them 1 to 5 left the
reader translating between two labellings of the same five lines. The tutor is
told the same letters, and the answer key is marked against its own option and
declared authoritative, so a model that would have answered differently cannot
tell a student the marked answer is wrong.
"Preferred response" and "Source page 518" are gone: the first labelled a block
that is obviously the answer, the second named a page of a book the learner
does not have. The clocks moved out of a grey strip across the explanation,
where they read as part of the answer, to the foot of the rail with everything
else about the session.
**AI Mode.** Sources are headed and counted at the end, where evidence belongs,
with the practise button after them rather than above. That button appears only
when there is something to build from and says what it will build — it used to
sit under "how can I help you today?" offering to make a session out of
nothing. A cited question opens in place: `/questions/:id` is the editor, so
following one dropped a learner into a form for changing the question they had
just been told about. And a session built from a chat is named like every other
session, rather than after the chat — asking "hi" produced "hi — practice".
Also: two test questions with raw `<p> </p>` in their stems were live in
the bank; retired. And 36 article summaries were written as a table of contents
with the colons filed off — "Peanut allergy prevention and management: LEAP
guidelines by risk tier, risk stratification, and anaphylaxis treatment" — every
noun phrase sounding informative and none of them saying anything. Rewritten as
claims, with the rule added to the prompt that produced them.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
A library holds both articles and questions now. It held only questions, so the
bookmark on an article had nowhere to write and stood in for the questions
filed under the topic instead — which is not what a reader who saved the
reading asked for, and left a topic with no questions unsaveable. Its own
table rather than a nullable column beside `question_id`: that shape allows a
row with both or neither, and every read then has to say which kind it is
looking at.
Which libraries already hold an article is now asked of the server, as one
question. It was kept on the device because the API could not answer, which was
wrong on the second machine and silently so. Putting one back is the same
control rather than an undo somewhere else.
"Short" is called Summary, because that is what the section is called, and it
is a toggle rather than one tab of three — the whole topic, or the part of it
worth revising, which is a different kind of choice from Long versus Clinical.
It names its own state, so a reader can tell why two thirds of the contents are
not there. The stored variant stays `short`: renaming it would be a data
migration to change a word on a button.
Also: `litellm==1.28.13` has been withdrawn from PyPI, so requirements.txt
could not be edited at all without the pip layer failing to rebuild — which is
what blocked pinning Pillow. Repinned to 1.53.1, the nearest still published;
the three things we use are unchanged in it, and both suites pass on the new
set. Pillow is pinned properly now rather than arriving through PyMuPDF.
One consequence, handled: `litellm.utils.get_valid_models()` now returns
nothing unless a provider's own API key is in the environment, and ours is a
proxy. That branch is only reached when no proxy is configured, and it now says
so instead of answering with an empty list that reads as "this site has no
models".
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
A question's stem image is two to four megabytes of scanned radiograph, and a
media grid is forty of those pulled at full size to draw forty postage stamps.
`?w=256` and `?w=640` now serve a WebP copy instead, made on the first ask and
kept beside the original under `thumbs/{width}/{key}` — same bucket, so nothing
new has to be configured for them to be backed up or thrown away.
Three rules, all about not making this a way to spend the server's afternoon.
Those two widths and no others: any other `?w=` is refused with a 400, because
an endpoint that resizes to whatever the query string asks for is a CPU sink
anybody can point at. Never enlarged: a 180px image asked for at 640 is served
as it is, since scaling up invents detail and charges bytes for it. And best
effort throughout — a PDF, an SVG, a truncated upload or a file that is not the
image its name claims all serve their original rather than failing, because a
preview must never take down the page that wanted it.
Authorisation is unchanged and still runs first: a thumbnail of a file you may
not read is a file you may not read. They stay `private, no-store` like
everything else here — they are behind authentication, so there is nothing for
a shared cache to do with them, and the win is the byte count.
EXIF rotation is read before anything measures the image. Every phone stores a
portrait photograph sideways with a flag; a thumbnail made without reading it
is a sideways thumbnail.
Pillow rather than sharp, which is Node. It is not pinned in requirements: the
pin invalidates the pip layer, and that layer no longer builds because
litellm==1.28.13 has been withdrawn from PyPI. Re-pinning litellm is a
deliberate upgrade of the AI layer, not something to slip into this. Noted in
the TODO.
Also: the article hover-card excerpt was printing `[[288|eczema]]` at readers.
The generic markdown-link rule does not know our own cross-reference syntax, so
it left the brackets and the id behind.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
`Question.is_shared` defaulted to 1 and was only ever set by a route nothing
called, so in practice it divided the bank into "everything" and "everything,
plus your own private ones" — a distinction that cost every recommendation
denominator a join and never changed an answer. Who may reach the bank is the
site's own access rules; who may manage a question is the category grant tree.
So the two predicates the whole bank was built on are now the same thing, and
say what they actually mean: a question is out of reach if it has been deleted
or belongs to a course. Nothing else. The column is dropped, the route that set
it is gone, the bulk "share" action with it, and the Private tile and pill go
from the question manager.
The tests that turned on it have been rewritten rather than deleted, because
the rule they were really about survives: revoking a question still revokes
every session carrying it — by deleting it, which is the only revocation left.
Several others named a category holding exactly two reachable questions and
then answered two particular ids; that category holds four now, so they name
the pair instead. A session's own sharing flag is untouched — that is a
different thing, and it is still how a session is handed to somebody.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
The exam player takes the window. The shell was sized against the header with
a number that did not include the navbar's own 32px of margin, so the block bar
— the one thing on the screen that must always be reachable — sat below the
fold and had to be scrolled to. Exam mode now hides the site chrome entirely
and is the viewport, which makes the arithmetic honest and matches what a board
looks like: item and block in a box at the left, the two arrows in the middle,
the tools at the right, the question-status rail down the side, and the clock,
Pause and End Block along the bottom.
Shortcuts is gone from the bar, and the labs open into the column beside the
question in both modes rather than a box over it.
Nothing is handed in behind the learner's back. The clock reaching zero stops
the block and says so; closing Time's Up submits, and the player stays put
showing the answers, which is the review. The server no longer settles an
expired attempt at all — listing sessions used to mark any paper whose clock
had run out, so opening a page could score a block the learner had walked away
from, and the first they knew of it was a result.
Reviewing an attempt is now the player with the answers in, not a dropdown and
a card. Same rail, same layout, same labs, same way out — and on a phone the
same burger opens the same question list, from one shared rule about which
routes are a session.
Also: the rule-out toggle sits beside its option instead of pinned to the far
edge of the card, so an option box is as wide as its own words; the voice
picker leaves the player, since a reader's voice is a setting and not a
decision to retake every session; figures carry no invented "Figure 1" — a
label is what prose refers to, and the backfill knew of no prose, so 346 of
them said only that an image was an image; and the landing page shows the two
modes happening rather than promising six things in a sentence.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
landing page describes this product
An objective is now required to build a session, not only asked for in
the interface — the interface asks, and this is the same rule where it
cannot be walked past. Only where there is something to choose: a
deployment with no exams, and the first administrator of a fresh one,
must still be able to build a session. A rule that locks an empty site is
not a rule, it is a fault.
Saving a question into a folder is one box that searches what you have
and offers to make what you do not. It used to say "make one in the
question bank" and leave you to go and do it, which means leaving the
question you were reading and coming back to find your place. A name
that already exists exactly is not offered twice; a partial match offers
both, because wanting a narrower folder called "cardio" is not the same
as wanting the one called "Cardiology misses".
The landing page is rebuilt. Its copy described a product from months ago
— "upload a PDF, AI extracts questions", which is one feature of many now
— and it was 568 lines of inline style objects, which cannot express a
hover, a media query or a keyframe. The figures come from
/api/public/stats and count up; a failed fetch renders the section
without them rather than showing noughts, which would be a lie about an
empty bank. Motion is CSS and SVG, and prefers-reduced-motion turns all
of it off — including forcing the scroll-revealed elements visible,
since a hidden element with its animation removed is how respecting that
setting turns into a blank page.
Two smaller ones from the screenshots: the collections shelf is boxed
rather than scrolling past everything else on the page, and its rows no
longer carry the entire stem — lab tables and all — in a native tooltip
that covered half the screen and could not be dismissed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
Choosing what you are studying for has no way past it now but to answer.
It decides which questions exist, how relevance is weighted and what
readiness measures against, so an account that never answered it was
being shown the whole bank by accident rather than by choice.
What is guarded instead is asking a question that cannot be answered: if
the list of objectives fails to load, or there are none, nothing is shown
at all. A modal with no options in it is not a question, it is a locked
door.
GET /api/public/stats, unauthenticated, so the landing page can state what
there is rather than what someone typed into the markup months ago — a
number written into a page goes stale the week after and nothing breaks
to say so. Counts only, and only of published material: how much there
is, never what it is, so there is nothing here to walk.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
Cap moved from /cap/ under this app to cap.pedshub.com, so anything else
on this machine can use the same instance. Caddy terminates it, the
backend keeps verifying over the compose network rather than going out
and back, and the widget endpoint is configuration rather than a path
baked into the component. Verified: a challenge is issued on the
subdomain, and a token that was never issued is still refused.
"Correct using hints" is now a per-topic figure. The knowledge profile's
accuracy bar was two-tone because /study-tools/recommendations carried
only `answered` and `correct`; the hint count existed lifetime-wide but
never per topic, and inferring one from the other would have been a
different set of answers drawn as though it were this one. The column
was already on attempt_answers, so it is a group-by, and the bar is
three-tone as the reference has it.
And the objective is asked for. It decides which questions exist, how
relevance is weighted, and what readiness measures against — and it was
possible to sit a whole board paper without ever being asked, because no
objective quietly means the entire bank. That is a reasonable default and
a poor thing to arrive at by accident. Five of six accounts here had
never set one.
It can be declined: "everything" is a real answer, and trapping somebody
behind a modal because a list failed to load would be worse than the gap
it closes. Declining is still a choice made, which is the point.
Also in this commit, from the exam-player work: Show answer in study mode
that reveals without recording an answer, review keyed on the attempt
being closed rather than every question being answered — a block that
timed out with nothing answered is over too — and the exam top and bottom
bars. That work found something worth knowing: the exam player is *served*
questions with no correct answer and no explanation, so review cannot
un-hide what it never had, and the player refetches the marked version
once the attempt closes. Nothing is revealed while a block is running.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
Proof-of-work rather than a puzzle, and — the reason for it — nothing
about the person signing up is described to a third party in order to let
them in. Turnstile and then hCaptcha were both here; both told Cloudflare
who was at the door.
The `cap` service runs on the compose network with its own Redis
database, kept apart from the app's so a flush of one cannot clear the
other's challenges. The widget talks to /cap/ on this origin, proxied by
the frontend's nginx, so the browser reaches nobody else either. Caddy
passes the whole host through to that container, so it needed no change.
Two things that had to be found rather than read:
Cap's key API is undocumented. The routes are `/auth/login` and
`/server/keys`, and the Bearer value is base64 JSON of `{token, hash}` —
not the session token itself, which is why the obvious call returns
"Malformed session token". The site key and secret were created that way
rather than by hand in a dashboard.
And an nginx proxy_pass whose target is a variable passes the URI through
untouched: the trailing slash that strips a location prefix on a literal
target does nothing. Cap was being asked for /cap/<key>/challenge and
answering NOT_FOUND until the prefix was stripped by an explicit rewrite.
Verified end to end against the running service: a challenge is issued
through the public path, and a token that was never issued is refused
rather than waved through.
Also here: the register modal's Name and Email were bare labels that
neither wrapped their input nor named it, so a screen reader met two
boxes with no names and clicking the word did nothing.
And the knowledge profile paginates ten to a page and expands each row to
its two bars beside the next step. "Correct using hints" is missing from
that bar because /study-tools/recommendations does not carry it per
topic — inferring it from the lifetime figure would be a different set of
answers, so the bar is honestly two-tone until the backend offers it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
analysis is a real tab
Adaptive selection knew what you were weak at and nothing about what the
exam is made of, so being weak at something worth 5% of the paper ranked
the same as being weak at something worth 1%. Every score is now
multiplied by the weight the board publishes for that topic's domain —
the same `exam_blueprints.weight` behind the Relevance column.
A topic the blueprint does not cover takes the median published weight. A
zero would make unmapped material unreachable and the highest would make
it the priority; neither is a claim the blueprint supports. With no study
objective the multiplier is absent and selection is about weakness alone,
exactly as before.
Weight scales weakness, it does not replace it: a topic you are certain of
does not surface because it is worth 5% of the paper, because (1 −
accuracy) is near zero and no multiplier rescues that. docs/adaptive-
sessions.md says all of this, including what is still open.
Session analysis is the third tab rather than a link out of the page —
two of the three used to change what you were reading and the third took
you somewhere else. The tab bar is one component both routes wear,
AnalysisSessionPage's body is a component the tab renders in place, and
the tab lives in the address so a link opens where it says.
Two things that were wrong turned up in that work: a session nobody had
sat showed 0% in the figures and "0% correct" in the donut — two separate
statements of a score on a session that had none — and the old third tab
disappeared entirely for anyone with no attempts, so the strip silently
changed shape.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
The incomplete-block warning is the one from the screenshot: a red
heading that says the block is incomplete, the count of unanswered items,
the sentence about resuming not matching exam day, and End Block against
Remain in Block. My version asked the question in my own words and led
with the wrong button.
Pausing says "Exam Paused" and offers Return to exam. Nothing else — the
warning about real exams is somebody else's disclaimer, not ours.
Exit session asks "Are you sure you want to end this session?" before it
goes, rather than going.
Time's Up says what it is and the button says Close, which is the only
thing left to do: it is already handed in and marked, and Close lands on
the session's analysis.
One name for one action: the bottom button read Skip on an unanswered
question and Next on an answered one, while the arrow an inch above it
said Next for both.
And the rail shows stems again once the block is handed in. Numbers while
it is being sat — reading ahead is not something the exam being rehearsed
allows — but there is nothing left to protect afterwards, so the review
reads like study mode.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
Retiring three vocabularies at once was my call and the wrong one. Keyword
had to go — it was the old route to an organ system, which a topic now
carries, and that took Systems from half the bank to all of it. Subject
and disease went with it on the argument that the topic tree says the same
thing. It mostly does, and "mostly" is not a reason to remove the
vocabulary people had learned to filter by.
203 subjects and 2,275 diseases are back, with their 14,029 links, and the
Disciplines and Diseases pickers with them. Keywords stay retired.
The backup I wrote before deleting was not where I said it was:
`./backups` is mounted on db-backup, not on backend, so the file went with
the next container rebuild. The rows came from the nightly dump instead,
which is what that dump is for. scripts/restore_subject_disease_tags reads
a pg_dump extract, is idempotent, and resets the sequence afterwards so
the next tag created by hand does not collide with a restored one.
/tags serves subjects and diseases from their own links again, and systems
through the topics that carry them.
The Performance tab is ordered as the reference has it: the trend beside
the split it is a trend in, and Completion's four figures underneath.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
Deleting a user failed with a not-null violation from quiz_attempts.
Every foreign key to users is already CASCADE or SET NULL in Postgres,
but the ORM relationships had no passive_deletes, so SQLAlchemy insisted
on emptying each one itself by writing NULL into columns that refuse it.
passive_deletes leaves it to the database, which knows. The two tables
that genuinely cannot forget a user — question_categories and
quiz_categories are NOT NULL and NO ACTION — hand their rows to the
administrator doing the deleting: the taxonomy is the site's, not the
author's.
/tags counted through question_tag_links, which is now empty, so every
organ system read zero and an active exam hid them entirely. It counts
through the topics that carry them instead: 2,919 of 2,924 questions, all
sixteen systems with real numbers. The session builder's Disciplines and
Symptoms pickers were over the retired vocabularies and are gone —
Topics is the same axis said once and said better, 673 against 203.
The end-block dialog offered one button. A confirmation with one button
is not a confirmation: it now leads with the way back into the block,
says how many are unanswered as a sentence rather than a grid to count
by eye, and the unanswered are numbers you can press to go there.
Registration asks for the password twice, on both forms — a password you
cannot see is one you can mistype into an account you then cannot open.
The public pages had no footer, so signing in meant losing the way to
About, Contact and the clinical disclaimer. They sit in a plain layout
that keeps it.
Draft batches can be filed from the workbench: the topic they file into
is a picker at the top, and nothing crosses over until it is set.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
claims to hold them
The scaffolding is down. 203 subject, 2,275 disease and 4,281 keyword
tags, and 25,356 links, deleted — backed up first to a 1.9MB JSON of
replayable rows, because "we can always put it back" should be true
rather than said. The 16 system rows stay: categories point at them.
With them go the things that only existed to feed them — the
classify_questions task, its snapshot helpers, POST /tags/classify and
its status poll — and the three Taxonomy tabs that would now always read
zero. A tab showing 0 forever teaches people the page is broken.
The organ-system filter in the session builder moved onto categories with
the rest, including everything beneath a matched topic, so it groups the
way the analysis does.
Registration: `settings:registration_enabled` was set to false, and there
was no switch anywhere on the site to set it back. The API had always
accepted it; the Site policy page had never shown it. So the site could
be closed to new members with the admin looking at three switches, all
correct, and no way to see the one that was actually refusing them. It is
now the first switch on that page, and says plainly that the ones below
it have nothing to act on while it is off. The SSO-only flag was hidden
the same way and is shown when SSO is configured.
Deleting a topic no longer silently unfiles its questions. It asks where
they go, and says how many are waiting, unless the topic is empty — the
same rule promotion now follows. Its extra category links move too,
minus any that would duplicate a pair the destination already has.
Back links: Trash, Extraction jobs, Taxonomy and the Handbook had none at
all, and Access pointed at the wrong section. They are one component now,
each returning one step to the section it was opened from. Editorial has
its own entry in the section bar, so its Tools card is gone rather than
being a second door to the same room.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
that stops answering the wrong question
Invite-only was set and the sign-up form had nowhere to type a code.
There are two registration forms — /register and the modal on the landing
page — and only the first had been taught about invite codes. The modal
is the one most people meet, so turning the gate on failed everybody with
"an invite code is required" and no field to satisfy it. It now asks the
same signup-policy question and shows the same field.
Dictation records to our own transcriber first and falls back to the
browser's recogniser only where recording is unavailable. It was the
other way round for speed, but the browser's speech stack announces
itself to the user in ways we do not control — Firefox interrupts the
page with a warning about a missing Speech Dispatcher library, which is
alarming and is not about us.
A draft question could be promoted into the bank with no category. That
question would reach nothing: no discipline, no organ system, no
relevance, no row on any tab of the analysis — in the bank and invisible
to every page that counts. Promotion now refuses, before an id is spent.
And "Your overall analysis" is out of the session rail. It put lifetime
figures one click away while you were standing in front of a single
session, which is the thing that was supposed to have moved to the
Performance tab.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
A question reached a system through a symptom keyword it happened to
mention — question → keyword → parent system — and only 726 of 4,281
keywords had ever been given a parent. The Systems tab saw 1,492 of 2,924
questions while Disciplines saw all of them.
The system now sits on the category: question_categories.system_id. Every
question has a category, so every question reaches a system. 2,919 of
2,924, and all sixteen buckets have real content.
It stays a third way of asking rather than the discipline tree relabelled
because a topic's system is assigned separately from where it sits in the
tree. scripts/assign_category_systems takes the discipline as a default
and lets the topic's own name overrule it, which is exactly the case that
makes the axis worth having: conjunctivitis is filed under Infectious
Disease and is an eye, osteomyelitis is filed there and is a bone. 110 of
660 topics were decided that way.
Two regex traps caught in the dry run and fixed before applying:
"adRENAL" matched the kidney rule, and "Abnormal Uterine Bleeding" matched
the bleeding rule. Both now have a specific rule above the general one.
I first tried to fix this by parenting the orphan keywords to systems,
deriving each keyword's system from the questions carrying it. The dry run
showed why that was the wrong shape: it reached only 534 of 3,555 orphans,
and inherited every coarse edge of the discipline map — conjunctivitis came
out as Multisystem because conjunctivitis questions are filed under
Infectious Disease. That script is left in place, unapplied, as the record
of a measurement worth keeping.
No ForeignKey on system_id in the model: question_tags is a raw-SQL table
with no ORM class, and declaring one leaves every metadata build unable to
resolve it. The constraint is real in Postgres.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
Four things, all from one screenshot pair.
The Review button in a study session was inherited from the exam player.
Reviewing a block before handing it in is an exam idea; a study session
has nothing to hand in — it keeps going until every question is answered
and at that point it *is* the review. The review link, the top-bar
button and the rail button are exam-only now, and a study session whose
questions are all answered says "Finish session" and submits rather than
opening a dialog to ask a second time.
The drawer's "Qbank" pointed at /questions, which has never been a route
— /questions/:id is the editor. It went nowhere. It points at
/question-bank, and Collections and AI Mode join the list.
While a session is open on a narrow screen, the navbar burger now opens
that session's questions instead of the site menu, which is a tab inside
the same drawer. Two menu buttons an inch apart, one of which leaves the
session you are sitting, is the wrong offer. The player claims the button
only while it has no rail, and hands it back when it leaves.
And the drawer says what AMBOSS's does: a Review badge once everything is
answered, the mode in the title, a progress bar under the count, and the
session and question clocks pinned beneath the list.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
AI Mode could cite an article and link to it; it could not do the other
half of the job. POST /ai/conversations/{id}/practice turns an answer
into a study session, built from what that answer actually cited: a
question it named first, then questions filed under the category of an
article it named, then retrieval on the learner's own words. Everything
goes through the bank's visibility rules on the way out — a chat is not a
route to questions a learner could not otherwise reach. Study mode, never
exam: this is reading followed by practice, not a paper.
Two false alarms on the Settings page, both visible in a screenshot:
The STT test called /model/info on the LiteLLM proxy. Our virtual key is
scoped to llm_api_routes and cannot, so a working transcription model
reported a red 403. It now falls back to /v1/models, which the key may
call, and says plainly that the proxy would not confirm what the model is
for — presence, not suitability.
And the TTS test raised a 400 carrying an instruction ("use the Preview
button"), which the page rendered in red with a ✗. That is not a failure.
It answers, and Preview stays the way to hear a voice.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
Favorites and the question libraries in one place. Card and Table views
with the choice remembered, sort by last used / created / name / size
with a direction control, a count line, and a search over name and date.
Favorites leads as a fixed row: it is the one shelf nobody made and
everybody has, so it cannot be renamed or deleted.
Sorted by when each was last used, not when it was made — the order
things were created in is nobody's mental model of their own shelf. A
library nobody has opened falls back to its age, because it is newer to
the learner than it is to the database. That needed
`user_collections.last_used_at`: null on every existing row, since
backfilling from created_at would invent a use that never happened.
A shelf opens in place rather than linking away. The obvious link would
have been /questions?collection=N, and there is no page there that reads
it — the old bank browser was dismantled — so the card would have led
nowhere. Questions can be taken back out from the open shelf, and any
shelf can be sat as a session through the existing explicit_ids builder.
The ⋯ menu moved out of QuizPage into components/MoreMenu; the player
keeps its own look and its own children through className props. It no
longer closes on any click inside, which the player's feedback form and
share dialog were relying on by accident.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
The session analysis ranked its weakest topics by primary category only,
while the Analysis page asked the same question three ways and rolled
answers up the category tree. Two sets of rules for "where does this
question belong" is two pages that can disagree about a learner and
neither able to explain why.
So the rules moved to services/knowledge_groups.py: ancestor roll-up,
article reached through its category, organ system reached through the
symptom keyword. study_tools now asks that service instead of building
the lookups inline, and GET /attempts/{id}/recommendations gives one
session the same Articles / Disciplines / Systems switch. Grouping is its
own call, so changing it does not re-read the question table and the peer
statistics beside it. A running exam ranks nothing — marking it there
would answer the question the exam is asking.
The ungrouped `recommendations` key is gone from the analysis payload
along with the code that built it.
And the document page had no way back. It is reached from the Tools
workbench, which by design has no menu of its own, so leaving it meant
the browser button. It opens onto Tools now, as Tools opens onto
Settings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
Your score is the share of questions right at your most recent answer to
each. It is deliberately not called an equated score: AMBOSS's EPC rests
on psychometrics we do not have, and a number dressed up as one would be
a claim we cannot support. The card says so.
Against everyone else compares you with other learners on the questions
you have in common — not with their scores on whatever they happened to
sit. A percentile over different question sets reads someone who worked
through the hardest fifty in the bank as weaker than someone who did
fifty easy ones, which is the opposite of true.
Neither appears before it means anything, and each says which half is
missing: more questions of your own, more questions shared with others,
or more learners. The cohort reported is the most any one shared question
saw — distinct learners cannot be summed across questions without
counting the same person once per question.
The "readiness is still locked" note sat above the tab switch and so
appeared on Performance, where it described a table that is on the other
tab. Moved down to the table it is about.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
`GET /study-tools/performance-over-time` returns a point per completed
session with two figures: that session's percentage, and the running
score across everything answered up to that day. The chart draws the
running line and marks the sessions along it — a single session of twelve
questions swings too far to say anything about whether a learner is
improving.
It stays shut below 40 answers or 3 sessions and says which of the two it
is waiting for, rather than drawing a line through two points and letting
the shape suggest a trend that is not there.
LineChart was in the tree unused, with a hardcoded slate palette that
vanishes on a dark page. Rewritten against the theme tokens.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
Opening a tip before answering is a nudge. The answer that follows is
still right — it is counted as right, and the percentage is not docked —
but it is not the same as right, so it keeps its own arc on the donut and
its own line in the legend: "3 correct after a tip".
attempt_answers.used_hint records it. The player reports which questions
had a tip opened before the answer went in; a tip read afterwards is
revision and does not count, which is the difference two of the tests
turn on. Both endings agree about it — an explicit submit carries the
list, and an exam that runs out takes it from the saved progress, so a
tab closing cannot launder a score.
Found while wiring this: RichText declared its component overrides inline
in the render, so every one was a fresh component type and React
remounted the whole rendered tree on each render. An open tip closed
itself every time the exam clock ticked. The map is memoised now.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
A question got wrong in March and right in September is 50% by one count
and 100% by another, and both are true. The Performance tab now says
which it is answering: All attempts is every answer ever given — how much
work has been done — and Latest attempt keeps only the most recent answer
to each question — what is known now.
GET /study-tools/answer-split returns both splits plus the session and
unique-question counts, under the same exclusions as everything else that
measures: no repetitions, no course quizzes, no expired attempts. A blank
is its own slice, never folded into incorrect.
The ring itself moves out of AnalysisSessionPage into components/Donut so
the session view and the lifetime view cannot drift apart. Its legend
gains .is-answered, which the session page had been asking for without
anything defining it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
"How am I doing" and "how was I doing last month" are different questions,
and a single lifetime figure cannot answer both. Analysis now carries a
Completion panel on the Performance tab: questions answered against the
bank, how many were right, time per question, total time — over 7 days,
30 days, 3 months, or everything.
GET /study-tools/completion?days=N does the counting. It leaves out what
would not be a measurement: repetitions (you already know that answer),
course quizzes (they belong to their course), and expired attempts. A
question left blank is not a wrong answer, so the percentage is out of
what was answered, not out of what was set. Nothing answered reports
nothing rather than 0%.
The Tools workbench has no menu of its own by design, which left no way
back; it now opens onto Settings where it was reached from.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
The phone had a dot grid dropped under the top bar — a different thing
in a different place doing the rail's job worse. It is a drawer holding
the same rail the desktop has, with the site's own menu on the other
tab, because the alternative is a second hamburger elsewhere for the
same purpose. The dot grid and its styles are gone.
And the extraction pipeline was run end to end against a three-question
PDF rather than reasoned about. It works: three questions, stems,
options, correct answers and explanations, landing in a draft batch and
not in the bank. But the run found a real bug on the way.
A document's text is read from the search index, not from the file. When
that index is missing — never processed, or lost to a restart — every
page is skipped and the job fails with "the AI could not find questions
with correct answers in this page range". That is the wrong diagnosis,
and it sends people to change the model, the prompt and the page range,
none of which is the problem. The two failures are now counted apart and
named apart: no stored text says so and says to re-process; a model that
found nothing says that instead.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
Sitting the same questions again is practice, not a new measurement. You
have already seen the answers, so getting them right the second time
says nothing about whether you knew them — and it cannot be allowed to
raise a figure that means "how much of this do you know". A repeated
session is titled "(repetition)", analysed in full on its own page, and
left out of every aggregate: the overall accuracy, the per-quiz history,
the averages, and the readiness that drives recommendations.
Deleting a single session is gone — control, endpoint, tests and all. A
session is a record of work done, and removing one edits the history
every figure on the analysis is computed from, which turns a measurement
into a number somebody chose. Starting again is still offered whole,
under Settings, Your data, which takes everything rather than the parts
that flatter.
Two layout bugs behind that. The category tree kept its appearance in
QuestionBankPage.css, so it looked right on the bank and took whatever
the host page did to a label everywhere else — in the question editor
that centred the name, leaving it adrift with the count at the far
right; it owns its own stylesheet now. And the editor's grid collapsed
to `1fr` below 900px, whose automatic minimum lets one unshrinkable
child push the column past the window: the page had padding down its
left and none down its right because the right was off the screen.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
The knowledge profile ranked topics by how much of *our* bank sat under
each one, which is a fact about us rather than about the exam. It made
cardiology and rheumatology equally worth an evening whenever we happened
to hold the same number of each. The ABP publishes that one is 5% of the
paper and the other 2%, and exam_blueprints.weight has held that since
the blueprint landed.
A domain's weight is divided among the topics beneath it in proportion
to the material each holds, so the topics under a domain add up to its
published share. 672 of our categories now carry one. A topic the
outline does not cover keeps the bank-share figure rather than reporting
nothing — and the row says which it is, because the two numbers mean
different things and should not be read as the same one.
Session analysis is a link to the last session rather than a third tab
with nothing behind it — a session's analysis is a session, and the rail
beside this page is the list of them.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
Study mode held a choice as a draft and waited for "Submit response" — a
second press to confirm something already decided, on every question.
Clicking an option marks it now, green or red, with the explanation.
Free text is the exception and keeps Enter, because typing is not
choosing.
Figures carried a generated caption: "Figure from question #3360 (from
images/doc_23/page_704_img_0.jpeg)". That describes the database, not
the picture, and showed a learner an internal file path. 341 of them are
cleared, the indexer no longer writes them, and an unlabelled figure now
says nothing rather than "Figure 1". A screen reader still gets the
label and caption when there are any, and the position when there are
not.
Suspend, Restart and Edit are gone from above the question. Three
buttons over a question nobody was looking away from to press them; Exit
is in the bar at the bottom with the session's own controls, and
restarting and editing belong to the session list and the editor.
And iOS Safari's zoom-on-focus is fixed once rather than per field.
Safari zooms the whole page in when a control smaller than 16px takes
focus and never zooms back out, leaving the layout scaled and broken. It
was being remembered at each individual field, which meant it was
forgotten at most of them — a dozen were still under 16px. One rule for
every control on a coarse pointer.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
Unanswered questions were counted as wrong in every percentage the site
reports. That made leaving an exam early look like failing it, and made
the figure say more about how far you got than about how well you did —
and how far you got is already the number sitting beside it.
An unanswered question is not a wrong answer. It is not an answer.
score_percent() and answered_counts() give the rule one definition, used
by all seven places that reported a percentage: submission, attempt
history, per-quiz history, the overall average, per-quiz stats, one
attempt's detail, and the session analysis. The list endpoints count in
one query rather than one per row.
The review dialog said unanswered questions count as incorrect, which
was true and is not any more. It now says they will not be marked wrong,
and will not be marked.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
The card offered Review answers and Resume session at once on a session
still in progress, which is the muddle: there is nothing to review yet
and nothing to resume once it is done. It is one or the other now, and
what decides it is whether anything is left to answer — not whether it
was an exam or a study session, which have the same two states as each
other. A study session keeps going until every question is answered and
becomes the review at that point, without waiting to be handed in.
Repeat is offered either way. The questions worth sitting again are
worth sitting again now.
"Skipped" meant gone past, and was shown for questions in a session
still running that had not been reached. Those read "not yet answered".
And a timed block is now ninety seconds a question, set from the count
rather than asked for. Choosing a limit is a decision nobody has the
information to make — the pace belongs to the exam being rehearsed, not
to a preference — and a block sat at the wrong pace teaches the wrong
pace. Forty questions is an hour. An explicit limit is still honoured.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
It ran on a wall clock. An hour away from the tab spent an hour of the
exam on questions that were never shown, and every per-question figure
was a fiction — which is the number the whole analysis is built on.
Three things stop it now. The tab being hidden, which catches switching
away. An explicit pause. And, for the commonest case the other two miss
— the tab left open on the exam while the person is in another room —
an idle watch: three minutes with no mousemove, key, wheel, touch or
scroll and it asks "Still there?", with the clock already stopped by the
time the question appears. A stray pointer movement does not answer it;
somebody has to say they are there.
Three minutes, not one, and scrolling counts as activity: reading a long
vignette is minutes without a click, and interrupting genuine reading to
ask whether you are reading is worse than occasionally crediting a
minute nobody was there for.
The server was the other half. seconds_remaining computed from
started_at and total_time, so a paused client made no difference to what
the server thought was left. It reads the saved time_left now, which is
what the player decrements only while the exam is on screen, falling
back to the wall clock for progress saved before this existed.
And a five-minute warning, said once. An exam that ends without notice
is a scramble; one that nags is a distraction.
Reverts the exam-exit-submits rule from earlier in this branch, which
was built on the opposite premise and would have charged wall-clock time
and then graded an exam whose clock should simply have stopped. Leaving
suspends, in both modes, and the overview no longer promises a clock
that does not stop for a break.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
Opening the analysis of a live attempt graded it whatever the mode. In
an exam that is a way to answer, look at whether it was right, and go
back and change it — the exam defeated rather than analysed. It reports
progress now: how many are answered, how long it is taking, and each row
as answered or not. No score, no percentage, and the donut counts how
far through it is instead of how much of it is right.
Study mode still grades live, because study mode marks each answer as it
is given; there is nothing here it has not already said.
Recommendations are withheld too, which is stricter than AMBOSS — they
show a dash for correct and then list the topics to go back to, which
says which questions were wrong by another route. A recommendation is a
verdict.
The withholding stops the moment the exam is over, submitted or expired:
settle_if_expired grades through the same function a manual submit does
and sets completed_at, and everything opens from there.
Tested on both sides, because this is an integrity rule and would come
back quietly the next time the live-analysis path was touched.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
Extraction wrote straight into `questions`, so a machine's first attempt
took a permanent id the moment it was produced. Ids come from a sequence
and are never reissued: every rejected draft burned one, and every draft
that needed fixing was sitting in the bank while it was being fixed.
A run now lands in a batch of drafts with their own table and their own
sequence. They are read, corrected and decided there, and `accept` is
the only place a Question is created — a copy rather than a translation,
because every field a draft holds is a field a question has, so nothing
is lost at the moment of acceptance.
Accepting is all or nothing, and everything is checked before anything
is created: a call that reports failure must not leave questions behind
from the drafts it got through first. My own test caught that — the
first question existed before the second draft was refused.
Readiness is reported for every draft rather than only on the attempt to
accept it, so a reviewer sees what needs work before opening anything.
A decided draft keeps its row and records what it became, so a batch
reads as a history of what was decided rather than emptying as it is
worked through. An acceptance cannot be undone from here: the question
exists, and deciding twice would make a second one.
No embeddings for drafts. A vector is for finding a question in the
bank, and a draft is not in the bank.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
Two shapes, because a plan is asked to do two different things. Papers
are rehearsal: each block is drawn to the ABP's published weights, so
sitting one says something about how you would do on the day. Domains
are study: the board's twenty-four content areas in its own order and
carrying its own titles, each given the share of the plan the board
gives it on the exam.
Both were written, then run against the real bank, which found two bugs
a unit test on a clean fixture would not have. Domains 19 and 20 —
nephrology and genitourinary — both map to our "Nephrology & Urology",
so a question sat in two pools and was dealt twice; the deal now keeps a
record of what has gone. And chunking every question a domain has into
blocks of forty gave preventive care six blocks and the plan a hundred
and sixty, which is not a plan: blocks are shared out by weight, with at
least one per domain so nothing the board examines is left out.
Built on the live bank alongside what was already there: Boards: Full
Papers (12 × 40) and Boards: By Content Domain (27 blocks, 1069
questions). Nothing existing was touched.
Psychosocial Issues and Child Abuse and Neglect — 6% of the paper
between them — had no category of ours at all, so they could contribute
nothing. Both now exist, with sub-topics named from the board's own
subdomains, and all 24 domains map to categories.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
The number beside a folder in Topic reading was a question count while
the browser lists articles, so "Hyperinflammatory Sepsis 4" meant four
questions and opened onto no reading at all. It counts what it opens
now, rolled up over the subtree, and a branch with nothing to read in it
is not offered — a folder with a number on it is a promise.
The hover card could not be reached. Its body was pointer-events: none,
on the idea that a hint should not sit between the reader and the link —
but the card is offset below the link and never covered it, while the
pointer travelling down to Split view crossed a body it could not enter,
so no mouseenter fired and the hide timer closed it on the way. The card
takes the pointer now, with a bridge across the gap.
And clicking the words opens the card rather than the article. A
cross-reference is read mid-sentence, and navigating away to find out
whether it was worth following is the thing that breaks the thread; the
card's two controls — beside what you are reading, or a tab for later —
are how you go. That also gives touch a route, where hover has none.
Modified and middle clicks are still the browser's.
The listing sent content and sections for all 331 articles, 214KB of
prose a list never renders. It sends what a list needs, which is 21KB.
The footer sat wherever the content stopped, so a page still loading put
it halfway up the screen with background below it. The shell is a column
the height of the window.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
The tutor is handed the correct answer and the explanation and told it
may reveal them, which is why it has never been offered during a running
exam — require_question_access already refuses that, whatever anyone
sets. What was missing is the other half: an administrator can now
withhold it from study sessions too.
Enforced on the server rather than by hiding a button, because hiding a
button does not stop a request. Reviewing a finished attempt is not
"during" and is unaffected; the answers are shown by then anyway. If
Redis is unreachable the tutor stays on — nothing is revealed that study
mode does not already show, so the permissive direction is the safe one
here.
GET /teach/prompt renders the instructions against a stand-in question,
so an educator answering "why did the tutor say that?" can read them
rather than infer them.
And a handbook at /handbook, for anyone who maintains questions or
articles whatever access they hold. It answers the things that were only
in the code: that a question links to an article three different ways —
a further-reading row, a key point carrying an article and section, and
a [[id|label]] marker in prose keyed by id so renaming does not break it
— what the tutor is told, why a blueprint shapes a paper, why deleting a
question hides it, and why changing the embedding model invalidates
every vector.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
Four gaps, one change.
Articles could not belong to an exam at all — an article reached one only
by inference through its category, which cannot say that the same article
belongs to a basic-science step and a clinical one showing different
views in each. article_exam_links says whether it is in the group;
Exam.article_views already decided what is shown once you are there.
POST /exams/ wrote name, slug, sort order and active, and silently
dropped family, description and article views, so a new objective landed
in "Other" showing everything whatever was asked for. It writes what it
is given now, and PATCH can change it afterwards.
Membership was one link row at a time, which nobody would do for three
thousand questions. POST /exams/{id}/assign takes whole topics with
everything beneath them — questions and articles both — and is
idempotent, so widening a selection and running it again adds only what
is new.
And the point of all of it: a real paper is not a uniform draw. The ABP
publishes that 12% of a general paediatrics exam is preventive care and
2% is rheumatology; forty questions drawn evenly is forty coin flips.
exam_blueprints holds a board's published outline — its own numbering,
its headings, its weights — and blueprint_category_links maps it onto
our taxonomy rather than bending the tree to fit, because their outline
is arranged for examining and ours for studying.
The sampler uses largest-remainder, so twenty-four percentages still come
to forty questions, and a domain that cannot supply its share gives the
shortfall back to be spread over those that can — the paper keeps its
length and loses only accuracy, and the working is returned so the
shortfall is visible rather than silent.
Seeded from the ABP General Pediatrics Content Outline (Oct 2024):
structure and published weights only, no exam material. 120 lines, 22 of
24 domains mapped; Psychosocial Issues and Child Abuse and Neglect have
no category of ours and are reported rather than hidden.
Creating an objective is now an administrator's rather than a
moderator's: it appears in everyone's picker and scopes the whole bank,
which is site configuration, and it sits with the other site switches a
moderator cannot reach.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
Question ids come from a sequence and are never reissued, and fourteen
tables point at them — attempts, quiz membership, exam membership,
media, article links, notes, favourites, feedback. Deleting the row took
all of that with it, so "restore" could only ever have meant typing the
text in again as a different question.
DELETE now sets deleted_at. The question leaves the bank, the builder,
search and every share path at once, because the exclusion lives in
general_question_predicate rather than at each call site. Restoring puts
back the same id, so everything that pointed at it still does. Erasing
for real requires the trash first and a moderator, and the confirmation
says what goes with it.
The trash page holds questions instead of tests. A test is a selection
you can remake in a minute; nobody wanted those back.
Used and withdrawn invite codes can be removed — an unused one is still
withdrawn rather than deleted, so it stays visible as having been issued
and stopped.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
Settings linked to a second dashboard with its own tab bar and its own
visual language. The admin sections are rendered in Settings now, under
headings that say who they are for — You, Content, The site — and each
has its own address, so People, AI models, Safety and Search are links.
/admin redirects into Settings for anyone who bookmarked it. AdminPage
takes a `section` prop and drops its tab row when embedded; it is loaded
lazily, so it is not in a learner's download.
Comments are gone: router, model, table and the half of the test file
that covered them. They were a discussion thread nobody was obliged to
answer, and feedback replaced them with a message addressed to whoever
maintains the question. The table was empty, so nothing was lost —
verified before dropping it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
A copy button tells you nothing about what you are about to send. The
dialog names the session, counts its questions and shows the stem it
opens on, then offers the link with Copy and the places people actually
send one — email, WhatsApp, Telegram.
Sharing is the administrator's to allow. GET /quizzes/share-policy is
asked before the dialog offers to make a link, so a switch that has been
thrown reads as "not offered" rather than as a button that fails when
pressed. A link already issued keeps working either way.
Removes the second, lesser share block that sat inside the save panel.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
The quiz player is a box the height of the window. The question used to
scroll the whole page, which took the session rail and the navigation off
screen exactly when you wanted them; now each column scrolls on its own
and the bar — Exit session, Previous, Next, Review — stays put.
Two site-wide switches, together under Settings → Site policy because
both are the administrator's and both apply to everyone:
* Sharing can be turned off. That stops new links being made; one
already handed to somebody keeps working, since revoking it would
break something a learner has already given away.
* Sign-up can be made invite-only, with single-use codes carrying a
note of who each is for and, afterwards, who it let in. A spent code
is kept rather than deleted — that record is the point of invite-only.
The alphabet has no O/0 or I/1/l, because these get read aloud.
The registration form asks for a code only when the site needs one, via
an unauthenticated policy endpoint — it has to know before there is an
account to ask with. It never says whether a given code is valid before
the account exists, which would make it somewhere to guess them. The
first account is always allowed, or a new install would lock itself out
before an administrator existed to issue a code.
Flags fall back to their defaults when Redis is down, in the safe
direction each way: sharing keeps working, sign-up does not silently
open.
Found on the way: the registration form's three labels named nothing —
no `for`, no wrapping — so a screen reader announced unlabelled boxes.
Backend 261/261, frontend 328/328.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
Comments are gone. A thread under every question was a discussion nobody
moderated, and what it was used for was telling an educator something was
wrong. That is now feedback: a private report, carrying the question id,
that someone is expected to act on.
* Give feedback sits in the question bar's new "more" menu, beside Save
and Share — occasional actions, folded away rather than each taking a
slot in a bar read on every question.
* An educator gets a badge of what is outstanding. Each row names the
question and opens its editor, where the report sits beside the field
it is about; reply, resolve, reopen or delete from there.
* Resolving keeps the report. A question with a history of the same
complaint should visibly have one; deleting is for the ones that were
never about the question.
* A granted educator sees only their own branch. The badge answers
quietly with zero for someone with no access, so the header can ask
without first working out who is asking.
The question bank is now the Qbank: create a session, and the last three
with Resume. Its facets, tag tree and create-a-quiz were a second copy of
the custom-session page; marking and folders belong in the player while
you are sitting a question. Import and export moved to the question
manager, which is the one place questions are managed, and which now has
a Preview that opens over the list instead of a page you have to come
back from.
Fixed while there: a session in progress analysed as 0/0 with an empty
table, because the analysis read attempt_answers — written on submit —
while the session list counted the saved progress. They read the same
thing now. The category trail is gone from the player: it named the
answer's own topic and led out of a session part-way through. An option's
reasoning opens on click and closes on the next one.
Backend 253/253, frontend 323/323.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
`[[403|urethritis]]` became `<ArticleLink slug="403">`, which built the
href `/articles/s/403` — the slug route — and asked the preview endpoint
to resolve "403" as a slug. Neither exists, so the hover card never
appeared and the link 404'd. Every one of the 2,150 links is written by
id, because an id survives a rename and a slug does not, so this was the
whole library and not one article.
resolve_slug now takes an id as well as a current or historical slug,
and the link addresses the article directly when it is written by id.
Also: a view of one section no longer prints a heading repeating the tab
above it. "Short" over a heading reading "In short" says the same word
twice, and hid the only content behind a chevron. No collapse control
over a single section, and no contents list of one entry.
And a horizontal-overflow guard that only half worked: `overflow-x:
hidden` was on body but not html, so the browser could still propagate
the overflow to the viewport and scroll the whole page sideways — which
is how the navbar came to be clipped mid-word. The exam name now
truncates with an ellipsis instead of clipping.
Backend 242/242, frontend 316/316.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
question_media replaced the two filename columns months ago: any number
of figures per question, each with a role, a label the prose can refer
to, a caption and an order. Only the editor's own endpoint ever read
them. The editor showed the two legacy text fields, and the player and
the answer review rendered the legacy paths — so the model existed and
nothing used it.
- FigureManager in the question editor: add from the image bank, name,
caption, reorder, remove, per role. A figure with no caption is called
out, because a caption is how anyone finds it again. The image id is
shown, since that is what the link survives a rename by.
- FigureStrip on the player and the review. Explanation figures are
labelled thumbnails that open full size and page between them — a
stack of full-width radiographs between two paragraphs pushes the
explanation off the screen, and "as in Figure 2" needs Figure 2 to be
named where it sits. A stem figure stays full size: it is the question.
- question_figures.py is the single place rows become what a page
renders, so the three views cannot disagree.
- Explanation figures are withheld until answers are revealed, the same
rule the explanation itself follows.
The legacy paths still render where a question was never backfilled, so
nothing that worked before stops working.
Backend 242/242, frontend 290/290.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
Access lived in three screens over two tables: category grants in the
question manager, media-library grants in the image bank, and nothing at
all for articles. Nobody could see what one person actually held.
/access is one surface over the same tables. A person on the left,
everything they have on the right. A granted branch shows its children
as covered rather than as separately tickable — a checkbox that changes
nothing is where a permissions screen starts lying — and the count of
categories a grant actually reaches is stated, not implied.
"Everything" is the moderator role, and the page says so instead of
inventing a wildcard grant that would silently mean the same thing and
be impossible to audit. While it is on, the branches below are hidden,
because they no longer apply. Nobody can change their own access.
The gap this closes: an educator granted a branch could edit its
questions but not the articles filed under it — articles were
moderator-or-author only. An article is filed under a category, so a
grant over that branch now covers its reading too. No new table: the
inheritance that category grants already had does the work.
Backend 242/242, frontend 284/284.
Also: the split-view test now focuses the link rather than hovering it.
Hover starts a 350ms timer; focus reveals at once, because the component
does not make a keyboard reader wait. That takes the wall clock out of a
test about the split view. Earlier failures were it losing CPU to the
backend suite running alongside it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
Answers the two questions owed: the readiness shrinkage, the three
groupings, the priority ranking, and the four steps of adaptive
selection — including where it is weaker than it looks.
Writing it up surfaced two defects, both fixed here:
* adaptive_select took the first 2,000 candidate rows. The bank is
2,948, so about a third of it could never be selected, and which
third depended on database order. The cap is gone; two integer
columns per question is not a size worth protecting against.
* category lookup was a linear scan through every candidate for every
recorded answer — O(answers x candidates), the slowest part of
building a session. It is a dict now.
Left alone and documented instead, because changing them changes which
questions a learner is given and that is not a silent decision: adaptive
ordering uses raw category accuracy rather than the shrunk readiness the
recommendations page uses, and difficulty is a filter rather than
something the session moves along.
Backend 231/231.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
The same answers asked three ways, as AMBOSS does it: which reading to
go back to, which organ system is weak, which discipline is weak. It was
Systems/Subtopics, where "Systems" meant top-level categories — which
are disciplines, not systems — and "Subtopics" meant every category
below them.
* Articles (the default): rows are the published article behind a
category, so the row links straight to the reading.
* Systems: the 16 organ systems. No question is tagged with a system
directly — it carries a symptom keyword filed under one — so
membership rolls up through the keyword's parent.
* Disciplines: top-level categories, which is what the old "systems"
grouping actually was.
Only 1,502 of 2,948 questions carry a system tag, so the Systems tab
says so rather than showing half the bank as if it were the whole of it,
and relevance there is measured against what the grouping can see.
"Practise this topic" now practises the row you are looking at, on its
own axis. That needed system_ids on the builder — matched as "any tag
beneath this system", where the existing tag_ids is "every one of these
tags", so the two cannot be conflated.
Backend 228/228, frontend 266/266.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
Two bugs, one visible cause. `.an-page`, `.an-rail` and four more classes
were defined in both AnalysisPage.css and AnalysisSessionPage.css with
different values — one a 280px grid, the other 260px. Once the two pages
shared a rail both stylesheets loaded together, the later won, and the
content column collapsed to rail width: "General Pediatrics" wrapped one
letter per line and the table headers floated away from their rows.
AnalysisShell now owns the frame and the session list for both views.
The page stylesheets style their content and nothing else.
And a session nobody has sat is no longer a bespoke "nothing here" panel.
GET /attempts/quiz/{id}/analysis answers with the same shape at zero —
0%, 0/20, every row "skipped" — so it is visibly the same page the
learner will see filled in, with a line saying why the figures are zero
and Start below. A part-finished session says how many are outstanding
and offers Resume. Once an attempt exists the quiz address returns the
real analysis, so both ways in reach the same page.
Mobile: below 1000px the rail becomes a band above the content that
starts closed — on a phone the first thing on screen should be the
analysis asked for. Search field is 16px on touch so iOS does not zoom
the page in and refuse to zoom back out; rail rows are 44px targets.
Backend 226/226, frontend 265/265.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
An unsuspended exam keeps running. When its clock runs out it is
submitted with what was answered and the score counts — a learner who
ran out of time sat an exam, which is a result and not an accident to
hide. Previously it was graded, flagged expired=1, excluded from every
statistic, and the client was told the opposite ("submit manually").
- attempt_expiry.settle_if_expired: one path, used by resume and by the
sessions list, so an exam left open elsewhere shows its score rather
than "in progress" forever. Suspended attempts hold their clock and
never expire.
- resume returns {expired_submitted, attempt_id}; the client opens the
analysis. The suspend dialog and the leave warning now say what
actually happens.
- delete: saved progress and device lock cleared; a study-plan block
whose only completed attempt is deleted goes back to unfinished.
- POST /attempts/reset-all: typed RESET, removes attempts, answers,
in-progress state, plan progress, reading marks, saved questions and
question notes; leaves the account, authored content and AI chats.
Settings → Your data, with the counts reported afterwards.
Also fixed on the way: the first version of the sessions-list change
mutated the dict it was iterating; the test only passed because it had
one attempt. Now two.
Backend 223/223, frontend 258/258.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
From the three recordings and the AMBOSS screenshots.
Study plans
- Blocks of about 40, split evenly: 202 questions is six blocks of
33-34, not five of 50 and one of 2. Reseeded (no progress or reading
existed yet); the seeder now splits the same way.
- A block has its own page, laid out as a course module: the plan's
blocks down the left, this block's reading then its session in the
middle, back / previous / next along the bottom. Study or exam mode
is chosen there, before the session exists; afterwards the mode is
shown, not offered. The plan page is the table of contents and links
into blocks rather than starting anything.
- Progress on a block comes from the same /quizzes/sessions row the
Sessions page shows, so the two cannot disagree.
Sessions <-> plans
- A session started from a block carries its place in the plan: the
session list and the analysis both return `plan` (plan, block,
position, previous and next block). The analysis shows a strip with
the way back to the block and on to the next one.
- Submitting a session marks its block complete. Nothing ever set
completed_at before — every block read as unfinished forever.
Recommendations
- Framed by the learner's chosen study objective: answers and bank
material linked to a different exam are left out, and the page is
titled for the exam. Unlinked material stays in, as elsewhere.
Backend 216/216, frontend 257/258 (the one failure is
ArticleSplitView, which is timing-flaky under the full run and is
unrelated to this change; being checked separately).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
The page at /quizzes showed the same fifteen rows twice — once under a
"Sessions" tab as a list, once under a "Library" tab as cards — with
nothing distinguishing them. The navbar carried the duplication too,
with "Sessions" and "History" both pointing at the same page.
Board Review I-XII already exist as study plans. The Library tab was
showing the bulk quizzes those plans were built from, so the same twelve
titles appeared in both systems. Those quizzes are now origin='plan':
still real, still the parent of their questions via source_quiz_id, but
no longer offered as something to pick off a list. Once a learner has
actually sat one it is history, so the session list keeps it.
- QuizzesPage deleted; /sessions is the only listing
- /quizzes/* redirects to /sessions/*, preserving path and query
- submitting a session lands on its analysis, not the old score page
- the answer review drops its score hero, which the analysis owns and
stated differently; a course quiz keeps its card, having no analysis
- delete-attempt moves to the analysis page, where the session lives
Backend 208/208, frontend 246/246.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
The linking was the gap
The marker system was built weeks ago — resolves by id, survives a rename, shows
a preview on hover — and not one of 333 articles used it. Every article was
written in isolation, so a piece on croup named stridor and epiglottitis and
offered no way to reach either. `scripts/link_articles.py` reads what is written
and links it: 3,718 cross-references across 307 articles, by id, so a later
rename cannot break them.
Conservative on purpose, because a wrong link is worse than a missing one: only
the first mention in a section, whole words, longest title first so "Otitis media
with effusion" beats "Otitis media", never inside an existing link, marker,
heading, code span or table, and never an article to itself.
That exposed a second thing: the reading view had its own Markdown pipeline with
its own cross-reference regex, and it only understood the old slug form. It would
have printed every one of those 3,718 links as literal brackets. Article prose
now goes through the same renderer as the rest of the site.
Short and Clinical looked empty
Both are usually a single section, and everything starts collapsed, so the tab
showed one heading over blank space. A view of one section is not a contents
page; it opens.
Removed
Quiz reminders — emailed nudges to retake anything under 75%, with a scheduler
that existed solely to send them: the model, the service, the scheduler, the
email, the table. Article comments. The dashboard's in-progress list and its
stat cards, both of which the analysis page now answers better.
One mistake worth recording: the first pass at removing the reminder cleanup used
a regex that took 109 lines with it, including an unrelated endpoint. The test
suite caught it (`/attempts/quiz/{id}/in-progress` returning 404 instead of 403),
and the file was restored and edited by exact match instead.
208 backend, 249 frontend green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
Performance by category counted every row in an attempt, and an attempt holds a
row for each question including the ones never answered. A 360-question sitting
that was opened and abandoned therefore landed as 360 wrong answers, which is
why Emergency Medicine read 0% of 400 and Gastroenterology 1.1% of 277 — figures
that describe a sitting nobody worked through, not a learner who cannot do
emergency medicine.
Accuracy now counts only questions that were actually answered, and the note
under the heading says so. Coverage is a separate question from accuracy and
conflating them made both useless.
Also: the category performance block is gone from the dashboard, where it
duplicated the one on Analysis; and the nav says Sessions rather than Quizzes,
with History beside it — "quiz" describes the packaging, a learner sits a
session, and the two entries answer different questions: what can I sit, and
what have I sat.
208 backend, 249 frontend green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
Figures
A question could carry exactly one stem image and one explanation image, each a
bare path with no title, no legend, and no way for the prose to refer to it.
`question_media` makes a figure a row: it points at an image already in the bank,
carries a role, a label the text can name ("Figure 1"), a caption and an order,
and there can be as many as the question needs. The same radiograph can serve two
questions without being stored twice.
The 346 existing paths were backfilled into figure records and retitled —
`page_339_img_0.png` says where a file came from and nothing about what it shows,
so the filename moved into the caption where it is still searchable, and the
title became something a person can read.
On the editor question: no new platform needed. Milkdown is already installed —
ProseMirror-based, MIT, GFM tables, code blocks, LaTeX — and already used for
articles, courses and the quick question modal. Only the question *page* still
has plain textareas, and that swap is written down rather than rushed, because
the stem carries manual-highlight offsets and a WYSIWYG rewrite would move them.
Fewer hints during a quiz
The category trail and the difficulty pill were shown beside every stem. Being
told a question is filed under Neonatology, or that it is "hard", narrows the
answer before the stem has been read. Both now wait until the answer is in,
where the trail becomes a way to more of the same topic.
The dashboard is about questions
Quizzes and attempts describe how the material happens to be packaged. What a
learner is working through is questions: how many of the bank they have seen,
how many they have answered correctly, and their average. The old per-quiz
performance card — which needed two attempts before it showed anything — is
gone, superseded by the session analysis. The greeting sits above "continue your
study" rather than below it, where it read as a heading for the wrong section.
208 backend, 249 frontend green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
The results page showed a score and a wall of explanations. What a learner needs
afterwards is where the time went and what to go back to, so
/analysis/session/:attemptId gives them: a rail of recent sessions, the four
figures they act on — correct, completed, time per question, total time — a
donut, the weakest topics, and a paginated table of every question with its
status, difficulty, time and how peers did on it.
Time per question was not recorded at all, so it could not be reported. It is
now (`attempt_answers.seconds_spent`), banked when you leave a question and
including the one still open at submission — without that the last question of
every session would show nothing. Answers from before this read "—" rather than
claiming zero, and a question nobody else has answered has no peer rate rather
than 0%, which would read as everyone having failed it.
Also in this pass, from the review:
* quiz categories are gone from the library — a second taxonomy beside the
real one, putting a heading above every test;
* the board review sets are numbered rather than dated, in both the quizzes
and the study plans built from the same material, so a learner does not meet
2019 in one place and VII in another;
* the footer's standing note is one clause, and the gap above it no longer
looks like the page ended early.
Everything else asked for today is written down in docs/TODO.md rather than
half-built: resume instead of restart, an unsuspended exam that keeps running,
deleting a session's data, reset-all-data with a warning, recommendations split
by article/discipline/system, and the adaptive session. Two questions I owe
answers to are in there too.
208 backend, 249 frontend green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
The question toolbar now carries what a learner actually reaches for. An
attending tip — one sentence of the kind said at the bedside, stored separately
from the explanation because it is read before the answer is known and must not
give it away. A note of their own on that question, replacing a single global
note that was one page for everything and so was never about the question in
front of you. Saving to a folder, which the collections API has supported all
along with nothing in the player able to call it. And the share link, which
previously only appeared on the start screen.
Panels open one at a time under the toolbar; two at once would push the options
off screen.
Reset question resets one question, not the attempt: a misclick should cost the
answer you just gave, not the nineteen before it.
The clock shows session time, time on this question and the running average, in
study mode as well as exam mode — four minutes on one question is the number
that says whether you are learning or stuck, countdown or no countdown. It
pauses, because time spent making tea is not time spent thinking.
208 backend, 249 frontend green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
The objective did almost nothing
It scoped question counts and nothing else, which is why changing it appeared to
have no effect. An exam now carries a family (USMLE, COMLEX, boards), a
description, and the article views it offers, and `/exams/` reports what the
current objective actually changes rather than leaving the learner to guess.
Reading follows from it: an article returns only the views its objective allows,
so someone revising a basic-science step is never shown bedside dosing they must
not act on — a view you can open but must never use is worse than one you were
never offered. An editor still gets the whole article, because they cannot edit
what they cannot see. An objective configured to show nothing falls back to all
three; that is a configuration mistake, not a preference worth honouring.
Unused figures deleted, at the user's request
3,262 figures — 334 MB — that nothing had ever used. "Unused" was defined by
exclusion and every exclusion was checked rather than assumed: kept if any
question uses it as a stem or explanation image, if any question version
mentions it, or if it appears in article prose or a flashcard. 440 kept, and
five question figures spot-checked as still readable afterwards. MinIO is now
596 objects, 520 MB, down from 3,858 and 854 MB.
This is not reversible from the application; the nightly borg backup of the
volume is the only way back, and that is stated in the script rather than
assumed.
For the record, since it was asked: the extraction is PyMuPDF, with an MD5 skip
list for repeated branding images. It pulled every embedded image from all 18
source PDFs, which is why one 767-page document alone produced 908 of them.
208 backend tests green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
The links I put in the save bar are gone — that bar was right as it was, and a
row of navigation crammed above it was clutter in the one place a person is
trying to finish a question. The footer is where going somewhere else belongs.
`SiteFooter` replaces the copyright line: four columns — Study, Library, Find,
PedsHub — with About, Contact, Account and Settings among them, and the standing
note that this is revision material rather than clinical guidance, said once at
the bottom of every page. A test asserts every link points at a route that
actually exists, because a footer full of dead links is worse than a short one:
the reader learns not to trust any of them.
Two retrieval faults the writing found
A bare condition name is a thin query. "Rickets" alone retrieved five passages
about *Rickettsia* — an embedding has little to go on in one word, and the
nearest neighbours of a short string are whatever looks like it. Asking as
"Rickets in children: definition, causes, clinical features, diagnosis and
management" took the contamination from five passages to none, so both the
pipeline and the generated route now ask that way.
And a category that names a department rather than a condition retrieves chapter
headings and whatever sits near them. "Pediatric Nephrology" passed the material
check with entirely irrelevant passages, and an article called that is a
department, not something to revise. Those names are now excluded from the topic
list.
Both were found by an agent writing articles and reporting what looked wrong,
rather than by anything automated noticing.
247 frontend tests green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
"PREP" is the American Academy of Pediatrics' trademark for their own product.
The plans here are our own sets of questions grouped by year, so they are now
named for what they are: Board Review 2021, and Mixed Review for the plan that
draws from every year at once.
Renamed in the database as well as the code — 13 plans, 14 quizzes a learner had
already generated from a block, and the 12 year tags, which appear in the
question bank's filters and are as visible as the plans. The seeder matches both
the old and new names so a fresh import still finds its material, and the tagger
mints the new one so the next run cannot undo this. Prompts and comments that
described the source PDFs by that name now describe them by what they are.
The generation run's 377 failures were not a bug
Every call was reserving the model's full 64k output ceiling, and OpenRouter
refuses the whole request when the balance is below the reservation — "you
requested up to 64000 tokens, but can only afford 52017" — however short the
answer would actually be. `_call_model` now takes a max_tokens, and the article
writer asks for 4000, which is comfortable for three views of one topic and
keeps each request small enough to be affordable. 98 articles were written
before the balance ran down; 158 exist in total.
Generation is paused at the user's request while credits are topped up.
208 backend, 243 frontend green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
The generation run stalled at topic 28 with the process alive and the log
frozen. `_call_model` had no timeout — every other call in ai_service.py has
one — so a stalled connection to the proxy hung the caller indefinitely. An
interactive request survives that because the person gives up; an unattended run
of five hundred topics does not, it just stops quietly and looks busy.
It now takes a timeout, generous by default and 150s from the article writer:
long enough for a full article, short enough that a stall is noticed in minutes
rather than found hours later with nothing written since.
Separately, the AI Mode tests passed this morning and failed this evening with
no code between them. Not flakiness: they call the real Redis rate limiter, and
sixty-eight runs of the suite had exhausted a daily limit of sixty. A test that
depends on shared external state stops testing the code and starts reporting how
often it has been run, so the limiter is now patched out for those tests.
208 backend tests green, and the run is moving again.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
An admin could already edit an article's prose, but not which view a section
belonged to and not its references at all — generation attached those and
nothing could touch them. And the revisions the API had been writing since the
CMS landed were unreachable from the interface.
Editing one view at a time
Short, Long and Clinical are tabs, each showing only its own sections with a
count on the tab. All three live in one list because they are one article, but
editing them together made it impossible to tell which version you were
changing, and a stray edit to the clinical view while meaning to fix the long
one is a mistake nobody notices until a learner does. A parent can only be an
earlier top-level section of the same view, which is what the server enforces.
Deleting a section lifts its children rather than taking them with it: a
survivor pointing at a section that no longer exists is worse than an orphan.
References are editable and structured
Title, author and pages, so the editorial queue's "published without sources"
stays a truthful question. A save that does not mention references leaves them
alone rather than clearing them, or an older client would silently strip the
provenance generation attached.
Version history
Every save is listed with what it was, and any of them can be opened or put
back. Restoring is itself a save, so the version you are leaving is kept too — a
history you can only walk one way is not a safety net, it is a trapdoor. Someone
else's draft returns 403 rather than being readable through its history.
One bug this turned up: the page-number field was derived from the parsed array
on every keystroke, so typing "12, 14" became "1214" the moment the comma
landed. The field now holds what you are typing and the array holds what gets
saved, and on blur it shows what was actually stored so a dropped entry is
visible rather than a silent difference.
208 backend, 243 frontend green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
Renames the middle view to Short and puts it first: it is the quickest way to
tell whether this is the article you wanted, and the full text is one click
away. Existing generated articles were migrated in place.
The prompt now asks for bullets that each carry a fact, because "X is important
to recognise" is a bullet that survives revision and teaches nothing.
The retrieval bug that made the last run mostly skips
The prose filter — drop chunks under 200 characters, since they are headings and
index lines — ran *after* taking the top fourteen hits. A broad query like
"Immunodeficiency" or a specialty name matches chapter titles first, so all
fourteen were index lines and the filter left nothing: the topic was skipped as
having no source material when the library holds plenty. Retrieval now asks for
five times what it needs and keeps the first passages that are actually prose.
Immunodeficiency went from 0 passages to 14, Pediatric Cardiology 0 to 14.
That is the same mistake the folder filter has a comment warning about — filter
inside the ranking, not after it — made two functions later.
Two things I got wrong and corrected rather than worked around: a `LIKE
'%key_points%'` check reported the migration had failed, when `_` is a
single-character wildcard and it was matching the title "Key points"; and a
variant count showing no Short sections was taken against the old image, where
Short was not yet a known variant and was being coerced to Long.
203 backend, 234 frontend green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
MinIO was resolving to the wrong container
Putting the backend on danvics_milvus to reach the clinical index gave it a
second service called `minio`, and Docker resolved that one first. Every object
read failed with InvalidAccessKeyId while the bucket simply looked empty — all
435 stem images unservable, and nothing in the logs saying why. The quiz MinIO
now answers to `quiz-minio`, which nothing else on this host claims.
A topic named after a shelf retrieved headings, not prose
"Pediatric Pulmonology" returned ten chunks whose top hit was 29 characters —
`**270** Pediatric Pulmonology`, an index line. Chapter titles rank well against
a query that looks like a chapter title. The model was handed a prompt with
citations and no content and said so, which was the correct response and read as
a JSON failure.
Two gates, both stated in the code. A chunk under 200 characters is a heading or
a running header rather than something to write from. A topic whose passages
total under 3,000 characters is skipped with the count in the reason, rather than
asking a model to write a medical article out of fragments — it will either
refuse or invent, and only one of those is visible.
The 71 generated drafts are deleted at the user's request. Nothing linked to
them and generation is resumable, so the cost was model calls rather than work.
Question bank corrections, from the agent that ran alongside:
262 questions had OCR-mangled units repaired — `inEq/L`, `mrnol/L`, flattened
`10⁹` superscripts and the rest — each with a version snapshot written first, so
every edit is reversible from the existing question editor. 94 stem images that
belonged to the explanation were removed; PREP's own `Item Q37A` / `Item C37B`
labels turned out to be a far better signal than word cues, taking the confident
split from 69/58/308 to 300/81/54. 13 uncertain images are listed for a person.
203 backend, 223 frontend green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
Standardises cross-references the way we agreed, and puts a CMS around articles
so hundreds of generated drafts are reviewable rather than merely present.
Links, made rename-proof
`[[7|Febrile seizures]]` resolves by id and displays the text — the id is the
part that must not change, the text is what keeps prose readable while you write
it. `[[old-slug]]` still resolves and is rewritten to the id form on save, not in
a migration: an article nobody has touched is not broken, and rewriting prose no
one asked to change is how an editor stops trusting the editor. Every slug an
article has ever had is kept, so a rename redirects instead of 404ing, and a save
reports markers pointing at nothing — at the moment the person who wrote the link
is still looking at it.
Three views of one topic
The full article to study from, the key points to revise from, the clinical view
to act from, with doses. They are views of one article rather than three
articles, so the numbers cannot drift apart and a question linked to the topic
still means one thing. Each section carries its variant; articles written before
this are the long view, unchanged.
CMS
draft → in review → published, with an author able to submit and only a
moderator able to publish. Every save snapshots what was there, restorable, and
restoring is itself snapshotted or the way back from a mistaken restore is gone.
The editorial queue is work rather than inventory: waiting for review, generated
and unread, published without sources, published with nothing to practise,
barely written. An empty bucket is drawn as good news, not as an alert.
Articles from the clinical library
The library index is 1.8M chunks of reference texts embedded with bge-m3 — the
same model PedsHub already uses, so our query vectors are directly comparable and
nothing had to be re-indexed. Retrieval supplies the facts and the provenance;
the model supplies the prose. References are built from the metadata of the
passages actually retrieved, never from the model, so a reference cannot be
invented — the same property that makes an AI Mode citation trustworthy. A topic
with fewer than three grounding passages is skipped rather than written from
memory. Everything lands as a draft.
Two things worth naming. The generated text is original writing grounded in those
books, not extracts from them: their facts are usable, their sentences are their
publishers'. And there are two Milvus servers on this host — the collection with
the data is the one reached as `milvus`, not the similarly named one on the other
stack, which I wired up first and which silently refused.
Also fixed along the way: `litellm==1.28.13` has been withdrawn from PyPI, so
requirements.txt could no longer be resolved from scratch and the image only
built because of a cached layer. Later additions go in their own layer until the
pins are refreshed.
182 backend, 223 frontend green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
The design settled earlier, built as described: retrieval decides what may be
cited, and the server enforces it.
The model is handed a shortlist of at most fourteen sources from the learner's
own library and told to cite them by marker. Afterwards every citation it wrote
is checked against that shortlist and anything else is deleted before it is
stored or shown. A hallucinated citation is not unlikely here, it is impossible
— surviving is not a decision the model gets to make. A URL it invents is not a
citation either: only the marker form counts, so a plausible-looking link stays
in the prose citing nothing.
Retrieval reuses the hybrid search already in place, and each corpus keeps its
own visibility rules — the bank predicate and exam scope for questions, the
draft rule for articles, deck ownership for cards. A question source carries the
stem only: a chat that printed the answer would hand away the practice it exists
to prepare you for.
Curated links do the job they were built for. A retrieved row an educator tied
to another retrieved row is boosted, because two things somebody already linked
surfacing for one query is evidence rather than coincidence. Nothing is stored
for this; the boost lives only in that ordering, and the answer marks those
sources so the reader knows which claim rests on an educator's judgement rather
than on a ranking.
Citations are stored with the answer as filtered, so reopening a thread shows
the links it showed at the time rather than a fresh retrieval that may now rank
differently. In the page the markers become numbers and each number opens its
source; a section citation deep-links into that section.
Two smaller decisions worth naming: a question appears in the thread the moment
you send it and is handed back to the input if the answer fails, because typed
words are not something to lose on a 502; and someone else's thread returns 404
rather than 403, since whether it exists is not your business either.
182 backend, 206 frontend green — 16 of the backend tests are the citation
contract and the retrieval boundary.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
Thirteen plans were seeded with an API to serve them and nothing that called it,
so the whole feature existed only in the database. Two pages and the editing
endpoints it was missing.
/study-plans lists the plans with progress stated in blocks — "3 of 6 blocks"
is something you can act on, where "50%" only tells you how you feel about it.
/study-plans/:id is one plan: each block shows Articles, then Sessions, in that
order, because that is the order the block is meant to be done in.
Reading is now part of a block (migration f4a5b6c7d8e9). "Mark as read" is the
learner's own claim and reversible — someone who ticks the wrong row should be
able to fix it without an educator, and progress nobody can correct stops being
trusted and then stops being used. It is a separate table from `article_views`
on purpose: opening an article is not the same claim as having finished with it.
A draft article attached to a block is listed for the educator who can open it
and left out for everyone else, rather than offered as a dead link.
Editing is inline on the learner's own page rather than a separate builder, so
the thing being changed and the thing a learner sees are the same object.
Moderators create (as a draft — an empty plan is not something to put in front
of anyone), rename, publish, delete; add, rename, reorder and remove blocks;
move questions between blocks of one plan; attach reading found by searching
rather than by id.
Two places where the obvious implementation leaves the data wrong, both tested:
deleting a block out of the middle shuffles the survivors down, or the next
insert collides with a position nothing occupies; and reordering parks every row
outside the range before writing the real positions, because (plan_id, position)
is unique and the first move would otherwise collide with a position still held.
A partial order is refused rather than half-applied.
166 backend, 188 frontend green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XeFQJXJTfHKTfbfsdxv57Z
Five corpora were each searchable from their own page, which meant knowing which
of five pages held the thing you were looking for before you could look for it.
`GET /search` runs them together.
Visibility is never re-implemented here. Questions go through the same bank
predicate and exam scope as the question bank, articles through the same draft
rule, cards through deck ownership, images through library grants. A search page
with its own idea of who may see what is how private content leaks, so the tests
that matter are the boundary ones: a peer's search reaches neither another
user's unshared question nor their deck, and a draft is invisible to everyone
but the educator who wrote it.
A section hit is reported under its article, not beside it — ten sections of one
article are one result with ten places to start reading, not ten results burying
everything else. This is what the section index was backfilled for; each one
links straight to that section.
Results are grouped by kind rather than interleaved by score. A question and an
article are different kinds of answer, and a single ranked list makes you read
every row to work out which kind each one is. Snippets show the window around
the match rather than the opening of the document, because every document's
opening looks the same. A question found only by the semantic ranker says so.
The header box has two ways out: pick a suggestion and go straight to that
article, or press Enter and search everything. Suggestions are lexical and
prefix-first — a typeahead is finishing the word you are typing, and a semantic
neighbour of half a word is noise — and debounced 180ms so typing is not a
request per keystroke. One corpus failing is logged and returned as a gap in the
answer rather than a failed page.
154 backend, 163 frontend green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XeFQJXJTfHKTfbfsdxv57Z
Three things, all from how AMBOSS actually behaves rather than from a
description of it.
Row by row, two menus
The articles page is now the column browser itself rather than a grid behind a
"▼ All categories" toggle. Opening a topic opens its contents in the next
column, so the trail you took stays on screen and you can step back a level
without losing your place. Topics and the articles filed under them share a
column, because to a reader those are the same list — things this heading
contains — and only the icon separates a folder you can open from a page you can
read. Articles filed nowhere sit in the root column instead of being unreachable
for want of a heading. Under 720px it is one column plus a back button. Search
is a different question from browsing — you already know the name — so it still
answers with a flat list of matches.
Sections that collapse
An article is a reference you consult, so it opens as a contents page: headings
only, each expanding where it sits. A section may now sit under an earlier
top-level one (`parent_id` on the section JSON, absent on every article written
before this), which is how "ROS questionnaire" belongs to "Review of systems"
rather than standing alongside it. The contents rail nests the same way. Nesting
is refused where it could not render: its own parent, a parent later in the
article, a parent outside it, or a sub-section of a sub-section. A deep link
opens the target section and its parent — landing on a collapsed heading looks
like the link went nowhere. References are pinned last however they were
written; a reader scrolling for content should not hit the bibliography halfway
down.
Links that show where they go
`[[febrile-seizures]]` or `[[Febrile seizures|febrile-seizures]]` in article
prose becomes an in-app link that previews the target on hover: title, a couple
of sentences of actual prose with the markup taken out, and how much is there.
Following a link to find out whether it was worth following is the thing that
breaks a train of thought. Slugs, not ids, because that is what an educator
writes and it outlives a renumbering. One fetch per article for the life of the
page, a 350ms delay so crossing a link summons nothing, and no card at all on
touch, where a card would sit between the finger and the link.
Dead CSS for the old section modal and the always-open section block is gone —
nothing rendered those class names any more, and stale rules winning on source
order has bitten this page before.
146 backend, 152 frontend green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XeFQJXJTfHKTfbfsdxv57Z
Systems were never systems
The 27 top-level rows were disciplines and care settings — Cardiology,
Emergency Medicine, Neonatology, and a stray condition (Sepsis) — not organ
systems. Cardiology is a discipline; Cardiovascular System is a system. So the
facet was mislabelled, and there was no organ-system axis at all.
Both fixes, as asked:
* that tree is now the "Topics" facet, which is what it always was;
* "Systems" is a new flat axis of 16 organ systems, matching how AMBOSS keeps
Systems flat while nesting Disciplines and Symptoms.
Tags can nest (migration e3f4a5b6c7d8)
`question_tags` gains parent_id and sort_order. A tag may sit under one of the
same kind (Surgery > Hand surgery) or under a system, which is how symptoms are
grouped by where they present. 726 symptoms are now filed under the system they
appear in; the remaining 3,536 stay top-level rather than being forced into an
approximate bucket. A false positive the dry run caught: "vision" was matching
"Health Supervision" — the same trap as erythema/erythematosus earlier, fixed
with a word boundary.
Admin can grow the taxonomy without a migration
POST /tags creates a top-level entry or a child; PATCH renames, reorders and
reparents, refusing a cycle; DELETE reparents children to the deleted tag's
parent rather than orphaning them, and can move its questions elsewhere;
POST /tags/{id}/questions attaches questions. Everything appears in every picker
immediately, because they all read the same endpoint.
Article sections were indexed but empty — `_rebuild_section_index` only runs on
save, so articles written before it existed had no rows. Backfilled: 10 articles,
28 sections, now embedded and searchable. Section-scoped question links already
worked (7 of 34 links name a section).
Tests: 10 new backend covering the tree shape, adding top-level and child
entries, kind rules, duplicate refusal, cycle refusal, rename/reparent, question
attachment, delete-reparents-children, delete-with-move, and the moderator gate.
141 backend, 136 frontend green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017acfNLsJpnkvH3sCZSjMJM
Serving went straight to disk with FileResponse, so object storage was
effectively write-only: bytes went to the bucket and were still read from the
volume. `/uploads/{path}` now tries the local file first, then the object,
keeping the existing authorisation and path-confinement checks in front of both.
That is what makes the volume removable at all.
Migration (scripts/migrate_uploads_to_s3.py)
Every file is copied and read back with a SHA-256 comparison before anything is
deleted, and deletion is a separate opt-in flag that refuses to run if a single
file failed to verify. 3,852 files, 853.7 MB, all verified, then removed from the
volume — which now holds 0 files.
A bug this caught in its own first run: verification used `storage_service.load`,
which falls back to the volume, so it compared each local file against itself and
reported 3,852 perfect matches against an empty bucket. `s3_object` reads
strictly from S3 with no fallback, and verification uses that. The fallback is
right for serving and wrong for verifying, and the two now have separate calls.
Proven before deleting: a file removed from the volume still served correctly and
byte-identically from the bucket.
Backups, corrected: borgmatic already covers /var/lib/docker/volumes, so
quiz_minio_data is backed up nightly with 7/4/6 retention — my earlier claim that
MinIO was outside the backup routine was wrong, based on db-backup alone.
Existing archives still hold the old uploads volume, so there is no window in
which these files exist in only one place.
Tests: 131 backend green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017acfNLsJpnkvH3sCZSjMJM
Storage
Media now goes through `storage_service`, which has two backends: the container
volume, and S3/MinIO. A volume can only be mounted by one host, has no presigned
URLs and no lifecycle rules, none of which suits ~860 MB of media. Reads fall
back to the volume when an object is missing, so the existing uploads keep
working and files can migrate gradually rather than in one risky pass.
A row stores the object key, never a URL: a URL embeds the backend, so a row
holding `http://minio:9000/...` breaks the moment the backend changes.
MinIO publishes no host ports — the backend reaches it over the compose network,
and 9000/9001 are already taken on this host by other stacks.
Image libraries (migration d2e3f4a5b6c7)
An image belongs to a library, and a person is granted a library the way they are
granted a category, so access can be given to some images without giving away all
of them. Tags reuse the shared `question_tags` vocabulary rather than inventing a
media-only one. Uploads are type- and size-checked, stored through the service,
and embedded so an image can be found by what it shows.
Classification finished
The 316 questions the chooser had declined are now filed with `--force`, which
takes the nearest candidate from the same shortlist the chooser saw. 306 were
forced, 10 the chooser accepted on this pass. No question sits on a bare system
any more:
system only 2,730 -> 0
condition/subsystem 214 -> 1,782
full depth 4 -> 1,166
A forced match is a weaker signal than a chosen one, so expect more errors among
those 306 — but the original system stays as a cross-link, so nothing is lost and
they can be corrected by hand.
Tests: 8 new backend covering library scoping, edit confinement, shared-vocabulary
tags, storage indirection on upload, and type/size limits. 131 backend green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WgRcMaScVEL7TBLpnAoSV9
An admin can now say "you edit Step 1 Cardiology" rather than only "you edit this
category". Each dimension on a grant is nullable and means "any"; a grant covers
the questions matching all the dimensions it sets, and holding several grants is
the union of their coverage (migration c1d2e3f4a5b6). A check constraint refuses
a grant that names nothing, which would otherwise mean "everything".
Permission checks now run against a predicate over Question rather than a set of
category ids, so the exam and discipline dimensions actually take effect on edit,
delete and bulk actions instead of being silently ignored.
Pediatrics is unbound back to a global tag. With counts already scoped by the
learner's active exam, one global row gives the right number per exam, so
scoping the row bought nothing and duplicating a 6,740-tag vocabulary per exam
would have to be repeated for every rename and merge.
Tests: 123 backend green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01365DYKu14YtsBKv2ycW6eG
A discipline may now belong to one exam. `question_tags.exam_id` NULL keeps a tag
shared — Cardiology means the same thing whichever exam you sit — while a set
exam_id scopes it. Boards Pediatrics and a future Step 1 Pediatrics are therefore
separate rows over genuinely different bodies of content, not one label stretched
across both. Uniqueness moves from (name, type) to (name, type, exam) to allow it
(migration b0c1d2e3f4a5).
`scripts/bind_exam_tags.py` binds Pediatrics to Pediatrics Boards and tags the
884 questions in that exam that were missing it — the whole bank is paediatrics,
so it now reads 2,948.
Facet counts are computed within the learner's active exam, and a tag scoped to a
different exam is left out: an unscoped list offered disciplines that could not
match anything they were studying. With no exam chosen, everything is offered as
before.
Tests: 4 new backend (same name once per exam, unscoped list offers all, choosing
an exam scopes counts and hides other exams' tags, switching exam switches which
Pediatrics is offered). Full suite green: 123 backend.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PpfzbZ1QTLMeVYxM2kyq8m
Editing a question now snapshots its previous state. The last 5 are kept — the
value is undoing a recent mistake, not an audit trail, and an uncapped history of
full question bodies grows without bound (migration a9b0c1d2e3f4).
A restore snapshots the current state first, so the restore is itself undoable.
History is gated by the same per-category grant that gates editing, so it cannot
be read by someone who could not have made the edit. The question editor shows
the versions with their dates and a Restore action.
Also added docs/TODO.md tracking everything requested and not yet delivered:
AI Mode and its citation contract, global search, study-plan editing and
articles-in-blocks, admin settings revamp, image libraries and question folders,
media management, nested article sections with references and per-section notes
and feedback, per-question notes and feedback in the runner, tutorial mode, the
per-question performance table, the Overview dashboard, systems subsystems, and
dropping "Pediatrics" as a discipline.
Tests: 6 new backend (snapshot on edit, cap at five newest-first, restore,
restore is undoable, refused without edit rights, unknown version). Full suites
green: 119 backend, 136 frontend, build clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PpfzbZ1QTLMeVYxM2kyq8m
The PREP sets were loose admin-generated quizzes. They are now study plans: one
per year, split into blocks of 50 numbered "Block 1", "Block 2", plus a
"PREP Mixed" plan of 300 drawn at random across every year. 12 plans, 2,821
questions, applied to production.
Block membership is snapshotted rather than stored as a filter — a plan you are
part-way through must not reshuffle between visits. Re-running the seeder updates
years whose questions changed and leaves the mixed draw alone unless --reshuffle.
Starting a block reuses the learner's existing quiz for it; without that,
reopening a block would create a duplicate test each time and scatter the
attempts across them. Only questions the learner may see are included.
Tests: 113 backend green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PpfzbZ1QTLMeVYxM2kyq8m
Images were findable only by the filename someone typed. `media_assets` gives
them a title, caption, alt text, a category on the shared tree and tags, with a
weighted tsvector so they are searchable now (migration y7e8f9a0b1c2).
The embedding column is filled from the caption today. A vision-capable model can
fill it from the image itself later without another migration — and because
`embedding_model` stamps every vector, a text-embedded caption and a
vision-embedded image stay distinguishable instead of being silently mixed in one
index. Adding "media" to the embeddable kinds is all the retry task, the full
regeneration and the health report needed.
`media_tag_links.tag_id` carries no ORM-level foreign key: `question_tags` is
created by raw DDL rather than a model, so the constraint lives in the migration
where the table actually exists.
Tests: 113 backend green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PpfzbZ1QTLMeVYxM2kyq8m
An article embedded as a single vector, which finds the article but not the
paragraph — so a citation could only ever point at the top of a page. Sections
live in a JSON column and cannot carry a vector or a full-text index, so they are
now projected into `article_section_index`: one row per section with its own
embedding and weighted tsvector (migration x6d7e8f9a0b1).
- Rows are keyed by section id, so editing a section updates it, removing one
deletes it, and an unchanged section is not re-embedded on every save.
- `article_section` joins the embeddable kinds, so the retry task, the full
regeneration and the health report cover it without further changes.
- `hybrid_ids(db, query, "article_section")` searches it like any other corpus.
This is the groundwork for grouped search results (article, then the sections
that matched) and for AI citations that deep-link to the right section.
Tests: 113 backend green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PpfzbZ1QTLMeVYxM2kyq8m
Exams (migration v4b5c6d7e8f9)
"Pediatrics Boards" was a hardcoded checkbox that filtered nothing. Exams are now
rows: Pediatrics Boards and USMLE Step 2 CK ship seeded, and everything already
in the bank is linked to the boards. Membership is a link table, not a column,
because one paediatric cardiology question can count towards several exams.
The learner's choice lives on `users.active_exam_id`, so it follows them between
devices instead of sitting in one browser's storage. Choosing an exam scopes the
bank; a question with no exam links stays visible, since unlinked content is
unclassified rather than excluded. A switcher sits in the navbar.
AI mode — matching, never generating
Both entry points build a test from the educator-reviewed questions that already
exist, ranked against the request. Nothing is invented:
- POST /questions/builder/describe turns "what I want to study" into a test.
- POST /questions/builder/from-upload matches a document against the bank. The
file is read in memory and never stored — it is a search query, not a source
of questions, so there is nothing to retain or expire. 10 MB cap, 30 questions.
Handing a whole document to `websearch_to_tsquery` builds one enormous
conjunction that matches nothing, so text over 300 characters is reduced to its
most distinctive terms, OR-joined, before it reaches the lexical ranker.
Continue your study (migration w5c6d7e8f9a0)
A dashboard panel with the sessions in flight and the articles most recently
opened. `article_views` records one row per learner and article, written best
effort so a reading page never fails because a bookkeeping write did.
Tests: 5 new exam tests (active exams and counts, choice persisted and cleared,
unknown/inactive refused, bank scoping including unlinked questions, moderator-only
creation) and 7 for AI-mode matching (no questions created, invisible questions
excluded, no-match reported rather than an empty test, upload limits enforced).
Full suites green: 113 backend, 136 frontend, build clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PpfzbZ1QTLMeVYxM2kyq8m
Retrieval generalised beyond questions
`_text_for_question`, `embed_question` and `hybrid_question_ids` all hardcoded
the questions table, so there was nothing to call for an article or a card. That
layer is now corpus-agnostic:
- `Embeddable` mixin gives articles and flashcards the same embedding,
embedding_model and embedded_at columns questions have, plus a weighted
full-text vector (migration u3a4b5c6d7e8).
- `embed_record(row, kind)` is one code path for all three — they share an
embedding space, so they must share the model and provenance rules too.
- `hybrid_ids(db, query, kind)` ranks any corpus; `hybrid_question_ids` stays as
a thin alias for existing callers.
- Article and flashcard search moved off `ILIKE '%term%'`, which could not find
a jaundice article from "yellow newborn".
- The retry task and full regeneration now sweep every corpus, and the health
report breaks down current/stale/missing per kind.
- Articles embed on create and on edit, with failures left to the retry task.
Quoted phrases replace the keyword-only mode
`websearch_to_tsquery` already gives "absence seizure" exact-phrase semantics,
and the semantic ranker sits out a quoted query. That covers the one case a
keyword-only toggle was for — exact lookup — per query rather than as a sticky
setting whose every position returns a subset of the default.
Full-page question editor (/questions/new, /questions/:id)
Editing happened in a cramped modal. There is now a page with room for the stem,
per-option explanations, a searchable category picker with primary plus extras,
difficulty, and images. It shows the question's id with a copy button, and
Duplicate creates a variant without retyping the stem. `GET /questions/detail/{id}`
backs it, pathed under /detail/ so it cannot shadow the static routes.
Question bank filter bar restyled — the toggle and count read as one control
instead of two grey pills crowding the result count.
Tests: 101 backend green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PpfzbZ1QTLMeVYxM2kyq8m
Questions and articles both pointed at `question_categories`; decks had no
category at all, so the three content types could not be filtered together and a
topic's cards were unreachable from its category.
- `flashcard_decks.category_id` references the same tree (migration
t2f3a4b5c697), so one category now spans questions, articles and cards.
- `GET /flashcards/` takes `category_id` and includes descendants, so a parent
category picks up everything filed beneath it.
- `PATCH /flashcards/{id}` files or unfiles a deck, refusing a category id that
does not exist rather than storing a dangling reference.
- The cards page shows each deck's category as a selector.
Tests: 5 new backend (all three types resolve to the same id, descendant
filtering, file and unfile, unknown category refused, renaming leaves the
category alone) and 1 new frontend. Full suites green: 101 backend,
136 frontend, build clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PpfzbZ1QTLMeVYxM2kyq8m
Search
- Retrieval was hybrid in name only: the keyword filter was applied to the SQL
query, so results were the *intersection* of the two rankers. A question that
matched the meaning but not the literal string could never be returned. It is
now a union, fused with Reciprocal Rank Fusion (a text rank and a cosine
distance are not on comparable scales, so RRF uses only their orderings).
- Added a generated `search_vector` tsvector + GIN index, so the lexical half is
ranked full text rather than ILIKE substring matching.
- Chose Postgres + pgvector over OpenSearch/Elasticsearch: a search cluster
would add a second datastore to keep in sync and a JVM on this host, to
replace an index Postgres maintains inside the same transaction.
- Removed the keyword-only mode. It looks precise but silently drops the
question that asks the same thing in different words.
Embeddings — measured on 500 real questions, using each question's own
explanation as a paraphrase query (known answer, no hand labelling):
bge-small (local CPU, 384d) R@1 0.840 R@5 0.953 186ms/query
bge-m3 (LiteLLM proxy, 1024d) R@1 0.847 R@5 0.973 93ms/query
BGE-M3 wins on both quality and latency and needs no extra credential, since
llm.danvics.com already serves `openrouter-bge-m3`.
Three gaps this exposed, all fixed:
- Nothing recorded which model produced a stored vector, so changing models
silently mixed incomparable spaces. `embedding_model` / `embedded_at` now
stamp every vector, `GET /admin/embedding/health` reports current vs stale vs
missing, and regeneration defaults to stale-only.
- The generator read the model from env while the stamp read a Redis override,
so a vector could be labelled with a model that did not produce it. Both now
resolve through one function, with a regression test.
- Embedding at creation is best effort, and a failure left a question invisible
to semantic search forever. `retry_missing_embeddings` runs every 15 minutes
via Celery beat and backfills missing or stale rows.
- Query embeddings are cached in Redis per model, so typing is not a network
round-trip per keystroke.
`dimensions` is only sent to OpenAI's embedding-3 family; BGE-M3 rejects it.
Tests: 8 new backend tests (union not intersection, fusion ordering, per-ranker
failure degradation, provenance stamping, stale/missing accounting, generator
and stamp agreement). Full suites green: 95 backend, 127 frontend, build clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014yhHB8Pc7oQqyqn2Vo9DXA
Analysis / recommendations (AMBOSS parity, verified on next.amboss.com):
- GET /study-tools/recommendations ranks focus areas by the study time most
likely to raise the score. Readiness is the learner's accuracy in a category
shrunk toward their own overall accuracy in proportion to sample size, so two
unlucky answers do not read as a knowledge gap; it unlocks after 40 answers.
Relevance is the share of the bank a category holds. Counts roll up through
the category tree, so a system inherits its children's questions.
It is deliberately not called EPC and does not claim to predict an exam.
- New /analysis page: Performance and Recommendations tabs, readiness summary,
adaptive-session box, and expandable focus rows showing questions seen,
answered correctly, the linked article and a per-topic practice action.
Per-category educator grants:
- category_grants table (migration p8b9c0d1e253) plus utils/category_grants.py
resolving a grant to the category and all of its descendants.
- Question create, edit, delete, bulk and the manager summary now accept a
moderator OR an educator granted the affected categories, and refuse moves
that would push a question out of the holder's scope. Summary counts are
scoped to the grant.
- Moderator endpoints to list, add and revoke grants, plus /my-grants driving
the nav link and the manager's scope banner; grantable-users avoids handing
moderators the admin-only user list.
- GrantsPanel in the question manager: grant, list and revoke with inline
confirmation.
Question page:
- The category trail was a fixed 78px band that wrapped into several rows and
pushed the stem down the page, followed by three more stacked strips. It is
now one scrollable meta line (breadcrumb + difficulty + type) and a single
AMBOSS-style action bar (Mark / Listen / Listen through / Clear) between the
stem and the options. Difficulty is exposed on the runner payload.
Deploy fix: index.html shipped with no cache header, so browsers kept serving
the previous bundle references and a release looked like nothing had changed.
nginx now sends no-cache for HTML and immutable long-cache for hashed assets.
Tests: 16 new backend (recommendation shrinkage, roll-up, locking, grant scope
across create/edit/delete/bulk/summary, moderator gate) and 10 new frontend.
Full suites green: 88 backend, 116 frontend, build clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014yhHB8Pc7oQqyqn2Vo9DXA
Quiz management area (AMBOSS parity, verified against next.amboss.com):
- GET /quizzes/sessions returns one management row per accessible quiz —
attempt state, live answered/total from Redis, last score and activity —
so the page no longer fans out per-quiz requests.
- QuizzesPage rebuilt as a session list grouped by day with a progress bar,
a state-aware primary action (Start / Resume / Review) and an action menu
matching AMBOSS: Analysis, Repeat, Rename, Share, Edit, Category, Delete.
Rename and delete confirm inline; no browser popups.
- Sessions / Library / Categories tabs replace the flat card grid.
- QuizPage honours ?restart=1 so Repeat always begins a fresh attempt.
Question manager (new moderator page at /questions/manage):
- GET /questions/manage/summary counts editorial gaps; /questions/bank gains
a `needs` filter (category / explanation / difficulty / private) so the
health tiles double as one-click filters.
- POST /questions/bulk applies category, difficulty, sharing or delete to up
to 500 checked questions in one call, moderator-only.
- Question edit/create modals extracted to components/QuestionEditors.jsx and
shared by the question bank and the manager instead of being duplicated.
Showcase articles:
- scripts/seed_showcase_articles.py seeds eight short starter articles across
the main pediatric systems, each filed under a real category, with stable
hex section IDs and links to bank questions from the same category.
Mobile: dedicated stylesheets for both pages — rows stack, the action menu
becomes a bottom sheet and the bulk bar docks to the bottom edge.
Tests: 9 new backend tests (session feed states, ordering, Redis-outage
degradation, visibility; bulk actions, gap filters, moderator gate) and 9 new
frontend tests. Full suites green: 72 backend, 106 frontend, build clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014yhHB8Pc7oQqyqn2Vo9DXA
Create/bank pages use AMBOSS-style facets: Exams, Disciplines, Symptoms, Systems, Articles, Saved. Personal question libraries with add-to-library in study modal. Adaptive session shortcuts from performance (including weakest topics). Quiz restart with fresh attempt. Category counts computed with two grouped queries. Migration n7a8b9c0d142. 63 backend and 97 frontend tests pass.
Sample quiz demonstrates option explanations, key points with article links and linked cards. Bank study modal shows per-option explanations and key points. Performance shows main categories with an expandable hierarchy. AMBOSS-style picker polish (chevrons, search box, switch, auto title). Mobile spacing fixes. 97 frontend tests pass.
Key points on questions link into article sections (AMBOSS-style) with samples; difficulty tagging with builder/bank filters; adaptive session algorithm prefers unanswered questions then recycles older incorrect ones, weakest categories first with damping; question create/edit is now admin/educator only; expired exams no longer auto-submit on resume; exam suspend messaging updated. Migrations k4f5a6b7c819, l5a6b7c8d920, m6a7b8c9d031. 63 backend and 97 frontend tests pass.