Commit graph

22 commits

Author SHA1 Message Date
Daniel
e72cdd6716 feat: a versioned API, refresh tokens, and an end-to-end stack that found four bugs
**The API.** Every route now lives under `/api/v1`, with `/api/...` rewritten
onto it — one route, two spellings, so they cannot drift and the OpenAPI
document describes each endpoint once. Errors carry an `error` object with a
stable code, one human sentence and, for a validation failure, the fields that
were wrong; `detail` is untouched so nothing that reads it breaks. The whole
surface — 320 routes, their parameters and their status codes — is checked in
as `backend/tests/api-contract.json`, and a test fails on any difference,
naming the routes that moved. `docs/api.md` is the contract in prose.

**Refresh tokens**, so an app can stay signed in without keeping a password.
Rows rather than signatures: listable, withdrawable, stored as hashes, rotated
on every use. A spent token coming back ends the whole session, because a theft
and a replay look identical from the server and the safe reading is the unsafe
one. A browser is not given one — it has nowhere to put it and a person to ask.

**An end-to-end stack**: `docker-compose.test.yml` with its own Postgres and
Redis, `e2e/seed.py` for the smallest world the tests name, and Playwright with
five projects — desktop, iPhone, Pixel, iPad and a browserless API project.
Devices because every bug reported this week was a phone bug found by a person
looking at a screenshot; a desktop-only suite would have passed through all of
them. Forty tests, five clean runs.

It found four things in its first hour:

- **A fresh deploy could not start.** `create_all()` ran before
  `CREATE EXTENSION vector`, so any database that had never had pgvector
  installed died on the first table with a vector column. Invisible here
  because this one has had the extension for a year.
- **A figure in a published article was a 404 for everyone but an admin.**
  Media in the library is nobody's to read by default, and nothing made an
  exception for a drawing an article actually shows — so every illustration
  added this week was an empty box for every real user.
- **Every rate limit was one bucket for the whole site.** The backend saw
  nginx's address for every request, so ten bad passwords from anybody locked
  out everybody, and no log line could say who. nginx now takes the real
  address from the proxy and overwrites the header on the way in; uvicorn runs
  with --proxy-headers.
- **The reading page's breakpoints disagreed** — 1150px in the component,
  820px in the stylesheet. Between them the menu button claimed the contents
  drawer and then toggled a class on a rail that was still in the layout: the
  contents did not open and the site menu did not either. The button was dead
  on every tablet.

And two smaller ones: the login limiter counted successful sign-ins, so eleven
people behind one hospital NAT locked each other out — it is cleared by a
correct password now; and `/uploads/{path}` served GET and HEAD from one route
with one operation id, which makes every OpenAPI client generator refuse the
document.

The first admin's password is generated and printed once at first start when
`DEFAULT_ADMIN_PASSWORD` is blank, rather than the account not existing:
`docker compose logs backend | grep -A3 "FIRST ADMIN"`.

CI (`.forgejo/workflows/tests.yml`) runs the backend suite, the contract, the
frontend suite and the build on every push to dev, main or master, and the
end-to-end stack on those branches and on pull requests into them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
2026-09-13 01:23:38 +02:00
Daniel
3279e14bb2 refactor: remove the LMS
There will be no courses. What was there: one draft called "jk" with two empty
lessons, and 4,000 lines of code around it — courses, modules, lessons,
enrolments, per-lesson progress, SCORM, BigBlueButton, completion certificates,
three React pages, a router, two models.

Its real cost was everywhere else. Every query that measured practice had to
remember `Quiz.course_id.is_(None)`, and forgetting it in one place would have
silently mixed course attempts into a learner's analytics; the bank predicate
carried a subquery to exclude a course's own questions from every search,
recommendation and share; quiz access had a second, parallel rule about
enrolment. All of that is gone, so the remaining rules say what they mean.

`quizzes.allow_review` goes with it. It was only ever enforced for a course
quiz, so it had become a promise nothing keeps — the public session page was
still offering "no answer review" about sessions that review fine.

The fixtures' question 5 lived in a course quiz and stood for "a question that
exists but is not in your bank". There is no such thing now — a question is in
the bank unless it is deleted — so the counts it kept out of the numbers are
back in, and the tests that turned on it now turn on deletion or on the
attempt that actually holds a question.

Files the LMS uploaded stay on disk and stay protected: LEGACY_LMS_PREFIXES in
app/utils/upload_access.py is what keeps them unreachable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
2026-09-12 23:27:51 +02:00
Daniel
16aed6b6b0 feat: Cap on its own host, hints per topic, and an objective is asked for
Cap moved from /cap/ under this app to cap.pedshub.com, so anything else
on this machine can use the same instance. Caddy terminates it, the
backend keeps verifying over the compose network rather than going out
and back, and the widget endpoint is configuration rather than a path
baked into the component. Verified: a challenge is issued on the
subdomain, and a token that was never issued is still refused.

"Correct using hints" is now a per-topic figure. The knowledge profile's
accuracy bar was two-tone because /study-tools/recommendations carried
only `answered` and `correct`; the hint count existed lifetime-wide but
never per topic, and inferring one from the other would have been a
different set of answers drawn as though it were this one. The column
was already on attempt_answers, so it is a group-by, and the bar is
three-tone as the reference has it.

And the objective is asked for. It decides which questions exist, how
relevance is weighted, and what readiness measures against — and it was
possible to sit a whole board paper without ever being asked, because no
objective quietly means the entire bank. That is a reasonable default and
a poor thing to arrive at by accident. Five of six accounts here had
never set one.

It can be declined: "everything" is a real answer, and trapping somebody
behind a modal because a list failed to load would be worse than the gap
it closes. Declining is still a choice made, which is the point.

Also in this commit, from the exam-player work: Show answer in study mode
that reveals without recording an answer, review keyed on the attempt
being closed rather than every question being answered — a block that
timed out with nothing answered is over too — and the exam top and bottom
bars. That work found something worth knowing: the exam player is *served*
questions with no correct answer and no explanation, so review cannot
un-hide what it never had, and the player refetches the marked version
once the attempt closes. Nothing is revealed while a block is running.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
2026-09-12 06:24:23 +02:00
Daniel
b80e188eae feat: one session's topics, asked the same three ways — and a way back
The session analysis ranked its weakest topics by primary category only,
while the Analysis page asked the same question three ways and rolled
answers up the category tree. Two sets of rules for "where does this
question belong" is two pages that can disagree about a learner and
neither able to explain why.

So the rules moved to services/knowledge_groups.py: ancestor roll-up,
article reached through its category, organ system reached through the
symptom keyword. study_tools now asks that service instead of building
the lookups inline, and GET /attempts/{id}/recommendations gives one
session the same Articles / Disciplines / Systems switch. Grouping is its
own call, so changing it does not re-read the question table and the peer
statistics beside it. A running exam ranks nothing — marking it there
would answer the question the exam is asking.

The ungrouped `recommendations` key is gone from the analysis payload
along with the code that built it.

And the document page had no way back. It is reached from the Tools
workbench, which by design has no menu of its own, so leaving it meant
the browser button. It opens onto Tools now, as Tools opens onto
Settings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
2026-09-12 03:20:03 +02:00
Daniel
8a1b518502 feat: the two readiness cards
Your score is the share of questions right at your most recent answer to
each. It is deliberately not called an equated score: AMBOSS's EPC rests
on psychometrics we do not have, and a number dressed up as one would be
a claim we cannot support. The card says so.

Against everyone else compares you with other learners on the questions
you have in common — not with their scores on whatever they happened to
sit. A percentile over different question sets reads someone who worked
through the hardest fifty in the bank as weaker than someone who did
fifty easy ones, which is the opposite of true.

Neither appears before it means anything, and each says which half is
missing: more questions of your own, more questions shared with others,
or more learners. The cohort reported is the most any one shared question
saw — distinct learners cannot be summed across questions without
counting the same person once per question.

The "readiness is still locked" note sat above the tab switch and so
appeared on Performance, where it described a table that is on the other
tab. Moved down to the table it is about.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
2026-09-12 03:08:09 +02:00
Daniel
68d65ac782 feat: performance over time, locked until it means something
`GET /study-tools/performance-over-time` returns a point per completed
session with two figures: that session's percentage, and the running
score across everything answered up to that day. The chart draws the
running line and marks the sessions along it — a single session of twelve
questions swings too far to say anything about whether a learner is
improving.

It stays shut below 40 answers or 3 sessions and says which of the two it
is waiting for, rather than drawing a line through two points and letting
the shape suggest a trend that is not there.

LineChart was in the tree unused, with a hardcoded slate palette that
vanishes on a dark page. Rewritten against the theme tokens.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
2026-09-12 01:57:46 +02:00
Daniel
789cd1cc81 feat: right after a tip is its own slice
Opening a tip before answering is a nudge. The answer that follows is
still right — it is counted as right, and the percentage is not docked —
but it is not the same as right, so it keeps its own arc on the donut and
its own line in the legend: "3 correct after a tip".

attempt_answers.used_hint records it. The player reports which questions
had a tip opened before the answer went in; a tip read afterwards is
revision and does not count, which is the difference two of the tests
turn on. Both endings agree about it — an explicit submit carries the
list, and an exam that runs out takes it from the saved progress, so a
tab closing cannot launder a score.

Found while wiring this: RichText declared its component overrides inline
in the render, so every one was a fresh component type and React
remounted the whole rendered tree on each render. An open tip closed
itself every time the exam clock ticked. The map is memoised now.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
2026-09-12 01:51:03 +02:00
Daniel
8c28cc4e9b feat: all attempts vs latest attempt, with the donut shared
A question got wrong in March and right in September is 50% by one count
and 100% by another, and both are true. The Performance tab now says
which it is answering: All attempts is every answer ever given — how much
work has been done — and Latest attempt keeps only the most recent answer
to each question — what is known now.

GET /study-tools/answer-split returns both splits plus the session and
unique-question counts, under the same exclusions as everything else that
measures: no repetitions, no course quizzes, no expired attempts. A blank
is its own slice, never folded into incorrect.

The ring itself moves out of AnalysisSessionPage into components/Donut so
the session view and the lifetime view cannot drift apart. Its legend
gains .is-answered, which the session page had been asking for without
anything defining it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
2026-09-12 01:39:17 +02:00
Daniel
4e272e6ef0 feat: completion over a chosen time range, and a way back out of Tools
"How am I doing" and "how was I doing last month" are different questions,
and a single lifetime figure cannot answer both. Analysis now carries a
Completion panel on the Performance tab: questions answered against the
bank, how many were right, time per question, total time — over 7 days,
30 days, 3 months, or everything.

GET /study-tools/completion?days=N does the counting. It leaves out what
would not be a measurement: repetitions (you already know that answer),
course quizzes (they belong to their course), and expired attempts. A
question left blank is not a wrong answer, so the percentage is out of
what was answered, not out of what was set. Nothing answered reports
nothing rather than 0%.

The Tools workbench has no menu of its own by design, which left no way
back; it now opens onto Settings where it was reached from.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
2026-09-12 01:35:14 +02:00
Daniel
2d828c3c03 feat: a repetition does not raise your score; no more deleting a session
Sitting the same questions again is practice, not a new measurement. You
have already seen the answers, so getting them right the second time
says nothing about whether you knew them — and it cannot be allowed to
raise a figure that means "how much of this do you know". A repeated
session is titled "(repetition)", analysed in full on its own page, and
left out of every aggregate: the overall accuracy, the per-quiz history,
the averages, and the readiness that drives recommendations.

Deleting a single session is gone — control, endpoint, tests and all. A
session is a record of work done, and removing one edits the history
every figure on the analysis is computed from, which turns a measurement
into a number somebody chose. Starting again is still offered whole,
under Settings, Your data, which takes everything rather than the parts
that flatter.

Two layout bugs behind that. The category tree kept its appearance in
QuestionBankPage.css, so it looked right on the bank and took whatever
the host page did to a label everywhere else — in the question editor
that centred the name, leaving it adrift with the count at the far
right; it owns its own stylesheet now. And the editor's grid collapsed
to `1fr` below 900px, whose automatic minimum lets one unshrinkable
child push the column past the window: the page had padding down its
left and none down its right because the right was off the screen.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
2026-09-12 01:08:42 +02:00
Daniel
1b996b0a3d feat: relevance is the board's published share, not our bank's proportions
The knowledge profile ranked topics by how much of *our* bank sat under
each one, which is a fact about us rather than about the exam. It made
cardiology and rheumatology equally worth an evening whenever we happened
to hold the same number of each. The ABP publishes that one is 5% of the
paper and the other 2%, and exam_blueprints.weight has held that since
the blueprint landed.

A domain's weight is divided among the topics beneath it in proportion
to the material each holds, so the topics under a domain add up to its
published share. 672 of our categories now carry one. A topic the
outline does not cover keeps the bank-share figure rather than reporting
nothing — and the row says which it is, because the two numbers mean
different things and should not be read as the same one.

Session analysis is a link to the last session rather than a third tab
with nothing behind it — a session's analysis is a session, and the rail
beside this page is the list of them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
2026-09-12 00:53:54 +02:00
Daniel
b06da68f6b feat: knowledge profile grouped by Articles, Systems or Disciplines
The same answers asked three ways, as AMBOSS does it: which reading to
go back to, which organ system is weak, which discipline is weak. It was
Systems/Subtopics, where "Systems" meant top-level categories — which
are disciplines, not systems — and "Subtopics" meant every category
below them.

  * Articles (the default): rows are the published article behind a
    category, so the row links straight to the reading.
  * Systems: the 16 organ systems. No question is tagged with a system
    directly — it carries a symptom keyword filed under one — so
    membership rolls up through the keyword's parent.
  * Disciplines: top-level categories, which is what the old "systems"
    grouping actually was.

Only 1,502 of 2,948 questions carry a system tag, so the Systems tab
says so rather than showing half the bank as if it were the whole of it,
and relevance there is measured against what the grouping can see.

"Practise this topic" now practises the row you are looking at, on its
own axis. That needed system_ids on the builder — matched as "any tag
beneath this system", where the existing tag_ids is "every one of these
tags", so the two cannot be conflated.

Backend 228/228, frontend 266/266.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
2026-09-11 12:33:05 +02:00
Daniel
ac92793600 feat: study-plan blocks as modules, sessions that know their block
From the three recordings and the AMBOSS screenshots.

Study plans
- Blocks of about 40, split evenly: 202 questions is six blocks of
  33-34, not five of 50 and one of 2. Reseeded (no progress or reading
  existed yet); the seeder now splits the same way.
- A block has its own page, laid out as a course module: the plan's
  blocks down the left, this block's reading then its session in the
  middle, back / previous / next along the bottom. Study or exam mode
  is chosen there, before the session exists; afterwards the mode is
  shown, not offered. The plan page is the table of contents and links
  into blocks rather than starting anything.
- Progress on a block comes from the same /quizzes/sessions row the
  Sessions page shows, so the two cannot disagree.

Sessions <-> plans
- A session started from a block carries its place in the plan: the
  session list and the analysis both return `plan` (plan, block,
  position, previous and next block). The analysis shows a strip with
  the way back to the block and on to the next one.
- Submitting a session marks its block complete. Nothing ever set
  completed_at before — every block read as unfinished forever.

Recommendations
- Framed by the learner's chosen study objective: answers and bank
  material linked to a different exam are left out, and the page is
  titled for the exam. Unlinked material stays in, as elsewhere.

Backend 216/216, frontend 257/258 (the one failure is
ArticleSplitView, which is timing-flaky under the full run and is
unrelated to this change; being checked separately).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
2026-09-11 04:31:21 +02:00
Daniel
5add9f23dd fix: a skipped question is not a wrong answer
Performance by category counted every row in an attempt, and an attempt holds a
row for each question including the ones never answered. A 360-question sitting
that was opened and abandoned therefore landed as 360 wrong answers, which is
why Emergency Medicine read 0% of 400 and Gastroenterology 1.1% of 277 — figures
that describe a sitting nobody worked through, not a learner who cannot do
emergency medicine.

Accuracy now counts only questions that were actually answered, and the note
under the heading says so. Coverage is a separate question from accuracy and
conflating them made both useless.

Also: the category performance block is gone from the dashboard, where it
duplicated the one on Analysis; and the nav says Sessions rather than Quizzes,
with History beside it — "quiz" describes the packaging, a learner sits a
session, and the two entries answer different questions: what can I sit, and
what have I sat.

208 backend, 249 frontend green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
2026-09-11 03:21:13 +02:00
Daniel
16dc431066 feat: analysis recommendations, category grants, compact question header
Analysis / recommendations (AMBOSS parity, verified on next.amboss.com):
- GET /study-tools/recommendations ranks focus areas by the study time most
  likely to raise the score. Readiness is the learner's accuracy in a category
  shrunk toward their own overall accuracy in proportion to sample size, so two
  unlucky answers do not read as a knowledge gap; it unlocks after 40 answers.
  Relevance is the share of the bank a category holds. Counts roll up through
  the category tree, so a system inherits its children's questions.
  It is deliberately not called EPC and does not claim to predict an exam.
- New /analysis page: Performance and Recommendations tabs, readiness summary,
  adaptive-session box, and expandable focus rows showing questions seen,
  answered correctly, the linked article and a per-topic practice action.

Per-category educator grants:
- category_grants table (migration p8b9c0d1e253) plus utils/category_grants.py
  resolving a grant to the category and all of its descendants.
- Question create, edit, delete, bulk and the manager summary now accept a
  moderator OR an educator granted the affected categories, and refuse moves
  that would push a question out of the holder's scope. Summary counts are
  scoped to the grant.
- Moderator endpoints to list, add and revoke grants, plus /my-grants driving
  the nav link and the manager's scope banner; grantable-users avoids handing
  moderators the admin-only user list.
- GrantsPanel in the question manager: grant, list and revoke with inline
  confirmation.

Question page:
- The category trail was a fixed 78px band that wrapped into several rows and
  pushed the stem down the page, followed by three more stacked strips. It is
  now one scrollable meta line (breadcrumb + difficulty + type) and a single
  AMBOSS-style action bar (Mark / Listen / Listen through / Clear) between the
  stem and the options. Difficulty is exposed on the runner payload.

Deploy fix: index.html shipped with no cache header, so browsers kept serving
the previous bundle references and a release looked like nothing had changed.
nginx now sends no-cache for HTML and immutable long-cache for hashed assets.

Tests: 16 new backend (recommendation shrinkage, roll-up, locking, grant scope
across create/edit/delete/bulk/summary, moderator gate) and 10 new frontend.
Full suites green: 88 backend, 116 frontend, build clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014yhHB8Pc7oQqyqn2Vo9DXA
2026-09-09 18:47:32 +02:00
Daniel
236547e646 feat: sample smart-links quiz, bank feedback polish, performance hierarchy
Sample quiz demonstrates option explanations, key points with article links and linked cards. Bank study modal shows per-option explanations and key points. Performance shows main categories with an expandable hierarchy. AMBOSS-style picker polish (chevrons, search box, switch, auto title). Mobile spacing fixes. 97 frontend tests pass.
2026-09-09 03:31:54 +02:00
Daniel
ff1aee6fad fix: resolve latest review findings
Category quiz creation counts extra-linked questions; primary category is validated on edit; statistics dedupe collapses case variants. 60 backend tests pass.
2026-09-08 19:53:25 +02:00
Daniel
a1e9340004 fix: auto-start quizzes in their mode, drop quiz code and verbose stats note
Timed quizzes start as exams and learning quizzes as study without a second mode prompt; reopening resumes automatically. Removed quiz code display from in-progress list and the verbose statistics basis sentence. Lab rows keep logical age order per test. 93 frontend tests pass.
2026-09-08 19:07:27 +02:00
Daniel
f5a084e49a feat: Orthobullets-style lab panel with source deep links and card links
Lab references deep-link to article sections or external sources, show linked cards with study links, and educators can attach cards and article targets. Grouped panel layout. Migration i2d3e4f5a607. 57 backend and 93 frontend tests pass.
2026-09-08 16:20:55 +02:00
Daniel
01337c8c25 feat: performance by category dashboard
Accuracy per category from completed non-expired general-bank answers, counting each question in its primary and additional categories. 56 backend and 90 frontend tests pass.
2026-09-08 16:03:45 +02:00
Daniel
cdb1ab7468 feat: multi-subcategory questions, statistics hardening, seeded lab values
Questions can belong to additional subcategories (junction table, counts, builder/bank filters, edit UI chips); response statistics dedupe duplicate options, match case-insensitively and exclude obsolete answers; AI question classification is command-line only (UI trigger removed); lab reference seed script with cited public pediatric ranges. Migration h1c2d3e4f506. 55 backend and 88 frontend tests pass.
2026-09-08 14:43:22 +02:00
Daniel
a3a6ef7995 feat: redesign quiz runner and add study tools
Add Orthobullets-inspired numbered-answer UI, explicit study response confirmation, response statistics, review navigation, safe calculator, keyboard controls and sourced educator lab references. Persist attempt mode to prevent query-flag exam disclosure. Combined deployed-image backend suite (22), frontend suite (48), build and synthetic desktop/mobile browser checks pass. PostgreSQL round-trip and independent review remain release gates; no production deployment.
2026-09-07 03:10:23 +02:00