pdf-quiz-generator/CLAUDE.md
Daniel 1f2770257f
Some checks failed
Tests / backend (push) Failing after 9s
Tests / frontend (push) Successful in 29s
Tests / e2e (push) Failing after 27s
docs: a Help page for the people who use the site, a Handbook for whoever runs it
They were one page, and it served neither. Somebody wanting to know what
"Studying for" does had to scroll past what the tutor is prompted with
and what re-embedding breaks.

/help — every signed-in account. Signing in through PedsHub SSO, what
"Studying for" scopes, Qbank against Sessions against Collections, study
and exam mode, where performance comes from, Reading and Cards and study
plans, "Make a deck", and how to report a bad question. An educator also
gets a section of their own: what a moderator has, what a grant gives
and what it does not, that nothing in the bank belongs to anybody, how a
draft becomes an article, how a plan is built, and what a moderator
cannot do. A learner never sees that half.

/handbook — administrators only now, with the same FAQ plus the rest.
The Settings card that pointed at it is admin-only to match, and
RequireAuth learned an admin door, which it did not have.

Help is in the account menu rather than the footer: a question you have
while working is answered from where you are.

And the repo docs describe the site that exists. CLAUDE.md lost the LMS
section — courses, modules, lessons, enrolments, all removed months ago
and still documented — and gained the permission model and the sign-in
flow. ADMIN.md's role table said moderators create courses; it now says
what the three roles actually reach, where roles come from, and that the
bank has no owners.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
2026-09-13 14:12:47 +02:00

16 KiB
Raw Permalink Blame History

PedsHub — AI Synopsis for Codebase Work

What this is

PedsHub is a pediatric question bank and reading library. Educators upload PREP exam PDFs, AI proposes MCQs and flashcards from them, a person reviews what is worth keeping, and learners sit it as sessions with an AI tutor beside them. Sign-in is through Authentik at sso.pedshub.com; the companion scribe app is at app.pedshub.com. Everything runs in Docker.

The LMS — courses, modules, lessons, enrolments — was removed. Anything in this file or the code that still mentions a course is stale; say so rather than building against it.

Critical rules

  • Never restart services while a Celery task is running — check docker compose logs celery --tail=5 first
  • Backend and Celery share the same code but build separate images — after changing backend code, you must docker compose build --no-cache backend celery then docker compose up -d backend celery --force-recreate
  • Frontend is a Vite build inside Docker — source changes require docker compose build frontend then docker compose up -d frontend. Vite minifies function names, so grep FunctionName on the built JS won't work.
  • Never import inside a function body if the same name exists at module level — Python treats it as a local variable for the entire function scope, causing UnboundLocalError before the import line executes. This was a real bug with sa_text.
  • docker compose restart does NOT pick up code changes — it reuses the old image. Always build then up -d --force-recreate.
  • Pydantic schemas must match DB nullability — if a column allows NULL, the schema field must be type | None. A mismatch causes 500 on serialization.
  • The Question model uses source_quiz_id as the Python attribute but quiz_id as the DB column — use Question.source_quiz_id in SQLAlchemy filters, never Question.quiz_id.

Stack

  • Backend: FastAPI + SQLAlchemy + PostgreSQL 16 (pgvector) + Redis + Celery
  • Frontend: React 18 + Vite + React Router 6 + plain CSS + Milkdown (markdown WYSIWYG) + Nginx
  • AI: LiteLLM proxy routes to Claude/GPT/Gemini/Bedrock. _proxy_model() in ai_service.py adds openai/ prefix for the proxy.
  • Vectors: ChromaDB for document page chunks (RAG), pgvector for question embeddings (semantic search)
  • Config: Backend reads .env via pydantic-settings. Frontend uses runtime window.__APP_CONFIG__ injected by docker-entrypoint.sh (not Vite build-time env).

Architecture

Browser → Nginx (frontend) → FastAPI (4 uvicorn workers)
                                ├── PostgreSQL (users, quizzes, questions, flashcards, attempts + pgvector embeddings)
                                ├── ChromaDB (document page chunks for extraction context)
                                ├── Redis (Celery broker, rate limits, settings, job progress, session locks)
                                └── Celery (2 fork workers: PDF processing, quiz extraction, flashcard generation, classification, embedding regeneration)
                                └── db-backup (daily PostgreSQL dumps, 14 daily / 4 weekly / 6 monthly retention, ./backups/)

## Who may do what

Three roles, and one kind of grant.

- **admin** — everything, including Settings: models, people, site policy.
- **moderator** — the whole bank: every question, article, category, card
  deck, image library, and the editorial queue. Nothing that configures the
  site.
- **user** — their own sittings, notes, starred questions and folders. No
  bank content at all unless a grant says otherwise.

Roles come from the identity provider's groups, never from this app. With
`OIDC_ROLE_CLAIM` set, `services/sso_roles.apply()` runs on every SSO sign-in
and brings the role into line with the person's groups — so removal at the
provider takes the role away here too. The two in-app role endpoints answer 409
while that mapping is on, because a role set here would be silently overwritten
at the next sign-in. The one demotion the sync refuses is the last
administrator.

`CategoryGrant` is the per-user half, given inside this app by a moderator or
admin (`/access`). It names a category branch, an image library or a folder,
and makes the holder an editor of what is inside it — which also makes the
Questions and Images menu entries appear. It never comes from a claim and is
never self-assignable.

**Nothing in the bank has an owner.** Questions, articles, categories, decks,
documents, media and shared quizzes all carry `user_id = NULL`, and every
creation path writes NULL. Authorship confers no rights anywhere: `may_edit_question`
and `can_edit_article` ask the role and the grants and nothing else. What keeps
an owner is what is genuinely one person's — attempts, notes, favourites,
collections, folders, study-plan progress, and the unshared quizzes that are
somebody's own sittings. See migration `q6a7b8c9d0e1`.

## Signing in

SSO only, through Authentik at `sso.pedshub.com`. There is no sign-up form, no
invite codes and no email sign-in codes — all three were this app doing the
provider's job, and all three were removed (migration `r7b8c9d0e1f2` drops
their tables).

- `GET /auth/sso/login` → 302 to the provider. It must stay `async` and the
  redirect must be awaited; authlib's Starlette client returns a coroutine, and
  returning it unawaited is a 500 on every click.
- `GET /auth/sso/callback` matches on the email claim, refuses an explicit
  `email_verified: false`, applies the role from groups, then parks the access
  token in Redis under a one-time code (`sso:exchange:<code>`, 60s) and
  redirects with `?code=`. The token never travels in a URL — nginx logs the
  request line, so it would be written to disk on every sign-in.
- `POST /auth/sso/exchange` spends the code once (GETDEL) and returns the token.
- `settings:sso_only` in Redis closes every password door: login, register,
  forgot/reset password, resend verification, and setting a password through
  `PUT /auth/me`. `GET /auth/signup-policy` reports it so the forms decline to
  draw themselves rather than being refused after the fact.

## Key directories

backend/app/ main.py — App startup, DDL migrations (setup_pgvector), router mounting, singleton lock config.py — All settings from .env models/ — SQLAlchemy ORM (user, quiz, question, flashcard, attempt, section, pdf_document, ...) schemas/ — Pydantic request/response models routers/ — API endpoints (auth, quizzes, questions, flashcards, attempts, admin, teach, tts, tags, ...) services/ ai_service.py — LLM calls, _proxy_model(), get_model_for_task() fallback chain extraction_modes.py — 6 quiz extraction modes + flashcard generation prompt search_service.py — Hybrid retrieval (BM25 + pgvector, RRF), rerank_ids() rerank_service.py — Cross-encoder rerank via the proxy's /v1/rerank vector_service.py — ChromaDB: store/query page chunks, LiteLLMEmbeddingFunction embedding_service.py — pgvector: embed questions for semantic search pdf_service.py — PyMuPDF: text extraction, image extraction with MD5 hash skip list tasks/ quiz_tasks.py — Celery: extract_quiz, classify_questions, regenerate_embeddings, generate_flashcard_deck pdf_tasks.py — Celery: process_pdf (text extraction + vectorization)

frontend/src/ App.jsx — Routes (public, authenticated, moderator-only) context/AuthContext.jsx — Login/logout/register, JWT token management pages/ DocumentDetailPage.jsx — Section management, "Extract Quiz" / "Create Flashcards" buttons, job progress QuestionBankPage.jsx — Browse questions, multi-category + tag filtering, TagBrowser component FlashcardsPage.jsx — Browse decks + card browser with search FlashcardStudyPage.jsx — Flip cards, got-it/review, keyboard nav, progress QuizPage.jsx — Take quiz (exam/study mode), timer, progress save to Redis AdminPage.jsx — Model config, user management, settings components/ Navbar.jsx — Auth-aware nav with jobs badge TeachChat.jsx — AI tutor drawer (lazy loaded, markdown/GFM tables) RichEditor.jsx — Milkdown markdown WYSIWYG editor with toolbar, GFM tables, math, history


## Database tables (key ones)
| Table | Purpose | Key FKs |
|-------|---------|---------|
| users | Accounts with role (admin/moderator/user) | — |
| pdf_documents | Uploaded PDFs | user_id → users |
| sections | Page ranges within a document | document_id → pdf_documents |
| quizzes | A sitting or a shared bank test. `is_shared=1` is the bank's and ownerless; `is_shared=0` is somebody's own session | section_id → sections (nullable), user_id → users (nullable) |
| questions | MCQ questions with pgvector embedding. `user_id` is always NULL — the bank has no owners | source_quiz_id → quizzes (nullable) |
| quiz_question_links | Quiz ↔ Question many-to-many | quiz_id, question_id |
| flashcard_decks | Flashcard deck metadata | section_id → sections, user_id → users |
| flashcards | Individual cards (front/back) | deck_id → flashcard_decks |
| question_tags | Tag definitions (subject/disease/keyword) | — |
| question_tag_links | Question ↔ Tag | question_id, tag_id |
| flashcard_tag_links | Flashcard ↔ Tag | flashcard_id, tag_id |
| quiz_attempts | User quiz sessions with score | quiz_id, user_id |

## Common patterns
- **Tag filtering SQL**: `WHERE tag_id = ANY(:tag_ids) GROUP BY ... HAVING COUNT(DISTINCT tag_id) = :cnt` — AND logic across tags
- **Multi-category filtering**: `category_ids` param (comma-separated), uses `.in_()` — OR logic within categories
- **Job progress**: Celery tasks push steps to Redis lists (`extraction:steps:{job_id}`), frontend polls `GET /quizzes/job/{job_id}`
- **Model selection**: Admin configures models per task (extraction, teach, article, tts, stt, keyword, flashcard, tool). `get_model_for_task(db, task)` returns (model_id, api_key) with fallback to `settings.LITELLM_MODEL`; `get_configured_model(db, task)` returns None instead of falling back, for work that must not run on a model nobody chose.
- **Vision fallback**: `vision_service.image_context(db, images, model_id=...)` returns the message parts to splice into a user message. If the job's model is vision-capable it gets the images; if not, the `tool` model describes them and the primary reads the description. Capability comes from the proxy's `/model/info` `supports_vision` (cached in-process), and where that is absent from a one-off 8px probe (cached in Redis) — never from a hard-coded list of model names. No tool model configured raises `VisionUnavailable`.
- **Reranking**: `search_service.rerank_ids(db, q, kind, ids)` reorders the first 50 of a fused ranking with a cross-encoder (`LITELLM_RERANK_MODEL`, default `cohere-rerank-v4.0-pro`, Redis override `settings:rerank_model`). It is a *permutation* — an unset, unreachable or malformed reranker returns the ids untouched, never fewer. Applied to question and article search, the test builder's description path, and the AI Mode shortlist; deliberately not to the typeahead, flashcard/media search, or the AI Mode closeness thresholds. Cached in Redis for a day, keyed on model + query + document text. See docs/reranking.md.
- **Uploads**: `file_intake.read()` then `kind_of()` then `text_from()` — size, then type sniffed from the leading bytes, then text. PDF, DOCX and images only (2 MB); an extension is a claim, never the decision. Images are read by the `tool` model, and say so when none is configured.
- **Small talk in AI Mode**: `ai_mode_service.is_small_talk()` runs *before* retrieval. A greeting scores 0.460.51 against a clinical corpus, which is either side of the adjacency threshold, so the closeness gate cannot be the guard. Mode "chat" means no shortlist, no citations, and no thread name.
- **`==key points==`**: a remark plugin (`utils/keyPoints.js`) → `<mark class="key-point">`. Never rewrite the markdown string before parsing: highlights and the read-aloud cursor are offsets into the raw text.
- **useEffect dependencies**: Use `.join(',')` on arrays to create a stable string key (e.g., `tagIdsKey`, `catIdsKey`)
- **Admin page data refresh**: `loadData(false)` — the `false` param skips the loading spinner on re-fetch after actions

## Database migrations (Alembic)

**What's a migration?** A migration is a small, ordered change to the database schema — adding a column, renaming a table, changing a type. Each change lives in a Python file under `backend/alembic/versions/`. Alembic tracks which ones have been applied in an `alembic_version` table inside Postgres, so it knows what's new next time you run it.

**Why it exists here:** until now, schema was created via `Base.metadata.create_all()` in `main.py:478`, which only creates *missing tables* — it never alters existing ones. Every column change required manual `ALTER TABLE`. Alembic makes schema changes versioned, reversible, and reproducible across environments.

**Current setup**
- `alembic.ini` contains no hardcoded URL; `alembic/env.py` injects `DATABASE_URL` from the container's env.
- Live DB is stamped at revision `9bac7bf02e38` (the latest in `alembic/versions/`).
- `Base.metadata.create_all()` remains in place as a fallback for fresh deploys — **don't remove it** without first generating a complete baseline migration from the live schema.

**Developer workflow**
```bash
# where am I?
docker compose exec backend alembic current
docker compose exec backend alembic heads

# create a new migration (auto-diff model vs live DB)
docker compose exec backend alembic revision --autogenerate -m "add some column"
# ↑ review the generated file under backend/alembic/versions/ before applying

# apply pending migrations
docker compose exec backend alembic upgrade head

# roll back the last one
docker compose exec backend alembic downgrade -1

When to write one — any schema change: new column, dropped column, renamed field, new table, altered index, new FK. Model edit → migration → apply → commit both together.

Gotchas

  • Migrations run as a normal transaction. A failed migration rolls back cleanly.
  • --autogenerate doesn't catch: server_default changes, CHECK constraints, enum value additions, data migrations. Hand-edit the file when needed.
  • After applying a new migration in dev, rebuild the backend image (docker compose build backend celery) so it ships with the migration file baked in.
  • The alembic_version table should only ever have one row. If you see multiple, you have branched heads — run alembic merge to reconcile.

What NOT to do

  • Don't add from sqlalchemy import text as X inside functions — import at module top only
  • Don't use Question.quiz_id — it's Question.source_quiz_id
  • Don't set Content-Type: multipart/form-data manually on axios FormData uploads — axios handles it
  • Don't use [someValue === null] as a useEffect dependency — it evaluates to a constant boolean
  • Don't docker compose restart expecting code changes to apply — must rebuild
  • Don't use window.confirm() — user hates browser popups, use inline confirmation or the Dialog component
  • Don't use MDXEditor — it's an MDX parser that chokes on < and { in medical content. Milkdown (CommonMark) is used instead.
  • Don't put documents listing on the dashboard — it's in Settings page under Nextcloud.
  • Don't gate anything on x.user_id == current_user.id for bank content. Authorship confers nothing; ask the role or a grant.
  • Don't stamp a creator on new bank content — write user_id=None.
  • Don't reintroduce invite codes or email sign-in codes. The identity provider does that.
  • Don't put a token in a URL. nginx logs the request line.