They were one page, and it served neither. Somebody wanting to know what "Studying for" does had to scroll past what the tutor is prompted with and what re-embedding breaks. /help — every signed-in account. Signing in through PedsHub SSO, what "Studying for" scopes, Qbank against Sessions against Collections, study and exam mode, where performance comes from, Reading and Cards and study plans, "Make a deck", and how to report a bad question. An educator also gets a section of their own: what a moderator has, what a grant gives and what it does not, that nothing in the bank belongs to anybody, how a draft becomes an article, how a plan is built, and what a moderator cannot do. A learner never sees that half. /handbook — administrators only now, with the same FAQ plus the rest. The Settings card that pointed at it is admin-only to match, and RequireAuth learned an admin door, which it did not have. Help is in the account menu rather than the footer: a question you have while working is answered from where you are. And the repo docs describe the site that exists. CLAUDE.md lost the LMS section — courses, modules, lessons, enrolments, all removed months ago and still documented — and gained the permission model and the sign-in flow. ADMIN.md's role table said moderators create courses; it now says what the three roles actually reach, where roles come from, and that the bank has no owners. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TqXevQJhxFrM7jJg82cgZN
16 KiB
PedsHub — AI Synopsis for Codebase Work
What this is
PedsHub is a pediatric question bank and reading library. Educators upload PREP
exam PDFs, AI proposes MCQs and flashcards from them, a person reviews what is
worth keeping, and learners sit it as sessions with an AI tutor beside them.
Sign-in is through Authentik at sso.pedshub.com; the companion scribe app is
at app.pedshub.com. Everything runs in Docker.
The LMS — courses, modules, lessons, enrolments — was removed. Anything in this file or the code that still mentions a course is stale; say so rather than building against it.
Critical rules
- Never restart services while a Celery task is running — check
docker compose logs celery --tail=5first - Backend and Celery share the same code but build separate images — after changing backend code, you must
docker compose build --no-cache backend celerythendocker compose up -d backend celery --force-recreate - Frontend is a Vite build inside Docker — source changes require
docker compose build frontendthendocker compose up -d frontend. Vite minifies function names, sogrep FunctionNameon the built JS won't work. - Never import inside a function body if the same name exists at module level — Python treats it as a local variable for the entire function scope, causing
UnboundLocalErrorbefore the import line executes. This was a real bug withsa_text. docker compose restartdoes NOT pick up code changes — it reuses the old image. Alwaysbuildthenup -d --force-recreate.- Pydantic schemas must match DB nullability — if a column allows NULL, the schema field must be
type | None. A mismatch causes 500 on serialization. - The Question model uses
source_quiz_idas the Python attribute butquiz_idas the DB column — useQuestion.source_quiz_idin SQLAlchemy filters, neverQuestion.quiz_id.
Stack
- Backend: FastAPI + SQLAlchemy + PostgreSQL 16 (pgvector) + Redis + Celery
- Frontend: React 18 + Vite + React Router 6 + plain CSS + Milkdown (markdown WYSIWYG) + Nginx
- AI: LiteLLM proxy routes to Claude/GPT/Gemini/Bedrock.
_proxy_model()in ai_service.py addsopenai/prefix for the proxy. - Vectors: ChromaDB for document page chunks (RAG), pgvector for question embeddings (semantic search)
- Config: Backend reads
.envvia pydantic-settings. Frontend uses runtimewindow.__APP_CONFIG__injected bydocker-entrypoint.sh(not Vite build-time env).
Architecture
Browser → Nginx (frontend) → FastAPI (4 uvicorn workers)
├── PostgreSQL (users, quizzes, questions, flashcards, attempts + pgvector embeddings)
├── ChromaDB (document page chunks for extraction context)
├── Redis (Celery broker, rate limits, settings, job progress, session locks)
└── Celery (2 fork workers: PDF processing, quiz extraction, flashcard generation, classification, embedding regeneration)
└── db-backup (daily PostgreSQL dumps, 14 daily / 4 weekly / 6 monthly retention, ./backups/)
## Who may do what
Three roles, and one kind of grant.
- **admin** — everything, including Settings: models, people, site policy.
- **moderator** — the whole bank: every question, article, category, card
deck, image library, and the editorial queue. Nothing that configures the
site.
- **user** — their own sittings, notes, starred questions and folders. No
bank content at all unless a grant says otherwise.
Roles come from the identity provider's groups, never from this app. With
`OIDC_ROLE_CLAIM` set, `services/sso_roles.apply()` runs on every SSO sign-in
and brings the role into line with the person's groups — so removal at the
provider takes the role away here too. The two in-app role endpoints answer 409
while that mapping is on, because a role set here would be silently overwritten
at the next sign-in. The one demotion the sync refuses is the last
administrator.
`CategoryGrant` is the per-user half, given inside this app by a moderator or
admin (`/access`). It names a category branch, an image library or a folder,
and makes the holder an editor of what is inside it — which also makes the
Questions and Images menu entries appear. It never comes from a claim and is
never self-assignable.
**Nothing in the bank has an owner.** Questions, articles, categories, decks,
documents, media and shared quizzes all carry `user_id = NULL`, and every
creation path writes NULL. Authorship confers no rights anywhere: `may_edit_question`
and `can_edit_article` ask the role and the grants and nothing else. What keeps
an owner is what is genuinely one person's — attempts, notes, favourites,
collections, folders, study-plan progress, and the unshared quizzes that are
somebody's own sittings. See migration `q6a7b8c9d0e1`.
## Signing in
SSO only, through Authentik at `sso.pedshub.com`. There is no sign-up form, no
invite codes and no email sign-in codes — all three were this app doing the
provider's job, and all three were removed (migration `r7b8c9d0e1f2` drops
their tables).
- `GET /auth/sso/login` → 302 to the provider. It must stay `async` and the
redirect must be awaited; authlib's Starlette client returns a coroutine, and
returning it unawaited is a 500 on every click.
- `GET /auth/sso/callback` matches on the email claim, refuses an explicit
`email_verified: false`, applies the role from groups, then parks the access
token in Redis under a one-time code (`sso:exchange:<code>`, 60s) and
redirects with `?code=`. The token never travels in a URL — nginx logs the
request line, so it would be written to disk on every sign-in.
- `POST /auth/sso/exchange` spends the code once (GETDEL) and returns the token.
- `settings:sso_only` in Redis closes every password door: login, register,
forgot/reset password, resend verification, and setting a password through
`PUT /auth/me`. `GET /auth/signup-policy` reports it so the forms decline to
draw themselves rather than being refused after the fact.
## Key directories
backend/app/ main.py — App startup, DDL migrations (setup_pgvector), router mounting, singleton lock config.py — All settings from .env models/ — SQLAlchemy ORM (user, quiz, question, flashcard, attempt, section, pdf_document, ...) schemas/ — Pydantic request/response models routers/ — API endpoints (auth, quizzes, questions, flashcards, attempts, admin, teach, tts, tags, ...) services/ ai_service.py — LLM calls, _proxy_model(), get_model_for_task() fallback chain extraction_modes.py — 6 quiz extraction modes + flashcard generation prompt search_service.py — Hybrid retrieval (BM25 + pgvector, RRF), rerank_ids() rerank_service.py — Cross-encoder rerank via the proxy's /v1/rerank vector_service.py — ChromaDB: store/query page chunks, LiteLLMEmbeddingFunction embedding_service.py — pgvector: embed questions for semantic search pdf_service.py — PyMuPDF: text extraction, image extraction with MD5 hash skip list tasks/ quiz_tasks.py — Celery: extract_quiz, classify_questions, regenerate_embeddings, generate_flashcard_deck pdf_tasks.py — Celery: process_pdf (text extraction + vectorization)
frontend/src/ App.jsx — Routes (public, authenticated, moderator-only) context/AuthContext.jsx — Login/logout/register, JWT token management pages/ DocumentDetailPage.jsx — Section management, "Extract Quiz" / "Create Flashcards" buttons, job progress QuestionBankPage.jsx — Browse questions, multi-category + tag filtering, TagBrowser component FlashcardsPage.jsx — Browse decks + card browser with search FlashcardStudyPage.jsx — Flip cards, got-it/review, keyboard nav, progress QuizPage.jsx — Take quiz (exam/study mode), timer, progress save to Redis AdminPage.jsx — Model config, user management, settings components/ Navbar.jsx — Auth-aware nav with jobs badge TeachChat.jsx — AI tutor drawer (lazy loaded, markdown/GFM tables) RichEditor.jsx — Milkdown markdown WYSIWYG editor with toolbar, GFM tables, math, history
## Database tables (key ones)
| Table | Purpose | Key FKs |
|-------|---------|---------|
| users | Accounts with role (admin/moderator/user) | — |
| pdf_documents | Uploaded PDFs | user_id → users |
| sections | Page ranges within a document | document_id → pdf_documents |
| quizzes | A sitting or a shared bank test. `is_shared=1` is the bank's and ownerless; `is_shared=0` is somebody's own session | section_id → sections (nullable), user_id → users (nullable) |
| questions | MCQ questions with pgvector embedding. `user_id` is always NULL — the bank has no owners | source_quiz_id → quizzes (nullable) |
| quiz_question_links | Quiz ↔ Question many-to-many | quiz_id, question_id |
| flashcard_decks | Flashcard deck metadata | section_id → sections, user_id → users |
| flashcards | Individual cards (front/back) | deck_id → flashcard_decks |
| question_tags | Tag definitions (subject/disease/keyword) | — |
| question_tag_links | Question ↔ Tag | question_id, tag_id |
| flashcard_tag_links | Flashcard ↔ Tag | flashcard_id, tag_id |
| quiz_attempts | User quiz sessions with score | quiz_id, user_id |
## Common patterns
- **Tag filtering SQL**: `WHERE tag_id = ANY(:tag_ids) GROUP BY ... HAVING COUNT(DISTINCT tag_id) = :cnt` — AND logic across tags
- **Multi-category filtering**: `category_ids` param (comma-separated), uses `.in_()` — OR logic within categories
- **Job progress**: Celery tasks push steps to Redis lists (`extraction:steps:{job_id}`), frontend polls `GET /quizzes/job/{job_id}`
- **Model selection**: Admin configures models per task (extraction, teach, article, tts, stt, keyword, flashcard, tool). `get_model_for_task(db, task)` returns (model_id, api_key) with fallback to `settings.LITELLM_MODEL`; `get_configured_model(db, task)` returns None instead of falling back, for work that must not run on a model nobody chose.
- **Vision fallback**: `vision_service.image_context(db, images, model_id=...)` returns the message parts to splice into a user message. If the job's model is vision-capable it gets the images; if not, the `tool` model describes them and the primary reads the description. Capability comes from the proxy's `/model/info` `supports_vision` (cached in-process), and where that is absent from a one-off 8px probe (cached in Redis) — never from a hard-coded list of model names. No tool model configured raises `VisionUnavailable`.
- **Reranking**: `search_service.rerank_ids(db, q, kind, ids)` reorders the first 50 of a fused ranking with a cross-encoder (`LITELLM_RERANK_MODEL`, default `cohere-rerank-v4.0-pro`, Redis override `settings:rerank_model`). It is a *permutation* — an unset, unreachable or malformed reranker returns the ids untouched, never fewer. Applied to question and article search, the test builder's description path, and the AI Mode shortlist; deliberately not to the typeahead, flashcard/media search, or the AI Mode closeness thresholds. Cached in Redis for a day, keyed on model + query + document text. See docs/reranking.md.
- **Uploads**: `file_intake.read()` then `kind_of()` then `text_from()` — size, then type sniffed from the leading bytes, then text. PDF, DOCX and images only (2 MB); an extension is a claim, never the decision. Images are read by the `tool` model, and say so when none is configured.
- **Small talk in AI Mode**: `ai_mode_service.is_small_talk()` runs *before* retrieval. A greeting scores 0.46–0.51 against a clinical corpus, which is either side of the adjacency threshold, so the closeness gate cannot be the guard. Mode "chat" means no shortlist, no citations, and no thread name.
- **`==key points==`**: a remark plugin (`utils/keyPoints.js`) → `<mark class="key-point">`. Never rewrite the markdown string before parsing: highlights and the read-aloud cursor are offsets into the raw text.
- **useEffect dependencies**: Use `.join(',')` on arrays to create a stable string key (e.g., `tagIdsKey`, `catIdsKey`)
- **Admin page data refresh**: `loadData(false)` — the `false` param skips the loading spinner on re-fetch after actions
## Database migrations (Alembic)
**What's a migration?** A migration is a small, ordered change to the database schema — adding a column, renaming a table, changing a type. Each change lives in a Python file under `backend/alembic/versions/`. Alembic tracks which ones have been applied in an `alembic_version` table inside Postgres, so it knows what's new next time you run it.
**Why it exists here:** until now, schema was created via `Base.metadata.create_all()` in `main.py:478`, which only creates *missing tables* — it never alters existing ones. Every column change required manual `ALTER TABLE`. Alembic makes schema changes versioned, reversible, and reproducible across environments.
**Current setup**
- `alembic.ini` contains no hardcoded URL; `alembic/env.py` injects `DATABASE_URL` from the container's env.
- Live DB is stamped at revision `9bac7bf02e38` (the latest in `alembic/versions/`).
- `Base.metadata.create_all()` remains in place as a fallback for fresh deploys — **don't remove it** without first generating a complete baseline migration from the live schema.
**Developer workflow**
```bash
# where am I?
docker compose exec backend alembic current
docker compose exec backend alembic heads
# create a new migration (auto-diff model vs live DB)
docker compose exec backend alembic revision --autogenerate -m "add some column"
# ↑ review the generated file under backend/alembic/versions/ before applying
# apply pending migrations
docker compose exec backend alembic upgrade head
# roll back the last one
docker compose exec backend alembic downgrade -1
When to write one — any schema change: new column, dropped column, renamed field, new table, altered index, new FK. Model edit → migration → apply → commit both together.
Gotchas
- Migrations run as a normal transaction. A failed migration rolls back cleanly.
--autogeneratedoesn't catch: server_default changes, CHECK constraints, enum value additions, data migrations. Hand-edit the file when needed.- After applying a new migration in dev, rebuild the backend image (
docker compose build backend celery) so it ships with the migration file baked in. - The
alembic_versiontable should only ever have one row. If you see multiple, you have branched heads — runalembic mergeto reconcile.
What NOT to do
- Don't add
from sqlalchemy import text as Xinside functions — import at module top only - Don't use
Question.quiz_id— it'sQuestion.source_quiz_id - Don't set
Content-Type: multipart/form-datamanually on axios FormData uploads — axios handles it - Don't use
[someValue === null]as a useEffect dependency — it evaluates to a constant boolean - Don't
docker compose restartexpecting code changes to apply — must rebuild - Don't use
window.confirm()— user hates browser popups, use inline confirmation or the Dialog component - Don't use MDXEditor — it's an MDX parser that chokes on
<and{in medical content. Milkdown (CommonMark) is used instead. - Don't put documents listing on the dashboard — it's in Settings page under Nextcloud.
- Don't gate anything on
x.user_id == current_user.idfor bank content. Authorship confers nothing; ask the role or a grant. - Don't stamp a creator on new bank content — write
user_id=None. - Don't reintroduce invite codes or email sign-in codes. The identity provider does that.
- Don't put a token in a URL. nginx logs the request line.