diff --git a/CLAUDE.md b/CLAUDE.md new file mode 100644 index 0000000..6ebe2a6 --- /dev/null +++ b/CLAUDE.md @@ -0,0 +1,94 @@ +# PedsHub — AI Synopsis for Codebase Work + +## What this is +PedsHub is a pediatric medical learning platform. Admins upload PREP exam PDFs, AI extracts MCQ questions (or generates flashcards), and users study them with an AI tutor. Everything runs in Docker. + +## Critical rules +- **Never restart services while a Celery task is running** — check `docker compose logs celery --tail=5` first +- **Backend and Celery share the same code but build separate images** — after changing backend code, you must `docker compose build --no-cache backend celery` then `docker compose up -d backend celery --force-recreate` +- **Frontend is a Vite build inside Docker** — source changes require `docker compose build frontend` then `docker compose up -d frontend`. Vite minifies function names, so `grep FunctionName` on the built JS won't work. +- **Never import inside a function body if the same name exists at module level** — Python treats it as a local variable for the entire function scope, causing `UnboundLocalError` before the import line executes. This was a real bug with `sa_text`. +- **`docker compose restart` does NOT pick up code changes** — it reuses the old image. Always `build` then `up -d --force-recreate`. +- **Pydantic schemas must match DB nullability** — if a column allows NULL, the schema field must be `type | None`. A mismatch causes 500 on serialization. +- **The Question model uses `source_quiz_id` as the Python attribute but `quiz_id` as the DB column** — use `Question.source_quiz_id` in SQLAlchemy filters, never `Question.quiz_id`. + +## Stack +- **Backend**: FastAPI + SQLAlchemy + PostgreSQL 16 (pgvector) + Redis + Celery +- **Frontend**: React 18 + Vite + React Router 6 + plain CSS + Nginx +- **AI**: LiteLLM proxy routes to Claude/GPT/Gemini/Bedrock. `_proxy_model()` in ai_service.py adds `openai/` prefix for the proxy. +- **Vectors**: ChromaDB for document page chunks (RAG), pgvector for question embeddings (semantic search) +- **Config**: Backend reads `.env` via pydantic-settings. Frontend uses runtime `window.__APP_CONFIG__` injected by `docker-entrypoint.sh` (not Vite build-time env). + +## Architecture +``` +Browser → Nginx (frontend) → FastAPI (4 uvicorn workers) + ├── PostgreSQL (users, quizzes, questions, flashcards, attempts + pgvector embeddings) + ├── ChromaDB (document page chunks for extraction context) + ├── Redis (Celery broker, rate limits, settings, job progress, session locks) + └── Celery (2 fork workers: PDF processing, quiz extraction, flashcard generation, classification, embedding regeneration) +``` + +## Key directories +``` +backend/app/ + main.py — App startup, DDL migrations (setup_pgvector), router mounting, singleton lock + config.py — All settings from .env + models/ — SQLAlchemy ORM (user, quiz, question, flashcard, attempt, section, pdf_document, ...) + schemas/ — Pydantic request/response models + routers/ — API endpoints (auth, quizzes, questions, flashcards, attempts, admin, teach, tts, tags, ...) + services/ + ai_service.py — LLM calls, _proxy_model(), get_model_for_task() fallback chain + extraction_modes.py — 6 quiz extraction modes + flashcard generation prompt + vector_service.py — ChromaDB: store/query page chunks, LiteLLMEmbeddingFunction + embedding_service.py — pgvector: embed questions for semantic search + pdf_service.py — PyMuPDF: text extraction, image extraction with MD5 hash skip list + tasks/ + quiz_tasks.py — Celery: extract_quiz, classify_questions, regenerate_embeddings, generate_flashcard_deck + pdf_tasks.py — Celery: process_pdf (text extraction + vectorization) + +frontend/src/ + App.jsx — Routes (public, authenticated, moderator-only) + context/AuthContext.jsx — Login/logout/register, JWT token management + pages/ + DocumentDetailPage.jsx — Section management, "Extract Quiz" / "Create Flashcards" buttons, job progress + QuestionBankPage.jsx — Browse questions, multi-category + tag filtering, TagBrowser component + FlashcardsPage.jsx — Browse decks + card browser with search + FlashcardStudyPage.jsx — Flip cards, got-it/review, keyboard nav, progress + QuizPage.jsx — Take quiz (exam/study mode), timer, progress save to Redis + AdminPage.jsx — Model config, user management, settings + components/ + Navbar.jsx — Auth-aware nav with jobs badge + TeachChat.jsx — AI tutor drawer (lazy loaded, markdown/GFM tables) +``` + +## Database tables (key ones) +| Table | Purpose | Key FKs | +|-------|---------|---------| +| users | Accounts with role (admin/moderator/user) | — | +| pdf_documents | Uploaded PDFs | user_id → users | +| sections | Page ranges within a document | document_id → pdf_documents | +| quizzes | Quiz metadata | section_id → sections (nullable), user_id → users | +| questions | MCQ questions with pgvector embedding | source_quiz_id → quizzes (nullable) | +| quiz_question_links | Quiz ↔ Question many-to-many | quiz_id, question_id | +| flashcard_decks | Flashcard deck metadata | section_id → sections, user_id → users | +| flashcards | Individual cards (front/back) | deck_id → flashcard_decks | +| question_tags | Tag definitions (subject/disease/keyword) | — | +| question_tag_links | Question ↔ Tag | question_id, tag_id | +| flashcard_tag_links | Flashcard ↔ Tag | flashcard_id, tag_id | +| quiz_attempts | User quiz sessions with score | quiz_id, user_id | + +## Common patterns +- **Tag filtering SQL**: `WHERE tag_id = ANY(:tag_ids) GROUP BY ... HAVING COUNT(DISTINCT tag_id) = :cnt` — AND logic across tags +- **Multi-category filtering**: `category_ids` param (comma-separated), uses `.in_()` — OR logic within categories +- **Job progress**: Celery tasks push steps to Redis lists (`extraction:steps:{job_id}`), frontend polls `GET /quizzes/job/{job_id}` +- **Model selection**: Admin configures models per task (extraction, teach, tts, keyword, flashcard). `get_model_for_task(db, task)` returns (model_id, api_key) with fallback to `settings.LITELLM_MODEL`. +- **useEffect dependencies**: Use `.join(',')` on arrays to create a stable string key (e.g., `tagIdsKey`, `catIdsKey`) +- **Admin page data refresh**: `loadData(false)` — the `false` param skips the loading spinner on re-fetch after actions + +## What NOT to do +- Don't add `from sqlalchemy import text as X` inside functions — import at module top only +- Don't use `Question.quiz_id` — it's `Question.source_quiz_id` +- Don't set `Content-Type: multipart/form-data` manually on axios FormData uploads — axios handles it +- Don't use `[someValue === null]` as a useEffect dependency — it evaluates to a constant boolean +- Don't `docker compose restart` expecting code changes to apply — must rebuild +- Don't use `window.confirm()` — user hates browser popups, use inline confirmation or the Dialog component diff --git a/frontend/src/pages/DocumentDetailPage.jsx b/frontend/src/pages/DocumentDetailPage.jsx index e8ebca4..5a8e72b 100644 --- a/frontend/src/pages/DocumentDetailPage.jsx +++ b/frontend/src/pages/DocumentDetailPage.jsx @@ -4,7 +4,7 @@ import { useAuth } from '../context/AuthContext' import api from '../api/client' import ConfirmButton from '../components/ConfirmButton' -function ExtractionProgress({ jobId, onDone, onClose }) { +function ExtractionProgress({ jobId, onDone, onClose, label = 'Extracting Questions' }) { const [steps, setSteps] = useState([]) const [status, setStatus] = useState('pending') const [quizId, setQuizId] = useState(null) @@ -60,7 +60,7 @@ function ExtractionProgress({ jobId, onDone, onClose }) {

- {status === 'completed' ? '✓ Extraction Complete' : status === 'failed' ? '✗ Extraction Failed' : '🤖 Extracting Questions…'} + {status === 'completed' ? `✓ ${label} Complete` : status === 'failed' ? `✗ ${label} Failed` : `🤖 ${label}…`}

@@ -194,7 +194,7 @@ export default function DocumentDetailPage() { const stored = JSON.parse(localStorage.getItem('pedquiz_jobs') || '[]') stored.unshift({ jobId: res.data.job_id, title, status: 'running', lastStep: 'Starting…', ts: Date.now() }) localStorage.setItem('pedquiz_jobs', JSON.stringify(stored.slice(0, 10))) - setActiveJob({ jobId: res.data.job_id, sectionName }) + setActiveJob({ jobId: res.data.job_id, sectionName, type: 'flashcard' }) } } catch (err) { setError(err.response?.data?.detail || 'Failed to start flashcard generation. Check AI model config.') @@ -284,7 +284,8 @@ export default function DocumentDetailPage() { {activeJob && ( { setActiveJob(null); navigate(`/quizzes/${quizId}`) }} + label={activeJob.type === 'flashcard' ? 'Generating Flashcards' : 'Extracting Questions'} + onDone={(quizId) => { setActiveJob(null); if (activeJob.type === 'flashcard') navigate('/flashcards'); else navigate(`/quizzes/${quizId}`) }} onClose={() => setActiveJob(null)} /> )}