pediatric-ai-scribe-v3/docs/clinical-assistant.md
Daniel 67e416c6d9
Some checks failed
Forgejo Android APK / Root app tests (push) Successful in 47s
Forgejo Docker Build / Root app tests (push) Successful in 48s
Forgejo Android APK / Build signed APK (push) Successful in 2m38s
Forgejo Docker Build / Build Docker image (push) Successful in 12s
Forgejo Docker Build / Deploy to the host (push) Failing after 2s
docs: merge the duplicate pairs and correct them against the running app
Three pairs of docs described the same thing twice, and the copies had drifted
apart. Merged each into one file, keeping the unique content from both:

- ARCHITECTURE.md -> architecture.md (its operational map: ownership, request
  flow, runtime boundaries, source of truth, deployment shape)
- DEVELOPMENT.md -> developer-guide.md (change workflow, Clinical Assistant
  high-risk areas, frontend rendering rules, deployment checks)
- transcription-options.md -> speech.md (the clinic setup table, and the list
  of browser-Whisper paths that must stay removed)

Then audited what remained against the code and the live database rather than
against the previous docs. Corrected:

- Google Vertex was still documented as a provider across nine files. The SDK
  is gone; AI_PROVIDER=vertex now logs an advisory and falls back to
  OpenRouter, and Gemini is reached through LiteLLM. Fixed the provider
  selection order to match src/utils/ai.js, which starts from LITELLM_API_BASE.
- promptSafe was documented on 8 routes; it is on 13.
- Node 20 -> 24, "24 vanilla JS modules" -> no fixed count, and
  transcribe.js/tts.js -> sttProvider.js/ttsProvider.js, which is what exists.
- STT/TTS are LiteLLM-only; README listed direct Google, AWS Transcribe and
  ElevenLabs paths that are not in the runtime.
- Learning Hub PPTX export was documented as pptxgenjs, which is not a
  dependency. It is pandoc against a reference deck.
- POST /api/admin/milestones/seed does not exist; it is /bulk-import.
- NEXTCLOUD_URL and NTFY_TOPIC are not read anywhere. Nextcloud is per-user in
  the users table, and the ntfy topic is derived as pedscribe-{userId}.
- A prose paragraph sat inside the Clinical Assistant settings table, so half
  the rows rendered as text.

Filled the gaps the audit exposed:

- database.md was missing 12 of 29 tables, including user_resources,
  personal_notes, login_codes, registration_invites and generated_image_jobs.
- developer-guide.md was missing 11 routers and 10 frontend modules.
- api-reference.md detailed 121 of 244 endpoints and said so, but whole
  features were absent. Added an endpoint index covering Clinical Assistant,
  My Resources, Notes, Diagrams, ED Encounters, invites and sign-in codes.
- configuration.md was missing METRICS_TOKEN, REDIS_URL, API_RATE_LIMIT_MAX,
  the LITELLM_* model variables, the DB_* ones maintenance.js reads, and the
  per-purpose S3 resolution scheme.
- clinical-assistant.md documented 2 of its 17 environment variables.
- features-explained.md had no entry for My Resources or Clinical Assistant.

Renamed the three remaining SHOUTING filenames to kebab-case, which is what the
docs viewer's prettyName() was working around, and rewrote README's index,
which listed architecture.md twice and omitted nine files.

Noted but not changed: the Turnstile site key is hardcoded in index.html rather
than read from TURNSTILE_SITE_KEY, and /api/health/detailed can report
tts: 'elevenlabs' though no ElevenLabs path exists.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dv6sqaY6Vq3ChZHMem3cnU
2026-09-12 04:57:35 +02:00

11 KiB

Clinical Assistant

The Clinical Assistant is a retrieval-grounded assistant for pediatric clinical reference questions. It is not the same as the app's note-generation/HPI workflow.

Responsibilities

Component Responsibility
Browser UI question input, source display, markdown/citation rendering, export
Ped-AI backend settings, MCP search call, answer prompt construction, model call
MCP server Nextcloud access, indexing, vector search, rerank, source metadata
LiteLLM model routing and provider abstraction

Request Flow

User asks a question
  -> browser posts to Ped-AI
  -> Ped-AI calls MCP `clinical_semantic_search`
  -> MCP returns source excerpts and metadata
  -> Ped-AI builds an answer prompt with source constraints
  -> LiteLLM model returns answer text
  -> browser renders answer and source cards

Source Rules

  • Prefer MCP file_path basename for displayed source titles when present.
  • Do not relabel one source as another requested source.
  • If the user names a source and retrieval does not return it, say that before using other sources.
  • Use citations only for returned source numbers.
  • Unknown citation numbers should remain plain text instead of being guessed.

Table And Markdown Rendering

LLM output is not guaranteed to be valid markdown. The browser renderer defensively handles common problems:

  • adjacent citation clusters,
  • missing closing bracket in narrow citation cases,
  • smashed bullet lists,
  • inline headings,
  • malformed pipe tables,
  • bare source numbers in source/citation table columns,
  • orphan markdown emphasis markers,
  • code blocks that must not be modified.

Renderer fixes must be narrow. Do not add broad repairs that turn arbitrary clinical numbers into citations.

Image Routing

Table lookup requests should stay in retrieval flow.

Examples that should use retrieval:

show me the table
show me Table 13.1
summarize the developmental table

Explicit visual creation/display requests can use image flow.

Examples:

create an infographic
generate a diagram
show me the image/figure

Caching Policy

Clinical answer response caching is intentionally disabled. Redis can support prompt suggestions and operational metadata, but final answers should be generated from current retrieval context.

Image Attachments

Users can attach up to 4 images (PNG, JPEG, WebP) to an outgoing clinical question. Attachments ride the outgoing question only for inference and persist with the saved chat once the question is sent:

  • They are validated client-side and authoritatively on the server (MIME allowlist, canonical base64, ≤ 5 MiB per image, ≤ 4 images, ≤ 10 MiB decoded total). Invalid input is rejected with 400 before any retrieval or provider call.
  • They are sent only with the outgoing clinical question for inference. Attaching images never disables retrieval: RAG/includeContext runs exactly as without images.
  • The conversation budget counts text only: images are excluded from the UTF-16 code-unit count. The server still validates every request.
  • Once sent, the message's attachments are stored in the saved chat payload (same bounded limits, re-validated on every save) and restored as thumbnails on load.
  • Only OpenAI-compatible providers (LiteLLM, OpenRouter, Azure) receive them as multimodal content parts (text + image_url data URIs) on the latest user message; the system/retrieval/history structure is unchanged. The direct Bedrock adapter refuses with a clear 400 before contacting the provider.
  • Attachments clear on a successful send and on New chat; a rejected send keeps them for correction.

Autosave, titles and saved-chat updates

After each completed assistant turn (and on any change to the conversation), the chat is autosaved with an 800 ms debounce to POST /api/clinical-assistant/chats. New chats get a title derived from the first user message (first 60 characters); later saves include the chat id and update the same row in place. Failures surface once per change and never block chat flow; oversized saves keep the 8 MiB / 400 / 413 semantics and are retried only on the next change, never truncated. The raw transcript stays canonical. The generated sidebar image and per-message image jobs persist with the chat again.

Translation

Every message offers Translate with a target-language picker. Translation is the local LibreTranslate container (LIBRETRANSLATE_URL, default http://libretranslate:5000), which is the only provider there is. clinical_assistant.translate_provider is read but any unrecognised value silently falls back to LibreTranslate, and no DeepL client exists in the code at all. Responses are cached per provider+message+lang. Patient text therefore never leaves the local network.

Settings

Important settings include:

All are stored in settings and edited under Admin → Clinical Assistant / Learning, except the image roster, which is written by the Image Generation card. Every one is read through getSetting, so an unset key falls back to the default in the right-hand column.

Setting Purpose
clinical_assistant.chat_model Chat model for answers; falls back to models.default
clinical_assistant.image_model Image model for explicit image generation; falls back to CLINICAL_ASSISTANT_IMAGE_MODEL, then openai-gpt-image-1
clinical_assistant.fallback_image_model Single retry target when the image model fails
clinical_assistant.allowed_models Comma-separated chat models a user may pick. Empty means no choice: the configured model is used. A non-empty list always includes the configured model; anything else is rejected with 400 model_not_allowed
clinical_assistant.allowed_image_models The same, for image models
clinical_assistant.image_model_roster Image models an admin added from Admin → Image Generation (+ Add). This is the pool the Image models tick-list offers; it is not itself an allowlist. Validated as up to 100 ids
clinical_assistant.search_limit Number of MCP results requested
clinical_assistant.context_chars Context characters requested from MCP
clinical_assistant.conversation_chars Input budget in UTF-16 code units. Empty means use CLINICAL_ASSISTANT_CONVERSATION_CHARS; a value must be 1000-1000000
clinical_assistant.show_sources true/false. Display only: hides the Sources panel and the citation markers. The prompt, the retrieval and the stored answer are byte-for-byte identical either way, so it cannot bias an answer; turning it back on restores the citations
clinical_assistant.preview_enabled true/false. Lets signed-out visitors try the assistant read-only; anything needing an account asks them to sign in
clinical_assistant.system_behavior Admin-editable assistant behavior guidance
clinical_assistant.image_behavior Guidance for the generate_image tool
clinical_assistant.patient_takehome_behavior Guidance for patient take-home text
clinical_assistant.prompt_model Model that generates the starter prompt pool
clinical_assistant.translate_provider Translation provider. libretranslate is the only value the server accepts
clinical_assistant.citations_enabled Legacy key, read only as a fallback for show_sources

search_limit and context_chars are capped by RERANKER_TOP_K in the MCP deployment, which is the real ceiling on every search. See retrieval-tuning.md.

Environment variables

Settings above are the normal way to configure the assistant. These environment variables sit underneath them — connection details, timeouts, and the defaults a setting falls back to.

Variable Default Purpose
CLINICAL_ASSISTANT_MCP_URL MCP endpoint. MCP_SERVER_URL is accepted as an older name.
CLINICAL_ASSISTANT_MCP_URLS Comma-separated list, tried in order, ahead of the single-URL variable.
CLINICAL_ASSISTANT_SEARCH_TOOL clinical_semantic_search Tool name to call on the MCP server. Only this value is accepted; the nc_semantic_search alias was removed, and anything else throws at startup rather than failing per request.
CLINICAL_ASSISTANT_MCP_INITIALIZE_TIMEOUT_MS 30000 Session handshake timeout.
CLINICAL_ASSISTANT_MCP_REQUEST_TIMEOUT_MS 90000 Per-search timeout.
CLINICAL_ASSISTANT_MCP_SESSION_TTL_MS 600000 How long an MCP session is reused.
CLINICAL_ASSISTANT_MCP_WARMUP on Set to false to skip opening an MCP session at boot. Tests set this.
CLINICAL_ASSISTANT_MCP_WARMUP_DELAY_MS 5000 Delay before that warmup.
CLINICAL_ASSISTANT_CONVERSATION_CHARS 120000 Input budget in UTF-16 code units, when the setting is empty.
CLINICAL_ASSISTANT_IMAGE_MODEL openai-gpt-image-1 Image model, when the setting is empty.
CLINICAL_ASSISTANT_PROMPT_MODEL Model for the starter prompt pool, when the setting is empty.
CLINICAL_ASSISTANT_PROMPT_POOL_TARGET 1000 How many example prompts to generate.
CLINICAL_ASSISTANT_PROMPT_POOL_REFRESH_MS 7 days How often the pool regenerates. 0 disables refresh.
CLINICAL_ASSISTANT_PROMPT_POOL_KEY clinical-assistant:prompt-pool:v2 Redis key holding the pool.
CLINICAL_ASSISTANT_PROMPT_POOL_WARMUP_DELAY_MS 15000 Delay before the pool warms at boot.
CLINICAL_ASSISTANT_EXAMPLE_CACHE_MS 600000 How long the examples endpoint caches its answer.

Choosing a model

The composer shows a Model button rather than the model id, which can be as long as openrouter-gemini-3.1-flash-image-preview; clicking it opens the list. The button is a face for #assistant-chat-model-select, which stays in the DOM as the state holder — so a choice made in the popup is saved by the same delegated change listener as before, under an account-scoped storage key. The whole control is hidden unless the allowlist offers more than one model.

For an image model to reach a user, an admin does two things: + Add it under Admin → Image Generation (which puts it in image_model_roster), then tick it in the Clinical Assistant's Image models list (which puts it in allowed_image_models). Discovery lists what the gateway advertises with mode image_generation; it never adds anything on its own.

Testing Priorities

Add or update tests when changing:

  • citation rendering,
  • source title cleanup,
  • named-source provenance behavior,
  • table rendering and table copy/CSV actions,
  • image intent routing,
  • image attachment validation, multimodal payload shape and saved-chat roundtrips,
  • autosave debounce, title derivation and saved-chat updates,
  • translation validation, caching and provider fallback,
  • MCP result normalization,
  • model discovery or settings behavior.