pediatric-ai-scribe-v3/TODO.md
Daniel 596fd897f6 docs: record the mail fixes and the collection rename
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-10 17:42:30 +02:00

9.6 KiB

TODO

Live state as of 2026-09-11. Everything not listed under Open is deployed and green (674 tests, three consecutive clean runs).

Open

Needs your decision

  • Replace Cloudflare Turnstile. Used on registration and password reset only (src/routes/auth.js); login is not gated, it relies on a 10-per-15-min limit and a constant-time credential check. Recommended replacement: ALTCHA — open source, self-hosted, proof-of-work, no third-party calls and no tracking, which also lets three CSP entries and frameSrc go away. Alternatives: mCaptcha (open source, self-hosted, heavier to run) and Cap (newer, smaller). hCaptcha is neither Google nor open source, so it trades one third party for another.
  • Kubernetes / CI-CD hardening. Details under Deployment readiness.
  • Finish pointing audio backups at MinIO — one blocked step. The audio-backups bucket now exists on the same MinIO the app already uses (assets:9000). The app's key is scoped by a MinIO user policy, so a bucket policy alone grants it nothing — verified, it returns AccessDenied. The remaining step attaches a second policy covering only that bucket, leaving the generated-images grant untouched. It needs the MinIO root credentials, and my permission layer blocked me from running it. Run: RU=$(cat ../ped-ai-storage/secrets/minio-root-user) RP=$(cat ../ped-ai-storage/secrets/minio-root-password) then docker compose exec -T -e MINIO_ROOT_USER="$RU" -e MINIO_ROOT_PASSWORD="$RP" pediatric-scribe node scripts/enable-audio-backup-bucket.js, set AUDIO_BACKUPS_S3_* (the script prints them), and restart. Until then recordings go to the encrypted Postgres column, which works; rows already there keep working afterwards.
  • Basic index has no reader. MilvusVectorStore.search() exists, but no tool calls it. Decide where the query path lives: pymilvus inside the deliberately-lean nextcloud-basic-mcp image, or a query API from the indexer container. Nothing can read that index until this is settled. Source: /home/danvics/docker/nextcloud-basic-mcp
  • Apply the restored clinical vector-store compose. Written, committed and validated at /home/danvics/docker/nextcloud-mcp-server, deliberately NOT applied — up -d recreates the live clinical Milvus.

Known gaps

  • Multi-collection, ped-ai half. The MCP side is deployed (clinical_semantic_search(collection=…) + clinical_list_collections, allowlisted by MILVUS_COLLECTIONS). ped-ai still searches one collection per request. Needs an admin setting for which collections to search, then fan-out and merge — dedupeSources in src/utils/clinicalRetrieval.js already merges and renumbers. See clinical-assist/COLLECTIONS.md.
  • Mail indexing works. It was never reached: mail ran last, after nine other sources, and Tables alone is thousands of rows at about a second each. Mail leads now — it is the only bounded source (identities only, capped by BASIC_INDEXING_MAIL_MAX_MESSAGES, bodies left to the processor), so it cannot starve the others the way they starved it. Messages are indexing.
  • Mail attachments were never indexed. An attachment's id is its index within its message, so /api/attachments/{id} meant nothing and answered 500 every time. Fixed to /api/messages/{id}/attachment/{id}, verified live against a real message.
  • The basic collection is renamed personal_assistant_bge_m3_1024 (was basic_bge_m3_1024_v2), matching mcp_bge_m3_1024 on the clinical side. Milvus grants name the collection, so the rename revoked basic_reader/basic_writer; bootstrap_basic.py restored them, but it must be bind-mounted because the operator image ships an older copy.
  • The indexed folder is Personal assistant, and the setting now takes a comma-separated list (Personal assistant,Clinical Notes), each walked recursively. Note the file reconciliation removed the chunks of the 68 Documents files, since a complete listing is the deletion authority and they are no longer under an indexed root. Entities went 27,930 -> ~9,700. Re-add those files under an indexed folder if they are still wanted.

Deployment readiness (CI/CD and Kubernetes)

What already exists: ci.yml (tests on PR and main), security.yml (weekly audit), docker-publish.yml and Android/versioning workflows, a Dockerfile HEALTHCHECK, and /api/health.

Worth doing before Kubernetes, roughly in order:

  • Fail CI on vulnerabilities. security.yml runs weekly and does not gate merges. npm audit --audit-level=high in ci.yml would have caught the nodemailer advisories at the PR.
  • Build the image on PRs too. docker-publish.yml only runs on tags, so a Dockerfile break is found at release time.
  • Run the e2e suite in CI. scripts/e2e.sh and docker-compose.e2e.yml exist but nothing calls them.
  • Separate liveness from readiness. /api/health is one endpoint; Kubernetes wants liveness (process up) apart from readiness (database, gateway and MinIO reachable), or rollouts take traffic too early.
  • Graceful shutdown. No SIGTERM handler, so a rolling update can cut off an in-flight transcription or image job.
  • Externalise state. Uploads and audio backups assume local paths and a single instance; more than one replica needs them all in MinIO/Postgres.
  • Config as secrets. Everything is env vars in compose today, which maps to ConfigMap/Secret cleanly, but JWT_SECRET, gateway keys and database credentials should be a Secret from the start.
  • Pin the base image by digest and keep the SBOM the build already has.

Done since this file was written

2026-09-11

  • Live transcription: proved working end to end against the live gateway — local-kokoro-tts produced 92KB of speech and mistral-voxtral-mini-transcribe returned the sentence back verbatim. The Settings picker offered six hardcoded ids that do not exist on this gateway (local-whisper-large-v3-turbo → 400 Invalid model name); it now lists the nine the gateway advertises, cached, with the admin default marked.

  • Assistant voice mode is wired end to end: record → browser recognition, falling back to server transcription → send → spoken answer. All four helpers it needs exist.

  • Recordings can be exported (server and local copies) and a recorder that dies — an error, or the microphone taken by another app, unplugged or revoked — now says so instead of appearing to record silence. No wake lock, deliberately: stopping on sleep or sign-out is the behaviour you want.

  • nodemailer 9.0.1 → 9.1.1, clearing four high advisories, two of them delivery bugs that can route mail to an attacker-controlled domain.

  • Admin routers state their own authentication. adminMilestones relied on adminConfig being mounted first on /api/admin; it failed closed, but on mount order rather than intent.

  • Test suite made deterministic. A file failed about one run in four with "Unable to deserialize cloned data": node:test parses each child's stdout, and page/server logging was landing inside those frames. Every test child's stdout is now pure TAP.

  • iOS: text fields are 16px on phones, so Safari no longer zooms the page on focus — which was also why fixed chrome (the menu button) scrolled away.

  • Settings claim corrected: there is no "use my normal physical exam" trigger, and the prompt forbids copying template content.

  • Storage stack is under version control (/home/danvics/docker/ped-ai-storage), secrets verified excluded, with a README recording the misleading project names and the MinIO/separate-etcd requirements.

  • Both Milvus stores keep objects in MinIO; verified end to end on the basic side (16 objects in the bucket, rows queryable, collection Loaded).

  • The operator image is built from source (Dockerfile.operator), so the drifted check.py that broke image generation can no longer be run by accident.

Worth knowing

  • Four repos were rescued from container images today: nextcloud-basic-mcp, clinical-assist, the deleted nextcloud-mcp-server compose file, and the operator's check.py drift. Prefer building from a repo over a live container.
  • Two Milvus instances, historically misleading names. Clinical index = nextcloud-mcp-server-milvus-1 (MinIO-backed, collection mcp_bge_m3_1024). Nextcloud basic index = ped-ai-storage-basic-milvus-1 (db basic, collection basic_bge_m3_1024_v2).
  • Embedded etcd is unusable with authorization on. Every non-root Milvus user failed etcdserver: invalid auth token. Both stores now run a separate etcd container, matching the profile that always worked.
  • Both Milvus stores now keep objects in MinIO, matching what clinical always did. COMMON_STORAGETYPE=local wrote segment files relative to the working directory, so a recreate destroyed them while etcd kept referencing them and the collection hung at Loading forever. S3 semantics also mean either store can be pointed at a managed bucket without touching Milvus — which is what makes a Terraform-managed deployment straightforward.
  • Milvus object-store credentials live in milvus-user.yaml in the protected secrets dir, not in the compose, because Milvus has no file-based option for them and the compose should stay reviewable.

Co-Authored-By: Claude Opus 5 noreply@anthropic.com