pediatric-ai-scribe-v3/TODO.md
Daniel 713ed830a3 chore: script to move audio backups onto MinIO; record where things stand
Same MinIO, its own bucket, as asked. The audio-backups bucket is created;
what remains is a MinIO IAM change, which needs root credentials — my
permission layer blocked that call, so it is a script to run rather than
something I applied.

The script is additive and reversible: it attaches a second policy covering
only the new bucket and carries the existing generated-images grant over
rather than replacing it (attaching only the new one would break image
storage). It prints the AUDIO_BACKUPS_S3_* values to set, and how to undo.

A bucket policy alone does not work here: MinIO evaluates the user policy
first and it denies by default. Verified — the app key gets AccessDenied on
the new bucket until its own policy allows it.

TODO records that, and the indexer being repointed from Documents to
Personal assistant so mail is finally reached.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dv6sqaY6Vq3ChZHMem3cnU
2026-09-10 17:00:02 +02:00

153 lines
9.6 KiB
Markdown

# TODO
Live state as of 2026-09-11. Everything not listed under **Open** is deployed
and green (674 tests, three consecutive clean runs).
## Open
### Needs your decision
- [ ] **Replace Cloudflare Turnstile.** Used on registration and password reset
only (`src/routes/auth.js`); login is not gated, it relies on a
10-per-15-min limit and a constant-time credential check. Recommended
replacement: **ALTCHA** — open source, self-hosted, proof-of-work, no
third-party calls and no tracking, which also lets three CSP entries and
`frameSrc` go away. Alternatives: **mCaptcha** (open source, self-hosted,
heavier to run) and **Cap** (newer, smaller). hCaptcha is neither Google
nor open source, so it trades one third party for another.
- [ ] **Kubernetes / CI-CD hardening.** Details under *Deployment readiness*.
- [ ] **Finish pointing audio backups at MinIO — one blocked step.** The
`audio-backups` bucket now exists on the same MinIO the app already uses
(`assets:9000`). The app's key is scoped by a MinIO *user* policy, so a
bucket policy alone grants it nothing — verified, it returns AccessDenied.
The remaining step attaches a second policy covering only that bucket,
leaving the generated-images grant untouched. It needs the MinIO root
credentials, and my permission layer blocked me from running it. Run:
`RU=$(cat ../ped-ai-storage/secrets/minio-root-user) RP=$(cat ../ped-ai-storage/secrets/minio-root-password)`
then `docker compose exec -T -e MINIO_ROOT_USER="$RU" -e MINIO_ROOT_PASSWORD="$RP" pediatric-scribe node scripts/enable-audio-backup-bucket.js`,
set `AUDIO_BACKUPS_S3_*` (the script prints them), and restart. Until then
recordings go to the encrypted Postgres column, which works; rows already
there keep working afterwards.
- [ ] **Basic index has no reader.** `MilvusVectorStore.search()` exists, but no
tool calls it. Decide where the query path lives: pymilvus inside the
deliberately-lean `nextcloud-basic-mcp` image, or a query API from the
indexer container. Nothing can read that index until this is settled.
Source: `/home/danvics/docker/nextcloud-basic-mcp`
- [ ] **Apply the restored clinical vector-store compose.** Written, committed
and validated at `/home/danvics/docker/nextcloud-mcp-server`, deliberately
NOT applied — `up -d` recreates the live clinical Milvus.
### Known gaps
- [ ] **Multi-collection, ped-ai half.** The MCP side is deployed
(`clinical_semantic_search(collection=…)` + `clinical_list_collections`,
allowlisted by `MILVUS_COLLECTIONS`). ped-ai still searches one collection
per request. Needs an admin setting for which collections to search, then
fan-out and merge — `dedupeSources` in `src/utils/clinicalRetrieval.js`
already merges and renumbers. See `clinical-assist/COLLECTIONS.md`.
- [ ] **Confirm mail indexing end to end.** Six separate breaks are fixed, and
mail is enabled (`BASIC_INDEXING_MAIL_ENABLED=true`), but no mail has been
indexed yet: the file pass was still working through 472 large PDFs (68
done in six hours, some over 1,500 chunks each) and mail runs after it.
On 2026-09-11 the indexed folder was changed from `Documents` to
`Personal assistant`, which currently lists 0 files, so the file pass is
now trivial and mail should be reached on a scan cycle (every 300s).
Watch for `[SCAN-*] Mail messages: N seen, M queued`.
Note: the 68 files already indexed from `Documents` are still in the
collection. Say the word if they should be cleared.
- [ ] **`image-size` DoS advisory (high), no upstream fix.** Every published
version up to 2.0.2 is affected; `npm audit fix` only offers a breaking
downgrade of pptxgenjs. Not reachable here: the only `addImage` call
(`src/routes/learningAI.js`) is fed PNGs the app generated itself and
fetched with an ownership check, never an uploaded file. Uploads are
allowlisted to jpeg/png/gif and checked against their magic bytes. Revisit
when pptxgenjs ships a patched dependency.
- [ ] **`uuid` advisory (moderate) via gaxios via Google auth.** Not reachable:
the flaw needs v3/v5/v6 with an explicit buffer, and gaxios calls only
`v4()` for a multipart boundary. Forcing an override risks Vertex auth.
## Deployment readiness (CI/CD and Kubernetes)
What already exists: `ci.yml` (tests on PR and main), `security.yml` (weekly
audit), `docker-publish.yml` and Android/versioning workflows, a Dockerfile
`HEALTHCHECK`, and `/api/health`.
Worth doing before Kubernetes, roughly in order:
- [ ] **Fail CI on vulnerabilities.** `security.yml` runs weekly and does not
gate merges. `npm audit --audit-level=high` in `ci.yml` would have caught
the nodemailer advisories at the PR.
- [ ] **Build the image on PRs too.** `docker-publish.yml` only runs on tags, so
a Dockerfile break is found at release time.
- [ ] **Run the e2e suite in CI.** `scripts/e2e.sh` and
`docker-compose.e2e.yml` exist but nothing calls them.
- [ ] **Separate liveness from readiness.** `/api/health` is one endpoint;
Kubernetes wants liveness (process up) apart from readiness (database,
gateway and MinIO reachable), or rollouts take traffic too early.
- [ ] **Graceful shutdown.** No SIGTERM handler, so a rolling update can cut off
an in-flight transcription or image job.
- [ ] **Externalise state.** Uploads and audio backups assume local paths and a
single instance; more than one replica needs them all in MinIO/Postgres.
- [ ] **Config as secrets.** Everything is env vars in compose today, which maps
to ConfigMap/Secret cleanly, but `JWT_SECRET`, gateway keys and database
credentials should be a Secret from the start.
- [ ] **Pin the base image by digest** and keep the SBOM the build already has.
## Done since this file was written
### 2026-09-11
- **Live transcription**: proved working end to end against the live gateway —
`local-kokoro-tts` produced 92KB of speech and
`mistral-voxtral-mini-transcribe` returned the sentence back verbatim. The
Settings picker offered six hardcoded ids that do not exist on this gateway
(`local-whisper-large-v3-turbo` → 400 Invalid model name); it now lists the
nine the gateway advertises, cached, with the admin default marked.
- **Assistant voice mode** is wired end to end: record → browser recognition,
falling back to server transcription → send → spoken answer. All four helpers
it needs exist.
- **Recordings can be exported** (server and local copies) and a recorder that
dies — an error, or the microphone taken by another app, unplugged or
revoked — now says so instead of appearing to record silence. No wake lock,
deliberately: stopping on sleep or sign-out is the behaviour you want.
- **nodemailer 9.0.1 → 9.1.1**, clearing four high advisories, two of them
delivery bugs that can route mail to an attacker-controlled domain.
- **Admin routers state their own authentication.** `adminMilestones` relied on
`adminConfig` being mounted first on `/api/admin`; it failed closed, but on
mount order rather than intent.
- **Test suite made deterministic.** A file failed about one run in four with
"Unable to deserialize cloned data": node:test parses each child's stdout, and
page/server logging was landing inside those frames. Every test child's stdout
is now pure TAP.
- **iOS**: text fields are 16px on phones, so Safari no longer zooms the page on
focus — which was also why fixed chrome (the menu button) scrolled away.
- **Settings claim corrected**: there is no "use my normal physical exam"
trigger, and the prompt forbids copying template content.
- Storage stack is under version control (`/home/danvics/docker/ped-ai-storage`),
secrets verified excluded, with a README recording the misleading project names
and the MinIO/separate-etcd requirements.
- Both Milvus stores keep objects in MinIO; verified end to end on the basic side
(16 objects in the bucket, rows queryable, collection Loaded).
- The operator image is built from source (`Dockerfile.operator`), so the drifted
`check.py` that broke image generation can no longer be run by accident.
## Worth knowing
- **Four repos were rescued from container images today**: `nextcloud-basic-mcp`,
`clinical-assist`, the deleted `nextcloud-mcp-server` compose file, and the
operator's `check.py` drift. Prefer building from a repo over a live container.
- **Two Milvus instances, historically misleading names.**
Clinical index = `nextcloud-mcp-server-milvus-1` (MinIO-backed, collection
`mcp_bge_m3_1024`). Nextcloud basic index = `ped-ai-storage-basic-milvus-1`
(db `basic`, collection `basic_bge_m3_1024_v2`).
- **Embedded etcd is unusable with authorization on.** Every non-root Milvus user
failed `etcdserver: invalid auth token`. Both stores now run a separate etcd
container, matching the profile that always worked.
- **Both Milvus stores now keep objects in MinIO**, matching what clinical always
did. `COMMON_STORAGETYPE=local` wrote segment files relative to the working
directory, so a recreate destroyed them while etcd kept referencing them and the
collection hung at Loading forever. S3 semantics also mean either store can be
pointed at a managed bucket without touching Milvus — which is what makes a
Terraform-managed deployment straightforward.
- **Milvus object-store credentials live in `milvus-user.yaml`** in the protected
secrets dir, not in the compose, because Milvus has no file-based option for
them and the compose should stay reviewable.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>