pediatric-ai-scribe-v3/docs/deployment.md
Daniel 36cb742ce7
All checks were successful
Forgejo Docker Build / Root app tests (push) Successful in 54s
Forgejo Docker Build / Build Docker image (push) Successful in 7s
ci: fix the failing job, split deploy out, and drop the Android build
Three things, one subject: making CI say the truth about this repo.

## The red on every run was ours, not the runners'

Every docker-build run came back success, success, failure — the same
shape for weeks. The failing job was `deploy`, and it was failing to
*not run*:

    if: ${{ github.event.inputs.deploy == 'true' }}

On a push there is no github.event.inputs at all. This Forgejo does not
treat that as false and skip; it dispatches the job, the runner cannot
resolve it, and the task ends in "Early termination". The runners were
never at fault, and nothing about them needed changing.

The `'runs-on' key not defined` line is a red herring: the `build` job
prints it too and succeeds. It names the job's *needs* target, not the
job, and the old android-apk workflow used `needs:` happily for months.

Deploy is now its own workflow with only workflow_dispatch — no
condition to evaluate, so nothing can be dispatched by mistake. No job
in either file now carries a job-level `if`. The one conditional left is
a *step* (push to registry), and step conditions are evaluated by the
runner once the job is already running, which is why that one has always
worked.

## dev and main

docker-build now runs on `dev` as well. Both branches prove the same two
things — tests pass, image builds — and only `main` publishes the image,
so nothing on `dev` can be mistaken for something deployable. Deploying
stays a person pressing a button after looking at the change.
CONTRIBUTING.md documents the flow.

## Android

Removed: the mobile/ Capacitor project, docs/mobile-build.md, and the
Android bits of scripts/release.sh. All of it is in git history — 4613a278
is the last commit that had it — for when it is rebuilt.

src/utils/platform.js stays. isMobileClient only decides token lifetime,
it is twelve lines, and it is the contract a future app would come back
to; deleting it would be a change to auth for no gain.

.github/workflows/ went too — all five. There is no GitHub remote on
this repository, so none of them has ever run, and two of them wrote
into mobile/ paths that no longer exist.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dv6sqaY6Vq3ChZHMem3cnU
2026-09-12 23:30:52 +02:00

10 KiB

Deployment

Prerequisites

  • Docker + Docker Compose
  • Reverse proxy (Caddy, Nginx, Traefik) for TLS termination
  • At least one configured AI provider (LiteLLM / OpenRouter / Bedrock / Azure)

What the image carries

Beyond Node, the runtime image installs a few tools that document export depends on. They are in Dockerfile and worth knowing about before trimming it:

For
pandoc-cli the fallback for Word export when the renderer cannot run
python3, py3-lxml, py3-pillow the slide renderer. Both libraries are C extensions with no Alpine wheels, so they come from apk rather than pip — installing them from source would mean carrying a compiler in the runtime image
python-pptx==1.0.2, python-docx==1.1.2 (pip) build the decks and the documents. Pinned: unpinned, a rebuild from the same commit could produce different output
poppler-utils pdftoppm, which turns a rendered deck into one image per slide so a vision model can see it. Only needed when slide review is switched on
ffmpeg, curl, jq audio handling and entrypoint scripting

Roughly 58MB of that is Python. PDF conversion is not in the image — it goes to Gotenberg over the network (GOTENBERG_URL, default http://gotenberg:3000), so PowerPoint and Word still work when Gotenberg is down and only PDF fails.

See my-resources.md for what the renderer does.

Images

Image Role
danielonyejesi/pediatric-ai-scribe-v3:latest App container. Published by CI on every tag push where configured. Pull directly or build from source.
pgvector/pgvector:pg16 Database.
redis:7-alpine Operational Redis cache/state.

Build from source

git clone https://github.com/ifedan-ed/pediatric-ai-scribe-v3.git
cd pediatric-ai-scribe-v3
cp .env.example .env
# edit .env — required: APP_URL, JWT_SECRET, DATA_ENCRYPTION_KEY, DB_PASSWORD, an AI provider
./scripts/build-image.sh
REV=$(git rev-parse HEAD)
scripts/deploy.sh "ped-ai-local:$REV" "$REV"

Use scripts/deploy.sh. Do not run docker compose up by hand.

Building is not deploying. docker compose takes its image from PED_AI_IMAGE in .env, and build-image.sh does not move that pin — naming a revision is also how a rollback is done. So a pin left behind by an earlier deploy starts that image, and every signal still reports success: the build completes, up says the container started, and /api/health returns {ok:true} from the wrong revision. This has happened: a stale pin silently reverted the app by 31 commits, removing a feature, and the missing feature was reported as a new bug.

scripts/deploy.sh <image-ref> [expected-revision] is what closes that gap:

  1. pulls the image if it is not local, and refuses to tear anything down until it exists;
  2. records what is serving now, so there is something to go back to;
  3. moves the PED_AI_IMAGE pin, so a later plain docker compose up brings up the same image rather than reverting;
  4. waits for the container to become healthy;
  5. asks /api/build which revision is actually serving and compares it to the expected one — catching a stale tag, a cached layer, or a rollback that never took;
  6. rolls back to the previous image if either check fails.

build-image.sh prints the exact deploy.sh line to run whenever the pin does not match the revision it just built.

The build uses Node 24 LTS and npm ci --omit=dev from the root lockfile. ./scripts/build-image.sh resolves the full checkout Git commit (including worktrees/packed refs) and passes GIT_REVISION through Compose. It only builds; starting or replacing production services remains a separate reviewed step. Use COMPOSE_FILE=docker-compose.local.yml ./scripts/build-image.sh for the local variant. For direct Docker builds:

docker build --build-arg GIT_REVISION="$(git rev-parse --verify 'HEAD^{commit}')" -t ped-ai-local:latest .

All Compose variants accept the same GIT_REVISION environment variable. A build without one is explicitly unknown (unversioned development), not a release provenance claim.

The Dockerfile rejects malformed revisions and writes the same full SHA to /app/BUILD_ID and org.opencontainers.image.revision. /api/build, the X-Build-Id header and asset query strings use that baked value. Git identifies the source commit, not local uncommitted changes: release from a clean checkout; a local dirty test image is not an exact representation of that commit.

Forgejo's existing trusted push/manual release workflows run a Node 24 root npm ci / npm test job on forgejo-local; APK and Docker jobs require it via needs. No untrusted pull-request code may run on that privileged runner. An isolated, unprivileged Forgejo PR runner is separate future provisioning, not an assumed label in these workflows. GitHub-hosted PR CI uses Node 24; GitHub release workflows also gate builds on root tests. Mobile dependency versions and signing/publishing gates are unchanged.

The default compose starts pediatric-ai-scribe on 127.0.0.1:3552, pedscribe-db internally, and ped-ai-redis internally.

Minimum .env

APP_URL=https://scribe.example.com
JWT_SECRET=<openssl rand -hex 32>
DATA_ENCRYPTION_KEY=<openssl rand -hex 32>
DB_PASSWORD=<strong password>

AI_PROVIDER=litellm
LITELLM_API_BASE=https://llm.example.com
LITELLM_API_KEY=sk-...

Full variable reference: docs/configuration.md.

Reverse proxy

App binds to 127.0.0.1:3552 only. TLS termination + host routing is the proxy's job.

Caddy

scribe.example.com {
    reverse_proxy localhost:3552
}

Nginx

server {
    listen 443 ssl http2;
    server_name scribe.example.com;
    ssl_certificate     /etc/ssl/certs/scribe.example.com.pem;
    ssl_certificate_key /etc/ssl/private/scribe.example.com.key;
    client_max_body_size 100M;
    location / {
        proxy_pass http://127.0.0.1:3552;
        proxy_set_header Host $host;
        proxy_set_header X-Real-IP $remote_addr;
        proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
        proxy_set_header X-Forwarded-Proto $scheme;
    }
}

App sets trust proxy: 1 so rate limiting uses the original client IP.

Volumes

Volume Contents Backup priority
pgdata All user data, encounters, memories, audit logs, settings Critical
scribe-logs Filesystem audit log files (JSONL by day) High for compliance evidence; Postgres also has audit/API/access tables

Postgres backup / restore

# Backup
docker exec pedscribe-db pg_dump -U pedscribe pedscribe > backup.sql

# Restore
cat backup.sql | docker exec -i pedscribe-db psql -U pedscribe pedscribe

Updating

From a Docker Hub pull

docker compose pull
docker compose up -d

Building from source

git pull
./scripts/build-image.sh --no-cache
docker compose up -d

On startup the container runs initDatabase() (idempotent baseline), then node-pg-migrate applies any new migration files. Collation-drift check auto- REINDEXes if the ICU library version changed between image builds.

Health

Endpoint Purpose
GET /api/health {ok:true} — public, used by Docker health check
GET /api/health/detailed Provider status — admin-auth required
GET /api/build Build ID (short git SHA) — useful for debugging cache invalidation
GET /metrics Prometheus metrics in text exposition format

Docker health check in Dockerfile: every 30 s, wget-spiders /api/health. Container marked unhealthy after 5 failures.

Resource footprint

  • RAM: 256 MB minimum, 512 MB recommended for one instance with a handful of concurrent users.
  • Disk: Postgres size scales with audit log retention, saved encounters, and documents.
  • CPU: idle load negligible; AI calls are network-bound on the LLM provider side.

Production checklist

  • JWT_SECRET ≥ 32 bytes (openssl rand -hex 32)
  • DATA_ENCRYPTION_KEY exactly 64 hex chars
  • DB_PASSWORD non-default
  • APP_URL = public URL (enables fail-closed CORS + HSTS + secure cookies)
  • HIPAA workload → use Bedrock or Azure OpenAI directly, or a LiteLLM gateway pointed at a BAA-eligible upstream. Not OpenRouter.
  • SMTP configured for verification + reset emails
  • Turnstile keys set for public-facing deployments
  • Reverse proxy serves valid TLS certs
  • Postgres dump scheduled off-host
  • Log retention and backup policy covers audit_log, api_log, access_log, and filesystem scribe-logs

CI / CD

Forgejo Actions only; there is no GitHub remote on this repository.

Workflow Trigger What it does
.forgejo/workflows/docker-build.yml push to dev or main test suite, then build the image. On main only, push it to git.danvics.com/danvics/pediatric-ai-scribe-v3:{revision,latest}
.forgejo/workflows/deploy.yml manual dispatch scripts/deploy.sh against the host: pin the image, wait for health, verify /api/build, roll back on disagreement

Deploying is never automatic — see "Branches" in CONTRIBUTING.md. Versioning is manual: scripts/release.sh X.Y.Z --push.

Ports

Service Internal External default
App 3000 127.0.0.1:3552
Postgres 5432 not exposed
Redis 6379 not exposed

Change the app's external port by editing the ports: mapping in docker-compose.yml.

Log destinations

  1. Container stdout (docker compose logs -f pediatric-scribe).
  2. Filesystem data/logs/YYYY-MM-DD.log (JSONL, one line per event).
  3. Postgres tables audit_log, api_log, access_log — batched writes via src/utils/auditQueue.js, drained on SIGTERM.
  4. Loki (if LOKI_URL set) — pushed fire-and-forget per event.

A central Prometheus/Loki/Grafana stack can also scrape GET /metrics and collect Docker logs with Promtail. Keep direct Loki push enabled only for structured application events that are useful for compliance and operations.

Auto-cleanup

Target Policy Frequency
saved_encounters Delete where expires_at < NOW(). Default 7 days (configurable via site.auto_delete_days). Hourly + 10 s after startup
audio_backups Delete where expires_at < NOW() (24 h default). Same schedule

Graceful shutdown

server.js handles SIGTERM and SIGINT:

  1. Close HTTP listener (new connections refused, in-flight finish).
  2. Drain src/utils/auditQueue.js (flush any pending audit/api/access writes).
  3. pool.end() — close Postgres pool cleanly.

9-second hard deadline — Docker sends SIGKILL after 10 s by default. Prevents in-flight note writes from being truncated on docker restart.