pediatric-ai-scribe-v3/Dockerfile
Daniel 154b896d5b
Some checks failed
Forgejo Docker Build / Build Docker image (push) Blocked by required conditions
Forgejo Docker Build / Deploy to the host (push) Blocked by required conditions
Forgejo Android APK / Root app tests (push) Successful in 58s
Forgejo Docker Build / Root app tests (push) Successful in 49s
Forgejo Android APK / Build signed APK (push) Has been cancelled
feat: Word is built by python-docx from the same typed source as the deck
Pandoc reads markdown, so every Word export had to flatten the resource to
markdown first — and a deck flattened to markdown stops being one. A comparison
became two headings and two lists, a callout became bold text, and a figure
became nothing at all, because markdown has nowhere to put it.

src/utils/docSpec.js reduces either source to the same blocks: a stored deck
where there is one, the markdown where there is not. scripts/render_docx.py
draws them. A comparison comes out as a labelled two-column table, a callout as
a shaded box, a table as a real table, a figure embedded at its own aspect ratio
with its caption, and speaker notes as muted indented text.

The deck wins over the markdown beside it, because that markdown is a
serialisation of the deck and reading it instead would be reading a lossy copy of
what is right there.

Word now carries the figures too. The export route skipped fetching them for
docx, which was correct when pandoc could not place them and wrong the moment
this could.

Pandoc stays installed and stays the fallback: a plainer document beats a failed
download. Both renderers now share one spawn helper.

Verified end to end: a deck with two figures exported as a six-page Word document
with both images embedded (537KB, two files in word/media), rendered to PDF and
looked at — the comparison is a labelled table, the figure sits at its true
aspect ratio, and the notes read as notes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dv6sqaY6Vq3ChZHMem3cnU
2026-09-11 21:58:53 +02:00

76 lines
4 KiB
Docker

# ─── OpenBao CLI, copied from upstream image (multi-arch automatic) ───
# Update the tag here to adopt a newer OpenBao. Binary is statically linked,
# safe to drop into the Node alpine image as-is.
# Pinned by digest, not by tag: a tag is a moving pointer, so two builds of the
# same commit could otherwise produce different images. These are manifest-list
# digests, so buildx still selects the right per-architecture variant.
FROM openbao/openbao:2.5.3@sha256:fdc6da21ca6963560c32336fd7feb9cf2d5e52668f1a1647205a4b41171f0806 AS bao-src
FROM node:24-alpine@sha256:e67514e5d0f6c46656005e1b693b2ec9d52e80b641307de684d4a015ba7a4eaf
WORKDIR /app
# ffmpeg: audio conversion for AWS Transcribe (WebM → PCM)
# curl: HTTP helper used by the OpenBao entrypoint and health/debug tooling
# jq: JSON parsing for the entrypoint's OpenBao secret-fetch step
# pandoc: Markdown → PPTX for Learning resources. It is large (~230MB), and it
# is here rather than in a sidecar because a sidecar would add a
# cross-stack network dependency to an export that must not fail for
# reasons outside this container. It also measures images, which
# pptxgenjs cannot: that library emits the target box verbatim with
# <a:stretch/>, so every image in every generated deck was distorted.
RUN apk add --no-cache ffmpeg curl jq pandoc-cli
# python-pptx builds the slide decks. pandoc still writes Word, where its output
# is good, but its pptx writer can only map markdown onto a handful of reference
# layouts: no per-slide layout, no positioning, no control over where an image
# lands or how large it is. That ceiling is the renderer's, not the model's — a
# better-written deck still came out as bullets on a template, and slides
# overflowed until autofit was injected into the emitted OOXML by hand.
#
# py3-lxml and py3-pillow come from apk rather than pip because both are C
# extensions and Alpine has no wheels for them; installing from source here
# would mean carrying a compiler in the runtime image. Adds ~58MB.
# poppler-utils supplies pdftoppm, which turns a rendered deck into one image
# per slide. That is the only way to let a vision model see what a deck actually
# looks like — Gotenberg converts to PDF and stops there.
RUN apk add --no-cache poppler-utils
RUN apk add --no-cache python3 py3-pip py3-lxml py3-pillow \
&& pip install --break-system-packages --no-cache-dir python-pptx==1.0.2 python-docx==1.1.2 \
&& python3 -c 'import pptx, docx'
# Pull the bao CLI out of the upstream image — matches host arch because
# buildx pulls the right manifest-list variant per build.
COPY --from=bao-src /bin/bao /usr/local/bin/bao
RUN /usr/local/bin/bao version
COPY package.json package-lock.json ./
# argon2 compiles native code via node-gyp — needs python3/make/g++ at build time
RUN apk add --no-cache --virtual .build-deps python3 make g++ \
&& npm ci --omit=dev \
&& apk del .build-deps
COPY . .
# One validated source revision for both runtime cache busting and OCI provenance.
# Direct development builds without an explicit revision remain visibly unversioned.
ARG GIT_REVISION=unknown
RUN node -e 'const r=process.argv[1]; if (r !== "unknown" && !require("./src/utils/buildId").isGitRevision(r)) throw new Error("GIT_REVISION must be a full lowercase Git SHA"); require("node:fs").writeFileSync("BUILD_ID", r + "\n");' -- "$GIT_REVISION"
LABEL org.opencontainers.image.revision=$GIT_REVISION
# Ensure the entrypoint is executable regardless of host file permissions
RUN chmod +x /app/docker-entrypoint.sh
RUN mkdir -p /app/data/logs
EXPOSE 3000
HEALTHCHECK --interval=30s --timeout=5s --start-period=20s \
CMD wget --no-verbose --tries=1 --spider http://localhost:3000/api/health || exit 1
# Entrypoint wrapper handles optional OpenBao secret fetch before exec'ing CMD.
# See docker-entrypoint.sh for the logic — it is a no-op if OPENBAO_ADDR is
# unset, so legacy .env-only deployments continue to work unchanged.
ENTRYPOINT ["/app/docker-entrypoint.sh"]
CMD ["node", "server.js"]