The e2e stack shared production's Postgres — same server, same database,
same table. Seeded robots sat in `users` beside real clinicians, and
anything a test wrote, or a migration under test changed, landed on real
data. Nothing about "run the tests" should be able to reach an account
belonging to a person.
Now it has a Postgres and a Redis of its own, both on tmpfs: created
empty on every run, held in RAM, gone on teardown. scripts/e2e.sh is one
command that recreates the stack, seeds it, runs the browser and leaves
the app up at 127.0.0.1:3553 so it can be clicked around in, with the
report served at :3554.
Two bugs fell out of it immediately, both of which only a database that
did not already exist could have found:
The schema could not be built from nothing. The entrypoint migrated
before the app created its baseline tables, so the first migration
failed on saved_encounters not existing. It never showed because every
database this has ever run against already had the baseline. Then, one
layer down, 1777800000000_generated-images creates a table with a
foreign key to learning_content — which the baseline stopped creating
when Learning Hub was removed. Restoring into a brand-new database could
not have booted. The entrypoint now stands aside when the database is
empty and lets the app do it in the order it already gets right, and the
foreign key is only created where its target is. All 20 migrations
replay from empty, producing the same 23 tables production has.
Configuration lives in the database, so a throwaway one starts at
defaults — 14 settings against production's 49. That is why every model
picker was empty: models.custom did not exist. The tests were right and
the environment was incomplete, so the seed now states what the suite
depends on, with fictional model ids: a test should not pass because of
a setting somebody changed on the live system last week.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dv6sqaY6Vq3ChZHMem3cnU
Adds the admin fixture the Search Sources screen needed, and repairs the reason
no browser-driving e2e test could log in at all.
The sign-in failure first. The suite drove the app over http on a container
hostname, which is not a secure context, so the browser provides no
crypto.randomUUID. AccountBoundary calls it to mint a session generation on
every sign-in; the call threw, the boot handler's catch swallowed it, and every
test landed on the login screen holding a perfectly valid session. Measured:
isSecureContext false and randomUUID undefined on
http://pediatric-ai-scribe-e2e:3000, both true on http://127.0.0.1:3553, where
boundary.enter() returns true and the app enters.
Chrome's --unsafely-treat-insecure-origin-as-secure was tried first and does not
work: Playwright rejects the --user-data-dir it must be paired with, and the
flag alone leaves isSecureContext false. Loopback needs no flags, so the runner
now uses the host network and the published port.
The seed is new. The e2e user was a registration someone did by hand once that
the shared Postgres happened to keep — enough to log in and no more. There was
no admin account, so nothing under /api/admin could be tested through a real
request, which is how the Search Sources card came to be verified by reading its
markup. e2e/seed.js creates both accounts and reconciles an existing one, so a
leftover with the wrong role cannot fail the suite for a reason unrelated to the
code. It resets passwords and grants admin, so it refuses any address outside
@ped-ai.test. The runner seeds before it tests.
The new spec covers what markup-reading could not: that an ordinary account is
refused the settings and never offered the Admin menu item, that no API key
comes back readable, that the Test button reports each source separately, and
that every control the save handler reads exists in a real render. Each account
gets its own browser context, because AccountBoundary allows one owner per
document and freezing the page on a second is the behaviour, not a bug.
10/10 pass on both projects. Two unit tests pin the loopback requirement and the
seed's domain guard so neither can be undone quietly.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dv6sqaY6Vq3ChZHMem3cnU
The pathway existed but was reachable only by API. It now has a tab of its own
next to the Learning Hub — related, not the same thing, and sitting together is
how someone discovers the difference — visible to every signed-in user with no
role gate in the markup.
Generate a deck or an article, see everything you have made, download each as
PowerPoint, Word or PDF, delete what you no longer want. The screen says
"Private to you" and "Nobody else sees these", because the distinction from
published Learning content is the thing a person needs to understand before
typing a patient's condition into it.
Downloads are fetched rather than linked: an <a href> cannot carry the
Authorization header. The blob is saved under the filename the server chose and
the object URL is revoked afterwards. Resource titles come from a model, so rows
are built as elements and a title is only ever assigned to textContent.
The e2e stack now joins danvics_convert too. It could previously reach only
Postgres and Redis, so a PDF download failed there in a way production would
not — which did at least prove the degradation path works: with Gotenberg
unreachable the response is "PDF conversion is unavailable right now. PowerPoint
and Word still work", and the other two formats download unaffected.
Verified in a browser as an ordinary user: the tab appears and opens, the form
swaps slide count for word count when the format changes, the library lists
their own work, and pptx, docx and pdf all download with sensible filenames
(36360, 13285 and 68310 bytes).
Also documents retrieval sizing in docs/retrieval-tuning.md — the per-feature
budgets, and RERANKER_TOP_K, which caps all of them and had until now appeared
in no configuration file at all.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dv6sqaY6Vq3ChZHMem3cnU
8 new spec files covering sections previously only smoke-tested:
- ai-endpoints-contract.spec.js — hits 8 real AI endpoints via request
context and fails if the response leaks TypeError / ReferenceError /
'Cannot read properties of undefined' / 'is not defined' / 'is not a
function'. This is the class of bug that shipped the PE-narrative
regression to prod because every page-level mock prevented the real
handler from running.
- encounter-workflow.spec.js — generate HPI, refine, clear transcript.
- encounter-save-load.spec.js — save draft, load popover, repopulate.
- wellvisit-workflow.spec.js — byvisit, milestones, SSHADESS (12+
reveal), visit note.
- vaxschedule-content.spec.js — schedule + catch-up panels populate
beyond "Loading".
- chart-review-workflow.spec.js — generate + load popover.
- learning-tab.spec.js — search filter, category pills, feed.
- settings-faq-dictation.spec.js — voice/password/2FA/Nextcloud
sections, FAQ expand/collapse, dictation generate flow.
Baseline fixes:
- Added CORS_ORIGINS + API_RATE_LIMIT_MAX env overrides so the e2e
container accepts the browser's Origin header and can absorb the
full suite's API traffic without tripping the 200/min guard.
- Server's /api/ rate limit is now configurable via
API_RATE_LIMIT_MAX (default stays 200).
- extensions-crud: replaced native page.on('dialog') listeners with
#confirm-modal-ok clicks (we moved off native confirm()).
- pe-guide-smoke + extensions-crud: mobile viewport opens the hamburger
before clicking sidebar tabs.
- fixtures.js: /api/refine mock uses 'refined' (real API shape), not
'content'. /api/chart-review replaced with /api/generate-chart-review.
Suite: 230 passed / 0 failed in 4m36s.
User-visible changes:
- Removed the REAL / SYNTH badge from each sound card. User feedback:
"ridiculous". Cards now just show the sound title + description +
player, no distinction beyond the player type (native <audio controls>
for recordings, play/stop/progress bar for synth).
- Removed the "real recordings; synth labelled SYNTH" subtitle from
the sounds library header.
- Single-playback policy across the whole PE Guide: when any sound
starts (audio or synth), every other playing sound stops. Covers:
audio → audio (pause the previous), audio → synth, synth → audio,
synth → synth. Listeners attach on each .pe-audio 'play' event and
on every synth play-button click.
E2E infrastructure fixes (for the test failures we hit):
- auth-gated-smoke.spec.js now imports test + loginAs from the shared
fixtures.js so the token cache is unified across every spec. Without
this, each spec file's module-scoped _tokenCache multiplied logins
and hit the 10/15min rate limit.
- server.js: /api/auth/login rate limit is now configurable via
LOGIN_RATE_LIMIT_MAX env var (default 10, prod unchanged).
- docker-compose.e2e.yml: LOGIN_RATE_LIMIT_MAX="500" so Playwright's
two-project (chromium + mobile-chrome) multi-worker runs can do
their logins without tripping the cap. Prod container unaffected.
- fixtures.js console.error allowlist expanded to suppress known
non-bugs: Cross-Origin-Opener-Policy warnings on http:// e2e
server, transient 401/403/404/503 resource loads, ERR_BLOCKED_BY_CLIENT.
Adds a second containerized instance of the app with Turnstile + SMTP
disabled so Playwright can log in without a bot challenge.
- docker-compose.e2e.yml: pediatric-ai-scribe-e2e on port 3553. Shares
postgres + pgdata with main so seeded test users (*@ped-ai.test) persist.
- Test user: e2e-user@ped-ai.test (created once via /api/auth/register
against the e2e container — SMTP is off so register auto-verifies).
- Tests log in once per worker via /api/auth/login (module-scoped token
cache) then inject the ped_auth cookie into each test's browser context.
This avoids the 10-per-15-min login rate-limit.
- Mobile viewport opens the sidebar via #btn-menu-toggle before clicking
tab buttons (which are hidden behind the hamburger <=768px).
Coverage: encounter, wellvisit, chart, vaxschedule, catchup, learning,
dictation, settings, calculators, faq + landing-page-after-login. Each
test clicks the tab, waits for the lazy component to render (>100 chars),
and asserts a known anchor string is present.
Total suite: 128 tests passing (53 desktop + 53 mobile Bedside/top-calcs +
11 desktop + 11 mobile auth-gated).