The unit tests for this feature read source files and assert patterns. They
prove the code says the right thing, not that the screen does it, and nothing
exercised the browser at all — so a mismatch between what the form sends and
what the route reads passed all of them.
Three real bugs shipped through that gap in one session: a modification that
updated the markdown but not the deck, generation that failed whenever the slide
reviewer was off, and a figure generated for a slide that never referenced it.
Every one was found by driving the running server by hand.
So these assert the request bodies, not only the rendering: that Generate sends
topic, kind, slideCount, refinement, model and all four options as the strings
the route compares against; that unticking the library sends 'false' rather than
omitting the field, which the route would read as on; and that Modify posts to
the right resource with every source option. Plus the screen's own behaviour —
availability gating on both cards, the illustration hint switching on and
staying off once overruled, the bounded searchable library, the two different
empty states, an article never being offered as slides, a local refusal that
spends no round trip, and a refused modification surfacing its reason. Fourteen
tests, both viewports.
The API is stubbed. This is the contract between the screen and the route, and
stubbing keeps it fast, free and deterministic.
Proven to catch regressions rather than merely pass: renaming useCorpus in the
form failed two tests, breaking the availability gating failed one, and
truncating the modify picker failed another.
Two flakes of my own were fixed rather than retried. openTab slept 400ms for the
library and picker instead of waiting for them, which made Modify report
"nothing to modify yet" under load. And the console-error guard failed on
net::ERR_ABORTED and net::ERR_NETWORK_CHANGED — a request in flight when the
context closes, and the host network reconfiguring under a browser that runs on
it. Both are the harness, not the page: anything genuinely failing carries a
status code and is still caught. Five consecutive clean full runs after.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dv6sqaY6Vq3ChZHMem3cnU
Three separate reasons tests were failing, none of them a defect in the app.
The calculators. e2e-harness.html loaded calculators.js and drugs-loader.js with
`defer` after they were split into ES modules; index.html was updated at the
time and this page was not. A module parsed as a classic script throws "Cannot
use import statement outside a module" before a line runs, so no click handler
was ever attached: the pills rendered from static HTML and did nothing. Only the
first calculator appeared to pass, because it carries `active` in the markup and
needs no click. That was 52 failures.
Settings and FAQ. Both moved from the tab rail into the account-card menu; the
helper still clicked button.tab-btn[data-tab=…] and timed out. Ten more.
The AI mocks, which had stopped intercepting for two independent reasons and so
were calling the real model on every run — spending credits and comparing
genuine output against strings like "MOCK HPI from dictation". A '**/api/x' glob
matches no URL on Playwright 1.50, and page.route fails silently when nothing
matches; measured against a real URL, that glob and '*/**/api/x' both matched
zero times where a regex matched. Fixing that alone was not enough: the app
registers a service worker that answers every /api/ request with its own
fetch(), and a request made inside a service worker never reaches page.route.
Blocking registration in the config puts them back in the page. The mocked
dictation test now finishes in 1.6s rather than 7.5s, which is what a real model
call costs.
Whole suite: 204 passed / 96 failed in 15.8 minutes, now 289 passed / 11 failed
in 6.8. The remaining eleven are spread across nine specs with no shared cause
and are not touched here.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dv6sqaY6Vq3ChZHMem3cnU
Adds the admin fixture the Search Sources screen needed, and repairs the reason
no browser-driving e2e test could log in at all.
The sign-in failure first. The suite drove the app over http on a container
hostname, which is not a secure context, so the browser provides no
crypto.randomUUID. AccountBoundary calls it to mint a session generation on
every sign-in; the call threw, the boot handler's catch swallowed it, and every
test landed on the login screen holding a perfectly valid session. Measured:
isSecureContext false and randomUUID undefined on
http://pediatric-ai-scribe-e2e:3000, both true on http://127.0.0.1:3553, where
boundary.enter() returns true and the app enters.
Chrome's --unsafely-treat-insecure-origin-as-secure was tried first and does not
work: Playwright rejects the --user-data-dir it must be paired with, and the
flag alone leaves isSecureContext false. Loopback needs no flags, so the runner
now uses the host network and the published port.
The seed is new. The e2e user was a registration someone did by hand once that
the shared Postgres happened to keep — enough to log in and no more. There was
no admin account, so nothing under /api/admin could be tested through a real
request, which is how the Search Sources card came to be verified by reading its
markup. e2e/seed.js creates both accounts and reconciles an existing one, so a
leftover with the wrong role cannot fail the suite for a reason unrelated to the
code. It resets passwords and grants admin, so it refuses any address outside
@ped-ai.test. The runner seeds before it tests.
The new spec covers what markup-reading could not: that an ordinary account is
refused the settings and never offered the Admin menu item, that no API key
comes back readable, that the Test button reports each source separately, and
that every control the save handler reads exists in a real render. Each account
gets its own browser context, because AccountBoundary allows one owner per
document and freezing the page on a second is the behaviour, not a bug.
10/10 pass on both projects. Two unit tests pin the loopback requirement and the
seed's domain guard so neither can be undone quietly.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dv6sqaY6Vq3ChZHMem3cnU
session persistence, per-tab model selector
Brings detailed coverage to sections that were previously only smoke-
tested:
- sickvisit-workflow.spec.js: generate → mocked note, refine bar
round-trip, load popover, New button clears demographics + transcript.
- hospitalcourse-workflow.spec.js: fill H&P → generate → mocked
narrative renders via /api/generate-hospital-course, refine updates
text, load popover open/close.
- auth-screen.spec.js: unauthenticated landing visible, login form
structure, forgot-password swap + return, register link is
intentionally display:none on this instance (pinned), register form
DOM still wired correctly if the link is manually unhidden, reg-
password has minlength=8 + type=password.
- session-persistence.spec.js: clear cookie simulates logout (UI login
is gated by Turnstile which can't be completed in the e2e container);
re-login via loginAs restores the last tab + sub-pill via localStorage.
- model-selector.spec.js: tab-model-select dropdowns render with >0
options across 7 tabs.
Fixture corrections:
- /api/sick-visit/note and /api/well-visit/note were not being mocked
at all — the wrong /api/generate-sick-visit pattern was intercepting
nothing, so tests fell through to the real AI backend. Both now have
correct patterns and response shapes (sick-visit returns `note`,
well-visit note returns `note`).
- /api/generate-hospital-course response key updated from `narrative`
to `hospitalCourse` to match what the frontend actually reads.
wellvisit-workflow: the Visit Note test now asserts a concrete
waitForResponse on /api/well-visit/note + text render, instead of the
previous "hit either endpoint" fallback.
Suite: 294 passed / 0 failed in 5m30s.
8 new spec files covering sections previously only smoke-tested:
- ai-endpoints-contract.spec.js — hits 8 real AI endpoints via request
context and fails if the response leaks TypeError / ReferenceError /
'Cannot read properties of undefined' / 'is not defined' / 'is not a
function'. This is the class of bug that shipped the PE-narrative
regression to prod because every page-level mock prevented the real
handler from running.
- encounter-workflow.spec.js — generate HPI, refine, clear transcript.
- encounter-save-load.spec.js — save draft, load popover, repopulate.
- wellvisit-workflow.spec.js — byvisit, milestones, SSHADESS (12+
reveal), visit note.
- vaxschedule-content.spec.js — schedule + catch-up panels populate
beyond "Loading".
- chart-review-workflow.spec.js — generate + load popover.
- learning-tab.spec.js — search filter, category pills, feed.
- settings-faq-dictation.spec.js — voice/password/2FA/Nextcloud
sections, FAQ expand/collapse, dictation generate flow.
Baseline fixes:
- Added CORS_ORIGINS + API_RATE_LIMIT_MAX env overrides so the e2e
container accepts the browser's Origin header and can absorb the
full suite's API traffic without tripping the 200/min guard.
- Server's /api/ rate limit is now configurable via
API_RATE_LIMIT_MAX (default stays 200).
- extensions-crud: replaced native page.on('dialog') listeners with
#confirm-modal-ok clicks (we moved off native confirm()).
- pe-guide-smoke + extensions-crud: mobile viewport opens the hamburger
before clicking sidebar tabs.
- fixtures.js: /api/refine mock uses 'refined' (real API shape), not
'content'. /api/chart-review replaced with /api/generate-chart-review.
Suite: 230 passed / 0 failed in 4m36s.
User-visible changes:
- Removed the REAL / SYNTH badge from each sound card. User feedback:
"ridiculous". Cards now just show the sound title + description +
player, no distinction beyond the player type (native <audio controls>
for recordings, play/stop/progress bar for synth).
- Removed the "real recordings; synth labelled SYNTH" subtitle from
the sounds library header.
- Single-playback policy across the whole PE Guide: when any sound
starts (audio or synth), every other playing sound stops. Covers:
audio → audio (pause the previous), audio → synth, synth → audio,
synth → synth. Listeners attach on each .pe-audio 'play' event and
on every synth play-button click.
E2E infrastructure fixes (for the test failures we hit):
- auth-gated-smoke.spec.js now imports test + loginAs from the shared
fixtures.js so the token cache is unified across every spec. Without
this, each spec file's module-scoped _tokenCache multiplied logins
and hit the 10/15min rate limit.
- server.js: /api/auth/login rate limit is now configurable via
LOGIN_RATE_LIMIT_MAX env var (default 10, prod unchanged).
- docker-compose.e2e.yml: LOGIN_RATE_LIMIT_MAX="500" so Playwright's
two-project (chromium + mobile-chrome) multi-worker runs can do
their logins without tripping the cap. Prod container unaffected.
- fixtures.js console.error allowlist expanded to suppress known
non-bugs: Cross-Origin-Opener-Policy warnings on http:// e2e
server, transient 401/403/404/503 resource loads, ERR_BLOCKED_BY_CLIENT.
Two additions:
1. e2e/fixtures.js — shared test infrastructure
- Custom `test` extending @playwright/test with two auto-fixtures on
every page: page.on('pageerror') and page.on('console') of type
'error'. Any uncaught JS error fails the test. This is the SSO-
bug-class safety net: if we'd had this earlier, the silent
admin.js ReferenceError would have failed CI instead of shipping.
- `authedPage` fixture — logs in via API, injects session cookie,
provides pre-authed page ready to drive.
- `mockAI(page)` helper — intercepts generate-* and transcribe
endpoints with canned JSON responses. Enables fast deterministic
CI runs. Opt-out via E2E_USE_REAL_AI=1 to hit real LiteLLM.
- Console-error allowlist for known noise (favicon 404, lazy-loaded
Whisper models, etc.).
2. e2e/tests/extensions-crud.spec.js — 11 tests covering the new
Pagers & Extensions feature end-to-end:
- empty state, add extension, add pager (correct grouping)
- edit persists, search by location + by number
- soft-delete with confirm → moves to trash
- restore from trash → reappears in active
- purge from trash → permanent
- cancel dialog keeps item, cancel form keeps nothing, validation
on required fields
Known follow-up: the e2e container (pediatric-ai-scribe-e2e) is still
on the pre-entrypoint image from before the OpenBao migration, so it
doesn't have the Extensions routes yet. Rebuilding it needs a small
entrypoint enhancement to honor docker-compose-level env overrides
vs OpenBao-fetched values (e.g. TURNSTILE_SECRET_KEY="" for the e2e
instance). That's separate work — this commit just lays in the tests.