Commit graph

14 commits

Author SHA1 Message Date
Daniel
b3d66caaca feat: a retried transcript goes back to the tab the audio came from
Retrying a kept recording transcribed it and put the text on the clipboard,
leaving you to find the right tab and paste. The app already had the answer: the
module was recorded with the audio. Retry now opens that tab and puts the text
in its transcript box.

Two things had to be true first.

The module was not actually being recorded. transcribeAudio never sent one, so
the server stored its default for every upload — all 28 rows in audio_backups
said "recording", and a retry had nowhere to send anything back to. Each
module's call now tags its own upload.

And the names disagreed. The recorders tagged 'encounter', 'soap', 'dictation'
while the recording-started events said 'enc', 'sick', 'dict'. One table now
holds the mapping and resolves the aliases, so the recorder that tags the
upload, the backup row that labels it and the retry that delivers it cannot
drift apart again.

Existing text is appended to, never replaced: a retry usually recovers
something on top of a live transcript, and overwriting would lose the words the
browser did hear. The box only exists once its tab's markup has been fetched, so
delivery polls briefly rather than guessing a delay, and falls back to the
clipboard if the tab never opens. An empty result says so rather than claiming
success. Backup rows now name their source and the button reads "Retry into
SOAP Note" instead of "Retry".

Verified in a browser: enc resolves to encounter, delivery switched tabs and
produced 'existing live transcript\n\nRECOVERED TEXT'.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dv6sqaY6Vq3ChZHMem3cnU
2026-09-11 12:39:50 +02:00
Daniel
fed4bd154f fix: recordings that produced nothing, and the boxes that zoomed on iOS
All checks were successful
Forgejo Android APK / Root app tests (push) Successful in 49s
Forgejo Android APK / Build signed APK (push) Successful in 2m8s
Measured in a real browser against the app rather than reasoned about.

The transcript boxes are contenteditable divs, and an editable div zooms on
focus exactly like an <input>. The earlier 16px sweep covered input, textarea
and select, so every workspace tab still zoomed while the calculators did not —
which is exactly what was reported. Every focusable text control in every tab
now measures 16px at phone width; the count of ones below it is zero.

Three ways a recording could end with nothing to show for it:

  - Safari supports none of the audio/webm types and throws NotSupportedError
    when handed one. Six modules built their own recorder on resume with
    "opus, else audio/webm", so resuming threw there and the recording stopped.
    There is now one codec chain in the app, and no module constructs a
    MediaRecorder of its own.

  - audio-recorder-failed is dispatched on document, and the encounter tab
    stopped its recording on any of them. The assistant's microphone failing
    ended a consultation being recorded in another tab. The recorder now
    travels with the event and the listener checks it is its own.

  - The server answers {success:true, text:''} for silence, and five modules
    assigned that straight into the transcript — emptying the box the browser
    had been filling live. It reads as a recording that vanished. Text is now
    required before overwriting, and a recording that captured nothing says so
    instead of resetting the button over an empty box.

Also: the citation counters were registered on prom-client's default registry
while the app serves its own, so they were never scraped. They read zero at
/metrics now instead of being absent, which is what the Grafana panels need.

And the reference linter passes for the first time, so scripts/e2e.sh gets past
its preflight: KaTeX is vendored (it was referenced by the assistant's LaTeX
rendering but never shipped — three 404s a page load and no math), and the
JavaScript left behind by the removed image picker, saved-chats toggle, image
gallery and visual-output panel is gone.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dv6sqaY6Vq3ChZHMem3cnU
2026-09-11 00:16:19 +02:00
Daniel
f89dc01729 refactor: one place decides which bucket, on which S3, with which credentials
Three S3 configurations had grown separately — S3_* for documents,
GENERATED_IMAGES_S3_* for images, and AUDIO_BACKUPS_S3_* after them — with
different key names and their own client construction. That is why moving
storage meant hunting through several files.

src/utils/objectStorage.js now resolves settings for any purpose: its own
variables first, then the shared S3_* ones, with a per-purpose bucket name
(S3_BUCKET_AUDIO_BACKUPS). One endpoint plus three bucket names is enough
for the whole app, and a purpose that needs its own account still overrides
everything. Audio backups and documents use it; generated images keeps its
own tested storage module, whose variable names the resolver already
understands.

Nothing existing has to change: S3_ACCESS_KEY_ID, S3_SECRET_ACCESS_KEY and
the AWS_* fallbacks still resolve, and path-style addressing keeps each
purpose's previous default — off for documents, so a Backblaze endpoint
behaves as before, on where a custom endpoint implies MinIO. A _FILE
credential now always beats an inline one, so a mounted secret cannot be
shadowed by an inherited environment variable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dv6sqaY6Vq3ChZHMem3cnU
2026-09-10 17:09:13 +02:00
Daniel
80d468a85d feat: a recording survives switching to the Assistant, and signing out keeps it
Switching to the Assistant set window.location, which reloads the document
and silently ended any running recording. The assistant is a tab in the same
page and activateTab already rewrites the URL to /assistant, so the switch
now happens in place; the reload stays as a fallback. Verified in Chromium:
recorder still running, no reload, URL /assistant, assistant visible.

Signing out mid-recording used to end it with nothing kept. It now says so
first — "the audio will be saved for 24 hours so you can transcribe it
later" — and stores the audio either way, tagged with the module that
produced it ('encounter', 'soap', 'dictation'), which is what makes it
findable in Settings afterwards. Sign-out completes whether or not the save
worked, and the rescue never rejects, because the caller is on its way out.
Verified: the warning appears and the upload carries module=encounter.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dv6sqaY6Vq3ChZHMem3cnU
2026-09-10 16:54:01 +02:00
Daniel
523926ab17 feat: keep the screen awake while recording, and keep every recording 24h
Recording
- A screen wake lock is held for as long as a recording runs. Browsers drop
  the lock whenever the page is hidden, so it is taken again on return —
  without that, one glance away ended it for the session. The lock is
  reference counted (two recorders cannot release each other's), never
  requested while hidden (the request would just be rejected), and a denial
  or an unsupported browser leaves the recording running.
- Signing out releases it and stops the recording; nothing is sent, because
  the session that owned the audio is gone.
- start() on an already-running recorder is now a no-op instead of replacing
  the MediaRecorder and silently dropping everything captured so far.
- A recording that ends by itself — recorder error, or the microphone taken
  by another app, unplugged or revoked — takes the same path as pressing
  Stop, so it is transcribed and stored rather than left in a tab that still
  says "recording". Moving around the workspace already kept recording.

Retention
- Every recording is kept for 24 hours now, not only the ones whose
  transcription failed. /api/transcribe already has the audio, so this costs
  no second upload, and a storage failure is logged rather than thrown: it
  must never lose the transcription someone is waiting for.
- One store (src/utils/audioBackupStore.js) is shared by /api/transcribe and
  /api/audio-backups so the two cannot drift. Payload goes to object storage
  when AUDIO_BACKUPS_S3_* is set and to the encrypted Postgres column
  otherwise; metadata always stays in Postgres, so listing, ownership and
  expiry behave the same either way. Object keys are scoped by owner, and
  the expiry sweep deletes the object with the row.

Verified against the live database: round trip byte-identical, another user
reads null, 950 -> 48 bytes compressed, expired rows take their objects.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dv6sqaY6Vq3ChZHMem3cnU
2026-09-10 16:42:31 +02:00
Daniel
48a3b06ebb feat: export a recording; report a recorder that has silently died
Export
- Both stores already held the audio (server /audio-backups/:id/audio and
  the local IndexedDB record) but nothing exposed it, so a recording could
  not be taken out of the app. Each backup row now has a download, naming
  the file by its timestamp and using the extension actually recorded
  (webm, or m4a on iOS). A local record is only handed over to the account
  that owns it.

Robustness
- MediaRecorder had no onerror and nothing watched the audio track, so a
  recorder that failed, or a microphone claimed by another app, unplugged,
  or revoked, left the tab saying "recording" while capturing nothing.
  Both are now reported once, with the chunks captured so far kept, so
  stopping still returns the audio up to the failure.

Deliberately not added: a wake lock. Stopping when the screen sleeps or the
session ends is the intended behaviour — recording is meant to be
deliberate, and nothing is left behind.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dv6sqaY6Vq3ChZHMem3cnU
2026-09-10 15:58:58 +02:00
Daniel
f0f48a3578 fix: patch nodemailer; Settings offers STT models the gateway really has
Security
- nodemailer 9.0.1 -> 9.1.1, clearing four high advisories, two of which
  are delivery bugs that matter for an app that sends mail: recipient-domain
  validation bypass via RFC 5322 comments, and an IDN/punycode allow-list
  bypass, both of which can route mail to an attacker-controlled domain.

Live transcription
- The Settings picker was a hardcoded list of six ids
  (local-whisper-*, local-parakeet-v3, gemini-*). None of them resolve on
  this gateway, and /api/transcribe prefers the user's choice over the admin
  default, so picking one broke every recording with "Invalid model name".
  Verified against the live gateway: local-whisper-large-v3-turbo -> 400.
- The picker now lists what /model/info advertises as audio_transcription,
  cached for five minutes, with the built-in list kept only as a fallback
  and the admin default marked.
- The pipeline itself is healthy: local-kokoro-tts produced 92KB of speech
  and mistral-voxtral-mini-transcribe returned the sentence back verbatim.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dv6sqaY6Vq3ChZHMem3cnU
2026-09-10 15:54:57 +02:00
Daniel
a08524d95f harden logging and observability 2026-05-08 19:08:19 +02:00
Daniel
6da5565f89 harden audio backup settings rendering 2026-05-08 08:23:38 +02:00
Daniel
ea52890908 gate browser speech recognition setting 2026-05-08 08:05:23 +02:00
Daniel
bc0b43151d convert settings helpers to modules 2026-05-08 07:54:08 +02:00
Daniel
1756043125 test note refine correction policy 2026-05-08 07:46:39 +02:00
Daniel
03ee07f92f clarify template-only AI memory context 2026-05-08 07:38:27 +02:00
Daniel
20fc6798a9 remove browser whisper transcription 2026-05-08 07:33:12 +02:00