pediatric-ai-scribe-v3/docs/speech.md
Daniel f89dc01729 refactor: one place decides which bucket, on which S3, with which credentials
Three S3 configurations had grown separately — S3_* for documents,
GENERATED_IMAGES_S3_* for images, and AUDIO_BACKUPS_S3_* after them — with
different key names and their own client construction. That is why moving
storage meant hunting through several files.

src/utils/objectStorage.js now resolves settings for any purpose: its own
variables first, then the shared S3_* ones, with a per-purpose bucket name
(S3_BUCKET_AUDIO_BACKUPS). One endpoint plus three bucket names is enough
for the whole app, and a purpose that needs its own account still overrides
everything. Audio backups and documents use it; generated images keeps its
own tested storage module, whose variable names the resolver already
understands.

Nothing existing has to change: S3_ACCESS_KEY_ID, S3_SECRET_ACCESS_KEY and
the AWS_* fallbacks still resolve, and path-style addressing keeps each
purpose's previous default — off for documents, so a Backblaze endpoint
behaves as before, on where a custom endpoint implies MinIO. A _FILE
credential now always beats an inline one, so a mounted secret cannot be
shadowed by an inherited environment variable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dv6sqaY6Vq3ChZHMem3cnU
2026-09-10 17:09:13 +02:00

7.2 KiB

Speech: STT, TTS, Audio Backup

Transcription

POST /api/transcribe accepts multipart/form-data with one audio file up to 25 MB. Server STT is routed through LiteLLM.

Set TRANSCRIBE_PROVIDER=litellm, LITELLM_API_BASE, and LITELLM_STT_MODEL. Auto mode also uses LiteLLM when the gateway is configured.

The model for a request is users.stt_model (the person's own choice), then the stt.model setting, then LITELLM_STT_MODEL. A user's choice wins, so the Settings picker must never offer a model the gateway does not have: it lists what /model/info advertises as audio_transcription (discoverSTTModels() in src/utils/sttProvider.js, cached five minutes). LITELLM_STT_MODELS is a fallback for when discovery fails and is not a list of models known to work — its ids do not resolve on every gateway.

Provider Notes HIPAA posture
LiteLLM Sends audio through the configured LiteLLM /audio/transcriptions backend. Depends on the selected upstream.

Browser Whisper and browser-local Whisper workers are not part of the runtime. Do not add browser model downloads or Transformers.js STT back into the public app.

Web Speech Preview

Browser-native Web Speech can show interim text when the user explicitly enables it. It is browser/vendor dependent, may send audio to browser-provider cloud services, and should not be treated as the final clinical transcript.

Text To Speech

POST /api/text-to-speech returns audio from LiteLLM /audio/speech. The X-TTS-Provider response header identifies the LiteLLM model used. Requests are limited to 5000 characters.

Provider Notes
LiteLLM Uses LITELLM_TTS_MODEL and LITELLM_TTS_VOICE.

The admin/user voice pickers read available LiteLLM-compatible voices from LITELLM_TTS_VOICES.

Recording

A recording holds a screen wake lock for as long as it runs, so the screen going to sleep cannot suspend it. Browsers release the lock whenever the page is hidden, so it is requested again when the page becomes visible; a lock is never requested while hidden, because that request is rejected. The count is shared, so two recorders cannot release each other's lock. A browser that denies or lacks the API keeps recording without one.

Recording continues while the user moves around the workspace — the app is a single page and switching tabs does not touch the recorder. Signing out stops it and releases the lock; nothing is sent, because the session that owned the audio is gone.

AudioRecorder.start() on a running recorder is a no-op. Calling it again would replace the MediaRecorder and drop everything captured so far.

A recording can stop without anyone pressing Stop: the recorder can error, and the microphone can be claimed by another app, unplugged, or revoked. Both are reported once and dispatch audio-recorder-failed, which runs the same path as Stop — so the audio is transcribed and stored rather than left in a tab that still claims to be recording. Whatever was captured before the failure is kept.

Audio Backup

Every recording is kept for 24 hours, whether its transcription succeeded or not. POST /api/transcribe already holds the audio, so keeping it costs no second upload; a storage failure there is logged and the transcription still returns, because losing the transcript someone is waiting for would be worse than losing the copy.

src/utils/audioBackupStore.js is the only place that knows how a recording is kept, shared by /api/transcribe and /api/audio-backups so the two cannot drift apart.

  • Audio is gzipped, then AES-256-GCM encrypted (DATA_ENCRYPTION_KEY), before it is stored — in either backend.
  • Object storage is used when AUDIO_BACKUPS_S3_* is set; otherwise the payload goes in the audio_backups.audio_data column. Metadata (owner, module, sizes, expiry) is always in Postgres, so listing, ownership and expiry behave the same either way.
  • Object keys are recordings/<user id>/<timestamp>-<random>, scoped to their owner, so a leaked id cannot address someone else's audio.
  • Reads carry user_id and expires_at > NOW() in the query, so an expired or borrowed id reads as missing rather than as another person's recording.
  • The expiry sweep (cleanupExpired in src/db/database.js) deletes each object with its row, so audio cannot outlive its 24 hours in the bucket.
  • Rows written before encryption was added are passed through unencrypted on read, so old backups still play.
  • The Settings list can retry, download or delete a recording. Download is how a copy leaves the app — onto a phone's Files, a shared drive, an external recorder. It works for both server-side and browser-fallback copies, and a browser copy is only handed to the account that owns it.
  • Browser fallback storage (IndexedDB) is used only when the server cannot be reached at all.

Treat audio backups as sensitive clinical data even when encrypted.

Where object storage settings come from

src/utils/objectStorage.js resolves them for every purpose the same way, so moving the app to a different MinIO — or to a real S3 — is one set of variables rather than three schemes. For a purpose (documents, generated-images, audio-backups) it reads, in order:

  1. that purpose's own variables — AUDIO_BACKUPS_S3_ENDPOINT, ..._BUCKET, ..._REGION, ..._ACCESS_KEY[_FILE], ..._SECRET_KEY[_FILE];
  2. the shared ones — S3_ENDPOINT, S3_REGION, S3_ACCESS_KEY[_FILE], S3_SECRET_KEY[_FILE];
  3. a per-purpose bucket — S3_BUCKET_AUDIO_BACKUPS, S3_BUCKET_GENERATED_IMAGES.

So one endpoint plus three bucket names covers everything, while a purpose that needs its own account overrides all of it. A _FILE variant always beats an inline value, because a mounted secret should not be shadowed by an inherited environment variable. No bucket means "not configured", which is never an error — all three are optional.

Every name previously accepted still works, including S3_ACCESS_KEY_ID, S3_SECRET_ACCESS_KEY and the AWS_* fallbacks for documents, and GENERATED_IMAGES_S3_*. Path-style addressing keeps each purpose's old default (off for documents, so Backblaze keeps working) unless S3_FORCE_PATH_STYLE says otherwise; a custom endpoint turns it on where there was no older default, because that is nearly always MinIO.

Switching audio backups to MinIO

Storing every recording, rather than only the failures, makes object storage the better home. It is off by default because it needs a bucket and its own credentials; the existing generated-images key is scoped to that bucket and cannot reach another.

  1. Create an audio-backups bucket and a user with read/write/delete on it.
  2. Give the bucket a 24-hour expiry lifecycle rule, as a backstop for objects the sweep could not delete.
  3. Set AUDIO_BACKUPS_S3_ENDPOINT, AUDIO_BACKUPS_S3_BUCKET, AUDIO_BACKUPS_S3_REGION, and either AUDIO_BACKUPS_S3_ACCESS_KEY_FILE / AUDIO_BACKUPS_S3_SECRET_KEY_FILE (preferred — credentials stay out of the process environment) or AUDIO_BACKUPS_S3_ACCESS_KEY / AUDIO_BACKUPS_S3_SECRET_KEY.

New recordings then go to the bucket. Rows already in Postgres keep working: a row without storage_key is read from the column.