pediatric-ai-scribe-v3/docs/browser-whisper-troubleshooting.md
Daniel 503f5afaad feat: ED multi-stage UX, extensions polish, docs viewer + application-logic docs
Three concurrent themes from this session:

═══════════════════════════════════════════════════════════════════
ED ENCOUNTERS — per-stage cards + consolidate→MDM finalize
═══════════════════════════════════════════════════════════════════

UX redesign per Daniel's feedback ("every stage note should be shown,
if AI is told to modify that particular note then the modified version
is used in final mdm"):

- Each generated stage stays on screen as its own editable card with
  its own embedded "Don't Miss" panel. No more single rolling note
  element that gets replaced on each generation.
- gatherCurrentNotes() reads contenteditable text from each stage card
  before any operation (advance, finalize, persist) so inline edits
  flow into the next AI call and the final consolidate.
- Stage badge is now state-accurate. "Stage N (recording)" with yellow
  background after Add-more before generation; "Stage N" with gray
  after generation. Fixes the bug where the badge flipped to Stage 2
  the moment Add-more was clicked.
- Save & Done now runs TWO server-side AI calls in /finalize:
  1. edConsolidate (new prompt) → polished single final note that
     integrates every stage chronologically (HPI / ROS / PE / ED Course /
     A&P with disposition).
  2. edFinalize (rewritten with full inline 2023 AMA E/M element
     rubric — problems / data / risk definitions, level mapping with
     concrete examples) → MDM JSON.
- Two new cards render after finalize: blue-bordered Final Consolidated
  Note + green-bordered MDM. Stage cards become read-only.
- partial_data on the saved row now stores {stages, finalNote, mdm,
  finalized} so resume re-renders the full state.

Why two-call finalize: a single combined prompt makes the model cut
corners on one task. Two focused calls cost ~2× latency at the very end
of an encounter — acceptable since finalize is a one-time terminal
action, not a per-stage hot path.

Files: public/components/ed-encounter.html, public/js/ed-encounters.js,
src/routes/edEncounters.js, src/utils/prompts.js (edConsolidate added,
edFinalize rewritten).

═══════════════════════════════════════════════════════════════════
EXTENSIONS / PAGERS — visual polish
═══════════════════════════════════════════════════════════════════

Multiple iterations based on Daniel's feedback:

- Layout: align-items:flex-start so action buttons stay pinned top-right
  when long numbers wrap (was align-items:center → buttons drifted into
  the text area, causing visible overlap).
- Number: word-break:break-all + min-width:0 + font-feature-settings:tnum
  so long numbers wrap within their column instead of pushing under the
  buttons. Click-to-copy with a 0.55s green flash + ✓ copied badge.
- Phone/pager Font Awesome icon next to the number in the type color —
  at-a-glance type signal (replacing an earlier 3px left stripe that
  Daniel found visually bulky).
- Name: font-weight 700, font-size 14.5px, color g900, letter-spacing
  -0.012em — scan-target headline typography for long lists.
- Alternating subtle backgrounds by index (white vs #fafbfc) so a long
  list reads as distinct rows.
- Hover: card lifts 1px with a soft shadow; action buttons fade from
  55% to 100% opacity. Cubic-bezier transition on transform.
- Entrance: staggered fade-up animation per card (35ms × index, capped
  at 12). prefers-reduced-motion media query disables motion.
- Empty state: 48px FA icon + heading instead of plain gray text.

Files: public/js/extensions.js, public/css/styles.css.

═══════════════════════════════════════════════════════════════════
DOCS REORGANIZATION + APPLICATION-LOGIC DOCS + ADMIN VIEWER
═══════════════════════════════════════════════════════════════════

Document moves (preserving git history via git mv):
  BROWSER_WHISPER_SETUP.md          → docs/browser-whisper-setup.md
  BROWSER_WHISPER_TROUBLESHOOTING.md → docs/browser-whisper-troubleshooting.md
  DEVELOPER_GUIDE.md                → docs/developer-guide-extended.md
  EMBEDDINGS_SETUP.md               → docs/embeddings-setup.md
  FEATURES_EXPLAINED.md             → docs/features-explained.md
  IMPROVEMENTS.md                   → docs/improvements.md
  OPENID_SETUP.md                   → docs/openid-setup.md
  TRANSCRIPTION_OPTIONS.md          → docs/transcription-options.md
README.md updated with the new paths + a Documentation section that
links to docs/logic/ at the top.

New application-logic doc series (~8,300 lines total) at docs/logic/.
Built with 5 parallel doc-writing agents per Daniel's "use multiple
agents" directive. Each doc explains how a part of the app actually
works — application logic, data flow, design decisions, sacred zones,
how-to-extend recipes — at a depth that lets a new dev (or an AI
assistant) modify the code confidently.

  docs/logic/README.md                — index + recommended reading order
  docs/logic/architecture.md (2166 L) — frontend IIFE pattern, lazy tab
                                         load, backend route convention,
                                         schema, encryption, deployment
  docs/logic/clinical-notes.md (1546L) — every note tab + helper trio
  docs/logic/bedside-and-calculators.md (1373L) — bedside ES module
                                         pocket + calculators + PE Guide
                                         + suture selector
  docs/logic/auth-admin-learning.md (1281L) — auth (local+OIDC+2FA) +
                                         admin panel + Learning Hub
                                         (Quiz engine logic at sub-detail
                                         only — TODO follow-up)
  docs/logic/ai-and-voice.md (1128 L) — callAI 5-provider routing,
                                         prompts, voice/STT, helper trio
  docs/logic/ed-encounters.md (821 L) — multi-stage ED + MDM (this
                                         session's worked example)

Admin-only docs viewer:
- New route /api/admin/docs/{tree,file}: recursively walks docs/, returns
  the tree as JSON; /file?path=X validates path stays inside docs/ and
  renders markdown via marked. Both gated by req.user.role==='admin'.
- New tab "Docs" (book icon) in the sidebar, hidden by default and
  revealed in auth.js when user.role==='admin' (same pattern as the
  existing Admin and CMS tabs).
- New component public/components/admin-docs.html: split-pane layout
  with a tree sidebar + filter input + a markdown reader pane.
- New module public/js/admin-docs.js: lazy-loads the tree on first tab
  activation, renders collapsible folders, persists expanded state and
  last-opened path via UIState. Server-rendered HTML so no client
  markdown parser needed.
- CSS for the viewer (responsive split-pane, code-block styling, table
  scrolling, etc.).
- Mounted at /api/admin/docs (NOT /api) — important: mounting a router
  with router.use(authMiddleware) at /api accidentally 401s every other
  /api/* path (caught and fixed during testing — /api/health was 401'ing).

Files: docs/* (moved + new), README.md, public/components/admin-docs.html
(new), public/js/admin-docs.js (new), src/routes/adminDocs.js (new),
public/index.html (tab + section + script), public/js/auth.js (admin
gate + logout cleanup), public/css/styles.css (viewer styles), server.js
(mount).

═══════════════════════════════════════════════════════════════════
KNOWN GAPS (TODO follow-ups)
═══════════════════════════════════════════════════════════════════

- Learning Hub quiz engine (MCQ / multi-select / T-F scoring + attempt
  tracking + progress dashboard) is covered at the architectural level
  in docs/logic/auth-admin-learning.md but not drilled into the quiz
  data model and scoring flow. Worth a focused follow-up doc.
- ED finalize: if MDM step JSON parse fails, server returns 502 with
  the consolidated finalNote in the error payload, but client doesn't
  surface the partial result. Add a "MDM failed, retry" affordance.
- No e2e Playwright coverage for ED encounters or the new docs viewer.
2026-04-28 03:09:38 +02:00

6.8 KiB

Browser Whisper Troubleshooting

🎙️ What is Browser Whisper?

Browser Whisper is an optional client-side transcription feature that runs entirely in your browser using WebAssembly. It provides:

  • Zero network transmission (HIPAA-safe)
  • No API costs
  • Works offline
  • Privacy-first (audio never leaves device)

However, it requires downloading AI models from CDN servers.


⚠️ Common Issue: CDN Blocked

Error Message:

NetworkError: Failed to execute 'importScripts' on 'WorkerGlobalScope':
The script at 'https://cdn.jsdelivr.net/npm/@xenova/transformers@2.17.2' failed to load.

What This Means:

Your network/firewall is blocking access to:

  • cdn.jsdelivr.net (JavaScript library CDN)
  • cdn-lfs.huggingface.co (AI model files)

Why It Happens:

  1. Corporate firewall - Many organizations block CDN domains
  2. Browser extensions - Ad blockers, privacy tools may block CDN
  3. Network proxy - Company proxy might filter JavaScript CDN
  4. CSP restrictions - Very strict Content Security Policy

Solutions

Browser Whisper is optional! The app works perfectly fine with server-side transcription.

Server transcription providers:

  • Google Gemini (via Vertex AI) - HIPAA-eligible
  • AWS Transcribe - HIPAA-eligible
  • OpenAI Whisper - Fast, accurate
  • LiteLLM - Routes to any provider

To use server transcription:

  1. Go to Settings → Browser Transcription
  2. Leave it disabled (or if stuck, disable it)
  3. Record audio normally - will use server

Advantages:

  • More accurate (larger models)
  • No download needed
  • Works immediately
  • Professional grade

Option 2: Whitelist CDN Domains

If you control your network/firewall, whitelist these domains:

cdn.jsdelivr.net
cdn-lfs.huggingface.co
cdn-lfs-us-1.huggingface.co
cdn-lfs-us-2.huggingface.co
huggingface.co

For corporate IT:

  • These are legitimate AI/JavaScript CDNs
  • Used by major companies worldwide
  • No security risk (public CDN content)
  • Required only for browser-based AI features

Option 3: Disable Browser Extensions

Try disabling:

  • Ad blockers (uBlock Origin, AdBlock Plus)
  • Privacy extensions (Privacy Badger, Ghostery)
  • Script blockers (NoScript, ScriptSafe)

Then refresh and try again.

Option 4: Try Different Browser

Some browsers have stricter security:

  • Chrome - Best compatibility
  • Edge - Works well
  • ⚠️ Firefox - May block CDN
  • Safari - Limited WebAssembly support

🧪 How to Test If It's Working

Test 1: Check CDN Access

# From your computer, run:
curl -I https://cdn.jsdelivr.net/npm/@xenova/transformers@2.17.2

# Should return: HTTP/2 200
# If 403 or timeout: CDN is blocked

Test 2: Browser Console

  1. Open DevTools (F12)
  2. Go to Console tab
  3. Settings → Browser Transcription
  4. Click "Pre-download model"
  5. Watch for:
    ✅ [WhisperWorker] Transformers library loaded successfully
    OR
    ❌ NetworkError: Failed to load
    

Test 3: Network Tab

  1. Open DevTools (F12)
  2. Go to Network tab
  3. Click "Pre-download model"
  4. Look for requests to:
    • cdn.jsdelivr.net (should be 200 OK)
    • cdn-lfs.huggingface.co (should be 200 OK)
  5. If blocked: Status will show "failed" or "blocked"

📊 When to Use Each Option

Scenario Recommendation Why
Corporate network Server transcription CDN likely blocked
Home network Browser Whisper Fast, free, private
Mobile device Server transcription Limited storage/memory
Offline use needed Browser Whisper Works without internet (after initial download)
High accuracy needed Server transcription Larger models available
Maximum privacy Browser Whisper Audio never leaves device
Can't access CDN Server transcription No choice - CDN blocked

🔧 Technical Details

What Gets Downloaded (First Time Only):

Tiny model (~39 MB):

  • onnx-runtime.wasm (~10 MB)
  • whisper-tiny.en model files (~29 MB)
  • Cached in browser IndexedDB (permanent)

Base model (~74 MB):

  • Larger model, better accuracy

Small model (~244 MB):

  • Best quality, slower processing

Where It's Stored:

  • Location: Browser IndexedDB
  • Persistence: Permanent (until you clear browser data)
  • Shared: Across all tabs/windows for this domain
  • Size: Selected model size (39/74/244 MB)

Performance:

  • Tiny: 2-3 seconds per 30-second clip
  • Base: 3-5 seconds per 30-second clip
  • Small: 6-10 seconds per 30-second clip

FAQ

Q: Is Browser Whisper required? A: No! It's completely optional. Server transcription works great.

Q: Why doesn't it work on my corporate network? A: Most corporate firewalls block CDN domains for security. Use server transcription instead.

Q: Can I download the models manually? A: Not easily - they're optimized for CDN delivery. Use server transcription if CDN is blocked.

Q: Will server transcription cost money? A: Depends on your provider:

  • Google Vertex AI: ~$0.005 per minute
  • AWS Transcribe: ~$0.024 per minute
  • OpenAI: $0.006 per minute
  • Very affordable for typical use

Q: Is server transcription HIPAA-safe? A: Yes, if using:

  • Google Vertex AI (with BAA)
  • AWS Transcribe (with BAA)
  • Azure OpenAI (with BAA)

OpenAI Whisper direct is NOT HIPAA-eligible.

Q: Can I use both? A: Yes! Enable Browser Whisper in Settings. If it fails (CDN blocked), it automatically falls back to server transcription.

Q: How do I know which one is being used? A: Check the toast notification after recording:

  • "Transcribed locally" = Browser Whisper
  • "Transcribed via google-gemini/aws/openai" = Server

For Maximum Privacy (Home Network):

  1. Enable Browser Whisper
  2. Choose "Tiny" model (fast, good enough for dictation)
  3. Pre-download model
  4. Use offline

For Corporate/Clinical Use:

  1. Keep Browser Whisper disabled
  2. Configure server transcription:
    # In .env:
    TRANSCRIBE_PROVIDER=google
    GOOGLE_VERTEX_PROJECT=your-project
    
  3. Use with BAA for HIPAA compliance

For Best Accuracy:

  1. Use server transcription
  2. Configure Google Gemini 2.0 Flash or AWS Transcribe Medical
  3. Audio quality + large models = best results

🛠️ Still Having Issues?

  1. Check console logs: DevTools → Console → Look for [BrowserWhisper] errors
  2. Check network logs: DevTools → Network → Filter by jsdelivr or huggingface
  3. Verify server transcription works: Just disable Browser Whisper and record
  4. Contact IT: Ask to whitelist CDN domains (if you need Browser Whisper)

Remember: Browser Whisper is a nice-to-have feature. Server transcription is the primary, production-ready method that works everywhere!