pediatric-ai-scribe-v3/TRANSCRIPTION_OPTIONS.md

8.2 KiB

Transcription Options Guide

Overview

Pediatric AI Scribe v2+ offers three transcription methods, allowing you to choose between privacy, speed, and real-time feedback.


📊 Comparison Table

Feature Browser Whisper Server Transcription Web Speech API
Privacy 100% offline (with BAA) Sends to cloud
Accuracy Whisper Gemini/AWS Browser-dependent
Speed 2-10s ~1s Instant
Real-time Batch mode Batch mode Live streaming
HIPAA Yes (Vertex/AWS) No
Cost Free ~$0.005/min Free
Internet Not required Required Required
Setup None (bundled) API keys None (built-in)

What It Is

  • Runs OpenAI Whisper entirely in your browser using WebAssembly
  • Audio never leaves your device - 100% offline after initial page load
  • Models bundled in Docker image (self-hosted, no CDN)

When to Use

  • Clinical documentation (HIPAA-compliant)
  • Maximum privacy required
  • Offline/air-gapped environments
  • No API costs
  • Zero vendor dependency

How to Enable

  1. Settings → Browser Transcription
  2. Toggle "Enable browser transcription" ON
  3. (Optional) Click "Pre-download model" if you want to cache it first
  4. Start recording - transcription happens automatically after recording

Models Available

  • Tiny (~39MB) - Fast, good for short clips (2-3 seconds)
  • Base (~74MB) - Balanced accuracy and speed (3-5 seconds)
  • Small (~244MB) - Best quality, slower (6-10 seconds)

Performance

  • Transcribes ~30-second clip in 2-10 seconds (depending on model)
  • First run may be slower (model loading)
  • Subsequent runs are instant (cached)

Privacy

  • Audio never transmitted
  • Models run locally in WASM
  • No network calls during transcription
  • HIPAA-compliant

Option 2: Server Transcription (Cloud, Fast)

What It Is

  • Sends audio to your configured AI provider
  • Uses Google Gemini, AWS Transcribe, OpenAI Whisper, or LiteLLM

When to Use

  • Maximum speed (~1 second for 30-second clip)
  • Best accuracy (cloud models)
  • Long recordings (Browser Whisper can be slow for 5+ minutes)
  • HIPAA-compliant with BAA providers

HIPAA-Eligible Providers

  • Google Vertex AI (with BAA)
  • AWS Transcribe (with BAA)
  • Azure OpenAI (with BAA)
  • OpenAI Whisper Direct Not HIPAA-eligible

How to Enable

  • Configured via environment variables (.env)
  • No user action needed - just works if API keys present
  • Falls back automatically if Browser Whisper fails

Cost

  • Google Gemini: ~$0.005/minute
  • AWS Transcribe: ~$0.024/minute
  • OpenAI: $0.006/minute

Option 3: Web Speech API (Real-Time, Experimental) ⚠️

What It Is

  • Uses your browser's built-in speech recognition
  • Shows transcription in real-time as you speak (streaming)
  • Chrome/Edge → Google Cloud Speech
  • Safari → Apple Speech Recognition

⚠️ PRIVACY WARNING

  • Audio IS sent to cloud servers (Google, Apple, etc.)
  • NOT HIPAA-compliant
  • Only use for non-clinical, personal use

When to Use

  • Personal notes (non-clinical)
  • Want real-time feedback while speaking
  • Demonstration/testing
  • NEVER for patient data

How to Enable

  1. Settings → Real-Time Streaming Transcription
  2. Read privacy warning carefully
  3. Toggle "Enable real-time streaming" ON
  4. Confirm warning dialog
  5. Grants microphone permission
  6. Start recording - see words appear live

Limitations

  • Not available in all browsers (requires Web Speech API)
  • Accuracy varies by browser
  • Requires internet connection
  • May have usage limits

Choosing the Right Option

For Clinical Use (HIPAA Required)

Use: Browser Whisper (offline) OR Server (Vertex AI/AWS with BAA)

  • Browser Whisper: Maximum privacy, no costs
  • Server: Faster, better for long recordings

For Personal Use (Non-HIPAA)

Use: Any option

  • Browser Whisper: Best balance of privacy and accuracy
  • Server: Fastest
  • Web Speech: Real-time feedback

Decision Tree

Is this clinical/patient data?
├─ YES → Use Browser Whisper or Server (Vertex/AWS)
│   ├─ Need offline? → Browser Whisper
│   ├─ Need speed? → Server (Vertex AI)
│   └─ Want free? → Browser Whisper
│
└─ NO → Any option
    ├─ Want real-time? → Web Speech API
    ├─ Want privacy? → Browser Whisper
    └─ Want speed? → Server

Configuration

Browser Whisper

# No configuration needed - bundled in Docker image
# Models at: /app/public/models/Xenova/whisper-tiny.en/

Server Transcription

# .env file
TRANSCRIBE_PROVIDER=google  # google, aws, openai, litellm

# Google Vertex AI
GOOGLE_VERTEX_PROJECT=your-project-id
GOOGLE_APPLICATION_CREDENTIALS=/path/to/key.json

# AWS Transcribe
AWS_BEDROCK_REGION=us-east-1
AWS_ACCESS_KEY_ID=your-key
AWS_SECRET_ACCESS_KEY=your-secret

# OpenAI
OPENAI_API_KEY=sk-...

# LiteLLM (proxy)
LITELLM_API_BASE=http://localhost:4000
LITELLM_API_KEY=optional

Web Speech API

# No configuration - uses browser built-in
# Privacy warning shown in Settings UI

FAQ

Q: Which is most accurate?

A: Browser Whisper and Server (Gemini/Whisper) are equally accurate. Web Speech is slightly less accurate.

Q: Which is fastest?

A: Server transcription (~1s) > Web Speech (real-time) > Browser Whisper (2-10s)

Q: Which is most private?

A: Browser Whisper (100% offline) > Server (with BAA) > Web Speech (not private)

Q: Can I use multiple at once?

A: No. Priority: Web Speech > Browser Whisper > Server (whichever is enabled first)

Q: What if transcription fails?

A: Automatic fallback chain:

  1. Browser Whisper (if enabled)
  2. Falls back to Server (if configured)
  3. Falls back to live transcript (if available)

Q: Is Browser Whisper really offline?

A: Yes! Models are bundled in the Docker image. After the page loads once, transcription works with zero network access.

Q: Does Web Speech work offline?

A: No. It requires internet to send audio to cloud servers.

Q: Can I train/customize the models?

A: No. Browser Whisper uses pre-trained models. Server transcription uses cloud models. No custom training available.


Troubleshooting

Browser Whisper stuck at "Initializing"

  • Cause: Models not loaded or network blocked during initial download
  • Fix: See BROWSER_WHISPER_TROUBLESHOOTING.md

Server transcription returns "No provider"

  • Cause: API keys not configured
  • Fix: Set environment variables in .env

Web Speech says "Not supported"

  • Cause: Browser doesn't support Web Speech API
  • Fix: Use Chrome, Edge, or Safari

Transcription is slow

  • Browser Whisper: Try switching to "Tiny" model
  • Server: Check API provider status
  • Web Speech: Check internet connection

Best Practices

Clinical Documentation

  1. Use Browser Whisper for all patient data
  2. Enable audio backups (automatic in v2)
  3. Keep recordings under 5 minutes for faster processing
  4. Use "Tiny" model for quick notes, "Base" for detailed documentation

Personal Use

  1. Web Speech for quick, informal notes
  2. Browser Whisper for anything you want private
  3. Server for long recordings

Performance Optimization

  1. Pre-download Browser Whisper model before first use
  2. Use shorter clips (30-60 seconds) for fastest results
  3. Clear browser cache if models seem corrupted

Summary

Need Recommendation
Clinical/HIPAA Browser Whisper (offline)
Fast transcription Server (Vertex AI)
Real-time feedback Web Speech (non-clinical only)
Maximum privacy Browser Whisper
Zero cost Browser Whisper
Long recordings Server (faster for 5+ min clips)
Offline use Browser Whisper

Default recommendation: Browser Whisper for 95% of use cases. It's private, accurate, free, and offline. Only use alternatives when you have specific needs for speed or real-time feedback.