8.2 KiB
Transcription Options Guide
Overview
Pediatric AI Scribe v2+ offers three transcription methods, allowing you to choose between privacy, speed, and real-time feedback.
📊 Comparison Table
| Feature | Browser Whisper | Server Transcription | Web Speech API |
|---|---|---|---|
| Privacy | ⭐⭐⭐⭐⭐ 100% offline | ⭐⭐⭐⭐ (with BAA) | ⭐ Sends to cloud |
| Accuracy | ⭐⭐⭐⭐⭐ Whisper | ⭐⭐⭐⭐⭐ Gemini/AWS | ⭐⭐⭐ Browser-dependent |
| Speed | ⭐⭐⭐ 2-10s | ⭐⭐⭐⭐⭐ ~1s | ⭐⭐⭐⭐⭐ Instant |
| Real-time | ❌ Batch mode | ❌ Batch mode | ✅ Live streaming |
| HIPAA | ✅ Yes | ✅ (Vertex/AWS) | ❌ No |
| Cost | Free | ~$0.005/min | Free |
| Internet | ❌ Not required | ✅ Required | ✅ Required |
| Setup | None (bundled) | API keys | None (built-in) |
Option 1: Browser Whisper (Offline, Private) ⭐ RECOMMENDED
What It Is
- Runs OpenAI Whisper entirely in your browser using WebAssembly
- Audio never leaves your device - 100% offline after initial page load
- Models bundled in Docker image (self-hosted, no CDN)
When to Use
- ✅ Clinical documentation (HIPAA-compliant)
- ✅ Maximum privacy required
- ✅ Offline/air-gapped environments
- ✅ No API costs
- ✅ Zero vendor dependency
How to Enable
- Settings → Browser Transcription
- Toggle "Enable browser transcription" ON
- (Optional) Click "Pre-download model" if you want to cache it first
- Start recording - transcription happens automatically after recording
Models Available
- Tiny (~39MB) - Fast, good for short clips (2-3 seconds)
- Base (~74MB) - Balanced accuracy and speed (3-5 seconds)
- Small (~244MB) - Best quality, slower (6-10 seconds)
Performance
- Transcribes ~30-second clip in 2-10 seconds (depending on model)
- First run may be slower (model loading)
- Subsequent runs are instant (cached)
Privacy
- ✅ Audio never transmitted
- ✅ Models run locally in WASM
- ✅ No network calls during transcription
- ✅ HIPAA-compliant
Option 2: Server Transcription (Cloud, Fast)
What It Is
- Sends audio to your configured AI provider
- Uses Google Gemini, AWS Transcribe, OpenAI Whisper, or LiteLLM
When to Use
- ✅ Maximum speed (~1 second for 30-second clip)
- ✅ Best accuracy (cloud models)
- ✅ Long recordings (Browser Whisper can be slow for 5+ minutes)
- ✅ HIPAA-compliant with BAA providers
HIPAA-Eligible Providers
- Google Vertex AI (with BAA) ✅
- AWS Transcribe (with BAA) ✅
- Azure OpenAI (with BAA) ✅
- OpenAI Whisper Direct ❌ Not HIPAA-eligible
How to Enable
- Configured via environment variables (
.env) - No user action needed - just works if API keys present
- Falls back automatically if Browser Whisper fails
Cost
- Google Gemini: ~$0.005/minute
- AWS Transcribe: ~$0.024/minute
- OpenAI: $0.006/minute
Option 3: Web Speech API (Real-Time, Experimental) ⚠️
What It Is
- Uses your browser's built-in speech recognition
- Shows transcription in real-time as you speak (streaming)
- Chrome/Edge → Google Cloud Speech
- Safari → Apple Speech Recognition
⚠️ PRIVACY WARNING
- Audio IS sent to cloud servers (Google, Apple, etc.)
- NOT HIPAA-compliant
- Only use for non-clinical, personal use
When to Use
- ✅ Personal notes (non-clinical)
- ✅ Want real-time feedback while speaking
- ✅ Demonstration/testing
- ❌ NEVER for patient data
How to Enable
- Settings → Real-Time Streaming Transcription
- Read privacy warning carefully
- Toggle "Enable real-time streaming" ON
- Confirm warning dialog
- Grants microphone permission
- Start recording - see words appear live
Limitations
- Not available in all browsers (requires Web Speech API)
- Accuracy varies by browser
- Requires internet connection
- May have usage limits
Choosing the Right Option
For Clinical Use (HIPAA Required)
Use: Browser Whisper (offline) OR Server (Vertex AI/AWS with BAA)
- Browser Whisper: Maximum privacy, no costs
- Server: Faster, better for long recordings
For Personal Use (Non-HIPAA)
Use: Any option
- Browser Whisper: Best balance of privacy and accuracy
- Server: Fastest
- Web Speech: Real-time feedback
Decision Tree
Is this clinical/patient data?
├─ YES → Use Browser Whisper or Server (Vertex/AWS)
│ ├─ Need offline? → Browser Whisper
│ ├─ Need speed? → Server (Vertex AI)
│ └─ Want free? → Browser Whisper
│
└─ NO → Any option
├─ Want real-time? → Web Speech API
├─ Want privacy? → Browser Whisper
└─ Want speed? → Server
Configuration
Browser Whisper
# No configuration needed - bundled in Docker image
# Models at: /app/public/models/Xenova/whisper-tiny.en/
Server Transcription
# .env file
TRANSCRIBE_PROVIDER=google # google, aws, openai, litellm
# Google Vertex AI
GOOGLE_VERTEX_PROJECT=your-project-id
GOOGLE_APPLICATION_CREDENTIALS=/path/to/key.json
# AWS Transcribe
AWS_BEDROCK_REGION=us-east-1
AWS_ACCESS_KEY_ID=your-key
AWS_SECRET_ACCESS_KEY=your-secret
# OpenAI
OPENAI_API_KEY=sk-...
# LiteLLM (proxy)
LITELLM_API_BASE=http://localhost:4000
LITELLM_API_KEY=optional
Web Speech API
# No configuration - uses browser built-in
# Privacy warning shown in Settings UI
FAQ
Q: Which is most accurate?
A: Browser Whisper and Server (Gemini/Whisper) are equally accurate. Web Speech is slightly less accurate.
Q: Which is fastest?
A: Server transcription (~1s) > Web Speech (real-time) > Browser Whisper (2-10s)
Q: Which is most private?
A: Browser Whisper (100% offline) > Server (with BAA) > Web Speech (not private)
Q: Can I use multiple at once?
A: No. Priority: Web Speech > Browser Whisper > Server (whichever is enabled first)
Q: What if transcription fails?
A: Automatic fallback chain:
- Browser Whisper (if enabled)
- Falls back to Server (if configured)
- Falls back to live transcript (if available)
Q: Is Browser Whisper really offline?
A: Yes! Models are bundled in the Docker image. After the page loads once, transcription works with zero network access.
Q: Does Web Speech work offline?
A: No. It requires internet to send audio to cloud servers.
Q: Can I train/customize the models?
A: No. Browser Whisper uses pre-trained models. Server transcription uses cloud models. No custom training available.
Troubleshooting
Browser Whisper stuck at "Initializing"
- Cause: Models not loaded or network blocked during initial download
- Fix: See BROWSER_WHISPER_TROUBLESHOOTING.md
Server transcription returns "No provider"
- Cause: API keys not configured
- Fix: Set environment variables in
.env
Web Speech says "Not supported"
- Cause: Browser doesn't support Web Speech API
- Fix: Use Chrome, Edge, or Safari
Transcription is slow
- Browser Whisper: Try switching to "Tiny" model
- Server: Check API provider status
- Web Speech: Check internet connection
Best Practices
Clinical Documentation
- Use Browser Whisper for all patient data
- Enable audio backups (automatic in v2)
- Keep recordings under 5 minutes for faster processing
- Use "Tiny" model for quick notes, "Base" for detailed documentation
Personal Use
- Web Speech for quick, informal notes
- Browser Whisper for anything you want private
- Server for long recordings
Performance Optimization
- Pre-download Browser Whisper model before first use
- Use shorter clips (30-60 seconds) for fastest results
- Clear browser cache if models seem corrupted
Summary
| Need | Recommendation |
|---|---|
| Clinical/HIPAA | Browser Whisper (offline) |
| Fast transcription | Server (Vertex AI) |
| Real-time feedback | Web Speech (non-clinical only) |
| Maximum privacy | Browser Whisper |
| Zero cost | Browser Whisper |
| Long recordings | Server (faster for 5+ min clips) |
| Offline use | Browser Whisper |
Default recommendation: Browser Whisper for 95% of use cases. It's private, accurate, free, and offline. Only use alternatives when you have specific needs for speed or real-time feedback.