279 lines
8.2 KiB
Markdown
279 lines
8.2 KiB
Markdown
# Transcription Options Guide
|
|
|
|
## Overview
|
|
|
|
Pediatric AI Scribe v2+ offers **three transcription methods**, allowing you to choose between **privacy**, **speed**, and **real-time feedback**.
|
|
|
|
---
|
|
|
|
## 📊 Comparison Table
|
|
|
|
| Feature | Browser Whisper | Server Transcription | Web Speech API |
|
|
|---------|----------------|---------------------|----------------|
|
|
| **Privacy** | ⭐⭐⭐⭐⭐ 100% offline | ⭐⭐⭐⭐ (with BAA) | ⭐ Sends to cloud |
|
|
| **Accuracy** | ⭐⭐⭐⭐⭐ Whisper | ⭐⭐⭐⭐⭐ Gemini/AWS | ⭐⭐⭐ Browser-dependent |
|
|
| **Speed** | ⭐⭐⭐ 2-10s | ⭐⭐⭐⭐⭐ ~1s | ⭐⭐⭐⭐⭐ Instant |
|
|
| **Real-time** | ❌ Batch mode | ❌ Batch mode | ✅ Live streaming |
|
|
| **HIPAA** | ✅ Yes | ✅ (Vertex/AWS) | ❌ No |
|
|
| **Cost** | Free | ~$0.005/min | Free |
|
|
| **Internet** | ❌ Not required | ✅ Required | ✅ Required |
|
|
| **Setup** | None (bundled) | API keys | None (built-in) |
|
|
|
|
---
|
|
|
|
## Option 1: Browser Whisper (Offline, Private) ⭐ RECOMMENDED
|
|
|
|
### What It Is
|
|
- Runs **OpenAI Whisper** entirely in your browser using WebAssembly
|
|
- Audio **never leaves your device** - 100% offline after initial page load
|
|
- Models bundled in Docker image (self-hosted, no CDN)
|
|
|
|
### When to Use
|
|
- ✅ Clinical documentation (HIPAA-compliant)
|
|
- ✅ Maximum privacy required
|
|
- ✅ Offline/air-gapped environments
|
|
- ✅ No API costs
|
|
- ✅ Zero vendor dependency
|
|
|
|
### How to Enable
|
|
1. Settings → Browser Transcription
|
|
2. Toggle "Enable browser transcription" ON
|
|
3. (Optional) Click "Pre-download model" if you want to cache it first
|
|
4. Start recording - transcription happens automatically after recording
|
|
|
|
### Models Available
|
|
- **Tiny** (~39MB) - Fast, good for short clips (2-3 seconds)
|
|
- **Base** (~74MB) - Balanced accuracy and speed (3-5 seconds)
|
|
- **Small** (~244MB) - Best quality, slower (6-10 seconds)
|
|
|
|
### Performance
|
|
- Transcribes ~30-second clip in 2-10 seconds (depending on model)
|
|
- First run may be slower (model loading)
|
|
- Subsequent runs are instant (cached)
|
|
|
|
### Privacy
|
|
- ✅ Audio never transmitted
|
|
- ✅ Models run locally in WASM
|
|
- ✅ No network calls during transcription
|
|
- ✅ HIPAA-compliant
|
|
|
|
---
|
|
|
|
## Option 2: Server Transcription (Cloud, Fast)
|
|
|
|
### What It Is
|
|
- Sends audio to your configured AI provider
|
|
- Uses Google Gemini, AWS Transcribe, OpenAI Whisper, or LiteLLM
|
|
|
|
### When to Use
|
|
- ✅ Maximum speed (~1 second for 30-second clip)
|
|
- ✅ Best accuracy (cloud models)
|
|
- ✅ Long recordings (Browser Whisper can be slow for 5+ minutes)
|
|
- ✅ HIPAA-compliant with BAA providers
|
|
|
|
### HIPAA-Eligible Providers
|
|
- **Google Vertex AI** (with BAA) ✅
|
|
- **AWS Transcribe** (with BAA) ✅
|
|
- **Azure OpenAI** (with BAA) ✅
|
|
- **OpenAI Whisper Direct** ❌ Not HIPAA-eligible
|
|
|
|
### How to Enable
|
|
- Configured via environment variables (`.env`)
|
|
- No user action needed - just works if API keys present
|
|
- Falls back automatically if Browser Whisper fails
|
|
|
|
### Cost
|
|
- Google Gemini: ~$0.005/minute
|
|
- AWS Transcribe: ~$0.024/minute
|
|
- OpenAI: $0.006/minute
|
|
|
|
---
|
|
|
|
## Option 3: Web Speech API (Real-Time, Experimental) ⚠️
|
|
|
|
### What It Is
|
|
- Uses your browser's built-in speech recognition
|
|
- Shows transcription **in real-time** as you speak (streaming)
|
|
- Chrome/Edge → Google Cloud Speech
|
|
- Safari → Apple Speech Recognition
|
|
|
|
### ⚠️ PRIVACY WARNING
|
|
- **Audio IS sent to cloud servers** (Google, Apple, etc.)
|
|
- **NOT HIPAA-compliant**
|
|
- Only use for non-clinical, personal use
|
|
|
|
### When to Use
|
|
- ✅ Personal notes (non-clinical)
|
|
- ✅ Want real-time feedback while speaking
|
|
- ✅ Demonstration/testing
|
|
- ❌ **NEVER for patient data**
|
|
|
|
### How to Enable
|
|
1. Settings → Real-Time Streaming Transcription
|
|
2. Read privacy warning carefully
|
|
3. Toggle "Enable real-time streaming" ON
|
|
4. Confirm warning dialog
|
|
5. Grants microphone permission
|
|
6. Start recording - see words appear live
|
|
|
|
### Limitations
|
|
- Not available in all browsers (requires Web Speech API)
|
|
- Accuracy varies by browser
|
|
- Requires internet connection
|
|
- May have usage limits
|
|
|
|
---
|
|
|
|
## Choosing the Right Option
|
|
|
|
### For Clinical Use (HIPAA Required)
|
|
**Use:** Browser Whisper (offline) OR Server (Vertex AI/AWS with BAA)
|
|
- Browser Whisper: Maximum privacy, no costs
|
|
- Server: Faster, better for long recordings
|
|
|
|
### For Personal Use (Non-HIPAA)
|
|
**Use:** Any option
|
|
- Browser Whisper: Best balance of privacy and accuracy
|
|
- Server: Fastest
|
|
- Web Speech: Real-time feedback
|
|
|
|
### Decision Tree
|
|
|
|
```
|
|
Is this clinical/patient data?
|
|
├─ YES → Use Browser Whisper or Server (Vertex/AWS)
|
|
│ ├─ Need offline? → Browser Whisper
|
|
│ ├─ Need speed? → Server (Vertex AI)
|
|
│ └─ Want free? → Browser Whisper
|
|
│
|
|
└─ NO → Any option
|
|
├─ Want real-time? → Web Speech API
|
|
├─ Want privacy? → Browser Whisper
|
|
└─ Want speed? → Server
|
|
```
|
|
|
|
---
|
|
|
|
## Configuration
|
|
|
|
### Browser Whisper
|
|
```bash
|
|
# No configuration needed - bundled in Docker image
|
|
# Models at: /app/public/models/Xenova/whisper-tiny.en/
|
|
```
|
|
|
|
### Server Transcription
|
|
```bash
|
|
# .env file
|
|
TRANSCRIBE_PROVIDER=google # google, aws, openai, litellm
|
|
|
|
# Google Vertex AI
|
|
GOOGLE_VERTEX_PROJECT=your-project-id
|
|
GOOGLE_APPLICATION_CREDENTIALS=/path/to/key.json
|
|
|
|
# AWS Transcribe
|
|
AWS_BEDROCK_REGION=us-east-1
|
|
AWS_ACCESS_KEY_ID=your-key
|
|
AWS_SECRET_ACCESS_KEY=your-secret
|
|
|
|
# OpenAI
|
|
OPENAI_API_KEY=sk-...
|
|
|
|
# LiteLLM (proxy)
|
|
LITELLM_API_BASE=http://localhost:4000
|
|
LITELLM_API_KEY=optional
|
|
```
|
|
|
|
### Web Speech API
|
|
```bash
|
|
# No configuration - uses browser built-in
|
|
# Privacy warning shown in Settings UI
|
|
```
|
|
|
|
---
|
|
|
|
## FAQ
|
|
|
|
### Q: Which is most accurate?
|
|
**A:** Browser Whisper and Server (Gemini/Whisper) are equally accurate. Web Speech is slightly less accurate.
|
|
|
|
### Q: Which is fastest?
|
|
**A:** Server transcription (~1s) > Web Speech (real-time) > Browser Whisper (2-10s)
|
|
|
|
### Q: Which is most private?
|
|
**A:** Browser Whisper (100% offline) > Server (with BAA) > Web Speech (not private)
|
|
|
|
### Q: Can I use multiple at once?
|
|
**A:** No. Priority: Web Speech > Browser Whisper > Server (whichever is enabled first)
|
|
|
|
### Q: What if transcription fails?
|
|
**A:** Automatic fallback chain:
|
|
1. Browser Whisper (if enabled)
|
|
2. Falls back to Server (if configured)
|
|
3. Falls back to live transcript (if available)
|
|
|
|
### Q: Is Browser Whisper really offline?
|
|
**A:** Yes! Models are bundled in the Docker image. After the page loads once, transcription works with zero network access.
|
|
|
|
### Q: Does Web Speech work offline?
|
|
**A:** No. It requires internet to send audio to cloud servers.
|
|
|
|
### Q: Can I train/customize the models?
|
|
**A:** No. Browser Whisper uses pre-trained models. Server transcription uses cloud models. No custom training available.
|
|
|
|
---
|
|
|
|
## Troubleshooting
|
|
|
|
### Browser Whisper stuck at "Initializing"
|
|
- **Cause:** Models not loaded or network blocked during initial download
|
|
- **Fix:** See BROWSER_WHISPER_TROUBLESHOOTING.md
|
|
|
|
### Server transcription returns "No provider"
|
|
- **Cause:** API keys not configured
|
|
- **Fix:** Set environment variables in `.env`
|
|
|
|
### Web Speech says "Not supported"
|
|
- **Cause:** Browser doesn't support Web Speech API
|
|
- **Fix:** Use Chrome, Edge, or Safari
|
|
|
|
### Transcription is slow
|
|
- **Browser Whisper:** Try switching to "Tiny" model
|
|
- **Server:** Check API provider status
|
|
- **Web Speech:** Check internet connection
|
|
|
|
---
|
|
|
|
## Best Practices
|
|
|
|
### Clinical Documentation
|
|
1. Use Browser Whisper for all patient data
|
|
2. Enable audio backups (automatic in v2)
|
|
3. Keep recordings under 5 minutes for faster processing
|
|
4. Use "Tiny" model for quick notes, "Base" for detailed documentation
|
|
|
|
### Personal Use
|
|
1. Web Speech for quick, informal notes
|
|
2. Browser Whisper for anything you want private
|
|
3. Server for long recordings
|
|
|
|
### Performance Optimization
|
|
1. Pre-download Browser Whisper model before first use
|
|
2. Use shorter clips (30-60 seconds) for fastest results
|
|
3. Clear browser cache if models seem corrupted
|
|
|
|
---
|
|
|
|
## Summary
|
|
|
|
| Need | Recommendation |
|
|
|------|---------------|
|
|
| Clinical/HIPAA | Browser Whisper (offline) |
|
|
| Fast transcription | Server (Vertex AI) |
|
|
| Real-time feedback | Web Speech (non-clinical only) |
|
|
| Maximum privacy | Browser Whisper |
|
|
| Zero cost | Browser Whisper |
|
|
| Long recordings | Server (faster for 5+ min clips) |
|
|
| Offline use | Browser Whisper |
|
|
|
|
**Default recommendation:** Browser Whisper for 95% of use cases. It's private, accurate, free, and offline. Only use alternatives when you have specific needs for speed or real-time feedback.
|