Add comprehensive transcription options documentation
This commit is contained in:
parent
ca14094c0a
commit
b8b9e8974b
1 changed files with 279 additions and 0 deletions
279
TRANSCRIPTION_OPTIONS.md
Normal file
279
TRANSCRIPTION_OPTIONS.md
Normal file
|
|
@ -0,0 +1,279 @@
|
|||
# Transcription Options Guide
|
||||
|
||||
## Overview
|
||||
|
||||
Pediatric AI Scribe v2+ offers **three transcription methods**, allowing you to choose between **privacy**, **speed**, and **real-time feedback**.
|
||||
|
||||
---
|
||||
|
||||
## 📊 Comparison Table
|
||||
|
||||
| Feature | Browser Whisper | Server Transcription | Web Speech API |
|
||||
|---------|----------------|---------------------|----------------|
|
||||
| **Privacy** | ⭐⭐⭐⭐⭐ 100% offline | ⭐⭐⭐⭐ (with BAA) | ⭐ Sends to cloud |
|
||||
| **Accuracy** | ⭐⭐⭐⭐⭐ Whisper | ⭐⭐⭐⭐⭐ Gemini/AWS | ⭐⭐⭐ Browser-dependent |
|
||||
| **Speed** | ⭐⭐⭐ 2-10s | ⭐⭐⭐⭐⭐ ~1s | ⭐⭐⭐⭐⭐ Instant |
|
||||
| **Real-time** | ❌ Batch mode | ❌ Batch mode | ✅ Live streaming |
|
||||
| **HIPAA** | ✅ Yes | ✅ (Vertex/AWS) | ❌ No |
|
||||
| **Cost** | Free | ~$0.005/min | Free |
|
||||
| **Internet** | ❌ Not required | ✅ Required | ✅ Required |
|
||||
| **Setup** | None (bundled) | API keys | None (built-in) |
|
||||
|
||||
---
|
||||
|
||||
## Option 1: Browser Whisper (Offline, Private) ⭐ RECOMMENDED
|
||||
|
||||
### What It Is
|
||||
- Runs **OpenAI Whisper** entirely in your browser using WebAssembly
|
||||
- Audio **never leaves your device** - 100% offline after initial page load
|
||||
- Models bundled in Docker image (self-hosted, no CDN)
|
||||
|
||||
### When to Use
|
||||
- ✅ Clinical documentation (HIPAA-compliant)
|
||||
- ✅ Maximum privacy required
|
||||
- ✅ Offline/air-gapped environments
|
||||
- ✅ No API costs
|
||||
- ✅ Zero vendor dependency
|
||||
|
||||
### How to Enable
|
||||
1. Settings → Browser Transcription
|
||||
2. Toggle "Enable browser transcription" ON
|
||||
3. (Optional) Click "Pre-download model" if you want to cache it first
|
||||
4. Start recording - transcription happens automatically after recording
|
||||
|
||||
### Models Available
|
||||
- **Tiny** (~39MB) - Fast, good for short clips (2-3 seconds)
|
||||
- **Base** (~74MB) - Balanced accuracy and speed (3-5 seconds)
|
||||
- **Small** (~244MB) - Best quality, slower (6-10 seconds)
|
||||
|
||||
### Performance
|
||||
- Transcribes ~30-second clip in 2-10 seconds (depending on model)
|
||||
- First run may be slower (model loading)
|
||||
- Subsequent runs are instant (cached)
|
||||
|
||||
### Privacy
|
||||
- ✅ Audio never transmitted
|
||||
- ✅ Models run locally in WASM
|
||||
- ✅ No network calls during transcription
|
||||
- ✅ HIPAA-compliant
|
||||
|
||||
---
|
||||
|
||||
## Option 2: Server Transcription (Cloud, Fast)
|
||||
|
||||
### What It Is
|
||||
- Sends audio to your configured AI provider
|
||||
- Uses Google Gemini, AWS Transcribe, OpenAI Whisper, or LiteLLM
|
||||
|
||||
### When to Use
|
||||
- ✅ Maximum speed (~1 second for 30-second clip)
|
||||
- ✅ Best accuracy (cloud models)
|
||||
- ✅ Long recordings (Browser Whisper can be slow for 5+ minutes)
|
||||
- ✅ HIPAA-compliant with BAA providers
|
||||
|
||||
### HIPAA-Eligible Providers
|
||||
- **Google Vertex AI** (with BAA) ✅
|
||||
- **AWS Transcribe** (with BAA) ✅
|
||||
- **Azure OpenAI** (with BAA) ✅
|
||||
- **OpenAI Whisper Direct** ❌ Not HIPAA-eligible
|
||||
|
||||
### How to Enable
|
||||
- Configured via environment variables (`.env`)
|
||||
- No user action needed - just works if API keys present
|
||||
- Falls back automatically if Browser Whisper fails
|
||||
|
||||
### Cost
|
||||
- Google Gemini: ~$0.005/minute
|
||||
- AWS Transcribe: ~$0.024/minute
|
||||
- OpenAI: $0.006/minute
|
||||
|
||||
---
|
||||
|
||||
## Option 3: Web Speech API (Real-Time, Experimental) ⚠️
|
||||
|
||||
### What It Is
|
||||
- Uses your browser's built-in speech recognition
|
||||
- Shows transcription **in real-time** as you speak (streaming)
|
||||
- Chrome/Edge → Google Cloud Speech
|
||||
- Safari → Apple Speech Recognition
|
||||
|
||||
### ⚠️ PRIVACY WARNING
|
||||
- **Audio IS sent to cloud servers** (Google, Apple, etc.)
|
||||
- **NOT HIPAA-compliant**
|
||||
- Only use for non-clinical, personal use
|
||||
|
||||
### When to Use
|
||||
- ✅ Personal notes (non-clinical)
|
||||
- ✅ Want real-time feedback while speaking
|
||||
- ✅ Demonstration/testing
|
||||
- ❌ **NEVER for patient data**
|
||||
|
||||
### How to Enable
|
||||
1. Settings → Real-Time Streaming Transcription
|
||||
2. Read privacy warning carefully
|
||||
3. Toggle "Enable real-time streaming" ON
|
||||
4. Confirm warning dialog
|
||||
5. Grants microphone permission
|
||||
6. Start recording - see words appear live
|
||||
|
||||
### Limitations
|
||||
- Not available in all browsers (requires Web Speech API)
|
||||
- Accuracy varies by browser
|
||||
- Requires internet connection
|
||||
- May have usage limits
|
||||
|
||||
---
|
||||
|
||||
## Choosing the Right Option
|
||||
|
||||
### For Clinical Use (HIPAA Required)
|
||||
**Use:** Browser Whisper (offline) OR Server (Vertex AI/AWS with BAA)
|
||||
- Browser Whisper: Maximum privacy, no costs
|
||||
- Server: Faster, better for long recordings
|
||||
|
||||
### For Personal Use (Non-HIPAA)
|
||||
**Use:** Any option
|
||||
- Browser Whisper: Best balance of privacy and accuracy
|
||||
- Server: Fastest
|
||||
- Web Speech: Real-time feedback
|
||||
|
||||
### Decision Tree
|
||||
|
||||
```
|
||||
Is this clinical/patient data?
|
||||
├─ YES → Use Browser Whisper or Server (Vertex/AWS)
|
||||
│ ├─ Need offline? → Browser Whisper
|
||||
│ ├─ Need speed? → Server (Vertex AI)
|
||||
│ └─ Want free? → Browser Whisper
|
||||
│
|
||||
└─ NO → Any option
|
||||
├─ Want real-time? → Web Speech API
|
||||
├─ Want privacy? → Browser Whisper
|
||||
└─ Want speed? → Server
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Configuration
|
||||
|
||||
### Browser Whisper
|
||||
```bash
|
||||
# No configuration needed - bundled in Docker image
|
||||
# Models at: /app/public/models/Xenova/whisper-tiny.en/
|
||||
```
|
||||
|
||||
### Server Transcription
|
||||
```bash
|
||||
# .env file
|
||||
TRANSCRIBE_PROVIDER=google # google, aws, openai, litellm
|
||||
|
||||
# Google Vertex AI
|
||||
GOOGLE_VERTEX_PROJECT=your-project-id
|
||||
GOOGLE_APPLICATION_CREDENTIALS=/path/to/key.json
|
||||
|
||||
# AWS Transcribe
|
||||
AWS_BEDROCK_REGION=us-east-1
|
||||
AWS_ACCESS_KEY_ID=your-key
|
||||
AWS_SECRET_ACCESS_KEY=your-secret
|
||||
|
||||
# OpenAI
|
||||
OPENAI_API_KEY=sk-...
|
||||
|
||||
# LiteLLM (proxy)
|
||||
LITELLM_API_BASE=http://localhost:4000
|
||||
LITELLM_API_KEY=optional
|
||||
```
|
||||
|
||||
### Web Speech API
|
||||
```bash
|
||||
# No configuration - uses browser built-in
|
||||
# Privacy warning shown in Settings UI
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## FAQ
|
||||
|
||||
### Q: Which is most accurate?
|
||||
**A:** Browser Whisper and Server (Gemini/Whisper) are equally accurate. Web Speech is slightly less accurate.
|
||||
|
||||
### Q: Which is fastest?
|
||||
**A:** Server transcription (~1s) > Web Speech (real-time) > Browser Whisper (2-10s)
|
||||
|
||||
### Q: Which is most private?
|
||||
**A:** Browser Whisper (100% offline) > Server (with BAA) > Web Speech (not private)
|
||||
|
||||
### Q: Can I use multiple at once?
|
||||
**A:** No. Priority: Web Speech > Browser Whisper > Server (whichever is enabled first)
|
||||
|
||||
### Q: What if transcription fails?
|
||||
**A:** Automatic fallback chain:
|
||||
1. Browser Whisper (if enabled)
|
||||
2. Falls back to Server (if configured)
|
||||
3. Falls back to live transcript (if available)
|
||||
|
||||
### Q: Is Browser Whisper really offline?
|
||||
**A:** Yes! Models are bundled in the Docker image. After the page loads once, transcription works with zero network access.
|
||||
|
||||
### Q: Does Web Speech work offline?
|
||||
**A:** No. It requires internet to send audio to cloud servers.
|
||||
|
||||
### Q: Can I train/customize the models?
|
||||
**A:** No. Browser Whisper uses pre-trained models. Server transcription uses cloud models. No custom training available.
|
||||
|
||||
---
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Browser Whisper stuck at "Initializing"
|
||||
- **Cause:** Models not loaded or network blocked during initial download
|
||||
- **Fix:** See BROWSER_WHISPER_TROUBLESHOOTING.md
|
||||
|
||||
### Server transcription returns "No provider"
|
||||
- **Cause:** API keys not configured
|
||||
- **Fix:** Set environment variables in `.env`
|
||||
|
||||
### Web Speech says "Not supported"
|
||||
- **Cause:** Browser doesn't support Web Speech API
|
||||
- **Fix:** Use Chrome, Edge, or Safari
|
||||
|
||||
### Transcription is slow
|
||||
- **Browser Whisper:** Try switching to "Tiny" model
|
||||
- **Server:** Check API provider status
|
||||
- **Web Speech:** Check internet connection
|
||||
|
||||
---
|
||||
|
||||
## Best Practices
|
||||
|
||||
### Clinical Documentation
|
||||
1. Use Browser Whisper for all patient data
|
||||
2. Enable audio backups (automatic in v2)
|
||||
3. Keep recordings under 5 minutes for faster processing
|
||||
4. Use "Tiny" model for quick notes, "Base" for detailed documentation
|
||||
|
||||
### Personal Use
|
||||
1. Web Speech for quick, informal notes
|
||||
2. Browser Whisper for anything you want private
|
||||
3. Server for long recordings
|
||||
|
||||
### Performance Optimization
|
||||
1. Pre-download Browser Whisper model before first use
|
||||
2. Use shorter clips (30-60 seconds) for fastest results
|
||||
3. Clear browser cache if models seem corrupted
|
||||
|
||||
---
|
||||
|
||||
## Summary
|
||||
|
||||
| Need | Recommendation |
|
||||
|------|---------------|
|
||||
| Clinical/HIPAA | Browser Whisper (offline) |
|
||||
| Fast transcription | Server (Vertex AI) |
|
||||
| Real-time feedback | Web Speech (non-clinical only) |
|
||||
| Maximum privacy | Browser Whisper |
|
||||
| Zero cost | Browser Whisper |
|
||||
| Long recordings | Server (faster for 5+ min clips) |
|
||||
| Offline use | Browser Whisper |
|
||||
|
||||
**Default recommendation:** Browser Whisper for 95% of use cases. It's private, accurate, free, and offline. Only use alternatives when you have specific needs for speed or real-time feedback.
|
||||
Loading…
Reference in a new issue