diff --git a/TRANSCRIPTION_OPTIONS.md b/TRANSCRIPTION_OPTIONS.md new file mode 100644 index 0000000..265c320 --- /dev/null +++ b/TRANSCRIPTION_OPTIONS.md @@ -0,0 +1,279 @@ +# Transcription Options Guide + +## Overview + +Pediatric AI Scribe v2+ offers **three transcription methods**, allowing you to choose between **privacy**, **speed**, and **real-time feedback**. + +--- + +## 📊 Comparison Table + +| Feature | Browser Whisper | Server Transcription | Web Speech API | +|---------|----------------|---------------------|----------------| +| **Privacy** | ⭐⭐⭐⭐⭐ 100% offline | ⭐⭐⭐⭐ (with BAA) | ⭐ Sends to cloud | +| **Accuracy** | ⭐⭐⭐⭐⭐ Whisper | ⭐⭐⭐⭐⭐ Gemini/AWS | ⭐⭐⭐ Browser-dependent | +| **Speed** | ⭐⭐⭐ 2-10s | ⭐⭐⭐⭐⭐ ~1s | ⭐⭐⭐⭐⭐ Instant | +| **Real-time** | ❌ Batch mode | ❌ Batch mode | ✅ Live streaming | +| **HIPAA** | ✅ Yes | ✅ (Vertex/AWS) | ❌ No | +| **Cost** | Free | ~$0.005/min | Free | +| **Internet** | ❌ Not required | ✅ Required | ✅ Required | +| **Setup** | None (bundled) | API keys | None (built-in) | + +--- + +## Option 1: Browser Whisper (Offline, Private) ⭐ RECOMMENDED + +### What It Is +- Runs **OpenAI Whisper** entirely in your browser using WebAssembly +- Audio **never leaves your device** - 100% offline after initial page load +- Models bundled in Docker image (self-hosted, no CDN) + +### When to Use +- ✅ Clinical documentation (HIPAA-compliant) +- ✅ Maximum privacy required +- ✅ Offline/air-gapped environments +- ✅ No API costs +- ✅ Zero vendor dependency + +### How to Enable +1. Settings → Browser Transcription +2. Toggle "Enable browser transcription" ON +3. (Optional) Click "Pre-download model" if you want to cache it first +4. Start recording - transcription happens automatically after recording + +### Models Available +- **Tiny** (~39MB) - Fast, good for short clips (2-3 seconds) +- **Base** (~74MB) - Balanced accuracy and speed (3-5 seconds) +- **Small** (~244MB) - Best quality, slower (6-10 seconds) + +### Performance +- Transcribes ~30-second clip in 2-10 seconds (depending on model) +- First run may be slower (model loading) +- Subsequent runs are instant (cached) + +### Privacy +- ✅ Audio never transmitted +- ✅ Models run locally in WASM +- ✅ No network calls during transcription +- ✅ HIPAA-compliant + +--- + +## Option 2: Server Transcription (Cloud, Fast) + +### What It Is +- Sends audio to your configured AI provider +- Uses Google Gemini, AWS Transcribe, OpenAI Whisper, or LiteLLM + +### When to Use +- ✅ Maximum speed (~1 second for 30-second clip) +- ✅ Best accuracy (cloud models) +- ✅ Long recordings (Browser Whisper can be slow for 5+ minutes) +- ✅ HIPAA-compliant with BAA providers + +### HIPAA-Eligible Providers +- **Google Vertex AI** (with BAA) ✅ +- **AWS Transcribe** (with BAA) ✅ +- **Azure OpenAI** (with BAA) ✅ +- **OpenAI Whisper Direct** ❌ Not HIPAA-eligible + +### How to Enable +- Configured via environment variables (`.env`) +- No user action needed - just works if API keys present +- Falls back automatically if Browser Whisper fails + +### Cost +- Google Gemini: ~$0.005/minute +- AWS Transcribe: ~$0.024/minute +- OpenAI: $0.006/minute + +--- + +## Option 3: Web Speech API (Real-Time, Experimental) ⚠️ + +### What It Is +- Uses your browser's built-in speech recognition +- Shows transcription **in real-time** as you speak (streaming) +- Chrome/Edge → Google Cloud Speech +- Safari → Apple Speech Recognition + +### ⚠️ PRIVACY WARNING +- **Audio IS sent to cloud servers** (Google, Apple, etc.) +- **NOT HIPAA-compliant** +- Only use for non-clinical, personal use + +### When to Use +- ✅ Personal notes (non-clinical) +- ✅ Want real-time feedback while speaking +- ✅ Demonstration/testing +- ❌ **NEVER for patient data** + +### How to Enable +1. Settings → Real-Time Streaming Transcription +2. Read privacy warning carefully +3. Toggle "Enable real-time streaming" ON +4. Confirm warning dialog +5. Grants microphone permission +6. Start recording - see words appear live + +### Limitations +- Not available in all browsers (requires Web Speech API) +- Accuracy varies by browser +- Requires internet connection +- May have usage limits + +--- + +## Choosing the Right Option + +### For Clinical Use (HIPAA Required) +**Use:** Browser Whisper (offline) OR Server (Vertex AI/AWS with BAA) +- Browser Whisper: Maximum privacy, no costs +- Server: Faster, better for long recordings + +### For Personal Use (Non-HIPAA) +**Use:** Any option +- Browser Whisper: Best balance of privacy and accuracy +- Server: Fastest +- Web Speech: Real-time feedback + +### Decision Tree + +``` +Is this clinical/patient data? +├─ YES → Use Browser Whisper or Server (Vertex/AWS) +│ ├─ Need offline? → Browser Whisper +│ ├─ Need speed? → Server (Vertex AI) +│ └─ Want free? → Browser Whisper +│ +└─ NO → Any option + ├─ Want real-time? → Web Speech API + ├─ Want privacy? → Browser Whisper + └─ Want speed? → Server +``` + +--- + +## Configuration + +### Browser Whisper +```bash +# No configuration needed - bundled in Docker image +# Models at: /app/public/models/Xenova/whisper-tiny.en/ +``` + +### Server Transcription +```bash +# .env file +TRANSCRIBE_PROVIDER=google # google, aws, openai, litellm + +# Google Vertex AI +GOOGLE_VERTEX_PROJECT=your-project-id +GOOGLE_APPLICATION_CREDENTIALS=/path/to/key.json + +# AWS Transcribe +AWS_BEDROCK_REGION=us-east-1 +AWS_ACCESS_KEY_ID=your-key +AWS_SECRET_ACCESS_KEY=your-secret + +# OpenAI +OPENAI_API_KEY=sk-... + +# LiteLLM (proxy) +LITELLM_API_BASE=http://localhost:4000 +LITELLM_API_KEY=optional +``` + +### Web Speech API +```bash +# No configuration - uses browser built-in +# Privacy warning shown in Settings UI +``` + +--- + +## FAQ + +### Q: Which is most accurate? +**A:** Browser Whisper and Server (Gemini/Whisper) are equally accurate. Web Speech is slightly less accurate. + +### Q: Which is fastest? +**A:** Server transcription (~1s) > Web Speech (real-time) > Browser Whisper (2-10s) + +### Q: Which is most private? +**A:** Browser Whisper (100% offline) > Server (with BAA) > Web Speech (not private) + +### Q: Can I use multiple at once? +**A:** No. Priority: Web Speech > Browser Whisper > Server (whichever is enabled first) + +### Q: What if transcription fails? +**A:** Automatic fallback chain: +1. Browser Whisper (if enabled) +2. Falls back to Server (if configured) +3. Falls back to live transcript (if available) + +### Q: Is Browser Whisper really offline? +**A:** Yes! Models are bundled in the Docker image. After the page loads once, transcription works with zero network access. + +### Q: Does Web Speech work offline? +**A:** No. It requires internet to send audio to cloud servers. + +### Q: Can I train/customize the models? +**A:** No. Browser Whisper uses pre-trained models. Server transcription uses cloud models. No custom training available. + +--- + +## Troubleshooting + +### Browser Whisper stuck at "Initializing" +- **Cause:** Models not loaded or network blocked during initial download +- **Fix:** See BROWSER_WHISPER_TROUBLESHOOTING.md + +### Server transcription returns "No provider" +- **Cause:** API keys not configured +- **Fix:** Set environment variables in `.env` + +### Web Speech says "Not supported" +- **Cause:** Browser doesn't support Web Speech API +- **Fix:** Use Chrome, Edge, or Safari + +### Transcription is slow +- **Browser Whisper:** Try switching to "Tiny" model +- **Server:** Check API provider status +- **Web Speech:** Check internet connection + +--- + +## Best Practices + +### Clinical Documentation +1. Use Browser Whisper for all patient data +2. Enable audio backups (automatic in v2) +3. Keep recordings under 5 minutes for faster processing +4. Use "Tiny" model for quick notes, "Base" for detailed documentation + +### Personal Use +1. Web Speech for quick, informal notes +2. Browser Whisper for anything you want private +3. Server for long recordings + +### Performance Optimization +1. Pre-download Browser Whisper model before first use +2. Use shorter clips (30-60 seconds) for fastest results +3. Clear browser cache if models seem corrupted + +--- + +## Summary + +| Need | Recommendation | +|------|---------------| +| Clinical/HIPAA | Browser Whisper (offline) | +| Fast transcription | Server (Vertex AI) | +| Real-time feedback | Web Speech (non-clinical only) | +| Maximum privacy | Browser Whisper | +| Zero cost | Browser Whisper | +| Long recordings | Server (faster for 5+ min clips) | +| Offline use | Browser Whisper | + +**Default recommendation:** Browser Whisper for 95% of use cases. It's private, accurate, free, and offline. Only use alternatives when you have specific needs for speed or real-time feedback.