# Transcription Options Guide ## Overview Pediatric AI Scribe v2+ offers **three transcription methods**, allowing you to choose between **privacy**, **speed**, and **real-time feedback**. --- ## 📊 Comparison Table | Feature | Browser Whisper | Server Transcription | Web Speech API | |---------|----------------|---------------------|----------------| | **Privacy** | ⭐⭐⭐⭐⭐ 100% offline | ⭐⭐⭐⭐ (with BAA) | ⭐ Sends to cloud | | **Accuracy** | ⭐⭐⭐⭐⭐ Whisper | ⭐⭐⭐⭐⭐ Gemini/AWS | ⭐⭐⭐ Browser-dependent | | **Speed** | ⭐⭐⭐ 2-10s | ⭐⭐⭐⭐⭐ ~1s | ⭐⭐⭐⭐⭐ Instant | | **Real-time** | ❌ Batch mode | ❌ Batch mode | ✅ Live streaming | | **HIPAA** | ✅ Yes | ✅ (Vertex/AWS) | ❌ No | | **Cost** | Free | ~$0.005/min | Free | | **Internet** | ❌ Not required | ✅ Required | ✅ Required | | **Setup** | None (bundled) | API keys | None (built-in) | --- ## Option 1: Browser Whisper (Offline, Private) ⭐ RECOMMENDED ### What It Is - Runs **OpenAI Whisper** entirely in your browser using WebAssembly - Audio **never leaves your device** - 100% offline after initial page load - Models bundled in Docker image (self-hosted, no CDN) ### When to Use - ✅ Clinical documentation (HIPAA-compliant) - ✅ Maximum privacy required - ✅ Offline/air-gapped environments - ✅ No API costs - ✅ Zero vendor dependency ### How to Enable 1. Settings → Browser Transcription 2. Toggle "Enable browser transcription" ON 3. (Optional) Click "Pre-download model" if you want to cache it first 4. Start recording - transcription happens automatically after recording ### Models Available - **Tiny** (~39MB) - Fast, good for short clips (2-3 seconds) - **Base** (~74MB) - Balanced accuracy and speed (3-5 seconds) - **Small** (~244MB) - Best quality, slower (6-10 seconds) ### Performance - Transcribes ~30-second clip in 2-10 seconds (depending on model) - First run may be slower (model loading) - Subsequent runs are instant (cached) ### Privacy - ✅ Audio never transmitted - ✅ Models run locally in WASM - ✅ No network calls during transcription - ✅ HIPAA-compliant --- ## Option 2: Server Transcription (Cloud, Fast) ### What It Is - Sends audio to your configured AI provider - Uses Google Gemini, AWS Transcribe, OpenAI Whisper, or LiteLLM ### When to Use - ✅ Maximum speed (~1 second for 30-second clip) - ✅ Best accuracy (cloud models) - ✅ Long recordings (Browser Whisper can be slow for 5+ minutes) - ✅ HIPAA-compliant with BAA providers ### HIPAA-Eligible Providers - **Google Vertex AI** (with BAA) ✅ - **AWS Transcribe** (with BAA) ✅ - **Azure OpenAI** (with BAA) ✅ - **OpenAI Whisper Direct** ❌ Not HIPAA-eligible ### How to Enable - Configured via environment variables (`.env`) - No user action needed - just works if API keys present - Falls back automatically if Browser Whisper fails ### Cost - Google Gemini: ~$0.005/minute - AWS Transcribe: ~$0.024/minute - OpenAI: $0.006/minute --- ## Option 3: Web Speech API (Real-Time, Experimental) ⚠️ ### What It Is - Uses your browser's built-in speech recognition - Shows transcription **in real-time** as you speak (streaming) - Chrome/Edge → Google Cloud Speech - Safari → Apple Speech Recognition ### ⚠️ PRIVACY WARNING - **Audio IS sent to cloud servers** (Google, Apple, etc.) - **NOT HIPAA-compliant** - Only use for non-clinical, personal use ### When to Use - ✅ Personal notes (non-clinical) - ✅ Want real-time feedback while speaking - ✅ Demonstration/testing - ❌ **NEVER for patient data** ### How to Enable 1. Settings → Real-Time Streaming Transcription 2. Read privacy warning carefully 3. Toggle "Enable real-time streaming" ON 4. Confirm warning dialog 5. Grants microphone permission 6. Start recording - see words appear live ### Limitations - Not available in all browsers (requires Web Speech API) - Accuracy varies by browser - Requires internet connection - May have usage limits --- ## Choosing the Right Option ### For Clinical Use (HIPAA Required) **Use:** Browser Whisper (offline) OR Server (Vertex AI/AWS with BAA) - Browser Whisper: Maximum privacy, no costs - Server: Faster, better for long recordings ### For Personal Use (Non-HIPAA) **Use:** Any option - Browser Whisper: Best balance of privacy and accuracy - Server: Fastest - Web Speech: Real-time feedback ### Decision Tree ``` Is this clinical/patient data? ├─ YES → Use Browser Whisper or Server (Vertex/AWS) │ ├─ Need offline? → Browser Whisper │ ├─ Need speed? → Server (Vertex AI) │ └─ Want free? → Browser Whisper │ └─ NO → Any option ├─ Want real-time? → Web Speech API ├─ Want privacy? → Browser Whisper └─ Want speed? → Server ``` --- ## Configuration ### Browser Whisper ```bash # No configuration needed - bundled in Docker image # Models at: /app/public/models/Xenova/whisper-tiny.en/ ``` ### Server Transcription ```bash # .env file TRANSCRIBE_PROVIDER=google # google, aws, openai, litellm # Google Vertex AI GOOGLE_VERTEX_PROJECT=your-project-id GOOGLE_APPLICATION_CREDENTIALS=/path/to/key.json # AWS Transcribe AWS_BEDROCK_REGION=us-east-1 AWS_ACCESS_KEY_ID=your-key AWS_SECRET_ACCESS_KEY=your-secret # OpenAI OPENAI_API_KEY=sk-... # LiteLLM (proxy) LITELLM_API_BASE=http://localhost:4000 LITELLM_API_KEY=optional ``` ### Web Speech API ```bash # No configuration - uses browser built-in # Privacy warning shown in Settings UI ``` --- ## FAQ ### Q: Which is most accurate? **A:** Browser Whisper and Server (Gemini/Whisper) are equally accurate. Web Speech is slightly less accurate. ### Q: Which is fastest? **A:** Server transcription (~1s) > Web Speech (real-time) > Browser Whisper (2-10s) ### Q: Which is most private? **A:** Browser Whisper (100% offline) > Server (with BAA) > Web Speech (not private) ### Q: Can I use multiple at once? **A:** No. Priority: Web Speech > Browser Whisper > Server (whichever is enabled first) ### Q: What if transcription fails? **A:** Automatic fallback chain: 1. Browser Whisper (if enabled) 2. Falls back to Server (if configured) 3. Falls back to live transcript (if available) ### Q: Is Browser Whisper really offline? **A:** Yes! Models are bundled in the Docker image. After the page loads once, transcription works with zero network access. ### Q: Does Web Speech work offline? **A:** No. It requires internet to send audio to cloud servers. ### Q: Can I train/customize the models? **A:** No. Browser Whisper uses pre-trained models. Server transcription uses cloud models. No custom training available. --- ## Troubleshooting ### Browser Whisper stuck at "Initializing" - **Cause:** Models not loaded or network blocked during initial download - **Fix:** See BROWSER_WHISPER_TROUBLESHOOTING.md ### Server transcription returns "No provider" - **Cause:** API keys not configured - **Fix:** Set environment variables in `.env` ### Web Speech says "Not supported" - **Cause:** Browser doesn't support Web Speech API - **Fix:** Use Chrome, Edge, or Safari ### Transcription is slow - **Browser Whisper:** Try switching to "Tiny" model - **Server:** Check API provider status - **Web Speech:** Check internet connection --- ## Best Practices ### Clinical Documentation 1. Use Browser Whisper for all patient data 2. Enable audio backups (automatic in v2) 3. Keep recordings under 5 minutes for faster processing 4. Use "Tiny" model for quick notes, "Base" for detailed documentation ### Personal Use 1. Web Speech for quick, informal notes 2. Browser Whisper for anything you want private 3. Server for long recordings ### Performance Optimization 1. Pre-download Browser Whisper model before first use 2. Use shorter clips (30-60 seconds) for fastest results 3. Clear browser cache if models seem corrupted --- ## Summary | Need | Recommendation | |------|---------------| | Clinical/HIPAA | Browser Whisper (offline) | | Fast transcription | Server (Vertex AI) | | Real-time feedback | Web Speech (non-clinical only) | | Maximum privacy | Browser Whisper | | Zero cost | Browser Whisper | | Long recordings | Server (faster for 5+ min clips) | | Offline use | Browser Whisper | **Default recommendation:** Browser Whisper for 95% of use cases. It's private, accurate, free, and offline. Only use alternatives when you have specific needs for speed or real-time feedback.