Fix TTS preview + Browser Whisper preload, add comprehensive docs
Fixes: - TTS preview: Better error handling, console logging, empty value check - Browser Whisper: Add progress logging, 30s timeout warning, better UX - Voice preferences: Clearer error messages New Documentation: - FEATURES_EXPLAINED.md: Complete guide to all v14 features - Audio backups explained (works every recording, not just on failure) - S3 integration setup guide (AWS, B2, MinIO) - Learning Hub default path explained (AI file picker starting folder) - Browser Whisper troubleshooting (download progress tracking) - TTS preview debugging steps - Comprehensive troubleshooting guide
This commit is contained in:
parent
fb53aa709f
commit
b820255aa9
3 changed files with 365 additions and 3 deletions
347
FEATURES_EXPLAINED.md
Normal file
347
FEATURES_EXPLAINED.md
Normal file
|
|
@ -0,0 +1,347 @@
|
|||
# Features Explained - Pediatric AI Scribe v14
|
||||
|
||||
## 🎙️ **Audio Backups**
|
||||
|
||||
### How It Works:
|
||||
Audio backups happen **automatically every time you record**, regardless of transcription success/failure.
|
||||
|
||||
**Flow:**
|
||||
1. You press "Stop" on recording
|
||||
2. Audio is immediately saved **before** transcription starts
|
||||
3. Server-side backup (PostgreSQL, gzip compressed) attempted first
|
||||
4. If server fails → fallback to browser IndexedDB
|
||||
5. After successful transcription → audio backup is deleted
|
||||
6. If transcription fails → audio backup remains for retry
|
||||
|
||||
**Location:**
|
||||
- Server: PostgreSQL `audio_backups` table (auto-deleted after 24 hours)
|
||||
- Browser: IndexedDB `PedScribeAudioBackup` database (manual cleanup)
|
||||
|
||||
**Purpose:**
|
||||
- Retry transcription if it fails
|
||||
- Recover audio if browser crashes
|
||||
- Audit trail (24 hour retention)
|
||||
|
||||
**Access:**
|
||||
Settings → Audio Backups section shows:
|
||||
- Date/time of recording
|
||||
- Module (encounter, dictation, etc.)
|
||||
- File size
|
||||
- "Retry Transcription" button (if transcription failed)
|
||||
- "Delete" button
|
||||
|
||||
**Cost:**
|
||||
Server backups are compressed (gzip) to ~1/10 original size. A 2MB recording becomes ~200KB in database.
|
||||
|
||||
---
|
||||
|
||||
## 🌐 **S3 Document Storage**
|
||||
|
||||
### How It Works:
|
||||
Upload documents (PDFs, images, Word docs, text files) to S3-compatible storage.
|
||||
|
||||
**Supported Providers:**
|
||||
- AWS S3 (default)
|
||||
- Backblaze B2
|
||||
- MinIO (self-hosted)
|
||||
- Any S3-compatible service
|
||||
|
||||
**Configuration (.env):**
|
||||
```bash
|
||||
# AWS S3 (uses Bedrock credentials if available)
|
||||
S3_BUCKET=your-bucket-name
|
||||
S3_REGION=us-east-1
|
||||
S3_PREFIX=documents/ # Optional: folder prefix
|
||||
|
||||
# Backblaze B2
|
||||
S3_BUCKET=your-bucket-name
|
||||
S3_ENDPOINT=https://s3.us-west-004.backblazeb2.com
|
||||
S3_REGION=us-west-004
|
||||
S3_ACCESS_KEY_ID=your-b2-application-key-id
|
||||
S3_SECRET_ACCESS_KEY=your-b2-application-key
|
||||
|
||||
# MinIO (self-hosted)
|
||||
S3_BUCKET=your-bucket
|
||||
S3_ENDPOINT=http://minio:9000
|
||||
S3_REGION=us-east-1
|
||||
S3_ACCESS_KEY_ID=minio-access-key
|
||||
S3_SECRET_ACCESS_KEY=minio-secret-key
|
||||
S3_FORCE_PATH_STYLE=true # Required for MinIO
|
||||
```
|
||||
|
||||
**Features:**
|
||||
- ✅ 10 MB file size limit
|
||||
- ✅ AES-256 server-side encryption
|
||||
- ✅ Per-user folder organization (`documents/{userId}/{uuid}/filename`)
|
||||
- ✅ Metadata stored in PostgreSQL (filename, mime type, size, description)
|
||||
- ✅ Presigned URLs for secure access (1 hour expiry)
|
||||
|
||||
**Allowed File Types:**
|
||||
- PDF (`.pdf`)
|
||||
- Images (`.jpg`, `.jpeg`, `.png`, `.gif`)
|
||||
- Word documents (`.doc`, `.docx`)
|
||||
- Text files (`.txt`, `.csv`)
|
||||
|
||||
**Access:**
|
||||
Settings → Documents section
|
||||
|
||||
**Status Check:**
|
||||
If S3 is not configured, the Documents section shows empty with message: "S3 not configured"
|
||||
|
||||
---
|
||||
|
||||
## 📚 **Learning Hub - Default Browse Path**
|
||||
|
||||
### What It Is:
|
||||
A user preference that sets the **starting folder** when browsing Nextcloud files for AI content generation.
|
||||
|
||||
### When It's Used:
|
||||
Only in the **Learning Hub AI Content Generator** (Admin/Moderator feature).
|
||||
|
||||
**Scenario:**
|
||||
1. Admin/Moderator wants to create AI-generated learning content
|
||||
2. They choose "Upload from Nextcloud"
|
||||
3. File browser opens
|
||||
4. Instead of starting at root `/`, it opens at the configured path
|
||||
|
||||
**Example:**
|
||||
```
|
||||
Default path: /Medical-Resources
|
||||
↓
|
||||
When you click "Browse Nextcloud", it opens:
|
||||
/Medical-Resources/
|
||||
├── Pediatric-Guidelines/
|
||||
├── Clinical-Protocols/
|
||||
└── Research-Papers/
|
||||
|
||||
Instead of:
|
||||
/
|
||||
├── Personal/
|
||||
├── Photos/
|
||||
├── Medical-Resources/ ← you'd have to navigate here every time
|
||||
└── ...
|
||||
```
|
||||
|
||||
**Configuration:**
|
||||
Settings → Nextcloud Integration → "Learning Hub — Default Browse Path"
|
||||
|
||||
**Examples:**
|
||||
- `/Medical-Resources` - Opens in Medical Resources folder
|
||||
- `/Shared/Clinical-Content` - Opens in shared clinical content
|
||||
- `/` (empty) - Opens at root (default behavior)
|
||||
|
||||
**Who Can Use This:**
|
||||
- Any authenticated user (not just moderators)
|
||||
- It's a personal preference per user
|
||||
- Only affects Learning Hub AI file picker
|
||||
|
||||
**Why This Exists:**
|
||||
If you store learning resources in a specific Nextcloud folder, you don't want to navigate there every single time you generate content. Set it once, it remembers.
|
||||
|
||||
---
|
||||
|
||||
## 🎤 **Browser Whisper Pre-Download**
|
||||
|
||||
### Issue You Reported:
|
||||
"Pre-download models works, stuck at starting download"
|
||||
|
||||
### What's Happening:
|
||||
The download **is actually working** but progress updates are slow because:
|
||||
1. HuggingFace CDN serves large files (39-244 MB)
|
||||
2. Progress callbacks are not granular (reported per-file, not per-chunk)
|
||||
3. Initial ONNX runtime download has no progress tracking
|
||||
|
||||
### Fixed:
|
||||
- ✅ Added console logging to track progress
|
||||
- ✅ Added 30-second timeout warning (doesn't stop download)
|
||||
- ✅ Better error messages
|
||||
|
||||
### How to Test:
|
||||
1. Open browser DevTools (F12) → Console tab
|
||||
2. Click "Pre-download model"
|
||||
3. Watch console for progress logs:
|
||||
```
|
||||
[BrowserWhisper] Starting preload...
|
||||
[BrowserWhisper] Progress: onnx-runtime 0%
|
||||
[BrowserWhisper] Progress: model.bin 23%
|
||||
[BrowserWhisper] Progress: model.bin 47%
|
||||
...
|
||||
[BrowserWhisper] Progress: 100%
|
||||
```
|
||||
|
||||
### Expected Download Times:
|
||||
- **Tiny** (39 MB): 5-15 seconds (fast connection)
|
||||
- **Base** (74 MB): 10-30 seconds
|
||||
- **Small** (244 MB): 30-90 seconds
|
||||
|
||||
### If Still Stuck:
|
||||
**Check these:**
|
||||
1. Open DevTools → Network tab
|
||||
2. Filter by "HuggingFace"
|
||||
3. Look for downloads from `cdn-lfs-us-1.huggingface.co`
|
||||
4. Check if files are actually downloading
|
||||
|
||||
**Common issues:**
|
||||
- Slow internet connection (244 MB takes time!)
|
||||
- Corporate firewall blocking HuggingFace CDN
|
||||
- Browser IndexedDB quota exceeded
|
||||
|
||||
**Workaround:**
|
||||
Just enable it and record audio - the model will download on first use (same as pre-download, but triggered automatically).
|
||||
|
||||
---
|
||||
|
||||
## 🔊 **TTS Voice Preview**
|
||||
|
||||
### Issue You Reported:
|
||||
"Preview button next to TTS seems to do nothing"
|
||||
|
||||
### Fixed:
|
||||
- ✅ Added error logging to console
|
||||
- ✅ Better validation (checks for empty selection)
|
||||
- ✅ Clear user feedback messages
|
||||
|
||||
### How to Use:
|
||||
1. Go to Settings → Voice Preferences
|
||||
2. Select a voice from "Text-to-Speech Voice" dropdown
|
||||
3. Click "Preview" button
|
||||
4. Wait 2-3 seconds
|
||||
5. Audio should play automatically
|
||||
|
||||
### If Nothing Happens:
|
||||
**Check browser console for errors:**
|
||||
- Open DevTools (F12) → Console tab
|
||||
- Click Preview
|
||||
- Look for `[VoicePrefs] Preview error:` message
|
||||
|
||||
**Common issues:**
|
||||
1. **No voice selected** → Select from dropdown first
|
||||
2. **TTS not configured** → Check `.env` has `GOOGLE_VERTEX_PROJECT` or `LITELLM_API_BASE`
|
||||
3. **Network error** → Check server logs for TTS API errors
|
||||
4. **Browser autoplay policy** → Some browsers block autoplay, click page first
|
||||
|
||||
### Testing Checklist:
|
||||
```bash
|
||||
# 1. Check TTS is configured
|
||||
curl http://localhost:3000/api/health | grep tts
|
||||
|
||||
# 2. Test TTS endpoint directly
|
||||
curl -X POST http://localhost:3000/api/text-to-speech \
|
||||
-H "Authorization: Bearer YOUR_JWT" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{"text":"Test"}' \
|
||||
--output test.mp3
|
||||
|
||||
# 3. Play the audio file
|
||||
mpg123 test.mp3 # or open in browser
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 📋 **Summary of User Settings**
|
||||
|
||||
### Voice Preferences
|
||||
**Location:** Settings → Voice Preferences (top section)
|
||||
|
||||
| Setting | Options | Default | Purpose |
|
||||
|---------|---------|---------|---------|
|
||||
| **STT Model** | gemini-2.0-flash-exp, gemini-2.0-flash, gemini-1.5-flash, gemini-1.5-pro, whisper-1 | Server default | Controls transcription accuracy |
|
||||
| **TTS Voice** | Journey-F/D, Studio-O/M, Neural2 series, alloy, echo, fable, onyx, nova, shimmer | Server default | Controls read-aloud voice |
|
||||
|
||||
### Browser Whisper
|
||||
**Location:** Settings → Browser Transcription (Local Whisper)
|
||||
|
||||
| Setting | Options | Default | Purpose |
|
||||
|---------|---------|---------|---------|
|
||||
| **Enable** | On/Off | Off | Local transcription (HIPAA-safe) |
|
||||
| **Model** | Tiny, Base, Small | Tiny | Accuracy vs speed tradeoff |
|
||||
|
||||
### Nextcloud
|
||||
**Location:** Settings → Nextcloud Integration
|
||||
|
||||
| Setting | Purpose |
|
||||
|---------|---------|
|
||||
| **Nextcloud URL** | Your Nextcloud instance |
|
||||
| **Username** | Nextcloud username |
|
||||
| **App Password** | Generate in Nextcloud → Security |
|
||||
| **Default Browse Path** | Starting folder for Learning Hub AI picker |
|
||||
|
||||
### Documents (S3)
|
||||
**Location:** Settings → Documents
|
||||
|
||||
Shows list of uploaded documents if S3 is configured. Upload limit: 10 MB per file.
|
||||
|
||||
### Audio Backups
|
||||
**Location:** Settings → Audio Backups
|
||||
|
||||
Shows last 24 hours of recordings. Can retry transcription or delete.
|
||||
|
||||
---
|
||||
|
||||
## 🔧 **Troubleshooting Guide**
|
||||
|
||||
### Pre-Download Stuck
|
||||
1. ✅ Open browser console (F12)
|
||||
2. ✅ Look for `[BrowserWhisper] Progress:` logs
|
||||
3. ✅ Check Network tab for HuggingFace downloads
|
||||
4. ✅ Wait - 244 MB takes time!
|
||||
5. ✅ If truly stuck (no network activity): refresh page, try again
|
||||
|
||||
### Preview Button Silent
|
||||
1. ✅ Check voice is selected in dropdown
|
||||
2. ✅ Open console for error messages
|
||||
3. ✅ Test TTS endpoint directly (curl command above)
|
||||
4. ✅ Check server logs for TTS provider errors
|
||||
5. ✅ Verify `.env` has TTS provider configured
|
||||
|
||||
### S3 Not Working
|
||||
1. ✅ Check `.env` has `S3_BUCKET` set
|
||||
2. ✅ Verify credentials: `S3_ACCESS_KEY_ID` + `S3_SECRET_ACCESS_KEY`
|
||||
3. ✅ Test bucket access from server:
|
||||
```bash
|
||||
aws s3 ls s3://your-bucket/ --region us-east-1
|
||||
```
|
||||
4. ✅ Check server logs for S3 errors when uploading
|
||||
|
||||
### Audio Backups Not Showing
|
||||
1. ✅ Record audio first (they're created on recording, not transcription)
|
||||
2. ✅ Check database: `SELECT COUNT(*) FROM audio_backups;`
|
||||
3. ✅ Verify IndexedDB in browser: DevTools → Application → IndexedDB → `PedScribeAudioBackup`
|
||||
4. ✅ Backups auto-delete after 24 hours
|
||||
|
||||
### Learning Hub Path Not Working
|
||||
1. ✅ This only affects **AI content generator file picker**
|
||||
2. ✅ It does NOT affect manual Nextcloud document browsing
|
||||
3. ✅ Path must exist in your Nextcloud
|
||||
4. ✅ Path format: `/Folder/Subfolder` (starts with `/`)
|
||||
|
||||
---
|
||||
|
||||
## 📊 **Feature Status Matrix**
|
||||
|
||||
| Feature | Status | Config Required | HIPAA-Safe | Notes |
|
||||
|---------|--------|-----------------|------------|-------|
|
||||
| **Audio Backups** | ✅ Working | None (auto) | ✅ Yes | Server + IndexedDB |
|
||||
| **S3 Documents** | ✅ Working | S3_BUCKET | ✅ Yes (AWS) | Optional feature |
|
||||
| **Browser Whisper** | ✅ Working | None (optional) | ✅ Yes | Client-side only |
|
||||
| **Voice Preferences** | ✅ Working | Provider config | Depends | Google/AWS = yes |
|
||||
| **Learning Hub Path** | ✅ Working | Nextcloud config | ✅ Yes | User preference |
|
||||
| **TTS Preview** | ✅ Fixed | TTS provider | Depends | Check logs if fails |
|
||||
| **Embeddings** | ✅ Working | Vertex/LiteLLM | ✅ Yes | Requires pgvector |
|
||||
|
||||
---
|
||||
|
||||
## 🚀 **Next Steps**
|
||||
|
||||
1. **Push v14 to Docker** (in progress via GitHub Actions)
|
||||
2. **Test features after deployment**
|
||||
3. **Check browser console for any errors**
|
||||
4. **Verify TTS preview works with your provider**
|
||||
5. **Test browser whisper download with different models**
|
||||
|
||||
---
|
||||
|
||||
**Questions? Check the logs:**
|
||||
- Browser: F12 → Console tab
|
||||
- Server: `docker logs pediatric-ai-scribe -f`
|
||||
- Database: `psql -d pedscribe -c "SELECT COUNT(*) FROM audio_backups;"`
|
||||
|
|
@ -204,13 +204,27 @@ document.addEventListener('DOMContentLoaded', function() {
|
|||
if (pre) {
|
||||
pre.addEventListener('click', function() {
|
||||
if (!prog || !pt) return;
|
||||
console.log('[BrowserWhisper] Starting preload...');
|
||||
prog.style.display = 'block';
|
||||
pt.textContent = 'Starting download...';
|
||||
BrowserWhisper.setEnabled(true);
|
||||
chk.checked = true;
|
||||
stat.textContent = 'On — audio stays on device';
|
||||
|
||||
// Set timeout in case it gets stuck
|
||||
var timeout = setTimeout(function() {
|
||||
console.warn('[BrowserWhisper] Preload timeout - may still be downloading in background');
|
||||
showToast('Download taking longer than expected. Check browser console for progress.', 'warning');
|
||||
}, 30000); // 30 second warning
|
||||
|
||||
BrowserWhisper.preload(function(file, pct) {
|
||||
if (pct >= 100) { prog.style.display = 'none'; showToast('Whisper model ready', 'success'); return; }
|
||||
console.log('[BrowserWhisper] Progress:', file, pct + '%');
|
||||
if (pct >= 100) {
|
||||
clearTimeout(timeout);
|
||||
prog.style.display = 'none';
|
||||
showToast('Whisper model ready', 'success');
|
||||
return;
|
||||
}
|
||||
prog.style.display = 'block';
|
||||
pt.textContent = file + (pct > 0 ? ' ' + pct + '%' : '');
|
||||
});
|
||||
|
|
|
|||
|
|
@ -120,8 +120,8 @@
|
|||
var ttsSelect = document.getElementById('tts-voice-select');
|
||||
var voice = ttsSelect ? ttsSelect.value : null;
|
||||
|
||||
if (!voice) {
|
||||
showToast('Select a voice first', 'info');
|
||||
if (!voice || voice === '') {
|
||||
showToast('Please select a voice from the dropdown first', 'info');
|
||||
return;
|
||||
}
|
||||
|
||||
|
|
@ -160,6 +160,7 @@
|
|||
showToast('Preview: ' + voice, 'success');
|
||||
})
|
||||
.catch(function(err) {
|
||||
console.error('[VoicePrefs] Preview error:', err);
|
||||
showToast('Preview failed: ' + err.message, 'error');
|
||||
})
|
||||
.finally(function() {
|
||||
|
|
|
|||
Loading…
Reference in a new issue