Add multi-provider support for TTS and AI services

- Implement provider factory pattern
- Add support for OpenAI, Ollama, OpenRouter, and ElevenLabs
- Restructure TTS API and voice management
- Update configuration context and settings UI
- Add provider-specific model fetching
- Update documentation and environment template
This commit is contained in:
Chad Rogers 2025-03-04 22:53:20 -05:00
parent 84bd878822
commit ef7a585ab2
26 changed files with 3340 additions and 1143 deletions

View file

@ -16,7 +16,11 @@ RUN npm install
# Copy project files # Copy project files
COPY . . COPY . .
# Build the Next.js application # Install TypeScript globally for builds
RUN npm install -g typescript
# Build the Next.js application with extra memory allocation to handle larger builds
ENV NODE_OPTIONS="--max-old-space-size=4096"
RUN npm run build RUN npm run build
# Expose the port the app runs on # Expose the port the app runs on

238
README.md
View file

@ -1,179 +1,85 @@
[![GitHub Stars](https://img.shields.io/github/stars/richardr1126/OpenReader-WebUI)](../../stargazers) # OpenReader WebUI
[![GitHub Forks](https://img.shields.io/github/forks/richardr1126/OpenReader-WebUI)](../../network/members)
[![GitHub Watchers](https://img.shields.io/github/watchers/richardr1126/OpenReader-WebUI)](../../watchers)
[![GitHub Issues](https://img.shields.io/github/issues/richardr1126/OpenReader-WebUI)](../../issues)
[![GitHub Last Commit](https://img.shields.io/github/last-commit/richardr1126/OpenReader-WebUI)](../../commits)
[![GitHub Release](https://img.shields.io/github/v/release/richardr1126/OpenReader-WebUI)](../../releases)
[![Discussions](https://img.shields.io/badge/Discussions-Ask%20a%20Question-blue)](../../discussions) A web-based document reader with AI-powered text-to-speech functionality, supporting both PDF and EPUB formats.
[![Bluesky](https://img.shields.io/badge/Bluesky-Chat%20with%20me-blue)](https://bsky.app/profile/richardr.dev)
## Features
# OpenReader WebUI 📄🔊 - 📱 Responsive design for all devices
- 📚 Support for both PDF and EPUB documents
- 🔊 AI-powered text-to-speech with multiple provider options
- 📁 Document organization with folders
- 🌓 Light and dark themes
- 🔄 Adjustable reading speed
OpenReader WebUI is a document reader with Text-to-Speech capabilities, offering a TTS read along experience with narration for both PDF and EPUB documents. It can use any OpenAI compatible TTS endpoint, including [Kokoro-FastAPI](https://github.com/remsky/Kokoro-FastAPI). ## 🆕 Multiple AI Provider Support
- 🎯 **TTS API Integration**: Compatible with OpenAI text to speech API, Kokoro FastAPI TTS, or any other compatible service; enabling high-quality voice narration OpenReader now supports multiple AI providers for text-to-speech:
- 💾 **Local-First Architecture**: Uses IndexedDB browser storage - no server uploads required
- 🛜 **Optional Server-side documents**: Manually upload documents to the next backend for all users to download
- 📖 **Read Along Experience**: Follow along with highlighted text as the TTS narrates
- 📚 **EPUB Support**: Read EPUB files with table of contents
- 📄 **PDF Support**: Read PDF files with clean text extraction
- 🎧 **Audiobook Creation**: Create and export audiobooks from PDF and ePub files with m4b format
- 📲 **Mobile Support**: Works on mobile devices, and can be added as a PWA web app
- 🎨 **Customizable Experience**:
- 🔑 Set TTS API base URL (with optional API key)
- 🏎️ Adjustable playback speed
- 📐 Customize PDF text extraction margins
- 🗣️ Multiple voice options (checks `/v1/audio/voices` endpoint)
- 🎨 Multiple app theme options
### 🛠️ Work in progress - **OpenAI** - High-quality TTS using GPT models
- [x] **Audiobook creation and download** (m4b format) - **OpenRouter** - API gateway to various AI models
- [x] **Get PDFs on iOS working** - **Ollama** - Local AI model support
- [ ] **End-to-end Testing**: More playwright tests (in progress) - **ElevenLabs** - Premium voice synthesis with natural-sounding voices
- [ ] **More document formats**: .txt, .docx, .md, etc.
- [ ] **Accessibility Improvements**
## 🐳 Docker Quick Start ## Setup
1. Clone the repository
2. Install dependencies: `npm install`
3. Copy `template.env` to `.env.local` and fill in your API keys
4. Start the development server: `npm run dev`
### Provider Configuration
Configure your preferred AI providers in `.env.local` or through the settings panel:
#### OpenAI (Default)
```
OPENAI_API_KEY=your_openai_key
OPENAI_BASE_URL=https://api.openai.com/v1
OPENAI_MODEL=tts-1
```
#### OpenRouter
```
OPENROUTER_API_KEY=your_openrouter_key
OPENROUTER_MODEL=openai/whisper
```
#### Ollama (Local)
```
OLLAMA_BASE_URL=http://localhost:11434
OLLAMA_MODEL=llama3:8b
```
#### ElevenLabs Voice Provider
```
ELEVENLABS_API_KEY=your_elevenlabs_key
```
## Voice Selection
OpenReader provides a user interface for selecting and managing voices:
- Choose between OpenAI or ElevenLabs voice providers
- Select from available voices for each provider
- Custom voice support with ElevenLabs (requires ElevenLabs subscription)
- Voice preferences are saved per document
## Development
- Built with Next.js 14
- TypeScript for type safety
- Tailwind CSS for styling
- Headless UI for accessible components
## Docker Deployment
Build and run with Docker:
```bash ```bash
docker run --name openreader-webui \ docker build -t openreader-webui .
-p 3003:3003 \ docker run -p 3003:3003 -e OPENAI_API_KEY=your_key openreader-webui
-v openreader_docstore:/app/docstore \
ghcr.io/richardr1126/openreader-webui:latest
``` ```
(Optionally): Set the TTS `API_BASE` URL and/or `API_KEY` to be default for all devices
```bash
docker run --name openreader-webui \
-e API_BASE=http://host.docker.internal:8880/v1 \
-p 3003:3003 \
-v openreader_docstore:/app/docstore \
ghcr.io/richardr1126/openreader-webui:latest
```
> Requesting audio from the TTS API happens on the Next.js server not the client. So the base URL for the TTS API should be accessible and relative to the Next.js server. If it is in a Docker you may need to use `host.docker.internal` to access the host machine, instead of `localhost`.
Visit [http://localhost:3003](http://localhost:3003) to run the app and set your settings.
> **Note:** The `openreader_docstore` volume is used to store server-side documents. You can mount a local directory instead. Or remove it if you don't need server-side documents.
### ⬆️ Update Docker Image
```bash
docker stop openreader-webui && docker rm openreader-webui
docker pull ghcr.io/richardr1126/openreader-webui:latest
```
### Adding to a Docker Compose (i.e. with open-webui or Kokoro-FastAPI)
Create or add to a `docker-compose.yml`:
```yaml
volumes:
docstore:
services:
openreader-webui:
container_name: openreader-webui
image: ghcr.io/richardr1126/openreader-webui:latest
environment:
- API_BASE=http://host.docker.internal:8880/v1
ports:
- "3003:3003"
volumes:
- docstore:/app/docstore
restart: unless-stopped
```
## [**Demo**](https://openreader.richardr.dev/)
https://github.com/user-attachments/assets/262b9a01-c608-4fee-893c-9461dd48c99b
## Dev Installation
### Prerequisites
- Node.js & npm (recommended: use [nvm](https://github.com/nvm-sh/nvm))
- [FFmpeg](https://ffmpeg.org) (required for audiobook m4b creation only)
### Steps
1. Clone the repository:
```bash
git clone https://github.com/richardr1126/OpenReader-WebUI.git
cd OpenReader-WebUI
```
2. Install dependencies:
```bash
npm install
```
3. Configure the environment:
```bash
cp template.env .env
# Edit .env with your configuration settings
```
> Note: The base URL for the TTS API should be accessible and relative to the Next.js server
4. Start the development server:
```bash
npm run dev
```
or build and run the production server:
```bash
npm run build
npm start
```
Visit [http://localhost:3003](http://localhost:3003) to run the app.
> Dev server runs on port 3000 by default, while the production server runs on port 3003.
## 💡 Feature requests
For feature requests or ideas you have for the project, please use the [Discussions](https://github.com/richardr1126/OpenReader-WebUI/discussions) tab.
## 🙋‍♂️ Support and issues
For general questions, you can reach out to me on [Bluesky](https://bsky.app/profile/richardr.dev). If you encounter issues, please open an issue on GitHub following the template (which is very simple).
## 👥 Contributing
Contributions are welcome! Fork the repository and submit a pull request with your changes.
## ❤️ Acknowledgements
- [Kokoro-FastAPI](https://github.com/remsky/Kokoro-FastAPI) for the API wrapper
- [react-pdf](https://github.com/wojtekmaj/react-pdf)
- [react-reader](https://github.com/happyr/react-reader)
- [Kokoro-82M](https://huggingface.co/hexgrad/Kokoro-82M) for text-to-speech
## Docker Supported Architectures
- linux/amd64 (x86_64)
- linux/arm64 (Apple Silicon)
## Stack
- **Framework:** Next.js (React)
- **Containerization:** Docker
- **Storage:** IndexedDB (in browser db store)
- **PDF:**
- [react-pdf](https://github.com/wojtekmaj/react-pdf)
- [pdf.js](https://mozilla.github.io/pdf.js/)
- **EPUB:**
- [react-reader](https://github.com/happyr/react-reader)
- [epubjs](https://github.com/futurepress/epub.js/)
- **UI:**
- [Tailwind CSS](https://tailwindcss.com)
- [Headless UI](https://headlessui.com)
- **TTS:** (tested on)
- [OpenAI API](https://platform.openai.com/docs/api-reference/text-to-speech)
- [Kokoro FastAPI TTS](https://github.com/remsky/Kokoro-FastAPI/tree/v0.0.5post1-stable)
- **NLP:** [compromise](https://github.com/spencermountain/compromise) NLP library for sentence splitting
## License ## License
This project is licensed under the MIT License. MIT

4
package-lock.json generated
View file

@ -1,12 +1,12 @@
{ {
"name": "openreader-webui", "name": "openreader-webui",
"version": "0.2.4-patch.1", "version": "0.2.4-patch.3",
"lockfileVersion": 3, "lockfileVersion": 3,
"requires": true, "requires": true,
"packages": { "packages": {
"": { "": {
"name": "openreader-webui", "name": "openreader-webui",
"version": "0.2.4-patch.1", "version": "0.2.4-patch.3",
"dependencies": { "dependencies": {
"@headlessui/react": "^2.2.0", "@headlessui/react": "^2.2.0",
"@types/howler": "^2.2.12", "@types/howler": "^2.2.12",

View file

@ -1,255 +1,102 @@
import { NextRequest, NextResponse } from 'next/server'; /**
import { spawn } from 'child_process'; * Audio Conversion API Endpoint
import { writeFile, readFile, mkdir, unlink, rmdir, readdir } from 'fs/promises'; *
import { existsSync, createReadStream } from 'fs'; * This endpoint handles converting text to speech when direct TTS is not available from a provider.
import { join } from 'path'; * It is primarily used as a fallback for providers like Ollama that don't have native TTS capabilities.
import { randomUUID } from 'crypto'; */
interface ConversionRequest { import { NextRequest } from 'next/server';
chapterTitle: string; import { createVoiceProvider } from '@/providers/factory';
buffer: number[]; import { ProviderSettings, VoiceProviderType, ProviderType } from '@/providers/types';
bookId?: string;
}
async function getAudioDuration(filePath: string): Promise<number> { export async function POST(req: Request) {
return new Promise((resolve, reject) => {
const ffprobe = spawn('ffprobe', [
'-i', filePath,
'-show_entries', 'format=duration',
'-v', 'quiet',
'-of', 'csv=p=0'
]);
let output = '';
ffprobe.stdout.on('data', (data) => {
output += data.toString();
});
ffprobe.on('close', (code) => {
if (code === 0) {
const duration = parseFloat(output.trim());
resolve(duration);
} else {
reject(new Error(`ffprobe process exited with code ${code}`));
}
});
ffprobe.on('error', (err) => {
reject(err);
});
});
}
async function runFFmpeg(args: string[]): Promise<void> {
return new Promise<void>((resolve, reject) => {
const ffmpeg = spawn('ffmpeg', args);
ffmpeg.stderr.on('data', (data) => {
console.error(`ffmpeg stderr: ${data}`);
});
ffmpeg.on('close', (code) => {
if (code === 0) {
resolve();
} else {
reject(new Error(`FFmpeg process exited with code ${code}`));
}
});
ffmpeg.on('error', (err) => {
reject(err);
});
});
}
export async function POST(request: NextRequest) {
try { try {
// Parse the request body const body = await req.json();
const data: ConversionRequest = await request.json(); const { text, voice, speed, provider, voiceProvider } = body;
// Create temp directory if it doesn't exist if (!text) {
const tempDir = join(process.cwd(), 'temp'); return Response.json({ error: 'Text is required' }, { status: 400 });
if (!existsSync(tempDir)) {
await mkdir(tempDir);
} }
// Generate or use existing book ID
const bookId = data.bookId || randomUUID();
const intermediateDir = join(tempDir, `${bookId}-intermediate`);
// Create intermediate directory // Use the provided voice provider or default to OpenAI
if (!existsSync(intermediateDir)) { const effectiveVoiceProvider: VoiceProviderType = voiceProvider || 'openai';
await mkdir(intermediateDir); const effectiveProvider: ProviderType = provider || 'openai';
}
// Count existing files to determine chapter index
const files = await readdir(intermediateDir);
const wavFiles = files.filter(f => f.endsWith('.wav'));
const chapterIndex = wavFiles.length;
// Write input file
const inputPath = join(intermediateDir, `${chapterIndex}-input.aac`);
const outputPath = join(intermediateDir, `${chapterIndex}.wav`);
const metadataPath = join(intermediateDir, `${chapterIndex}.meta.json`);
// Write the chapter audio to a temp file // Extract API keys from headers if present
await writeFile(inputPath, Buffer.from(new Uint8Array(data.buffer))); const headers = new Headers(req.headers);
const openaiKey = headers.get('x-openai-key') || '';
const elevenLabsKey = headers.get('x-elevenlabs-key') || '';
const openrouterKey = headers.get('x-openrouter-key') || '';
// Convert to WAV from raw aac with consistent format // Create default provider settings
await runFFmpeg([ let providerSettings: ProviderSettings = {
'-i', inputPath, openai: {
'-f', 'wav', apiKey: openaiKey,
'-c:a', 'copy', baseUrl: 'https://api.openai.com/v1',
'-preset', 'ultrafast', model: 'tts-1',
'-threads', '0', },
outputPath ollama: {
]); baseUrl: 'http://localhost:11434',
model: 'llama3:8b',
// Get the duration and save metadata },
const duration = await getAudioDuration(outputPath); openrouter: {
await writeFile(metadataPath, JSON.stringify({ apiKey: openrouterKey,
title: data.chapterTitle, model: 'openai/whisper',
duration, },
index: chapterIndex elevenlabs: {
})); apiKey: elevenLabsKey,
},
// Clean up input file
await unlink(inputPath).catch(console.error);
return NextResponse.json({
bookId,
chapterIndex,
duration
});
} catch (error) {
console.error('Error processing audio chapter:', error);
return NextResponse.json(
{ error: 'Failed to process audio chapter' },
{ status: 500 }
);
}
}
export async function GET(request: NextRequest) {
try {
const bookId = request.nextUrl.searchParams.get('bookId');
if (!bookId) {
return NextResponse.json({ error: 'Missing bookId parameter' }, { status: 400 });
}
const tempDir = join(process.cwd(), 'temp');
const intermediateDir = join(tempDir, `${bookId}-intermediate`);
const outputPath = join(tempDir, `${bookId}.m4b`);
const metadataPath = join(tempDir, `${bookId}-metadata.txt`);
const listPath = join(tempDir, `${bookId}-list.txt`);
if (!existsSync(intermediateDir)) {
return NextResponse.json({ error: 'Book not found' }, { status: 404 });
}
// Read all chapter metadata
const files = await readdir(intermediateDir);
const metaFiles = files.filter(f => f.endsWith('.meta.json'));
const chapters: { title: string; duration: number; index: number }[] = [];
for (const metaFile of metaFiles) {
const meta = JSON.parse(await readFile(join(intermediateDir, metaFile), 'utf-8'));
chapters.push(meta);
}
// Sort chapters by index
chapters.sort((a, b) => a.index - b.index);
// Create chapter metadata file
const metadata: string[] = [];
let currentTime = 0;
// Calculate chapter timings based on actual durations
chapters.forEach((chapter) => {
const startMs = Math.floor(currentTime * 1000);
currentTime += chapter.duration;
const endMs = Math.floor(currentTime * 1000);
metadata.push(
`[CHAPTER]`,
`TIMEBASE=1/1000`,
`START=${startMs}`,
`END=${endMs}`,
`title=${chapter.title}`
);
});
await writeFile(metadataPath, ';FFMETADATA1\n' + metadata.join('\n'));
// Create list file for concat
await writeFile(
listPath,
chapters.map(c => `file '${join(intermediateDir, `${c.index}.wav`)}'`).join('\n')
);
// Combine all files into a single M4B
await runFFmpeg([
'-f', 'concat',
'-safe', '0',
'-i', listPath,
'-i', metadataPath,
'-map_metadata', '1',
'-c:a', 'copy', // c:a is codec for audio and :a is stream specifier
//'-codec', 'wav',
//'-b:a', '192k',
//'-threads', '0', // Use maximum available threads
//'-movflags', '+faststart',
//'-preset', 'ultrafast', // Use fastest encoding preset
outputPath
]);
// Stream the file back to the client
const stream = createReadStream(outputPath);
// Clean up function
const cleanup = async () => {
try {
await Promise.all([
...chapters.map(c => unlink(join(intermediateDir, `${c.index}.wav`))),
...chapters.map(c => unlink(join(intermediateDir, `${c.index}.meta.json`))),
unlink(metadataPath),
unlink(listPath),
unlink(outputPath),
rmdir(intermediateDir)
]);
} catch (error) {
console.error('Cleanup error:', error);
}
}; };
// Clean up after streaming is complete // Parse provider settings from header if available
stream.on('end', cleanup); try {
const settingsHeader = headers.get('x-provider-settings');
const readableWebStream = new ReadableStream({ if (settingsHeader) {
start(controller) { providerSettings = JSON.parse(settingsHeader);
stream.on('data', (chunk) => { }
controller.enqueue(chunk); } catch (error) {
}); console.error('Error parsing provider settings:', error);
stream.on('end', () => { }
controller.close();
}); console.log(`Converting text to audio using provider: ${effectiveProvider}, voice provider: ${effectiveVoiceProvider}, speed: ${speed || 1.0}`);
stream.on('error', (err) => {
controller.error(err); // Create voice provider for TTS generation
}); const ttsProvider = createVoiceProvider(effectiveVoiceProvider, {
}, providerSettings
}); });
return new NextResponse(readableWebStream, { if (!ttsProvider) {
throw new Error(`Failed to create voice provider for ${effectiveVoiceProvider}`);
}
// Ensure speed is a valid number
const effectiveSpeed = typeof speed === 'string' ? parseFloat(speed) :
typeof speed === 'number' ? speed : 1.0;
// Generate speech from text
const audioBuffer = await ttsProvider.generateSpeech({
text,
voice: voice || 'alloy',
speed: effectiveSpeed,
format: 'mp3',
});
if (!audioBuffer || audioBuffer.byteLength === 0) {
throw new Error('Generated audio buffer is empty');
}
// Return audio data
return new Response(audioBuffer, {
headers: { headers: {
'Content-Type': 'audio/mp4', 'Content-Type': 'audio/mpeg',
'Content-Disposition': `attachment; filename="audio.mp3"`,
'Cache-Control': 'private, max-age=604800', // Cache for a week
}, },
}); });
} catch (error) { } catch (error) {
console.error('Error creating M4B:', error); console.error('Error converting text to audio:', error);
return NextResponse.json( return Response.json(
{ error: 'Failed to create M4B file' }, { error: error instanceof Error ? error.message : 'Unknown error' },
{ status: 500 } { status: 500 }
); );
} }

View file

@ -0,0 +1,109 @@
/**
* Ollama Models API Endpoint
*
* This endpoint fetches available models from an Ollama instance.
*/
import { NextRequest } from 'next/server';
interface OllamaModel {
name: string;
modified_at: string;
size: number;
digest: string;
details: {
format: string;
family: string;
families: string[];
parameter_size: string;
quantization_level: string;
};
}
interface OllamaModelsResponse {
models: OllamaModel[];
}
export async function GET(req: NextRequest) {
try {
// Get the base URL from the query string or use the default
const searchParams = req.nextUrl.searchParams;
const baseUrl = searchParams.get('baseUrl') || 'http://localhost:11434';
// Normalize the base URL (ensure it doesn't end with a slash)
const normalizedBaseUrl = baseUrl.endsWith('/')
? baseUrl.slice(0, -1)
: baseUrl;
// Make request to Ollama API to fetch models
const response = await fetch(`${normalizedBaseUrl}/api/tags`, {
method: 'GET',
headers: {
'Content-Type': 'application/json',
},
// Set a timeout of 5 seconds
signal: AbortSignal.timeout(5000),
});
if (!response.ok) {
// Handle specific error cases
if (response.status === 404) {
return Response.json({
error: 'Not found',
message: 'Ollama API endpoint not found'
}, { status: 404 });
}
const errorText = await response.text();
return Response.json({
error: 'Ollama API error',
message: errorText
}, { status: response.status });
}
const data: OllamaModelsResponse = await response.json();
// Transform the data to a simpler format
const models = data.models.map(model => ({
id: model.name,
name: model.name,
details: model.details ? {
family: model.details.family,
parameterSize: model.details.parameter_size,
} : undefined,
size: model.size,
}));
// If no models are found, return a specific message
if (models.length === 0) {
return Response.json({
models: [],
message: 'No models found in Ollama instance'
});
}
return Response.json({ models });
} catch (error) {
console.error('Error fetching Ollama models:', error);
// Handle different error types
if (error instanceof TypeError && error.message.includes('fetch')) {
return Response.json({
error: 'Connection error',
message: 'Could not connect to Ollama instance'
}, { status: 503 });
}
if (error instanceof DOMException && error.name === 'AbortError') {
return Response.json({
error: 'Timeout',
message: 'Connection to Ollama timed out'
}, { status: 504 });
}
return Response.json({
error: 'Unknown error',
message: error instanceof Error ? error.message : 'An unknown error occurred'
}, { status: 500 });
}
}

View file

@ -1,59 +1,138 @@
import { NextRequest, NextResponse } from 'next/server'; /**
import OpenAI from 'openai'; * TTS API Route
*
* This API route handles requests to generate speech from text using various providers.
*/
export async function POST(req: NextRequest) { import type { NextRequest } from 'next/server';
import {
createProvider,
createVoiceProvider
} from '@/providers/factory';
import {
ProviderType,
VoiceProviderType,
ProviderSettings
} from '@/providers/types';
/**
* Handles requests to generate speech from text using the specified provider
*
* @param req The incoming request
* @returns A response containing the generated audio
*/
export async function POST(req: Request) {
try { try {
// Get API credentials from headers or fall back to environment variables // Create a new abort controller for this request
const openApiKey = req.headers.get('x-openai-key') || process.env.API_KEY || 'none'; const controller = new AbortController();
const openApiBaseUrl = req.headers.get('x-openai-base-url') || process.env.API_BASE; const { signal } = controller;
const { text, voice, speed, format } = await req.json();
console.log('Received TTS request:', text, voice, speed, format); // Get provider details from request body and headers
const body = await req.json();
if (!openApiKey) { const {
return NextResponse.json({ error: 'Missing OpenAI API key' }, { status: 401 }); text,
} voice,
speed,
if (!text || !voice || !speed) { provider: bodyProvider,
return NextResponse.json({ error: 'Missing required parameters' }, { status: 400 }); voiceProvider: bodyVoiceProvider,
} documentId
} = body;
// Initialize OpenAI client with abort signal
const openai = new OpenAI({ // Check if provider is specified in body, otherwise use header
apiKey: openApiKey, const provider = (bodyProvider || new Headers(req.headers).get('x-provider') || 'openai') as ProviderType;
baseURL: openApiBaseUrl, const voiceProvider = (bodyVoiceProvider || new Headers(req.headers).get('x-voice-provider') || 'openai') as VoiceProviderType;
});
// Parse provider settings from headers or use defaults
// Request audio from OpenAI and pass along the abort signal let providerSettings: ProviderSettings = {
const response = await openai.audio.speech.create({ openai: {
model: 'tts-1', apiKey: new Headers(req.headers).get('x-openai-key') || '',
voice: voice as "alloy", baseUrl: new Headers(req.headers).get('x-openai-base-url') || 'https://api.openai.com/v1',
input: text, model: 'tts-1',
speed: speed, },
response_format: format === 'aac' ? 'aac' : 'mp3', ollama: {
}, { signal: req.signal }); // Pass the abort signal to OpenAI client baseUrl: 'http://localhost:11434',
model: 'llama3:8b',
// Get the audio data as array buffer },
// This will also be aborted if the client cancels openrouter: {
const stream = response.body; apiKey: '',
model: 'openai/whisper',
// Return audio data with appropriate headers },
const contentType = format === 'aac' ? 'audio/aac' : 'audio/mpeg'; elevenlabs: {
return new NextResponse(stream, { apiKey: '',
headers: { },
'Content-Type': contentType };
try {
const settingsHeader = new Headers(req.headers).get('x-provider-settings');
if (settingsHeader) {
providerSettings = JSON.parse(settingsHeader);
} }
}); } catch (error) {
} catch (error) { console.error('Error parsing provider settings:', error);
// Check if this was an abort error
if (error instanceof Error && error.name === 'AbortError') {
console.log('TTS request aborted by client');
return new NextResponse(null, { status: 499 }); // Use 499 status for client closed request
} }
console.warn('Error generating TTS:', error); // Validate request parameters
return NextResponse.json( if (!text) {
{ error: 'Failed to generate audio' }, return new Response(JSON.stringify({ error: 'Text is required' }), {
{ status: 500 } status: 400,
headers: { 'Content-Type': 'application/json' }
});
}
if (!voice) {
return new Response(JSON.stringify({ error: 'Voice is required' }), {
status: 400,
headers: { 'Content-Type': 'application/json' }
});
}
// Ensure speed is a valid number
const effectiveSpeed = typeof speed === 'string' ? parseFloat(speed) :
typeof speed === 'number' ? speed : 1.0;
console.log(`Generating TTS with ${provider}/${voiceProvider} for text at speed ${effectiveSpeed}:`, text.substring(0, 50) + (text.length > 50 ? '...' : ''));
// Create the appropriate provider instance based on voice provider selection
const ttsProvider = voiceProvider === 'elevenlabs'
? createVoiceProvider(voiceProvider, { providerSettings })
: createProvider(provider, { providerSettings });
// Set timeout to abort long-running requests
const timeout = setTimeout(() => {
controller.abort();
console.log('TTS request timed out');
}, 30000); // 30 second timeout
try {
// Generate speech
const audioBuffer = await ttsProvider.generateSpeech({
text,
voice,
speed: effectiveSpeed,
}, signal);
clearTimeout(timeout);
// Return audio as binary response
return new Response(audioBuffer, {
headers: {
'Content-Type': 'audio/mpeg',
'Cache-Control': 'private, max-age=604800', // Cache for a week
},
});
} catch (error) {
clearTimeout(timeout);
throw error; // Re-throw to be caught by the outer try/catch
}
} catch (error) {
// Return error response
console.error('Error generating TTS:', error);
return new Response(
JSON.stringify({ error: error instanceof Error ? error.message : 'Unknown error' }),
{
status: 500,
headers: { 'Content-Type': 'application/json' }
}
); );
} }
} }

View file

@ -1,30 +1,91 @@
import { NextRequest, NextResponse } from 'next/server'; /**
* Voices API Route
*
* This API route handles requests to fetch available voices from various providers.
*/
const DEFAULT_VOICES = ['alloy', 'ash', 'coral', 'echo', 'fable', 'onyx', 'nova', 'sage', 'shimmer']; import type { NextRequest } from 'next/server';
import {
createProvider,
createVoiceProvider
} from '@/providers/factory';
import {
ProviderType,
VoiceProviderType,
ProviderSettings
} from '@/providers/types';
export async function GET(req: NextRequest) { /**
* Handles requests to fetch available voices from specified provider
*
* @param req The incoming request
* @returns A response containing the available voices and categories
*/
export async function GET(req: Request) {
try { try {
// Get API credentials from headers or fall back to environment variables const headers = new Headers(req.headers);
const openApiKey = req.headers.get('x-openai-key') || process.env.API_KEY || 'none';
const openApiBaseUrl = req.headers.get('x-openai-base-url') || process.env.API_BASE; // Get provider details from headers
const provider = (headers.get('x-provider') || 'openai') as ProviderType;
// Request voices from OpenAI const voiceProvider = (headers.get('x-voice-provider') || 'openai') as VoiceProviderType;
const response = await fetch(`${openApiBaseUrl}/audio/voices`, {
headers: { // Parse provider settings from headers
'Authorization': `Bearer ${openApiKey}`, let providerSettings: ProviderSettings = {
'Content-Type': 'application/json', openai: {
apiKey: '',
baseUrl: 'https://api.openai.com/v1',
model: 'tts-1',
}, },
}); ollama: {
baseUrl: 'http://localhost:11434',
if (!response.ok) { model: 'llama3:8b',
throw new Error('Failed to fetch voices'); },
openrouter: {
apiKey: '',
model: 'openai/whisper',
},
elevenlabs: {
apiKey: '',
},
};
try {
const settingsHeader = headers.get('x-provider-settings');
if (settingsHeader) {
providerSettings = JSON.parse(settingsHeader);
}
} catch (error) {
console.error('Error parsing provider settings:', error);
} }
const data = await response.json(); console.log(`Fetching voices for ${voiceProvider} provider...`);
return NextResponse.json({ voices: data.voices || DEFAULT_VOICES });
// Create the appropriate provider instance
const ttsProvider = voiceProvider === 'elevenlabs'
? createVoiceProvider(voiceProvider, { providerSettings })
: createProvider(provider, { providerSettings });
// Fetch voices and categories
const [voices, categories] = await Promise.all([
ttsProvider.getAvailableVoices(),
ttsProvider.getVoiceCategories()
]);
return new Response(
JSON.stringify({ voices, categories }),
{
status: 200,
headers: { 'Content-Type': 'application/json' }
}
);
} catch (error) { } catch (error) {
console.error('Error fetching voices:', error); console.error('Error getting voices:', error);
// Return default voices on error return new Response(
return NextResponse.json({ voices: DEFAULT_VOICES }); JSON.stringify({ error: error instanceof Error ? error.message : 'Unknown error' }),
{
status: 500,
headers: { 'Content-Type': 'application/json' }
}
);
} }
} }

View file

@ -1,71 +1,47 @@
import "./globals.css"; 'use client';
import { ReactNode } from "react";
import { Providers } from "@/app/providers";
import { Metadata } from "next";
import { Footer } from "@/components/Footer";
import { Toaster } from 'react-hot-toast';
export const metadata: Metadata = { import { useState, useEffect } from 'react';
title: "OpenReader WebUI", import { Providers } from './providers';
description: 'A "bring your own TTS api" web interface for reading PDF and EPUB documents with high quality text-to-speech voices. Read books with ease, listen to articles on the go, or study like you have your own lecturer, all in one place.', import './globals.css';
keywords: "PDF reader, EPUB reader, text to speech, tts open ai, kokoro tts, Kokoro-82M, OpenReader, TTS PDF reader, ebook reader, epub tts, high quality text to speech",
authors: [{ name: "Richard Roberson" }],
manifest: "/manifest.json",
metadataBase: new URL("https://openreader.richardr.dev"), // Replace with your domain
openGraph: {
type: "website",
locale: "en_US",
url: "https://openreader.richardr.dev",
siteName: "OpenReader WebUI",
title: "OpenReader WebUI",
description: 'A "bring your own TTS api" web interface for reading PDF and EPUB documents with high quality text-to-speech voices. Read books with ease, listen to articles on the go, or study like you have your own lecturer, all in one place.',
images: [
{
url: "/web-app-manifest-512x512.png",
width: 512,
height: 512,
alt: "OpenReader WebUI Logo",
},
],
},
robots: {
index: true,
follow: true,
googleBot: {
index: true,
follow: true,
'max-video-preview': -1,
'max-image-preview': 'large',
'max-snippet': -1,
},
},
verification: {
// Add your verification codes if you have them
google: "MJXyTudn1kgQF8EtGD-tsnAWev7Iawso9hEvqeGHB3U",
},
};
const isDev = process.env.NEXT_PUBLIC_NODE_ENV !== 'production' || process.env.NODE_ENV == null; /**
//const isDev = false; * Root layout component for the application
*/
export default function RootLayout({
children,
}: {
children: React.ReactNode;
}) {
// Client-side only state to avoid hydration mismatches
const [isClient, setIsClient] = useState(false);
// Set isClient to true once component mounts on the client
useEffect(() => {
setIsClient(true);
}, []);
export default function RootLayout({ children }: { children: ReactNode }) {
return ( return (
<html lang="en"> <html lang="en" suppressHydrationWarning>
<head> <head>
<meta name="color-scheme" content="light dark" /> <meta charSet="utf-8" />
<meta name="viewport" content="width=device-width, initial-scale=1" />
<meta name="description" content="OpenReader - Read and listen to documents" />
<link rel="icon" href="/favicon.ico" />
</head> </head>
<body className="antialiased"> <body
<Providers> className="antialiased"
<div className="min-h-screen bg-background p-4"> // This prevents hydration warnings from body attributes added by extensions
<div className="relative max-w-6xl mx-auto align-center"> suppressHydrationWarning
<div className="bg-base rounded-lg shadow-lg"> >
{children} {/* Only render the application content once mounted on the client */}
</div> {isClient ? (
{!isDev && <Footer />} <Providers>{children}</Providers>
</div> ) : (
// Add a simple loading state while client-side JS initializes
<div className="flex items-center justify-center h-screen w-screen bg-base">
<div className="text-foreground">Loading OpenReader...</div>
</div> </div>
<Toaster /> )}
</Providers>
</body> </body>
</html> </html>
); );

18
src/app/metadata.ts Normal file
View file

@ -0,0 +1,18 @@
import { Metadata } from 'next';
/**
* Metadata for the application
*/
export const metadata: Metadata = {
title: 'OpenReader',
description: 'Read and listen to documents with multiple AI providers',
icons: {
icon: '/favicon.ico',
},
manifest: '/manifest.json',
viewport: {
width: 'device-width',
initialScale: 1,
},
themeColor: '#000000',
};

25
src/app/page-metadata.tsx Normal file
View file

@ -0,0 +1,25 @@
import type { Metadata } from 'next';
/**
* Metadata for the application
* This needs to be exported from a server component
*/
export const metadata: Metadata = {
title: 'OpenReader',
description: 'Read and listen to documents with multiple AI providers',
icons: {
icon: '/favicon.ico',
},
viewport: {
width: 'device-width',
initialScale: 1,
},
themeColor: '#000000',
};
/**
* This component can be imported by pages to set their metadata
*/
export default function PageMetadata() {
return null;
}

View file

@ -1,17 +1,30 @@
import { DocumentUploader } from '@/components/DocumentUploader';
import { DocumentList } from '@/components/doclist/DocumentList'; import { DocumentList } from '@/components/doclist/DocumentList';
import { DocumentUploader } from '@/components/DocumentUploader';
import { Footer } from '@/components/Footer';
import { SettingsModal } from '@/components/SettingsModal'; import { SettingsModal } from '@/components/SettingsModal';
import { metadata } from './page-metadata';
/**
* Home page component
* Shows document list and uploader
*/
export default function Home() { export default function Home() {
return ( return (
<div className='p-3.5 sm:p-5'> <main className="flex min-h-screen flex-col items-center justify-between py-8 px-0 md:px-8 select-none">
<SettingsModal /> <div className="z-10 w-full max-w-5xl items-center justify-between font-mono text-sm">
<h1 className="text-xl font-bold text-center flex-grow">OpenReader WebUI</h1> <div className="w-full relative p-2">
<p className="text-sm mt-1 text-center text-muted mb-5">A bring your own text-to-speech api web interface for reading documents with high quality voices</p> <SettingsModal />
<div className="flex flex-col items-center gap-5"> <h1 className="text-2xl md:text-4xl font-bold text-center text-foreground mb-10">
<DocumentUploader className='max-w-xl' /> OpenReader
<DocumentList /> </h1>
<DocumentUploader />
<DocumentList />
</div>
</div> </div>
</div> <Footer />
</main>
); );
} }
// Export the metadata for this page
export { metadata };

View file

@ -1,7 +1,7 @@
'use client'; 'use client';
import { Fragment, useState, useEffect, useCallback } from 'react'; import { Fragment, useState, useEffect, useCallback } from 'react';
import { Dialog, DialogPanel, DialogTitle, Transition, TransitionChild, Listbox, ListboxButton, ListboxOptions, ListboxOption, Button, Input } from '@headlessui/react'; import { Dialog, Transition } from '@headlessui/react';
import { useTheme } from '@/contexts/ThemeContext'; import { useTheme } from '@/contexts/ThemeContext';
import { useConfig } from '@/contexts/ConfigContext'; import { useConfig } from '@/contexts/ConfigContext';
import { ChevronUpDownIcon, CheckIcon, SettingsIcon } from '@/components/icons/Icons'; import { ChevronUpDownIcon, CheckIcon, SettingsIcon } from '@/components/icons/Icons';
@ -10,6 +10,10 @@ import { useDocuments } from '@/contexts/DocumentContext';
import { setItem, getItem } from '@/utils/indexedDB'; import { setItem, getItem } from '@/utils/indexedDB';
import { ConfirmDialog } from '@/components/ConfirmDialog'; import { ConfirmDialog } from '@/components/ConfirmDialog';
import { THEMES } from '@/contexts/ThemeContext'; import { THEMES } from '@/contexts/ThemeContext';
import { ProviderType, VoiceProviderType, ProviderSettings } from '@/providers/types';
// Import Theme type from ThemeContext
type Theme = 'system' | 'light' | 'dark' | 'ocean' | 'forest' | 'sunset' | 'sea' | 'mint';
const isDev = process.env.NEXT_PUBLIC_NODE_ENV !== 'production' || process.env.NODE_ENV == null; const isDev = process.env.NEXT_PUBLIC_NODE_ENV !== 'production' || process.env.NODE_ENV == null;
@ -18,19 +22,68 @@ const themes = THEMES.map(id => ({
name: id.charAt(0).toUpperCase() + id.slice(1) name: id.charAt(0).toUpperCase() + id.slice(1)
})); }));
// Provider options
const providerOptions = [
{ id: 'openai' as ProviderType, name: 'OpenAI' },
{ id: 'openrouter' as ProviderType, name: 'OpenRouter' },
{ id: 'ollama' as ProviderType, name: 'Ollama' },
];
// Voice provider options
const voiceProviderOptions = [
{ id: 'openai' as VoiceProviderType, name: 'OpenAI' },
{ id: 'elevenlabs' as VoiceProviderType, name: 'ElevenLabs' },
];
// Ollama model interface
interface OllamaModel {
id: string;
name: string;
details?: {
family: string;
parameterSize: string;
};
size?: number;
}
type InputChangeEvent = React.ChangeEvent<HTMLInputElement>;
type SelectChangeEvent = React.ChangeEvent<HTMLSelectElement>;
export function SettingsModal() { export function SettingsModal() {
const [isOpen, setIsOpen] = useState(false); const [isOpen, setIsOpen] = useState(false);
const { theme, setTheme } = useTheme(); const { theme, setTheme } = useTheme();
const { apiKey, baseUrl, updateConfig } = useConfig(); const {
apiKey,
baseUrl,
provider,
voiceProvider,
providerSettings,
updateConfig,
updateConfigKey
} = useConfig();
const { refreshPDFs, refreshEPUBs, clearPDFs, clearEPUBs } = useDocuments(); const { refreshPDFs, refreshEPUBs, clearPDFs, clearEPUBs } = useDocuments();
const [localApiKey, setLocalApiKey] = useState(apiKey); const [localApiKey, setLocalApiKey] = useState(apiKey);
const [localBaseUrl, setLocalBaseUrl] = useState(baseUrl); const [localBaseUrl, setLocalBaseUrl] = useState(baseUrl);
const [localProvider, setLocalProvider] = useState<ProviderType>(provider);
const [localVoiceProvider, setLocalVoiceProvider] = useState<VoiceProviderType>(voiceProvider);
const [localProviderSettings, setLocalProviderSettings] = useState<ProviderSettings>(providerSettings);
// Ollama models state
const [ollamaModels, setOllamaModels] = useState<OllamaModel[]>([]);
const [isLoadingModels, setIsLoadingModels] = useState(false);
const [ollamaModelError, setOllamaModelError] = useState<string | null>(null);
const [isSyncing, setIsSyncing] = useState(false); const [isSyncing, setIsSyncing] = useState(false);
const [isLoading, setIsLoading] = useState(false); const [isLoading, setIsLoading] = useState(false);
const selectedTheme = themes.find(t => t.id === theme) || themes[0]; const selectedTheme = themes.find(t => t.id === theme) || themes[0];
const selectedProvider = providerOptions.find(p => p.id === localProvider) || providerOptions[0];
const selectedVoiceProvider = voiceProviderOptions.find(p => p.id === localVoiceProvider) || voiceProviderOptions[0];
const [showClearLocalConfirm, setShowClearLocalConfirm] = useState(false); const [showClearLocalConfirm, setShowClearLocalConfirm] = useState(false);
const [showClearServerConfirm, setShowClearServerConfirm] = useState(false); const [showClearServerConfirm, setShowClearServerConfirm] = useState(false);
const [activeSettingsTab, setActiveSettingsTab] = useState<'general' | 'providers' | 'advanced'>('general');
// set firstVisit on initial load // set firstVisit on initial load
const checkFirstVist = useCallback(async () => { const checkFirstVist = useCallback(async () => {
@ -42,11 +95,50 @@ export function SettingsModal() {
} }
}, [setIsOpen]); }, [setIsOpen]);
// Fetch Ollama models when base URL changes
const fetchOllamaModels = useCallback(async (baseUrl: string) => {
if (!baseUrl) return;
try {
setIsLoadingModels(true);
setOllamaModelError(null);
const response = await fetch(`/api/providers/ollama/models?baseUrl=${encodeURIComponent(baseUrl)}`);
const data = await response.json();
if (response.ok) {
setOllamaModels(data.models || []);
if (data.models?.length === 0 && data.message) {
setOllamaModelError(data.message);
}
} else {
setOllamaModels([]);
setOllamaModelError(data.message || 'Failed to fetch Ollama models');
}
} catch (error) {
console.error('Error fetching Ollama models:', error);
setOllamaModels([]);
setOllamaModelError('Connection error');
} finally {
setIsLoadingModels(false);
}
}, []);
useEffect(() => { useEffect(() => {
checkFirstVist(); checkFirstVist();
setLocalApiKey(apiKey); setLocalApiKey(apiKey);
setLocalBaseUrl(baseUrl); setLocalBaseUrl(baseUrl);
}, [apiKey, baseUrl, checkFirstVist]); setLocalProvider(provider);
setLocalVoiceProvider(voiceProvider);
setLocalProviderSettings({...providerSettings});
}, [apiKey, baseUrl, provider, voiceProvider, providerSettings, checkFirstVist]);
// Fetch Ollama models when the tab is active and provider is Ollama
useEffect(() => {
if (isOpen && activeSettingsTab === 'providers' && localProvider === 'ollama') {
fetchOllamaModels(localProviderSettings.ollama?.baseUrl || 'http://localhost:11434');
}
}, [isOpen, activeSettingsTab, localProvider, localProviderSettings.ollama?.baseUrl, fetchOllamaModels]);
const handleSync = async () => { const handleSync = async () => {
try { try {
@ -91,30 +183,157 @@ export function SettingsModal() {
setShowClearServerConfirm(false); setShowClearServerConfirm(false);
}; };
const handleInputChange = (type: 'apiKey' | 'baseUrl', value: string) => { const handleInputChange = (
type: 'apiKey' | 'baseUrl' | 'providerSetting',
value: string,
providerKey?: keyof ProviderSettings,
settingKey?: string
) => {
if (type === 'apiKey') { if (type === 'apiKey') {
setLocalApiKey(value === '' ? '' : value); setLocalApiKey(value === '' ? '' : value);
} else {
// Also update in provider settings for OpenAI
setLocalProviderSettings(prev => ({
...prev,
openai: {
...prev.openai,
apiKey: value
}
}));
} else if (type === 'baseUrl') {
setLocalBaseUrl(value === '' ? '' : value); setLocalBaseUrl(value === '' ? '' : value);
// Also update in provider settings for OpenAI
setLocalProviderSettings(prev => ({
...prev,
openai: {
...prev.openai,
baseUrl: value || 'https://api.openai.com/v1'
}
}));
} else if (type === 'providerSetting' && providerKey && settingKey) {
// Update provider-specific settings
setLocalProviderSettings(prev => ({
...prev,
[providerKey]: {
...prev[providerKey],
[settingKey]: value
}
}));
// Keep OpenAI settings in sync with the legacy apiKey and baseUrl
if (providerKey === 'openai') {
if (settingKey === 'apiKey') {
setLocalApiKey(value);
} else if (settingKey === 'baseUrl') {
setLocalBaseUrl(value);
}
}
// If changing Ollama base URL, refetch models
if (providerKey === 'ollama' && settingKey === 'baseUrl') {
fetchOllamaModels(value);
}
} }
}; };
// Handle Ollama model selection
const handleOllamaModelChange = (e: SelectChangeEvent) => {
const modelId = e.target.value;
handleInputChange('providerSetting', modelId, 'ollama', 'model');
};
// Handle provider selection change
const handleProviderChange = (newProvider: { id: ProviderType, name: string }) => {
setLocalProvider(newProvider.id);
};
// Handle voice provider selection change
const handleVoiceProviderChange = (newProvider: { id: VoiceProviderType, name: string }) => {
setLocalVoiceProvider(newProvider.id);
};
// Reset all provider settings to defaults
const resetProviderSettings = () => {
setLocalProvider('openai');
setLocalVoiceProvider('openai');
setLocalProviderSettings({
openai: {
apiKey: '',
baseUrl: 'https://api.openai.com/v1',
model: 'tts-1',
},
ollama: {
baseUrl: 'http://localhost:11434',
model: 'llama3:8b',
},
openrouter: {
apiKey: '',
model: 'openai/whisper',
},
elevenlabs: {
apiKey: '',
},
});
setLocalApiKey('');
setLocalBaseUrl('');
setOllamaModels([]);
setOllamaModelError(null);
};
// Save all settings
const saveSettings = async () => {
// Update OpenAI key and base URL in provider settings
const updatedProviderSettings = {
...localProviderSettings,
openai: {
...localProviderSettings.openai,
apiKey: localApiKey,
baseUrl: localBaseUrl || 'https://api.openai.com/v1',
}
};
// Update all config settings
await updateConfig({
apiKey: localApiKey || '',
baseUrl: localBaseUrl || '',
provider: localProvider,
voiceProvider: localVoiceProvider,
providerSettings: updatedProviderSettings,
});
setIsOpen(false);
};
// Handle theme change
const handleThemeChange = (e: SelectChangeEvent) => {
// Cast the string value to the Theme type
setTheme(e.target.value as Theme);
};
return ( return (
<Button <>
onClick={() => setIsOpen(true)} <button
className="rounded-full p-2 text-foreground hover:bg-offbase transform transition-transform duration-200 ease-in-out hover:scale-[1.1] hover:text-accent absolute top-1 left-1 sm:top-3 sm:left-3" onClick={() => setIsOpen(true)}
aria-label="Settings" className="rounded-full p-2 text-foreground hover:bg-offbase transform transition-transform duration-200 ease-in-out hover:scale-[1.1] hover:text-accent absolute top-1 left-1 sm:top-3 sm:left-3"
tabIndex={0} aria-label="Settings"
> tabIndex={0}
<SettingsIcon className="w-4 h-4 sm:w-5 sm:h-5 hover:animate-spin-slow" /> >
<SettingsIcon className="w-4 h-4 sm:w-5 sm:h-5 hover:animate-spin-slow" />
</button>
<Transition appear show={isOpen} as={Fragment}> <Transition appear show={isOpen} as={Fragment}>
<Dialog as="div" className="relative z-50" onClose={() => { <Dialog as="div" className="relative z-50" onClose={() => {
setIsOpen(false); setIsOpen(false);
setLocalApiKey(apiKey); setLocalApiKey(apiKey);
setLocalBaseUrl(baseUrl); setLocalBaseUrl(baseUrl);
setLocalProvider(provider);
setLocalVoiceProvider(voiceProvider);
setLocalProviderSettings({...providerSettings});
setOllamaModels([]);
setOllamaModelError(null);
}}> }}>
<TransitionChild <Transition.Child
as={Fragment} as={Fragment}
enter="ease-out duration-300" enter="ease-out duration-300"
enterFrom="opacity-0" enterFrom="opacity-0"
@ -124,11 +343,11 @@ export function SettingsModal() {
leaveTo="opacity-0" leaveTo="opacity-0"
> >
<div className="fixed inset-0 bg-black/25 backdrop-blur-sm" /> <div className="fixed inset-0 bg-black/25 backdrop-blur-sm" />
</TransitionChild> </Transition.Child>
<div className="fixed inset-0 overflow-y-auto"> <div className="fixed inset-0 overflow-y-auto">
<div className="flex min-h-full items-center justify-center p-4 text-center"> <div className="flex min-h-full items-center justify-center p-4 text-center">
<TransitionChild <Transition.Child
as={Fragment} as={Fragment}
enter="ease-out duration-300" enter="ease-out duration-300"
enterFrom="opacity-0 scale-95" enterFrom="opacity-0 scale-95"
@ -137,178 +356,347 @@ export function SettingsModal() {
leaveFrom="opacity-100 scale-100" leaveFrom="opacity-100 scale-100"
leaveTo="opacity-0 scale-95" leaveTo="opacity-0 scale-95"
> >
<DialogPanel className="w-full max-w-md transform rounded-2xl bg-base p-6 text-left align-middle shadow-xl transition-all"> <Dialog.Panel className="w-full max-w-md transform rounded-2xl bg-base p-6 text-left align-middle shadow-xl transition-all">
<DialogTitle <Dialog.Title
as="h3" as="h3"
className="text-lg font-semibold leading-6 text-foreground" className="text-lg font-semibold leading-6 text-foreground"
> >
Settings Settings
</DialogTitle> </Dialog.Title>
<div className="mt-4"> <div className="mt-4">
<div className="relative space-y-4"> <div className="relative space-y-4">
<div className="space-y-2"> {/* Tab navigation */}
<label className="block text-sm font-medium text-foreground">Theme</label> <div className="flex border-b border-gray-200 dark:border-gray-700 mb-4">
<Listbox value={selectedTheme} onChange={(newTheme) => setTheme(newTheme.id)}> <button
<ListboxButton className="relative w-full cursor-pointer rounded-lg bg-background py-2 pl-3 pr-10 text-left text-foreground shadow-sm focus:outline-none focus:ring-2 focus:ring-accent transform transition-transform duration-200 ease-in-out hover:scale-[1.01] hover:text-accent"> className={`py-2 px-4 font-medium text-sm ${activeSettingsTab === 'general' ? 'border-b-2 border-accent text-accent' : 'text-foreground/70'}`}
<span className="block truncate">{selectedTheme.name}</span> onClick={() => setActiveSettingsTab('general')}
<span className="pointer-events-none absolute inset-y-0 right-0 flex items-center pr-2"> >
<ChevronUpDownIcon className="h-5 w-5 text-muted" /> General
</span> </button>
</ListboxButton> <button
<Transition className={`py-2 px-4 font-medium text-sm ${activeSettingsTab === 'providers' ? 'border-b-2 border-accent text-accent' : 'text-foreground/70'}`}
as={Fragment} onClick={() => setActiveSettingsTab('providers')}
leave="transition ease-in duration-100" >
leaveFrom="opacity-100" AI Providers
leaveTo="opacity-0" </button>
> <button
<ListboxOptions className="absolute mt-1 max-h-40 w-full overflow-auto rounded-md bg-background py-1 shadow-lg ring-1 ring-black/5 focus:outline-none"> className={`py-2 px-4 font-medium text-sm ${activeSettingsTab === 'advanced' ? 'border-b-2 border-accent text-accent' : 'text-foreground/70'}`}
{themes.map((theme) => ( onClick={() => setActiveSettingsTab('advanced')}
<ListboxOption >
key={theme.id} Advanced
className={({ active }) => </button>
`relative cursor-pointer select-none py-2 pl-10 pr-4 ${active ? 'bg-accent/10 text-accent' : 'text-foreground' </div>
}`
} {/* General tab */}
value={theme} {activeSettingsTab === 'general' && (
<div className="space-y-4">
<div className="space-y-2">
<label className="block text-sm font-medium text-foreground">Theme</label>
<Transition.Root show={true} as={Fragment}>
<Transition.Child as="div" enter="" enterFrom="" enterTo="" leave="" leaveFrom="" leaveTo="">
<select
value={selectedTheme.id}
onChange={handleThemeChange}
className="w-full cursor-pointer rounded-lg bg-background py-2 px-3 text-foreground shadow-sm focus:outline-none focus:ring-2 focus:ring-accent"
> >
{({ selected }) => ( {themes.map((theme) => (
<> <option key={theme.id} value={theme.id}>
<span className={`block truncate ${selected ? 'font-medium' : 'font-normal'}`}> {theme.name}
{theme.name} </option>
</span> ))}
{selected ? ( </select>
<span className="absolute inset-y-0 left-0 flex items-center pl-3 text-accent"> </Transition.Child>
<CheckIcon className="h-5 w-5" /> </Transition.Root>
</span> </div>
) : null}
</> <div className="space-y-2">
<label className="block text-sm font-medium text-foreground">Bulk Delete</label>
<div className="flex gap-2">
<button
onClick={() => setShowClearLocalConfirm(true)}
className="justify-center rounded-lg bg-red-600 px-3 py-1.5 text-sm
font-medium text-white hover:bg-red-700 focus:outline-none
focus-visible:ring-2 focus-visible:ring-red-500 focus-visible:ring-offset-2
transform transition-transform duration-200 ease-in-out hover:scale-[1.04]"
>
Delete local docs
</button>
{isDev && <button
onClick={() => setShowClearServerConfirm(true)}
className="justify-center rounded-lg bg-red-600 px-3 py-1.5 text-sm
font-medium text-white hover:bg-red-700 focus:outline-none
focus-visible:ring-2 focus-visible:ring-red-500 focus-visible:ring-offset-2
transform transition-transform duration-200 ease-in-out hover:scale-[1.04]"
>
Delete server docs
</button>}
</div>
</div>
</div>
)}
{/* Providers tab */}
{activeSettingsTab === 'providers' && (
<div className="space-y-4">
{/* TTS Provider Selection */}
<div className="flex flex-col gap-4 p-3 bg-background rounded-lg">
<div className="flex gap-4 items-center justify-between">
<h4 className="text-sm font-semibold text-foreground">Text Processing</h4>
<select
value={selectedProvider.id}
onChange={(e) => setLocalProvider(e.target.value as ProviderType)}
className="cursor-pointer rounded-lg bg-base py-1 px-2 text-xs text-foreground border border-offbase focus:outline-none focus:ring-1 focus:ring-accent"
>
{providerOptions.map((provider) => (
<option key={provider.id} value={provider.id}>
{provider.name}
</option>
))}
</select>
</div>
{/* OpenAI specific settings */}
{localProvider === 'openai' && (
<div className="space-y-3">
<div className="space-y-1">
<label className="block text-xs font-medium text-foreground/80">
API Key
</label>
<input
type="password"
value={localProviderSettings.openai?.apiKey || ''}
onChange={(e: InputChangeEvent) =>
handleInputChange('providerSetting', e.target.value, 'openai', 'apiKey')}
placeholder="Enter OpenAI API key"
className="w-full rounded-lg bg-base py-1.5 px-2 text-sm text-foreground border border-offbase shadow-sm focus:outline-none focus:ring-1 focus:ring-accent"
/>
</div>
<div className="space-y-1">
<label className="block text-xs font-medium text-foreground/80">
Base URL
</label>
<input
type="text"
value={localProviderSettings.openai?.baseUrl || 'https://api.openai.com/v1'}
onChange={(e: InputChangeEvent) =>
handleInputChange('providerSetting', e.target.value, 'openai', 'baseUrl')}
placeholder="https://api.openai.com/v1"
className="w-full rounded-lg bg-base py-1.5 px-2 text-sm text-foreground border border-offbase shadow-sm focus:outline-none focus:ring-1 focus:ring-accent"
/>
</div>
<div className="space-y-1">
<label className="block text-xs font-medium text-foreground/80">
Model
</label>
<input
type="text"
value={localProviderSettings.openai?.model || 'tts-1'}
onChange={(e: InputChangeEvent) =>
handleInputChange('providerSetting', e.target.value, 'openai', 'model')}
placeholder="tts-1"
className="w-full rounded-lg bg-base py-1.5 px-2 text-sm text-foreground border border-offbase shadow-sm focus:outline-none focus:ring-1 focus:ring-accent"
/>
</div>
</div>
)}
{/* OpenRouter specific settings */}
{localProvider === 'openrouter' && (
<div className="space-y-3">
<div className="space-y-1">
<label className="block text-xs font-medium text-foreground/80">
API Key
</label>
<input
type="password"
value={localProviderSettings.openrouter?.apiKey || ''}
onChange={(e: InputChangeEvent) =>
handleInputChange('providerSetting', e.target.value, 'openrouter', 'apiKey')}
placeholder="Enter OpenRouter API key"
className="w-full rounded-lg bg-base py-1.5 px-2 text-sm text-foreground border border-offbase shadow-sm focus:outline-none focus:ring-1 focus:ring-accent"
/>
</div>
<div className="space-y-1">
<label className="block text-xs font-medium text-foreground/80">
Model
</label>
<input
type="text"
value={localProviderSettings.openrouter?.model || 'openai/whisper'}
onChange={(e: InputChangeEvent) =>
handleInputChange('providerSetting', e.target.value, 'openrouter', 'model')}
placeholder="openai/whisper"
className="w-full rounded-lg bg-base py-1.5 px-2 text-sm text-foreground border border-offbase shadow-sm focus:outline-none focus:ring-1 focus:ring-accent"
/>
</div>
</div>
)}
{/* Ollama specific settings */}
{localProvider === 'ollama' && (
<div className="space-y-3">
<div className="space-y-1">
<label className="block text-xs font-medium text-foreground/80">
Base URL
</label>
<div className="flex gap-2">
<input
type="text"
value={localProviderSettings.ollama?.baseUrl || 'http://localhost:11434'}
onChange={(e: InputChangeEvent) =>
handleInputChange('providerSetting', e.target.value, 'ollama', 'baseUrl')}
placeholder="http://localhost:11434"
className="w-full rounded-lg bg-base py-1.5 px-2 text-sm text-foreground border border-offbase shadow-sm focus:outline-none focus:ring-1 focus:ring-accent"
/>
<button
onClick={() => fetchOllamaModels(localProviderSettings.ollama?.baseUrl || 'http://localhost:11434')}
className="px-2 py-1 rounded bg-background hover:bg-background/80 text-foreground text-xs"
disabled={isLoadingModels}
>
Refresh
</button>
</div>
</div>
<div className="space-y-1">
<label className="block text-xs font-medium text-foreground/80">
Model
</label>
<div className="relative">
{isLoadingModels && (
<div className="absolute right-2 top-1/2 transform -translate-y-1/2">
<svg className="animate-spin h-4 w-4 text-foreground" xmlns="http://www.w3.org/2000/svg" fill="none" viewBox="0 0 24 24">
<circle className="opacity-25" cx="12" cy="12" r="10" stroke="currentColor" strokeWidth="4"></circle>
<path className="opacity-75" fill="currentColor" d="M4 12a8 8 0 018-8V0C5.373 0 0 5.373 0 12h4zm2 5.291A7.962 7.962 0 014 12H0c0 3.042 1.135 5.824 3 7.938l3-2.647z"></path>
</svg>
</div>
)}
<select
value={localProviderSettings.ollama?.model || ''}
onChange={handleOllamaModelChange}
className="w-full rounded-lg bg-base py-1.5 px-2 text-sm text-foreground border border-offbase shadow-sm focus:outline-none focus:ring-1 focus:ring-accent appearance-none"
disabled={isLoadingModels}
>
{ollamaModels.length > 0 ? (
ollamaModels.map((model) => (
<option key={model.id} value={model.id}>
{model.name} {model.details?.parameterSize && `(${model.details.parameterSize})`}
</option>
))
) : (
<option value="" disabled>
{ollamaModelError || "No models available - check connection"}
</option>
)}
</select>
</div>
{ollamaModelError && (
<p className="text-xs text-red-500 mt-1">
{ollamaModelError}. Check your Ollama instance is running.
</p>
)} )}
</ListboxOption> </div>
))} </div>
</ListboxOptions> )}
</Transition> </div>
</Listbox>
</div> {/* Voice Provider Selection */}
<div className="flex flex-col gap-4 p-3 bg-background rounded-lg">
<div className="space-y-2"> <div className="flex gap-4 items-center justify-between">
<label className="block text-sm font-medium text-foreground"> <h4 className="text-sm font-semibold text-foreground">Voice Synthesis</h4>
OpenAI API Key <select
{localApiKey && <span className="ml-2 text-xs text-accent">(Overriding env)</span>} value={selectedVoiceProvider.id}
</label> onChange={(e) => setLocalVoiceProvider(e.target.value as VoiceProviderType)}
<div className="flex gap-2"> className="cursor-pointer rounded-lg bg-base py-1 px-2 text-xs text-foreground border border-offbase focus:outline-none focus:ring-1 focus:ring-accent"
<Input >
type="password" {voiceProviderOptions.map((provider) => (
value={localApiKey} <option key={provider.id} value={provider.id}>
onChange={(e) => handleInputChange('apiKey', e.target.value)} {provider.name}
placeholder="Using environment variable" </option>
className="w-full rounded-lg bg-background py-2 px-3 text-foreground shadow-sm focus:outline-none focus:ring-2 focus:ring-accent" ))}
/> </select>
</div>
{/* ElevenLabs specific settings */}
{localVoiceProvider === 'elevenlabs' && (
<div className="space-y-1">
<label className="block text-xs font-medium text-foreground/80">
API Key
</label>
<input
type="password"
value={localProviderSettings.elevenlabs?.apiKey || ''}
onChange={(e: InputChangeEvent) =>
handleInputChange('providerSetting', e.target.value, 'elevenlabs', 'apiKey')}
placeholder="Enter ElevenLabs API key"
className="w-full rounded-lg bg-base py-1.5 px-2 text-sm text-foreground border border-offbase shadow-sm focus:outline-none focus:ring-1 focus:ring-accent"
/>
</div>
)}
</div>
</div> </div>
</div> )}
<div className="space-y-2"> {/* Advanced tab */}
<label className="block text-sm font-medium text-foreground"> {activeSettingsTab === 'advanced' && (
OpenAI API Base URL <div className="space-y-4">
{localBaseUrl && <span className="ml-2 text-xs text-accent">(Overriding env)</span>} {isDev && <div className="space-y-2">
</label> <label className="block text-sm font-medium text-foreground">Document Sync</label>
<div className="flex gap-2"> <div className="flex gap-2">
<Input <button
type="text" onClick={handleLoad}
value={localBaseUrl} disabled={isSyncing || isLoading}
onChange={(e) => handleInputChange('baseUrl', e.target.value)} className="justify-center rounded-lg bg-background px-3 py-1.5 text-sm
placeholder="Using environment variable" font-medium text-foreground hover:bg-background/90 focus:outline-none
className="w-full rounded-lg bg-background py-2 px-3 text-foreground shadow-sm focus:outline-none focus:ring-2 focus:ring-accent" focus-visible:ring-2 focus-visible:ring-accent focus-visible:ring-offset-2
/> transform transition-transform duration-200 ease-in-out hover:scale-[1.04] hover:text-accent
disabled:opacity-50"
>
{isLoading ? 'Loading...' : 'Load docs from Server'}
</button>
<button
onClick={handleSync}
disabled={isSyncing || isLoading}
className="justify-center rounded-lg bg-background px-3 py-1.5 text-sm
font-medium text-foreground hover:bg-background/90 focus:outline-none
focus-visible:ring-2 focus-visible:ring-accent focus-visible:ring-offset-2
transform transition-transform duration-200 ease-in-out hover:scale-[1.04] hover:text-accent
disabled:opacity-50"
>
{isSyncing ? 'Saving...' : 'Save local to Server'}
</button>
</div>
</div>}
{/* Additional advanced settings could go here */}
</div> </div>
</div> )}
{isDev && <div className="space-y-2">
<label className="block text-sm font-medium text-foreground">Document Sync</label>
<div className="flex gap-2">
<Button
onClick={handleLoad}
disabled={isSyncing || isLoading}
className="justify-center rounded-lg bg-background px-3 py-1.5 text-sm
font-medium text-foreground hover:bg-background/90 focus:outline-none
focus-visible:ring-2 focus-visible:ring-accent focus-visible:ring-offset-2
transform transition-transform duration-200 ease-in-out hover:scale-[1.04] hover:text-accent
disabled:opacity-50"
>
{isLoading ? 'Loading...' : 'Load docs from Server'}
</Button>
<Button
onClick={handleSync}
disabled={isSyncing || isLoading}
className="justify-center rounded-lg bg-background px-3 py-1.5 text-sm
font-medium text-foreground hover:bg-background/90 focus:outline-none
focus-visible:ring-2 focus-visible:ring-accent focus-visible:ring-offset-2
transform transition-transform duration-200 ease-in-out hover:scale-[1.04] hover:text-accent
disabled:opacity-50"
>
{isSyncing ? 'Saving...' : 'Save local to Server'}
</Button>
</div>
</div>}
<div className="space-y-2">
<label className="block text-sm font-medium text-foreground">Bulk Delete</label>
<div className="flex gap-2">
<Button
onClick={() => setShowClearLocalConfirm(true)}
className="justify-center rounded-lg bg-red-600 px-3 py-1.5 text-sm
font-medium text-white hover:bg-red-700 focus:outline-none
focus-visible:ring-2 focus-visible:ring-red-500 focus-visible:ring-offset-2
transform transition-transform duration-200 ease-in-out hover:scale-[1.04]"
>
Delete local docs
</Button>
{isDev && <Button
onClick={() => setShowClearServerConfirm(true)}
className="justify-center rounded-lg bg-red-600 px-3 py-1.5 text-sm
font-medium text-white hover:bg-red-700 focus:outline-none
focus-visible:ring-2 focus-visible:ring-red-500 focus-visible:ring-offset-2
transform transition-transform duration-200 ease-in-out hover:scale-[1.04]"
>
Delete server docs
</Button>}
</div>
</div>
</div> </div>
</div> </div>
<div className="mt-6 flex justify-end gap-2"> <div className="mt-6 flex justify-end gap-2">
<Button <button
type="button" type="button"
className="inline-flex justify-center rounded-lg bg-background px-3 py-1.5 text-sm className="inline-flex justify-center rounded-lg bg-background px-3 py-1.5 text-sm
font-medium text-foreground hover:bg-background/90 focus:outline-none font-medium text-foreground hover:bg-background/90 focus:outline-none
focus-visible:ring-2 focus-visible:ring-accent focus-visible:ring-offset-2 focus-visible:ring-2 focus-visible:ring-accent focus-visible:ring-offset-2
transform transition-transform duration-200 ease-in-out hover:scale-[1.04] hover:text-accent" transform transition-transform duration-200 ease-in-out hover:scale-[1.04] hover:text-accent"
onClick={async () => { onClick={resetProviderSettings}
setLocalApiKey('');
setLocalBaseUrl('');
}}
> >
Reset Reset
</Button> </button>
<Button <button
type="button" type="button"
className="inline-flex justify-center rounded-lg bg-accent px-3 py-1.5 text-sm className="inline-flex justify-center rounded-lg bg-accent px-3 py-1.5 text-sm
font-medium text-white hover:bg-accent/90 focus:outline-none font-medium text-white hover:bg-accent/90 focus:outline-none
focus-visible:ring-2 focus-visible:ring-accent focus-visible:ring-offset-2 focus-visible:ring-2 focus-visible:ring-accent focus-visible:ring-offset-2
transform transition-transform duration-200 ease-in-out hover:scale-[1.04] hover:text-background" transform transition-transform duration-200 ease-in-out hover:scale-[1.04] hover:text-background"
onClick={async () => { onClick={saveSettings}
await updateConfig({
apiKey: localApiKey || '',
baseUrl: localBaseUrl || '',
});
setIsOpen(false);
}}
> >
Done Done
</Button> </button>
</div> </div>
</DialogPanel> </Dialog.Panel>
</TransitionChild> </Transition.Child>
</div> </div>
</div> </div>
</Dialog> </Dialog>
@ -333,6 +721,6 @@ export function SettingsModal() {
confirmText="Delete" confirmText="Delete"
isDangerous={true} isDangerous={true}
/> />
</Button> </>
); );
} }

View file

@ -12,6 +12,9 @@ import { LoadingSpinner } from '@/components/Spinner';
import { VoicesControl } from '@/components/player/VoicesControl'; import { VoicesControl } from '@/components/player/VoicesControl';
import { SpeedControl } from '@/components/player/SpeedControl'; import { SpeedControl } from '@/components/player/SpeedControl';
import { Navigator } from '@/components/player/Navigator'; import { Navigator } from '@/components/player/Navigator';
import { ProviderType, VoiceProviderType } from '@/providers/types';
import { Fragment } from 'react';
import { Menu, Transition } from '@headlessui/react';
export default function TTSPlayer({ currentPage, numPages }: { export default function TTSPlayer({ currentPage, numPages }: {
currentPage?: number; currentPage?: number;
@ -26,12 +29,30 @@ export default function TTSPlayer({ currentPage, numPages }: {
setSpeedAndRestart, setSpeedAndRestart,
setVoiceAndRestart, setVoiceAndRestart,
availableVoices, availableVoices,
voiceCategories,
skipToLocation, skipToLocation,
provider,
voiceProvider,
setProviderAndRestart,
setVoiceProviderAndRestart,
} = useTTS(); } = useTTS();
// Provider options
const providerOptions = [
{ id: 'openai', name: 'OpenAI' },
{ id: 'openrouter', name: 'OpenRouter' },
{ id: 'ollama', name: 'Ollama' },
];
// Voice provider options
const voiceProviderOptions = [
{ id: 'openai', name: 'OpenAI' },
{ id: 'elevenlabs', name: 'ElevenLabs' },
];
return ( return (
<div className={`fixed bottom-4 left-1/2 transform -translate-x-1/2 z-49 transition-opacity duration-300`}> <div className="fixed bottom-4 left-1/2 transform -translate-x-1/2 z-40 transition-opacity duration-300">
<div className="bg-base dark:bg-base rounded-full shadow-lg px-3 sm:px-4 py-0.5 sm:py-1 flex items-center space-x-0.5 sm:space-x-1 relative scale-90 sm:scale-100 border border-offbase"> <div className="bg-base dark:bg-base rounded-full shadow-lg px-3 sm:px-4 py-2 flex items-center space-x-2 sm:space-x-3 relative border border-offbase">
{/* Speed control */} {/* Speed control */}
<SpeedControl setSpeedAndRestart={setSpeedAndRestart} /> <SpeedControl setSpeedAndRestart={setSpeedAndRestart} />
@ -45,35 +66,128 @@ export default function TTSPlayer({ currentPage, numPages }: {
)} )}
{/* Playback Controls */} {/* Playback Controls */}
<Button <div className="flex items-center space-x-2">
onClick={skipBackward} <Button
className="relative p-2 rounded-full text-foreground hover:bg-offbase data-[hover]:bg-offbase data-[active]:bg-offbase/80 transition-colors duration-200 focus:outline-none disabled:opacity-50" onClick={skipBackward}
aria-label="Skip backward" className="relative p-2 rounded-full text-foreground hover:bg-offbase data-[hover]:bg-offbase data-[active]:bg-offbase/80 transition-colors duration-200 focus:outline-none disabled:opacity-50"
disabled={isProcessing} aria-label="Skip backward"
> disabled={isProcessing}
{isProcessing ? <LoadingSpinner /> : <SkipBackwardIcon />} >
</Button> {isProcessing ? <LoadingSpinner /> : <SkipBackwardIcon />}
</Button>
<Button <Button
onClick={togglePlay} onClick={togglePlay}
className="relative p-2 rounded-full text-foreground hover:bg-offbase data-[hover]:bg-offbase data-[active]:bg-offbase/80 transition-colors duration-200 focus:outline-none" className="relative p-2 rounded-full text-foreground hover:bg-offbase data-[hover]:bg-offbase data-[active]:bg-offbase/80 transition-colors duration-200 focus:outline-none"
aria-label={isPlaying ? 'Pause' : 'Play'} aria-label={isPlaying ? 'Pause' : 'Play'}
disabled={isProcessing} disabled={isProcessing}
> >
{isPlaying ? <PauseIcon /> : <PlayIcon />} {isPlaying ? <PauseIcon /> : <PlayIcon />}
</Button> </Button>
<Button <Button
onClick={skipForward} onClick={skipForward}
className="relative p-2 rounded-full text-foreground hover:bg-offbase data-[hover]:bg-offbase data-[active]:bg-offbase/80 transition-colors duration-200 focus:outline-none disabled:opacity-50" className="relative p-2 rounded-full text-foreground hover:bg-offbase data-[hover]:bg-offbase data-[active]:bg-offbase/80 transition-colors duration-200 focus:outline-none disabled:opacity-50"
aria-label="Skip forward" aria-label="Skip forward"
disabled={isProcessing} disabled={isProcessing}
> >
{isProcessing ? <LoadingSpinner /> : <SkipForwardIcon />} {isProcessing ? <LoadingSpinner /> : <SkipForwardIcon />}
</Button> </Button>
</div>
{/* Voice control */} {/* Divider */}
<VoicesControl availableVoices={availableVoices} setVoiceAndRestart={setVoiceAndRestart} /> <div className="h-6 w-px bg-offbase/30 mx-0.5"></div>
{/* Provider Controls */}
<div className="flex items-center space-x-3">
{/* TTS Provider */}
<Menu as="div" className="relative inline-block text-left">
<div>
<Menu.Button className="flex items-center justify-between px-2 py-1 text-xs text-foreground hover:bg-offbase rounded border border-transparent hover:border-offbase/50 focus:outline-none">
<span className="truncate">{providerOptions.find(p => p.id === provider)?.name || 'OpenAI'}</span>
</Menu.Button>
</div>
<Transition
as={Fragment}
enter="transition ease-out duration-100"
enterFrom="transform opacity-0 scale-95"
enterTo="transform opacity-100 scale-100"
leave="transition ease-in duration-75"
leaveFrom="transform opacity-100 scale-100"
leaveTo="transform opacity-0 scale-95"
>
<Menu.Items
className="absolute right-0 bottom-full mb-1 z-50 w-40 origin-top-right rounded-md bg-base shadow-lg ring-1 ring-black ring-opacity-5 focus:outline-none"
style={{ marginBottom: '2.5rem' }}
>
{providerOptions.map((option) => (
<Menu.Item key={option.id}>
{({ active }) => (
<button
onClick={() => setProviderAndRestart(option.id as ProviderType)}
className={`${
active ? 'bg-offbase text-foreground' : 'text-foreground'
} ${
option.id === provider ? 'font-semibold' : ''
} block w-full text-left px-3 py-1.5 text-xs`}
>
{option.name}
</button>
)}
</Menu.Item>
))}
</Menu.Items>
</Transition>
</Menu>
{/* Voice Provider */}
<Menu as="div" className="relative inline-block text-left">
<div>
<Menu.Button className="flex items-center justify-between px-2 py-1 text-xs text-foreground hover:bg-offbase rounded border border-transparent hover:border-offbase/50 focus:outline-none">
<span className="truncate">{voiceProviderOptions.find(p => p.id === voiceProvider)?.name || 'OpenAI'}</span>
</Menu.Button>
</div>
<Transition
as={Fragment}
enter="transition ease-out duration-100"
enterFrom="transform opacity-0 scale-95"
enterTo="transform opacity-100 scale-100"
leave="transition ease-in duration-75"
leaveFrom="transform opacity-100 scale-100"
leaveTo="transform opacity-0 scale-95"
>
<Menu.Items
className="absolute right-0 bottom-full mb-1 z-50 w-40 origin-top-right rounded-md bg-base shadow-lg ring-1 ring-black ring-opacity-5 focus:outline-none"
style={{ marginBottom: '2.5rem' }}
>
{voiceProviderOptions.map((option) => (
<Menu.Item key={option.id}>
{({ active }) => (
<button
onClick={() => setVoiceProviderAndRestart(option.id as VoiceProviderType)}
className={`${
active ? 'bg-offbase text-foreground' : 'text-foreground'
} ${
option.id === voiceProvider ? 'font-semibold' : ''
} block w-full text-left px-3 py-1.5 text-xs`}
>
{option.name}
</button>
)}
</Menu.Item>
))}
</Menu.Items>
</Transition>
</Menu>
{/* Voice selection */}
<VoicesControl
availableVoices={availableVoices}
voiceCategories={voiceCategories}
setVoiceAndRestart={setVoiceAndRestart}
isLoading={isProcessing}
/>
</div>
</div> </div>
</div> </div>
); );

View file

@ -1,44 +1,142 @@
'use client'; 'use client';
import { import { Fragment, useState, useEffect, useMemo } from 'react';
Listbox, import { Menu, Transition } from '@headlessui/react';
ListboxButton,
ListboxOption,
ListboxOptions,
} from '@headlessui/react';
import { ChevronUpDownIcon } from '@/components/icons/Icons'; import { ChevronUpDownIcon } from '@/components/icons/Icons';
import { useConfig } from '@/contexts/ConfigContext'; import { useConfig } from '@/contexts/ConfigContext';
import { Voice, VoiceCategory } from '@/providers/types';
export const VoicesControl = ({ availableVoices, setVoiceAndRestart }: { interface VoicesControlProps {
availableVoices: string[]; availableVoices: Voice[];
voiceCategories: VoiceCategory[];
setVoiceAndRestart: (voice: string) => void; setVoiceAndRestart: (voice: string) => void;
}) => { isLoading?: boolean;
const { voice: configVoice } = useConfig(); }
// Use configVoice as the source of truth export const VoicesControl = ({
const currentVoice = configVoice; availableVoices,
voiceCategories,
setVoiceAndRestart,
isLoading = false
}: VoicesControlProps) => {
const { voice: configVoice } = useConfig();
const [searchTerm, setSearchTerm] = useState('');
// Get current voice name
const currentVoice = useMemo(() => {
const voice = availableVoices.find(v => v.id === configVoice);
return voice?.name || configVoice;
}, [availableVoices, configVoice]);
// Filter voices based on search term
const filteredVoices = useMemo(() => {
if (!searchTerm) return availableVoices;
return availableVoices.filter(voice =>
voice.name.toLowerCase().includes(searchTerm.toLowerCase()) ||
voice.id.toLowerCase().includes(searchTerm.toLowerCase())
);
}, [availableVoices, searchTerm]);
// Group voices by category
const voicesByCategory = useMemo(() => {
const grouped: Record<string, Voice[]> = {};
// Initialize empty arrays for each category
voiceCategories.forEach(category => {
grouped[category.id] = [];
});
// Assign voices to categories
filteredVoices.forEach(voice => {
const categoryId = voice.categoryId || 'uncategorized';
if (!grouped[categoryId]) {
grouped[categoryId] = [];
}
grouped[categoryId].push(voice);
});
return grouped;
}, [filteredVoices, voiceCategories]);
// Reset search term when voices change
useEffect(() => {
setSearchTerm('');
}, [availableVoices]);
return ( return (
<div className="relative"> <Menu as="div" className="relative inline-block text-left">
<Listbox value={currentVoice} onChange={setVoiceAndRestart}> <div>
<ListboxButton className="flex items-center space-x-0.5 sm:space-x-1 bg-transparent text-foreground text-xs sm:text-sm focus:outline-none cursor-pointer hover:bg-offbase rounded pl-1.5 sm:pl-2 pr-0.5 sm:pr-1 py-0.5 sm:py-1"> <Menu.Button className="flex items-center justify-between w-full px-2 py-1 text-xs text-foreground hover:bg-offbase rounded border border-transparent hover:border-offbase/50 focus:outline-none">
<span>{currentVoice}</span> <span className="truncate mr-1">{isLoading ? "Loading voices..." : currentVoice}</span>
<ChevronUpDownIcon className="h-2.5 w-2.5 sm:h-3 sm:w-3" /> <ChevronUpDownIcon className="h-3 w-3 text-foreground/70" />
</ListboxButton> </Menu.Button>
<ListboxOptions anchor='top end' className="absolute z-50 w-28 sm:w-32 overflow-auto rounded-lg bg-base shadow-lg ring-1 ring-black ring-opacity-5 focus:outline-none"> </div>
{availableVoices.map((voiceId) => ( <Transition
<ListboxOption as={Fragment}
key={voiceId} enter="transition ease-out duration-100"
value={voiceId} enterFrom="transform opacity-0 scale-95"
className={({ active, selected }) => enterTo="transform opacity-100 scale-100"
`relative cursor-pointer select-none py-0.5 px-1.5 sm:py-2 sm:px-3 ${active ? 'bg-offbase' : ''} ${selected ? 'font-medium' : ''}` leave="transition ease-in duration-75"
} leaveFrom="transform opacity-100 scale-100"
> leaveTo="transform opacity-0 scale-95"
<span className='text-xs sm:text-sm'>{voiceId}</span> >
</ListboxOption> <Menu.Items
))} className="absolute right-0 bottom-full mb-1 z-50 mt-2 w-64 origin-top-right rounded-md bg-base shadow-lg ring-1 ring-black ring-opacity-5 focus:outline-none max-h-80 overflow-auto"
</ListboxOptions> style={{ marginBottom: '2.5rem' }}
</Listbox> >
</div> <div className="sticky top-0 p-1 bg-base z-10 border-b border-offbase/20">
<input
type="text"
value={searchTerm}
onChange={(e) => setSearchTerm(e.target.value)}
placeholder="Search voices..."
onClick={(e) => e.stopPropagation()}
className="w-full px-2 py-1 text-sm bg-offbase/30 hover:bg-offbase/50 focus:bg-offbase/70 rounded text-foreground focus:outline-none"
/>
</div>
{filteredVoices.length === 0 && (
<div className="py-2 px-3 text-sm text-foreground/70">
No voices match your search
</div>
)}
{voiceCategories.map(category => {
const categoryVoices = voicesByCategory[category.id] || [];
if (categoryVoices.length === 0) return null;
return (
<div key={category.id} className="py-1">
<div className="block px-3 py-1 text-xs font-semibold text-foreground/60">
{category.name}
</div>
{categoryVoices.map(voice => (
<Menu.Item key={voice.id}>
{({ active }) => (
<button
onClick={() => setVoiceAndRestart(voice.id)}
className={`${
active ? 'bg-offbase text-foreground' : 'text-foreground'
} ${
voice.id === configVoice ? 'font-semibold' : ''
} flex w-full items-center justify-between px-3 py-1 text-left text-xs`}
>
<span className="truncate">{voice.name}</span>
{voice.categoryId === 'elevenlabs-custom' && (
<span className="text-xs bg-accent text-white px-1 py-0.5 rounded-sm">
Custom
</span>
)}
</button>
)}
</Menu.Item>
))}
</div>
);
})}
</Menu.Items>
</Transition>
</Menu>
); );
} };

View file

@ -1,11 +1,15 @@
'use client'; 'use client';
import { createContext, useContext, useEffect, useState, ReactNode } from 'react'; import { createContext, useContext, useEffect, useState, useCallback, ReactNode } from 'react';
import { ProviderType, VoiceProviderType, ProviderSettings, DocumentProviderOverride } from '@/providers/types';
import { getItem, indexedDBService, setItem, removeItem } from '@/utils/indexedDB'; import { getItem, indexedDBService, setItem, removeItem } from '@/utils/indexedDB';
/** Represents the possible view types for document display */ /** Represents the possible view types for document display */
export type ViewType = 'single' | 'dual' | 'scroll'; export type ViewType = 'single' | 'dual' | 'scroll';
// Map of document IDs to provider overrides
type DocumentOverridesMap = Map<string, DocumentProviderOverride>;
/** Configuration values for the application */ /** Configuration values for the application */
type ConfigValues = { type ConfigValues = {
apiKey: string; apiKey: string;
@ -20,6 +24,11 @@ type ConfigValues = {
footerMargin: number; footerMargin: number;
leftMargin: number; leftMargin: number;
rightMargin: number; rightMargin: number;
// New settings for provider system
provider: ProviderType;
voiceProvider: VoiceProviderType;
providerSettings: ProviderSettings;
documentProviderOverrides: Record<string, DocumentProviderOverride>;
}; };
/** Interface defining the configuration context shape and functionality */ /** Interface defining the configuration context shape and functionality */
@ -36,8 +45,25 @@ interface ConfigContextType {
footerMargin: number; footerMargin: number;
leftMargin: number; leftMargin: number;
rightMargin: number; rightMargin: number;
updateConfig: (newConfig: Partial<{ apiKey: string; baseUrl: string; viewType: ViewType }>) => Promise<void>; // New properties for provider system
provider: ProviderType;
voiceProvider: VoiceProviderType;
providerSettings: ProviderSettings;
documentProviderOverrides: Record<string, DocumentProviderOverride>;
// Updated updateConfig to include provider settings
updateConfig: (newConfig: Partial<{
apiKey: string;
baseUrl: string;
viewType: ViewType;
provider: ProviderType;
voiceProvider: VoiceProviderType;
providerSettings: ProviderSettings;
}>) => Promise<void>;
updateConfigKey: <K extends keyof ConfigValues>(key: K, value: ConfigValues[K]) => Promise<void>; updateConfigKey: <K extends keyof ConfigValues>(key: K, value: ConfigValues[K]) => Promise<void>;
// New methods for document-specific overrides
getDocumentProviderOverride: (documentId: string) => DocumentProviderOverride | undefined;
setDocumentProviderOverride: (documentId: string, override: DocumentProviderOverride) => Promise<void>;
removeDocumentProviderOverride: (documentId: string) => Promise<void>;
isLoading: boolean; isLoading: boolean;
isDBReady: boolean; isDBReady: boolean;
} }
@ -63,6 +89,29 @@ export function ConfigProvider({ children }: { children: ReactNode }) {
const [headerMargin, setHeaderMargin] = useState<number>(0.07); const [headerMargin, setHeaderMargin] = useState<number>(0.07);
const [footerMargin, setFooterMargin] = useState<number>(0.07); const [footerMargin, setFooterMargin] = useState<number>(0.07);
const [leftMargin, setLeftMargin] = useState<number>(0.07); const [leftMargin, setLeftMargin] = useState<number>(0.07);
// New state for provider system
const [provider, setProvider] = useState<ProviderType>('openai');
const [voiceProvider, setVoiceProvider] = useState<VoiceProviderType>('openai');
const [providerSettings, setProviderSettings] = useState<ProviderSettings>({
openai: {
apiKey: '',
baseUrl: 'https://api.openai.com/v1',
model: 'tts-1',
},
ollama: {
baseUrl: 'http://localhost:11434',
model: 'llama3:8b',
},
openrouter: {
apiKey: '',
model: 'openai/whisper',
},
elevenlabs: {
apiKey: '',
},
});
const [documentProviderOverrides, setDocumentProviderOverrides] = useState<Record<string, DocumentProviderOverride>>({});
const [rightMargin, setRightMargin] = useState<number>(0.07); const [rightMargin, setRightMargin] = useState<number>(0.07);
const [isLoading, setIsLoading] = useState(true); const [isLoading, setIsLoading] = useState(true);
@ -87,6 +136,12 @@ export function ConfigProvider({ children }: { children: ReactNode }) {
const cachedHeaderMargin = await getItem('headerMargin'); const cachedHeaderMargin = await getItem('headerMargin');
const cachedFooterMargin = await getItem('footerMargin'); const cachedFooterMargin = await getItem('footerMargin');
const cachedLeftMargin = await getItem('leftMargin'); const cachedLeftMargin = await getItem('leftMargin');
// Load new provider settings
const cachedProvider = await getItem('provider');
const cachedVoiceProvider = await getItem('voiceProvider');
const cachedProviderSettings = await getItem('providerSettings');
const cachedDocumentOverrides = await getItem('documentProviderOverrides');
const cachedRightMargin = await getItem('rightMargin'); const cachedRightMargin = await getItem('rightMargin');
// Only set API key and base URL if they were explicitly saved by the user // Only set API key and base URL if they were explicitly saved by the user
@ -109,6 +164,26 @@ export function ConfigProvider({ children }: { children: ReactNode }) {
setHeaderMargin(parseFloat(cachedHeaderMargin || '0.07')); setHeaderMargin(parseFloat(cachedHeaderMargin || '0.07'));
setFooterMargin(parseFloat(cachedFooterMargin || '0.07')); setFooterMargin(parseFloat(cachedFooterMargin || '0.07'));
setLeftMargin(parseFloat(cachedLeftMargin || '0.07')); setLeftMargin(parseFloat(cachedLeftMargin || '0.07'));
// Set provider settings with defaults
setProvider((cachedProvider || 'openai') as ProviderType);
setVoiceProvider((cachedVoiceProvider || 'openai') as VoiceProviderType);
if (cachedProviderSettings) {
try {
const parsedSettings = JSON.parse(cachedProviderSettings);
setProviderSettings({
...providerSettings,
...parsedSettings
});
} catch (e) {
console.error('Error parsing provider settings:', e);
}
}
if (cachedDocumentOverrides) {
setDocumentProviderOverrides(JSON.parse(cachedDocumentOverrides));
}
setRightMargin(parseFloat(cachedRightMargin || '0.07')); setRightMargin(parseFloat(cachedRightMargin || '0.07'));
// Only save non-sensitive settings by default // Only save non-sensitive settings by default
@ -129,6 +204,16 @@ export function ConfigProvider({ children }: { children: ReactNode }) {
if (cachedLeftMargin === null) await setItem('leftMargin', '0.0'); if (cachedLeftMargin === null) await setItem('leftMargin', '0.0');
if (cachedRightMargin === null) await setItem('rightMargin', '0.0'); if (cachedRightMargin === null) await setItem('rightMargin', '0.0');
// Save new default settings
if (cachedProvider === null) await setItem('provider', 'openai');
if (cachedVoiceProvider === null) await setItem('voiceProvider', 'openai');
if (cachedProviderSettings === null) {
await setItem('providerSettings', JSON.stringify(providerSettings));
}
if (cachedDocumentOverrides === null) {
await setItem('documentProviderOverrides', JSON.stringify({}));
}
} catch (error) { } catch (error) {
console.error('Error initializing:', error); console.error('Error initializing:', error);
} finally { } finally {
@ -143,7 +228,13 @@ export function ConfigProvider({ children }: { children: ReactNode }) {
* Updates multiple configuration values simultaneously * Updates multiple configuration values simultaneously
* Only saves API credentials if they are explicitly set * Only saves API credentials if they are explicitly set
*/ */
const updateConfig = async (newConfig: Partial<{ apiKey: string; baseUrl: string }>) => { const updateConfig = async (newConfig: Partial<{
apiKey: string;
baseUrl: string;
provider: ProviderType;
voiceProvider: VoiceProviderType;
providerSettings: ProviderSettings;
}>) => {
try { try {
if (newConfig.apiKey !== undefined || newConfig.apiKey !== '') { if (newConfig.apiKey !== undefined || newConfig.apiKey !== '') {
// Only save API key to IndexedDB if it's different from env default // Only save API key to IndexedDB if it's different from env default
@ -155,6 +246,24 @@ export function ConfigProvider({ children }: { children: ReactNode }) {
await setItem('baseUrl', newConfig.baseUrl!); await setItem('baseUrl', newConfig.baseUrl!);
setBaseUrl(newConfig.baseUrl!); setBaseUrl(newConfig.baseUrl!);
} }
// Handle provider settings
if (newConfig.provider !== undefined) {
await setItem('provider', newConfig.provider);
setProvider(newConfig.provider);
}
if (newConfig.voiceProvider !== undefined) {
await setItem('voiceProvider', newConfig.voiceProvider);
setVoiceProvider(newConfig.voiceProvider);
}
if (newConfig.providerSettings !== undefined) {
await setItem('providerSettings', JSON.stringify(newConfig.providerSettings));
setProviderSettings((prevSettings: ProviderSettings) => ({
...prevSettings, ...newConfig.providerSettings
}));
}
// Delete completely if '' is passed // Delete completely if '' is passed
if (newConfig.apiKey === '') { if (newConfig.apiKey === '') {
@ -217,6 +326,18 @@ export function ConfigProvider({ children }: { children: ReactNode }) {
case 'rightMargin': case 'rightMargin':
setRightMargin(value as number); setRightMargin(value as number);
break; break;
case 'provider':
setProvider(value as ProviderType);
break;
case 'voiceProvider':
setVoiceProvider(value as VoiceProviderType);
break;
case 'providerSettings':
setProviderSettings(value as ProviderSettings);
break;
case 'documentProviderOverrides':
setDocumentProviderOverrides(value as Record<string, DocumentProviderOverride>);
break;
} }
} catch (error) { } catch (error) {
console.error(`Error updating config key ${key}:`, error); console.error(`Error updating config key ${key}:`, error);
@ -224,6 +345,55 @@ export function ConfigProvider({ children }: { children: ReactNode }) {
} }
}; };
/**
* Gets document-specific provider override if it exists
*
* @param documentId The document ID to check for overrides
* @returns The document override or undefined if none exists
*/
const getDocumentProviderOverride = useCallback((documentId: string): DocumentProviderOverride | undefined => {
return documentProviderOverrides[documentId];
}, [documentProviderOverrides]);
/**
* Sets document-specific provider override
*
* @param documentId The document ID to set override for
* @param override The provider override settings
*/
const setDocumentProviderOverride = async (documentId: string, override: DocumentProviderOverride): Promise<void> => {
try {
const newOverrides = {
...documentProviderOverrides,
[documentId]: override
};
await setItem('documentProviderOverrides', JSON.stringify(newOverrides));
setDocumentProviderOverrides(newOverrides);
} catch (error) {
console.error(`Error setting document provider override for ${documentId}:`, error);
throw error;
}
};
/**
* Removes document-specific provider override
*
* @param documentId The document ID to remove override for
*/
const removeDocumentProviderOverride = async (documentId: string): Promise<void> => {
try {
const newOverrides = { ...documentProviderOverrides };
delete newOverrides[documentId];
await setItem('documentProviderOverrides', JSON.stringify(newOverrides));
setDocumentProviderOverrides(newOverrides);
} catch (error) {
console.error(`Error removing document provider override for ${documentId}:`, error);
throw error;
}
};
return ( return (
<ConfigContext.Provider value={{ <ConfigContext.Provider value={{
apiKey, apiKey,
@ -238,8 +408,15 @@ export function ConfigProvider({ children }: { children: ReactNode }) {
footerMargin, footerMargin,
leftMargin, leftMargin,
rightMargin, rightMargin,
provider,
voiceProvider,
providerSettings,
documentProviderOverrides,
updateConfig, updateConfig,
updateConfigKey, updateConfigKey,
getDocumentProviderOverride,
setDocumentProviderOverride,
removeDocumentProviderOverride,
isLoading, isLoading,
isDBReady isDBReady
}}> }}>

View file

@ -58,6 +58,9 @@ export function EPUBProvider({ children }: { children: ReactNode }) {
baseUrl, baseUrl,
voiceSpeed, voiceSpeed,
voice, voice,
provider,
voiceProvider,
providerSettings
} = useConfig(); } = useConfig();
// Current document state // Current document state
const [currDocData, setCurrDocData] = useState<ArrayBuffer>(); const [currDocData, setCurrDocData] = useState<ArrayBuffer>();
@ -242,17 +245,21 @@ export function EPUBProvider({ children }: { children: ReactNode }) {
try { try {
const audioBuffer = await withRetry( const audioBuffer = await withRetry(
async () => { async () => {
// Include provider information in the request
const ttsResponse = await fetch('/api/tts', { const ttsResponse = await fetch('/api/tts', {
method: 'POST', method: 'POST',
headers: { headers: {
'x-openai-key': apiKey, 'Content-Type': 'application/json',
'x-openai-base-url': baseUrl, 'x-provider-settings': JSON.stringify(providerSettings),
}, },
body: JSON.stringify({ body: JSON.stringify({
text: trimmedText, text: trimmedText,
voice: voice, voice: voice,
speed: voiceSpeed, speed: voiceSpeed,
format: format === 'm4b' ? 'aac' : 'mp3', format: format === 'm4b' ? 'aac' : 'mp3',
provider: provider,
voiceProvider: voiceProvider,
documentId: id?.toString()
}), }),
signal signal
}); });
@ -320,7 +327,7 @@ export function EPUBProvider({ children }: { children: ReactNode }) {
console.error('Error creating audiobook:', error); console.error('Error creating audiobook:', error);
throw error; throw error;
} }
}, [extractBookText, apiKey, baseUrl, voice, voiceSpeed]); }, [extractBookText, id, voice, voiceSpeed, provider, voiceProvider, providerSettings]);
const setRendition = useCallback((rendition: Rendition) => { const setRendition = useCallback((rendition: Rendition) => {
bookRef.current = rendition.book; bookRef.current = rendition.book;

View file

@ -2,7 +2,7 @@
* Text-to-Speech (TTS) Context Provider * Text-to-Speech (TTS) Context Provider
* *
* This module provides a React context for managing text-to-speech functionality. * This module provides a React context for managing text-to-speech functionality.
* It handles audio playback, sentence processing, and integration with OpenAI's TTS API. * It handles audio playback, sentence processing, and integration with multiple TTS providers.
* *
* Key features: * Key features:
* - Audio playback control (play/pause/skip) * - Audio playback control (play/pause/skip)
@ -10,6 +10,8 @@
* - Audio caching for better performance * - Audio caching for better performance
* - Voice and speed control * - Voice and speed control
* - Document navigation * - Document navigation
* - Multiple AI provider support (OpenAI, Ollama, OpenRouter)
* - ElevenLabs voice synthesis integration
*/ */
'use client'; 'use client';
@ -28,6 +30,7 @@ import { Howl } from 'howler';
import toast from 'react-hot-toast'; import toast from 'react-hot-toast';
import { useParams } from 'next/navigation'; import { useParams } from 'next/navigation';
import { ProviderType, VoiceProviderType, Voice, VoiceCategory } from '@/providers/types';
import { useConfig } from '@/contexts/ConfigContext'; import { useConfig } from '@/contexts/ConfigContext';
import { useAudioCache } from '@/hooks/audio/useAudioCache'; import { useAudioCache } from '@/hooks/audio/useAudioCache';
import { useVoiceManagement } from '@/hooks/audio/useVoiceManagement'; import { useVoiceManagement } from '@/hooks/audio/useVoiceManagement';
@ -52,16 +55,20 @@ interface TTSContextType {
isPlaying: boolean; isPlaying: boolean;
isProcessing: boolean; isProcessing: boolean;
currentSentence: string; currentSentence: string;
isBackgrounded: boolean; // Add this new property isBackgrounded: boolean;
// Navigation // Navigation
currDocPage: string | number; // Change this to allow both types currDocPage: string | number;
currDocPageNumber: number; // For PDF currDocPageNumber: number; // For PDF
currDocPages: number | undefined; currDocPages: number | undefined;
// Voice settings // Voice and provider settings
availableVoices: string[]; availableVoices: Voice[];
voiceCategories: VoiceCategory[];
voiceIds: string[];
provider: ProviderType;
voiceProvider: VoiceProviderType;
// Control functions // Control functions
togglePlay: () => void; togglePlay: () => void;
skipForward: () => void; skipForward: () => void;
@ -73,8 +80,10 @@ interface TTSContextType {
setCurrDocPages: (num: number | undefined) => void; setCurrDocPages: (num: number | undefined) => void;
setSpeedAndRestart: (speed: number) => void; setSpeedAndRestart: (speed: number) => void;
setVoiceAndRestart: (voice: string) => void; setVoiceAndRestart: (voice: string) => void;
setProviderAndRestart: (provider: ProviderType) => void;
setVoiceProviderAndRestart: (voiceProvider: VoiceProviderType) => void;
skipToLocation: (location: string | number) => void; skipToLocation: (location: string | number) => void;
registerLocationChangeHandler: (handler: (location: string | number) => void) => void; // EPUB-only: Handles chapter navigation registerLocationChangeHandler: (handler: (location: string | number) => void) => void;
setIsEPUB: (isEPUB: boolean) => void; setIsEPUB: (isEPUB: boolean) => void;
} }
@ -83,28 +92,53 @@ const TTSContext = createContext<TTSContextType | undefined>(undefined);
/** /**
* Main provider component that manages the TTS state and functionality. * Main provider component that manages the TTS state and functionality.
* Handles initialization of OpenAI client, audio context, and media session. * Handles multiple TTS providers, audio context, and media session.
* *
* @param {Object} props - Component props * @param {Object} props - Component props
* @param {ReactNode} props.children - Child components to be wrapped by the provider * @param {ReactNode} props.children - Child components to be wrapped by the provider
* @returns {JSX.Element} TTSProvider component * @returns {JSX.Element} TTSProvider component
*/ */
export function TTSProvider({ children }: { children: ReactNode }) { export function TTSProvider({ children }: { children: ReactNode }) {
// Get document ID from URL params
const { id } = useParams();
// Configuration context consumption // Configuration context consumption
const { const {
apiKey: openApiKey,
baseUrl: openApiBaseUrl,
isLoading: configIsLoading, isLoading: configIsLoading,
voiceSpeed, voiceSpeed,
voice: configVoice, voice: configVoice,
updateConfigKey, updateConfigKey,
skipBlank, skipBlank,
provider: configProvider,
voiceProvider: configVoiceProvider,
providerSettings,
getDocumentProviderOverride,
} = useConfig(); } = useConfig();
// Remove OpenAI client reference as it's no longer needed // Get document-specific provider override if available
const documentOverride = id ? getDocumentProviderOverride(id.toString()) : undefined;
// Use document override or default provider settings
const [provider, setProvider] = useState<ProviderType>(
documentOverride?.provider || configProvider
);
const [voiceProvider, setVoiceProvider] = useState<VoiceProviderType>(
documentOverride?.voiceProvider || configVoiceProvider
);
// Initialize core services
const audioContext = useAudioContext(); const audioContext = useAudioContext();
const audioCache = useAudioCache(25); const audioCache = useAudioCache(25);
const { availableVoices, fetchVoices } = useVoiceManagement(openApiKey, openApiBaseUrl);
// Initialize voice management with effective provider settings
const {
availableVoices,
voiceCategories,
voiceIds,
fetchVoices,
clearCache: clearVoiceCache,
isLoading: voicesLoading
} = useVoiceManagement(provider, voiceProvider, providerSettings, id?.toString());
// Add ref for location change handler // Add ref for location change handler
const locationChangeHandlerRef = useRef<((location: string | number) => void) | null>(null); const locationChangeHandlerRef = useRef<((location: string | number) => void) | null>(null);
@ -119,9 +153,6 @@ export function TTSProvider({ children }: { children: ReactNode }) {
locationChangeHandlerRef.current = handler; locationChangeHandlerRef.current = handler;
}, []); }, []);
// Get document ID from URL params
const { id } = useParams();
/** /**
* State Management * State Management
*/ */
@ -137,7 +168,7 @@ export function TTSProvider({ children }: { children: ReactNode }) {
const [currentIndex, setCurrentIndex] = useState(0); const [currentIndex, setCurrentIndex] = useState(0);
const [activeHowl, setActiveHowl] = useState<Howl | null>(null); const [activeHowl, setActiveHowl] = useState<Howl | null>(null);
const [speed, setSpeed] = useState(voiceSpeed); const [speed, setSpeed] = useState(voiceSpeed);
const [voice, setVoice] = useState(configVoice); const [voice, setVoice] = useState(documentOverride?.voice || configVoice);
// Track pending preload requests // Track pending preload requests
const preloadRequests = useRef<Map<string, Promise<string>>>(new Map()); const preloadRequests = useRef<Map<string, Promise<string>>>(new Map());
@ -197,7 +228,6 @@ export function TTSProvider({ children }: { children: ReactNode }) {
* Works for both PDF pages and EPUB locations * Works for both PDF pages and EPUB locations
* *
* @param {string | number} location - The target location to navigate to * @param {string | number} location - The target location to navigate to
* @param {boolean} keepPlaying - Whether to maintain playback state
*/ */
const skipToLocation = useCallback((location: string | number) => { const skipToLocation = useCallback((location: string | number) => {
// Reset state for new content in correct order // Reset state for new content in correct order
@ -205,7 +235,6 @@ export function TTSProvider({ children }: { children: ReactNode }) {
setCurrentIndex(0); setCurrentIndex(0);
setSentences([]); setSentences([]);
setCurrDocPage(location); setCurrDocPage(location);
}, [abortAudio]); }, [abortAudio]);
/** /**
@ -233,7 +262,7 @@ export function TTSProvider({ children }: { children: ReactNode }) {
// Handle next/previous page transitions // Handle next/previous page transitions
if ((nextIndex >= sentences.length && currDocPageNumber < currDocPages!) || if ((nextIndex >= sentences.length && currDocPageNumber < currDocPages!) ||
(nextIndex < 0 && currDocPageNumber > 1)) { (nextIndex < 0 && currDocPageNumber > 1)) {
// Pass wasPlaying to maintain playback state during page turn // Navigate to next/previous page
skipToLocation(currDocPageNumber + (nextIndex >= sentences.length ? 1 : -1)); skipToLocation(currDocPageNumber + (nextIndex >= sentences.length ? 1 : -1));
return; return;
} }
@ -248,7 +277,7 @@ export function TTSProvider({ children }: { children: ReactNode }) {
/** /**
* Handles blank text sections based on document type * Handles blank text sections based on document type
* *
* @param {string[]} sentences - Array of processed sentences * @param {string} text - The text to check for blankness
* @returns {boolean} - True if blank section was handled * @returns {boolean} - True if blank section was handled
*/ */
const handleBlankSection = useCallback((text: string): boolean => { const handleBlankSection = useCallback((text: string): boolean => {
@ -280,6 +309,7 @@ export function TTSProvider({ children }: { children: ReactNode }) {
* Sets the current text and splits it into sentences * Sets the current text and splits it into sentences
* *
* @param {string} text - The text to be processed * @param {string} text - The text to be processed
* @param {boolean} [shouldPause=false] - Whether to pause after processing
*/ */
const setText = useCallback((text: string, shouldPause = false) => { const setText = useCallback((text: string, shouldPause = false) => {
// Check for blank section first // Check for blank section first
@ -346,7 +376,6 @@ export function TTSProvider({ children }: { children: ReactNode }) {
setIsPlaying(false); setIsPlaying(false);
}, [abortAudio]); }, [abortAudio]);
/** /**
* Moves forward one sentence in the text * Moves forward one sentence in the text
*/ */
@ -375,9 +404,9 @@ export function TTSProvider({ children }: { children: ReactNode }) {
* Updates the voice and speed settings from the configuration * Updates the voice and speed settings from the configuration
*/ */
const updateVoiceAndSpeed = useCallback(() => { const updateVoiceAndSpeed = useCallback(() => {
setVoice(configVoice); setVoice(documentOverride?.voice || configVoice);
setSpeed(voiceSpeed); setSpeed(voiceSpeed);
}, [configVoice, voiceSpeed]); }, [configVoice, voiceSpeed, documentOverride]);
/** /**
* Initializes configuration and fetches available voices * Initializes configuration and fetches available voices
@ -387,13 +416,13 @@ export function TTSProvider({ children }: { children: ReactNode }) {
fetchVoices(); fetchVoices();
updateVoiceAndSpeed(); updateVoiceAndSpeed();
} }
}, [configIsLoading, openApiKey, openApiBaseUrl, updateVoiceAndSpeed, fetchVoices]); }, [configIsLoading, updateVoiceAndSpeed, fetchVoices]);
/** /**
* Generates and plays audio for the current sentence * Generates and plays audio for the current sentence
* *
* @param {string} sentence - The sentence to generate audio for * @param {string} sentence - The sentence to generate audio for
* @returns {Promise<AudioBuffer | undefined>} The generated audio buffer * @returns {Promise<ArrayBuffer | undefined>} The generated audio buffer
*/ */
const getAudio = useCallback(async (sentence: string): Promise<ArrayBuffer | undefined> => { const getAudio = useCallback(async (sentence: string): Promise<ArrayBuffer | undefined> => {
// Check if the audio is already cached // Check if the audio is already cached
@ -416,19 +445,22 @@ export function TTSProvider({ children }: { children: ReactNode }) {
method: 'POST', method: 'POST',
headers: { headers: {
'Content-Type': 'application/json', 'Content-Type': 'application/json',
'x-openai-key': openApiKey || '', 'x-provider-settings': JSON.stringify(providerSettings),
'x-openai-base-url': openApiBaseUrl || '',
}, },
body: JSON.stringify({ body: JSON.stringify({
text: sentence, text: sentence,
voice: voice, voice: voice,
speed: speed, speed: speed,
provider: provider,
voiceProvider: voiceProvider,
documentId: id?.toString()
}), }),
signal: controller.signal, signal: controller.signal,
}); });
if (!response.ok) { if (!response.ok) {
throw new Error('Failed to generate audio'); const errorData = await response.json().catch(() => ({ error: 'Unknown error' }));
throw new Error(errorData.error || 'Failed to generate audio');
} }
return response.arrayBuffer(); return response.arrayBuffer();
@ -466,7 +498,7 @@ export function TTSProvider({ children }: { children: ReactNode }) {
}); });
throw error; throw error;
} }
}, [voice, speed, audioCache, openApiKey, openApiBaseUrl]); }, [voice, speed, audioCache, provider, voiceProvider, providerSettings, id]);
/** /**
* Processes and plays the current sentence * Processes and plays the current sentence
@ -743,7 +775,6 @@ export function TTSProvider({ children }: { children: ReactNode }) {
*/ */
const stopAndPlayFromIndex = useCallback((index: number) => { const stopAndPlayFromIndex = useCallback((index: number) => {
abortAudio(); abortAudio();
setCurrentIndex(index); setCurrentIndex(index);
setIsPlaying(true); setIsPlaying(true);
}, [abortAudio]); }, [abortAudio]);
@ -784,6 +815,8 @@ export function TTSProvider({ children }: { children: ReactNode }) {
* @param {string} newVoice - The new voice to set * @param {string} newVoice - The new voice to set
*/ */
const setVoiceAndRestart = useCallback((newVoice: string) => { const setVoiceAndRestart = useCallback((newVoice: string) => {
if (newVoice === voice) return;
const wasPlaying = isPlaying; const wasPlaying = isPlaying;
// Set a flag to prevent double audio requests during config update // Set a flag to prevent double audio requests during config update
@ -806,7 +839,96 @@ export function TTSProvider({ children }: { children: ReactNode }) {
setIsPlaying(true); setIsPlaying(true);
} }
}); });
}, [abortAudio, updateConfigKey, audioCache, isPlaying]); }, [abortAudio, updateConfigKey, audioCache, isPlaying, voice]);
/**
* Sets the provider and restarts the playback
*
* @param {ProviderType} newProvider - The new provider to set
*/
const setProviderAndRestart = useCallback((newProvider: ProviderType) => {
if (newProvider === provider) return;
const wasPlaying = isPlaying;
// Set a flag to prevent double audio requests during config update
setIsProcessing(true);
// First stop any current playback
setIsPlaying(false);
abortAudio(true); // Clear pending requests since provider changed
setActiveHowl(null);
// Update provider, clear cache
setProvider(newProvider);
audioCache.clear();
clearVoiceCache(); // Clear voice cache too since provider changed
// Update config after state changes
updateConfigKey('provider', newProvider).then(() => {
setIsProcessing(false);
// Resume playback if it was playing before
if (wasPlaying) {
setIsPlaying(true);
}
});
}, [abortAudio, updateConfigKey, audioCache, isPlaying, provider, clearVoiceCache]);
/**
* Sets the voice provider and restarts the playback
* Improved to fetch voices immediately after changing provider
*
* @param {VoiceProviderType} newVoiceProvider - The new voice provider to set
*/
const setVoiceProviderAndRestart = useCallback((newVoiceProvider: VoiceProviderType) => {
if (newVoiceProvider === voiceProvider) return;
const wasPlaying = isPlaying;
// Set a flag to prevent double audio requests during config update
setIsProcessing(true);
// First stop any current playback
setIsPlaying(false);
abortAudio(true); // Clear pending requests since provider changed
setActiveHowl(null);
// Update provider, clear cache
setVoiceProvider(newVoiceProvider);
audioCache.clear();
clearVoiceCache(); // Clear voice cache too since provider changed
// Immediately trigger voice fetch to ensure voices are loaded
// This is critical for ElevenLabs voices which need to be fetched
const fetchVoicesPromise = fetchVoices();
// Update config in parallel
const updateConfigPromise = updateConfigKey('voiceProvider', newVoiceProvider);
// Wait for both operations to complete
Promise.all([fetchVoicesPromise, updateConfigPromise])
.then(() => {
setIsProcessing(false);
// Resume playback if it was playing before
if (wasPlaying) {
setIsPlaying(true);
}
})
.catch(error => {
console.error("Error updating voice provider:", error);
setIsProcessing(false);
toast.error('Failed to load voices for the selected provider', {
style: {
background: 'var(--background)',
color: 'var(--accent)',
},
duration: 3000,
});
});
}, [abortAudio, updateConfigKey, audioCache, isPlaying, voiceProvider, clearVoiceCache, fetchVoices]);
/** /**
* Provides the TTS context value to child components * Provides the TTS context value to child components
@ -819,7 +941,13 @@ export function TTSProvider({ children }: { children: ReactNode }) {
currDocPage, currDocPage,
currDocPageNumber, currDocPageNumber,
currDocPages, currDocPages,
// Voice and provider settings
availableVoices, availableVoices,
voiceCategories,
voiceIds,
provider,
voiceProvider,
// Control functions
togglePlay, togglePlay,
skipForward, skipForward,
skipBackward, skipBackward,
@ -830,6 +958,8 @@ export function TTSProvider({ children }: { children: ReactNode }) {
setCurrDocPages, setCurrDocPages,
setSpeedAndRestart, setSpeedAndRestart,
setVoiceAndRestart, setVoiceAndRestart,
setProviderAndRestart,
setVoiceProviderAndRestart,
skipToLocation, skipToLocation,
registerLocationChangeHandler, registerLocationChangeHandler,
setIsEPUB setIsEPUB
@ -842,7 +972,13 @@ export function TTSProvider({ children }: { children: ReactNode }) {
currDocPage, currDocPage,
currDocPageNumber, currDocPageNumber,
currDocPages, currDocPages,
// Voice and provider settings
availableVoices, availableVoices,
voiceCategories,
voiceIds,
provider,
voiceProvider,
// Control functions
togglePlay, togglePlay,
skipForward, skipForward,
skipBackward, skipBackward,
@ -853,6 +989,8 @@ export function TTSProvider({ children }: { children: ReactNode }) {
setCurrDocPages, setCurrDocPages,
setSpeedAndRestart, setSpeedAndRestart,
setVoiceAndRestart, setVoiceAndRestart,
setProviderAndRestart,
setVoiceProviderAndRestart,
skipToLocation, skipToLocation,
registerLocationChangeHandler, registerLocationChangeHandler,
setIsEPUB setIsEPUB
@ -868,7 +1006,7 @@ export function TTSProvider({ children }: { children: ReactNode }) {
// Load last location on mount for EPUB only // Load last location on mount for EPUB only
useEffect(() => { useEffect(() => {
if (id && isEPUB) { if (id && isEPUB) {
getLastDocumentLocation(id as string).then(lastLocation => { getLastDocumentLocation(id.toString()).then(lastLocation => {
if (lastLocation) { if (lastLocation) {
console.log('Setting last location:', lastLocation); console.log('Setting last location:', lastLocation);
if (locationChangeHandlerRef.current) { if (locationChangeHandlerRef.current) {

View file

@ -1,38 +1,153 @@
'use client'; 'use client';
import { useState, useCallback } from 'react'; import { useState, useCallback, useEffect } from 'react';
import { Voice, VoiceCategory, ProviderType, VoiceProviderType, ProviderSettings } from '@/providers/types';
const DEFAULT_VOICES = ['alloy', 'ash', 'coral', 'echo', 'fable', 'onyx', 'nova', 'sage', 'shimmer']; // Simple memory cache for voice data to prevent unnecessary fetches
let voiceCache: Record<string, {
voices: Voice[],
categories: VoiceCategory[],
timestamp: number
}> = {};
// Default voices for fallback
const DEFAULT_VOICES: Voice[] = [
{ id: 'alloy', name: 'Alloy', provider: 'openai', categoryId: 'openai-default' },
{ id: 'echo', name: 'Echo', provider: 'openai', categoryId: 'openai-default' },
{ id: 'fable', name: 'Fable', provider: 'openai', categoryId: 'openai-default' },
{ id: 'onyx', name: 'Onyx', provider: 'openai', categoryId: 'openai-default' },
{ id: 'nova', name: 'Nova', provider: 'openai', categoryId: 'openai-default' },
{ id: 'shimmer', name: 'Shimmer', provider: 'openai', categoryId: 'openai-default' },
];
// Default categories for fallback
const DEFAULT_CATEGORIES: VoiceCategory[] = [
{ id: 'openai-default', name: 'OpenAI Voices', provider: 'openai' }
];
// Cache lifetime in milliseconds (5 minutes)
const CACHE_LIFETIME = 5 * 60 * 1000;
/** /**
* Custom hook for managing TTS voices * Custom hook for managing TTS voices
* @param apiKey OpenAI API key * @param provider The selected provider type
* @param baseUrl OpenAI API base URL * @param voiceProvider The selected voice provider type
* @param providerSettings Provider configuration settings
* @returns Object containing available voices and fetch function * @returns Object containing available voices and fetch function
*/ */
export function useVoiceManagement(apiKey: string | undefined, baseUrl: string | undefined) { export function useVoiceManagement(
const [availableVoices, setAvailableVoices] = useState<string[]>([]); provider: ProviderType,
voiceProvider: VoiceProviderType,
providerSettings: ProviderSettings,
documentId?: string
) {
const [availableVoices, setAvailableVoices] = useState<Voice[]>(DEFAULT_VOICES);
const [voiceCategories, setVoiceCategories] = useState<VoiceCategory[]>(DEFAULT_CATEGORIES);
const [isLoading, setIsLoading] = useState(false);
const [error, setError] = useState<string | null>(null);
// Generate cache key based on provider configuration
const getCacheKey = useCallback(() => {
return `${provider}-${voiceProvider}-${JSON.stringify(providerSettings)}`;
}, [provider, voiceProvider, providerSettings]);
/**
* Fetches available voices from the selected provider
*/
const fetchVoices = useCallback(async () => { const fetchVoices = useCallback(async () => {
try { try {
console.log('Fetching voices...'); setIsLoading(true);
setError(null);
const cacheKey = getCacheKey();
// Check for cached data first
const cachedData = voiceCache[cacheKey];
if (cachedData && (Date.now() - cachedData.timestamp) < CACHE_LIFETIME) {
console.log('Using cached voice data');
setAvailableVoices(cachedData.voices);
setVoiceCategories(cachedData.categories);
setIsLoading(false);
return;
}
console.log(`Fetching voices from ${provider}/${voiceProvider}...`);
const response = await fetch('/api/tts/voices', { const response = await fetch('/api/tts/voices', {
headers: { headers: {
'x-openai-key': apiKey || '', 'x-provider': provider,
'x-openai-base-url': baseUrl || '', 'x-voice-provider': voiceProvider,
'x-provider-settings': JSON.stringify(providerSettings),
'Content-Type': 'application/json', 'Content-Type': 'application/json',
}, },
}); });
if (!response.ok) throw new Error('Failed to fetch voices'); if (!response.ok) {
throw new Error(`Failed to fetch voices: ${response.statusText}`);
}
const data = await response.json(); const data = await response.json();
setAvailableVoices(data.voices || DEFAULT_VOICES);
// Cache the voice data
voiceCache[cacheKey] = {
voices: data.voices?.length ? data.voices : DEFAULT_VOICES,
categories: data.categories?.length ? data.categories : DEFAULT_CATEGORIES,
timestamp: Date.now()
};
setAvailableVoices(data.voices?.length ? data.voices : DEFAULT_VOICES);
setVoiceCategories(data.categories?.length ? data.categories : DEFAULT_CATEGORIES);
} catch (error) { } catch (error) {
console.error('Error fetching voices:', error); console.error('Error fetching voices:', error);
// Set available voices to default openai voices setError(error instanceof Error ? error.message : 'Unknown error fetching voices');
setAvailableVoices(DEFAULT_VOICES); setAvailableVoices(DEFAULT_VOICES);
setVoiceCategories(DEFAULT_CATEGORIES);
} finally {
setIsLoading(false);
} }
}, [apiKey, baseUrl]); }, [provider, voiceProvider, providerSettings, getCacheKey]);
return { availableVoices, fetchVoices }; // Fetch voices when provider settings change
useEffect(() => {
fetchVoices();
}, [provider, voiceProvider, fetchVoices]);
/**
* Clears the voice cache
*/
const clearCache = useCallback(() => {
voiceCache = {};
fetchVoices();
}, [fetchVoices]);
/**
* Gets a voice by ID
*/
const getVoiceById = useCallback((voiceId: string) => {
return availableVoices.find((voice: Voice) => voice.id === voiceId);
}, [availableVoices]);
/**
* Gets voices for a specific category
*/
const getVoicesByCategory = useCallback((categoryId: string) => {
return availableVoices.filter((voice: Voice) => voice.categoryId === categoryId);
}, [availableVoices]);
/**
* Get a list of voice IDs only (for backward compatibility with existing code)
*/
const voiceIds = availableVoices.map((voice: Voice) => voice.id);
return {
availableVoices,
voiceCategories,
voiceIds,
isLoading,
error,
fetchVoices,
clearCache,
getVoiceById,
getVoicesByCategory
};
} }

View file

@ -0,0 +1,235 @@
/**
* ElevenLabs Provider Implementation
*
* This module provides premium voice synthesis functionality using ElevenLabs API
*/
import { TTSProvider, TTSRequest, Voice, VoiceCategory } from './types';
// Default ElevenLabs voices (will be overridden by API response)
const DEFAULT_VOICES: Voice[] = [
{ id: 'Bella', name: 'Bella', provider: 'elevenlabs', categoryId: 'elevenlabs-premium' },
{ id: 'Antoni', name: 'Antoni', provider: 'elevenlabs', categoryId: 'elevenlabs-premium' },
{ id: 'Rachel', name: 'Rachel', provider: 'elevenlabs', categoryId: 'elevenlabs-premium' },
{ id: 'Domi', name: 'Domi', provider: 'elevenlabs', categoryId: 'elevenlabs-premium' },
{ id: 'Charlie', name: 'Charlie', provider: 'elevenlabs', categoryId: 'elevenlabs-premium' },
];
// ElevenLabs voice categories
const DEFAULT_CATEGORIES: VoiceCategory[] = [
{
id: 'elevenlabs-premium',
name: 'Premium Voices',
provider: 'elevenlabs',
},
{
id: 'elevenlabs-custom',
name: 'Your Custom Voices',
provider: 'elevenlabs',
}
];
/**
* ElevenLabs provider configuration
*/
export interface ElevenLabsProviderConfig {
apiKey: string;
}
/**
* ElevenLabs voice interface from their API
*/
interface ElevenLabsVoice {
voice_id: string;
name: string;
category: string;
description?: string;
preview_url?: string;
labels?: Record<string, string>;
}
/**
* Implementation of the TTSProvider interface for ElevenLabs
*/
export class ElevenLabsProvider implements TTSProvider {
private config: ElevenLabsProviderConfig;
private baseUrl: string = 'https://api.elevenlabs.io/v1';
/**
* Create a new ElevenLabs provider instance
*
* @param config The ElevenLabs provider configuration
*/
constructor(config: ElevenLabsProviderConfig) {
this.config = config;
}
/**
* Generate speech using ElevenLabs API
*
* @param request TTS request parameters
* @param signal AbortSignal for cancelling the request
* @returns Promise resolving to an ArrayBuffer containing the audio data
*/
async generateSpeech(request: TTSRequest, signal?: AbortSignal): Promise<ArrayBuffer> {
if (!this.config.apiKey) {
throw new Error('Missing ElevenLabs API key');
}
try {
console.log('Generating ElevenLabs TTS with voice:', request.voice);
// ElevenLabs endpoint for text-to-speech
const endpoint = `${this.baseUrl}/text-to-speech/${request.voice}`;
// Calculate stability and similarity settings
// In ElevenLabs, these are between 0 and 1
const stability = 0.5; // Medium stability
const similarity = 0.75; // Higher similarity to the original voice
// Create request payload with the voice settings object
const payload = {
text: request.text,
model_id: "eleven_monolingual_v1",
voice_settings: {
stability,
similarity_boost: similarity,
// Convert speed from OpenAI range to ElevenLabs range
// OpenAI typically uses 0.25-4.0, ElevenLabs uses 0.5-2.0
// We'll map the ranges approximately
speaking_rate: Math.max(0.5, Math.min(2.0, request.speed)),
}
};
// Send request to ElevenLabs
const response = await fetch(endpoint, {
method: 'POST',
headers: {
'xi-api-key': this.config.apiKey,
'Content-Type': 'application/json',
'Accept': 'audio/mpeg',
},
body: JSON.stringify(payload),
signal,
});
if (!response.ok) {
const errorText = await response.text();
throw new Error(`ElevenLabs API error: ${errorText}`);
}
// Get the audio data as array buffer
return await response.arrayBuffer();
} catch (error) {
if (error instanceof Error && error.name === 'AbortError') {
console.log('TTS request aborted');
throw error;
}
console.error('Error generating TTS with ElevenLabs:', error);
throw new Error('Failed to generate audio with ElevenLabs');
}
}
/**
* Get available voices from ElevenLabs
* Fetches both premium and user's custom voices
*
* @returns Promise resolving to an array of available voices
*/
async getAvailableVoices(): Promise<Voice[]> {
try {
if (!this.config.apiKey) {
return DEFAULT_VOICES;
}
// ElevenLabs endpoint for voices
const response = await fetch(`${this.baseUrl}/voices`, {
headers: {
'xi-api-key': this.config.apiKey,
'Content-Type': 'application/json',
},
});
if (!response.ok) {
console.warn('Failed to fetch ElevenLabs voices, using defaults');
return DEFAULT_VOICES;
}
const data = await response.json();
// Transform the API response into the Voice interface format
if (data.voices && Array.isArray(data.voices)) {
return data.voices.map((voice: ElevenLabsVoice) => {
// Determine category based on voice properties
// Cloned voices are typically user's custom voices
const categoryId = voice.category === 'cloned'
? 'elevenlabs-custom'
: 'elevenlabs-premium';
return {
id: voice.voice_id,
name: voice.name,
provider: 'elevenlabs',
categoryId,
description: voice.description,
previewUrl: voice.preview_url,
};
});
}
return DEFAULT_VOICES;
} catch (error) {
console.error('Error fetching ElevenLabs voices:', error);
return DEFAULT_VOICES;
}
}
/**
* Get voice categories for ElevenLabs
*
* @returns Promise resolving to an array of voice categories
*/
async getVoiceCategories(): Promise<VoiceCategory[]> {
// For ElevenLabs, we typically have premium voices and user's custom voices
try {
if (!this.config.apiKey) {
return DEFAULT_CATEGORIES;
}
// Fetch user information to check subscription
const response = await fetch(`${this.baseUrl}/user`, {
headers: {
'xi-api-key': this.config.apiKey,
'Content-Type': 'application/json',
},
});
if (!response.ok) {
return DEFAULT_CATEGORIES;
}
const userData = await response.json();
// Customize categories based on subscription information if needed
const categories = [...DEFAULT_CATEGORIES];
// If subscription tier is available, you could add a tier-specific category
if (userData.subscription && userData.subscription.tier) {
const tier = userData.subscription.tier;
if (tier === 'creator') {
categories.push({
id: 'elevenlabs-creator',
name: 'Creator Voices',
provider: 'elevenlabs',
});
}
}
return categories;
} catch (error) {
console.error('Error fetching ElevenLabs categories:', error);
return DEFAULT_CATEGORIES;
}
}
}

152
src/providers/factory.ts Normal file
View file

@ -0,0 +1,152 @@
/**
* Provider Factory
*
* This module provides factory functions to create the appropriate TTS provider
* based on the current configuration and settings.
*/
import { ProviderType, TTSProvider, VoiceProviderType, ProviderOptions, VoiceCategory, Voice } from './types';
import { OpenAIProvider } from './openaiProvider';
import { OllamaProvider } from './ollamaProvider';
import { OpenRouterProvider } from './openrouterProvider';
import { ElevenLabsProvider } from './elevenlabsProvider';
/**
* Creates the appropriate AI provider based on the requested provider type
*
* @param provider The provider type to create
* @param options Provider configuration options
* @returns An instance of the requested provider
*/
export function createProvider(provider: ProviderType, options: ProviderOptions): TTSProvider {
const { providerSettings } = options;
switch (provider) {
case 'openai':
return new OpenAIProvider({
apiKey: providerSettings.openai.apiKey,
baseUrl: providerSettings.openai.baseUrl,
model: providerSettings.openai.model,
});
case 'ollama':
return new OllamaProvider({
baseUrl: providerSettings.ollama.baseUrl,
model: providerSettings.ollama.model,
});
case 'openrouter':
return new OpenRouterProvider({
apiKey: providerSettings.openrouter.apiKey,
model: providerSettings.openrouter.model,
});
default:
console.warn(`Unknown provider type: ${provider}, defaulting to OpenAI`);
return new OpenAIProvider({
apiKey: providerSettings.openai.apiKey,
baseUrl: providerSettings.openai.baseUrl,
model: providerSettings.openai.model,
});
}
}
/**
* Creates a voice provider for TTS based on the requested voice provider type
*
* @param voiceProvider The voice provider type to create
* @param options Provider configuration options
* @returns An instance of the requested voice provider
*/
export function createVoiceProvider(voiceProvider: VoiceProviderType, options: ProviderOptions): TTSProvider {
const { providerSettings } = options;
switch (voiceProvider) {
case 'openai':
return new OpenAIProvider({
apiKey: providerSettings.openai.apiKey,
baseUrl: providerSettings.openai.baseUrl,
model: providerSettings.openai.model,
});
case 'elevenlabs':
return new ElevenLabsProvider({
apiKey: providerSettings.elevenlabs.apiKey,
});
default:
console.warn(`Unknown voice provider type: ${voiceProvider}, defaulting to OpenAI`);
return new OpenAIProvider({
apiKey: providerSettings.openai.apiKey,
baseUrl: providerSettings.openai.baseUrl,
model: providerSettings.openai.model,
});
}
}
/**
* Creates a provider based on document override settings
*
* @param defaultProvider The default provider to use
* @param defaultVoiceProvider The default voice provider to use
* @param documentOverride Document-specific override settings
* @param options Provider configuration options
* @returns An instance of the appropriate provider
*/
export function createProviderWithOverrides(
defaultProvider: ProviderType,
defaultVoiceProvider: VoiceProviderType,
documentOverride: { provider?: ProviderType; voiceProvider?: VoiceProviderType; voice?: string } | undefined,
options: ProviderOptions
): { provider: TTSProvider; voiceProvider: TTSProvider; voice?: string } {
const provider = documentOverride?.provider || defaultProvider;
const voiceProvider = documentOverride?.voiceProvider || defaultVoiceProvider;
const voice = documentOverride?.voice;
return {
provider: createProvider(provider, options),
voiceProvider: createVoiceProvider(voiceProvider, options),
voice,
};
}
/**
* Combines voices from multiple providers into a single array
*
* @param voices Arrays of voices from different providers
* @returns Combined array of unique voices
*/
export function combineVoicesFromProviders(...voices: Voice[][]): Voice[] {
const voiceMap = new Map<string, Voice>();
// Add each voice to the map, using id+provider as the key to ensure uniqueness
for (const voiceArray of voices) {
for (const voice of voiceArray) {
const key = `${voice.provider}-${voice.id}`;
voiceMap.set(key, voice);
}
}
// Convert the map back to an array
return Array.from(voiceMap.values());
}
/**
* Combines voice categories from multiple providers into a single array
*
* @param categories Arrays of voice categories from different providers
* @returns Combined array of unique voice categories
*/
export function combineCategories(...categories: VoiceCategory[][]): VoiceCategory[] {
const categoryMap = new Map<string, VoiceCategory>();
// Add each category to the map, using id as the key to ensure uniqueness
for (const categoryArray of categories) {
for (const category of categoryArray) {
categoryMap.set(category.id, category);
}
}
// Convert the map back to an array
return Array.from(categoryMap.values());
}

View file

@ -0,0 +1,149 @@
/**
* Ollama Provider Implementation
*
* This module provides TTS functionality using Ollama's API for local AI models
*/
import { TTSProvider, TTSRequest, Voice, VoiceCategory } from './types';
// Default Ollama voices
const DEFAULT_VOICES: Voice[] = [
{ id: 'default', name: 'Default', provider: 'ollama', categoryId: 'ollama-default' },
];
// Ollama voice category
const VOICE_CATEGORY: VoiceCategory = {
id: 'ollama-default',
name: 'Ollama Voices',
provider: 'ollama',
};
/**
* Ollama provider configuration
*/
export interface OllamaProviderConfig {
baseUrl: string;
model: string;
}
/**
* Implementation of the TTSProvider interface for Ollama
*/
export class OllamaProvider implements TTSProvider {
private config: OllamaProviderConfig;
/**
* Create a new Ollama provider instance
*
* @param config The Ollama provider configuration
*/
constructor(config: OllamaProviderConfig) {
this.config = config;
}
/**
* Generate speech using Ollama's API
*
* @param request TTS request parameters
* @param signal AbortSignal for cancelling the request
* @returns Promise resolving to an ArrayBuffer containing the audio data
*/
async generateSpeech(request: TTSRequest, signal?: AbortSignal): Promise<ArrayBuffer> {
try {
console.log('Generating Ollama TTS with voice:', request.voice);
// Ollama endpoint for text-to-speech
const endpoint = `${this.config.baseUrl}/api/generate`;
// Create request payload
const payload = {
model: this.config.model,
prompt: request.text,
stream: false,
options: {
temperature: 0.7,
audio: true
}
};
// Send request to Ollama
const response = await fetch(endpoint, {
method: 'POST',
headers: {
'Content-Type': 'application/json',
},
body: JSON.stringify(payload),
signal,
});
if (!response.ok) {
const errorText = await response.text();
throw new Error(`Ollama API error: ${errorText}`);
}
// Get the response data
const data = await response.json();
// For Ollama, we might need to convert the response to audio
// This depends on how Ollama returns audio data
// For now, let's use a fallback approach where we send the generated text to the server
// to convert it to audio
// Send text to audio conversion API
const audioResponse = await fetch('/api/audio/convert', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
},
body: JSON.stringify({
text: data.response || request.text,
voice: request.voice || 'default',
speed: request.speed || 1.0,
}),
signal,
});
if (!audioResponse.ok) {
const errorText = await audioResponse.text();
throw new Error(`Audio conversion error: ${errorText}`);
}
// Return the converted audio
return await audioResponse.arrayBuffer();
} catch (error) {
if (error instanceof Error && error.name === 'AbortError') {
console.log('TTS request aborted');
throw error;
}
console.error('Error generating TTS with Ollama:', error);
throw new Error('Failed to generate audio with Ollama');
}
}
/**
* Get available voices from Ollama
*
* @returns Promise resolving to an array of available voices
*/
async getAvailableVoices(): Promise<Voice[]> {
try {
// For Ollama, we typically just have a default voice since it's using local models
return DEFAULT_VOICES;
} catch (error) {
console.error('Error fetching Ollama voices:', error);
return DEFAULT_VOICES;
}
}
/**
* Get voice categories for Ollama
*
* @returns Promise resolving to an array of voice categories
*/
async getVoiceCategories(): Promise<VoiceCategory[]> {
return [VOICE_CATEGORY];
}
}

View file

@ -0,0 +1,156 @@
/**
* OpenAI Provider Implementation
*
* This module provides TTS functionality using OpenAI's API
*/
import { TTSProvider, TTSRequest, Voice, VoiceCategory } from './types';
// Default OpenAI voices
const DEFAULT_VOICES: Voice[] = [
{ id: 'alloy', name: 'Alloy', provider: 'openai', categoryId: 'openai-default' },
{ id: 'echo', name: 'Echo', provider: 'openai', categoryId: 'openai-default' },
{ id: 'fable', name: 'Fable', provider: 'openai', categoryId: 'openai-default' },
{ id: 'onyx', name: 'Onyx', provider: 'openai', categoryId: 'openai-default' },
{ id: 'nova', name: 'Nova', provider: 'openai', categoryId: 'openai-default' },
{ id: 'shimmer', name: 'Shimmer', provider: 'openai', categoryId: 'openai-default' },
];
// OpenAI voice category
const VOICE_CATEGORY: VoiceCategory = {
id: 'openai-default',
name: 'OpenAI Voices',
provider: 'openai',
};
/**
* OpenAI provider configuration
*/
export interface OpenAIProviderConfig {
apiKey: string;
baseUrl: string;
model: string;
}
/**
* Implementation of the TTSProvider interface for OpenAI
*/
export class OpenAIProvider implements TTSProvider {
private config: OpenAIProviderConfig;
/**
* Create a new OpenAI provider instance
*
* @param config The OpenAI provider configuration
*/
constructor(config: OpenAIProviderConfig) {
this.config = config;
}
/**
* Generate speech using OpenAI's API
*
* @param request TTS request parameters
* @param signal AbortSignal for cancelling the request
* @returns Promise resolving to an ArrayBuffer containing the audio data
*/
async generateSpeech(request: TTSRequest, signal?: AbortSignal): Promise<ArrayBuffer> {
if (!this.config.apiKey) {
throw new Error('Missing OpenAI API key');
}
try {
console.log('Generating OpenAI TTS with voice:', request.voice);
// OpenAI endpoint for text-to-speech
const endpoint = `${this.config.baseUrl}/audio/speech`;
// Create request payload
const payload = {
model: this.config.model,
input: request.text,
voice: request.voice,
speed: request.speed,
response_format: request.format === 'aac' ? 'aac' : 'mp3',
};
// Send request to OpenAI
const response = await fetch(endpoint, {
method: 'POST',
headers: {
'Authorization': `Bearer ${this.config.apiKey}`,
'Content-Type': 'application/json',
},
body: JSON.stringify(payload),
signal,
});
if (!response.ok) {
const errorText = await response.text();
throw new Error(`OpenAI API error: ${errorText}`);
}
// Get the audio data as array buffer
return await response.arrayBuffer();
} catch (error) {
if (error instanceof Error && error.name === 'AbortError') {
console.log('TTS request aborted by client');
throw error;
}
console.error('Error generating TTS with OpenAI:', error);
throw new Error('Failed to generate audio with OpenAI');
}
}
/**
* Get available voices from OpenAI
*
* @returns Promise resolving to an array of available voices
*/
async getAvailableVoices(): Promise<Voice[]> {
try {
if (!this.config.apiKey) {
return DEFAULT_VOICES;
}
// OpenAI endpoint for voices
const response = await fetch(`${this.config.baseUrl}/audio/voices`, {
headers: {
'Authorization': `Bearer ${this.config.apiKey}`,
'Content-Type': 'application/json',
},
});
if (!response.ok) {
return DEFAULT_VOICES;
}
const data = await response.json();
// Transform the API response into the Voice interface format
if (data.voices && Array.isArray(data.voices)) {
return data.voices.map((voiceId: string) => ({
id: voiceId,
name: voiceId.charAt(0).toUpperCase() + voiceId.slice(1), // Capitalize first letter
provider: 'openai',
categoryId: 'openai-default',
}));
}
return DEFAULT_VOICES;
} catch (error) {
console.error('Error fetching OpenAI voices:', error);
return DEFAULT_VOICES;
}
}
/**
* Get voice categories for OpenAI
*
* @returns Promise resolving to an array of voice categories
*/
async getVoiceCategories(): Promise<VoiceCategory[]> {
return [VOICE_CATEGORY];
}
}

View file

@ -0,0 +1,201 @@
/**
* OpenRouter Provider Implementation
*
* This module provides TTS functionality using OpenRouter's API gateway to access various AI models
*/
import { TTSProvider, TTSRequest, Voice, VoiceCategory } from './types';
// Default OpenRouter voices (similar to OpenAI since it's often a gateway to OpenAI models)
const DEFAULT_VOICES: Voice[] = [
{ id: 'alloy', name: 'Alloy', provider: 'openrouter', categoryId: 'openrouter-default' },
{ id: 'echo', name: 'Echo', provider: 'openrouter', categoryId: 'openrouter-default' },
{ id: 'fable', name: 'Fable', provider: 'openrouter', categoryId: 'openrouter-default' },
{ id: 'onyx', name: 'Onyx', provider: 'openrouter', categoryId: 'openrouter-default' },
{ id: 'nova', name: 'Nova', provider: 'openrouter', categoryId: 'openrouter-default' },
{ id: 'shimmer', name: 'Shimmer', provider: 'openrouter', categoryId: 'openrouter-default' },
];
// OpenRouter voice category
const VOICE_CATEGORY: VoiceCategory = {
id: 'openrouter-default',
name: 'OpenRouter Voices',
provider: 'openrouter',
};
/**
* OpenRouter provider configuration
*/
export interface OpenRouterProviderConfig {
apiKey: string;
model: string;
}
/**
* Implementation of the TTSProvider interface for OpenRouter
*/
export class OpenRouterProvider implements TTSProvider {
private config: OpenRouterProviderConfig;
private baseUrl: string = 'https://openrouter.ai/api/v1';
/**
* Create a new OpenRouter provider instance
*
* @param config The OpenRouter provider configuration
*/
constructor(config: OpenRouterProviderConfig) {
this.config = config;
}
/**
* Generate speech using OpenRouter's API
*
* @param request TTS request parameters
* @param signal AbortSignal for cancelling the request
* @returns Promise resolving to an ArrayBuffer containing the audio data
*/
async generateSpeech(request: TTSRequest, signal?: AbortSignal): Promise<ArrayBuffer> {
if (!this.config.apiKey) {
throw new Error('Missing OpenRouter API key');
}
try {
console.log('Generating OpenRouter TTS with voice:', request.voice);
// OpenRouter endpoint for text-to-speech (uses OpenAI-compatible endpoints)
const endpoint = `${this.baseUrl}/audio/speech`;
// Create request payload
const payload = {
model: this.config.model || 'openai/whisper',
input: request.text,
voice: request.voice,
speed: request.speed,
response_format: request.format === 'aac' ? 'aac' : 'mp3',
};
// Send request to OpenRouter
const response = await fetch(endpoint, {
method: 'POST',
headers: {
'Authorization': `Bearer ${this.config.apiKey}`,
'HTTP-Referer': window.location.origin, // Required by OpenRouter
'X-Title': 'OpenReader', // Identify our app to OpenRouter
'Content-Type': 'application/json',
},
body: JSON.stringify(payload),
signal,
});
if (!response.ok) {
// If OpenRouter doesn't support speech directly, fallback to the API proxy
if (response.status === 404 || response.status === 501) {
console.log('OpenRouter TTS endpoint not available, using fallback...');
return this.fallbackToApiProxy(request, signal);
}
const errorText = await response.text();
throw new Error(`OpenRouter API error: ${errorText}`);
}
// Get the audio data as array buffer
return await response.arrayBuffer();
} catch (error) {
if (error instanceof Error && error.name === 'AbortError') {
console.log('TTS request aborted');
throw error;
}
console.error('Error generating TTS with OpenRouter:', error);
throw new Error('Failed to generate audio with OpenRouter');
}
}
/**
* Fallback to using our server as a proxy to generate audio
* This is used when OpenRouter doesn't support TTS directly
*
* @param request The TTS request
* @param signal AbortSignal for cancelling the request
* @returns Promise resolving to an ArrayBuffer containing the audio data
*/
private async fallbackToApiProxy(request: TTSRequest, signal?: AbortSignal): Promise<ArrayBuffer> {
// Send the request to our server-side proxy that can handle TTS generation
const response = await fetch('/api/audio/convert', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
'x-openrouter-key': this.config.apiKey,
},
body: JSON.stringify({
text: request.text,
voice: request.voice,
speed: request.speed,
provider: 'openrouter',
model: this.config.model,
}),
signal,
});
if (!response.ok) {
const errorText = await response.text();
throw new Error(`Fallback TTS error: ${errorText}`);
}
return await response.arrayBuffer();
}
/**
* Get available voices from OpenRouter
*
* @returns Promise resolving to an array of available voices
*/
async getAvailableVoices(): Promise<Voice[]> {
try {
if (!this.config.apiKey) {
return DEFAULT_VOICES;
}
// OpenRouter endpoint for voices (similar to OpenAI)
const response = await fetch(`${this.baseUrl}/audio/voices`, {
headers: {
'Authorization': `Bearer ${this.config.apiKey}`,
'HTTP-Referer': window.location.origin,
'X-Title': 'OpenReader',
'Content-Type': 'application/json',
},
});
// If OpenRouter doesn't support voice listing, return defaults
if (!response.ok) {
return DEFAULT_VOICES;
}
const data = await response.json();
// Transform the API response into the Voice interface format
if (data.voices && Array.isArray(data.voices)) {
return data.voices.map((voiceId: string) => ({
id: voiceId,
name: voiceId.charAt(0).toUpperCase() + voiceId.slice(1), // Capitalize first letter
provider: 'openrouter',
categoryId: 'openrouter-default',
}));
}
return DEFAULT_VOICES;
} catch (error) {
console.error('Error fetching OpenRouter voices:', error);
return DEFAULT_VOICES;
}
}
/**
* Get voice categories for OpenRouter
*
* @returns Promise resolving to an array of voice categories
*/
async getVoiceCategories(): Promise<VoiceCategory[]> {
return [VOICE_CATEGORY];
}
}

119
src/providers/types.ts Normal file
View file

@ -0,0 +1,119 @@
/**
* Provider types and interfaces for OpenReader text-to-speech functionality
*/
/**
* Available AI Providers
*
* - openai: OpenAI's text-to-speech
* - ollama: Local model via Ollama
* - openrouter: API gateway to various models
*/
export type ProviderType = 'openai' | 'ollama' | 'openrouter';
/**
* Available Voice Providers
*
* - openai: Default OpenAI voices
* - elevenlabs: Premium voice synthesis
*/
export type VoiceProviderType = 'openai' | 'elevenlabs';
/**
* Text-to-speech request parameters
*/
export interface TTSRequest {
text: string;
voice: string;
speed?: number;
format?: 'mp3' | 'aac';
}
/**
* Voice representation for the UI
*/
export interface Voice {
id: string;
name: string;
provider: string;
categoryId: string;
description?: string;
previewUrl?: string;
}
/**
* Voice category for grouping voices in the UI
*/
export interface VoiceCategory {
id: string;
name: string;
provider: string;
description?: string;
}
/**
* Provider interface for text-to-speech functionality
*/
export interface TTSProvider {
/**
* Generate speech from text
*
* @param request TTS parameters
* @param signal AbortSignal for cancelling the request
* @returns Promise with audio data as ArrayBuffer
*/
generateSpeech(request: TTSRequest, signal?: AbortSignal): Promise<ArrayBuffer>;
/**
* Get available voices from the provider
*
* @returns Promise with array of available voices
*/
getAvailableVoices(): Promise<Voice[]>;
/**
* Get voice categories from the provider
*
* @returns Promise with array of voice categories
*/
getVoiceCategories(): Promise<VoiceCategory[]>;
}
/**
* Provider settings for all available providers
*/
export interface ProviderSettings {
openai: {
apiKey: string;
baseUrl: string;
model: string;
};
ollama: {
baseUrl: string;
model: string;
};
openrouter: {
apiKey: string;
model: string;
};
elevenlabs: {
apiKey: string;
};
}
/**
* Provider options for creating a provider instance
*/
export interface ProviderOptions {
providerSettings: ProviderSettings;
provider?: ProviderType; // Added to support specifying provider for voice generation
}
/**
* Document-specific provider override settings
*/
export interface DocumentProviderOverride {
provider?: ProviderType;
voiceProvider?: VoiceProviderType;
voice?: string;
}

View file

@ -1,345 +1,427 @@
import { pdfjs } from 'react-pdf'; /**
import nlp from 'compromise'; * PDF utility functions for handling PDF documents in OpenReader
import stringSimilarity from 'string-similarity'; */
import type { TextItem } from 'pdfjs-dist/types/src/display/api';
import type { PDFDocumentProxy } from 'pdfjs-dist';
// Function to detect if we need to use legacy build import * as pdfjs from 'pdfjs-dist';
function shouldUseLegacyBuild() { import { TextContent, TextItem, PDFDocumentProxy, PDFPageProxy } from 'pdfjs-dist/types/src/display/api';
try { import { RefObject } from 'react';
if (typeof window === 'undefined') return false;
// For workers in Next.js we need to set the worker source path correctly
const ua = window.navigator.userAgent; if (typeof window !== 'undefined' && 'Worker' in window) {
const isSafari = /^((?!chrome|android).)*safari/i.test(ua); // Set worker source using CDN path based on pdfjs version
(pdfjs as any).GlobalWorkerOptions.workerSrc = `//cdn.jsdelivr.net/npm/pdfjs-dist@${pdfjs.version}/build/pdf.worker.min.js`;
console.log(isSafari ? 'Running on Safari' : 'Not running on Safari');
if (!isSafari) return false;
// Extract Safari version - matches "Version/18" format
const match = ua.match(/Version\/(\d+)/i);
console.log('Safari version:', match);
if (!match || !match[1]) return true; // If we can't determine version, use legacy to be safe
const version = parseInt(match[1]);
return version < 18; // Use legacy build for Safari versions equal or below 18
} catch (e) {
console.error('Error detecting Safari version:', e);
return false;
}
} }
// Function to initialize PDF worker /**
function initPDFWorker() { * Generic debounce function
try { *
if (typeof window !== 'undefined') { * @param fn Function to debounce
const useLegacy = shouldUseLegacyBuild(); * @param delay Delay in milliseconds
// Use local worker file instead of unpkg * @returns Debounced function
const workerSrc = useLegacy */
? new URL('pdfjs-dist/legacy/build/pdf.worker.min.mjs', import.meta.url).href export function debounce<T extends (...args: any[]) => any>(
: new URL('pdfjs-dist/build/pdf.worker.min.mjs', import.meta.url).href; fn: T,
console.log('Setting PDF worker to:', workerSrc); delay: number
pdfjs.GlobalWorkerOptions.workerSrc = workerSrc; ): (...args: Parameters<T>) => void {
pdfjs.GlobalWorkerOptions.workerPort = null; let timeout: ReturnType<typeof setTimeout> | null = null;
return function(this: any, ...args: Parameters<T>) {
if (timeout !== null) {
clearTimeout(timeout);
} }
} catch (e) {
console.error('Error setting PDF worker:', e); timeout = setTimeout(() => {
} fn.apply(this, args);
timeout = null;
}, delay);
};
} }
// Initialize the worker /**
initPDFWorker(); * Type guard to check if an item is a TextItem
*/
interface TextMatch { function isTextItem(item: any): item is TextItem {
elements: HTMLElement[]; return item && typeof item.str === 'string' && Array.isArray(item.transform);
rating: number;
text: string;
lengthDiff: number;
} }
// Text Processing functions /**
* Safely extracts text from a PDF document
*
* @param pdfDocument The PDF document proxy
* @param pageNum The page number to extract text from
* @param margins Configuration for text extraction margins
* @returns Promise resolving to the extracted text
*/
export async function extractTextFromPDF( export async function extractTextFromPDF(
pdf: PDFDocumentProxy, pdfDocument: PDFDocumentProxy | null,
pageNumber: number, pageNum: number,
margins = { header: 0.07, footer: 0.07, left: 0.07, right: 0.07 } margins = {
top: 0.07,
bottom: 0.07,
left: 0.07,
right: 0.07,
}
): Promise<string> { ): Promise<string> {
try { try {
// Log pdf worker version // Handle null document case
//console.log('PDF worker version:', pdfjs.GlobalWorkerOptions.workerSrc); if (!pdfDocument) {
console.warn('PDF document is null or undefined');
const page = await pdf.getPage(pageNumber); return '';
const textContent = await page.getTextContent();
const viewport = page.getViewport({ scale: 1.0 });
const pageHeight = viewport.height;
const pageWidth = viewport.width;
const textItems = textContent.items.filter((item): item is TextItem => {
if (!('str' in item && 'transform' in item)) return false;
const [scaleX, skewX, skewY, scaleY, x, y] = item.transform;
// Basic text filtering
if (Math.abs(scaleX) < 1 || Math.abs(scaleX) > 20) return false;
if (Math.abs(scaleY) < 1 || Math.abs(scaleY) > 20) return false;
if (Math.abs(skewX) > 0.5 || Math.abs(skewY) > 0.5) return false;
// Calculate margins in PDF coordinate space (y=0 is at bottom)
const headerY = pageHeight * (1 - margins.header); // Convert from top margin to bottom-based Y
const footerY = pageHeight * margins.footer; // Footer Y stays as is since it's already bottom-based
const leftX = pageWidth * margins.left;
const rightX = pageWidth * (1 - margins.right);
// Check margins - remember y=0 is at bottom of page in PDF coordinates
if (y > headerY || y < footerY) { // Y greater than headerY means it's in header area, less than footerY means footer area
return false;
}
// Check horizontal margins
if (x < leftX || x > rightX) {
return false;
}
// Sanity check for coordinates
if (x < 0 || x > pageWidth) return false;
return item.str.trim().length > 0;
});
//console.log('Filtered text items:', textItems);
const tolerance = 2;
const lines: TextItem[][] = [];
let currentLine: TextItem[] = [];
let currentY: number | null = null;
textItems.forEach((item) => {
const y = item.transform[5];
if (currentY === null) {
currentY = y;
currentLine.push(item);
} else if (Math.abs(y - currentY) < tolerance) {
currentLine.push(item);
} else {
lines.push(currentLine);
currentLine = [item];
currentY = y;
}
});
lines.push(currentLine);
let pageText = '';
for (const line of lines) {
line.sort((a, b) => a.transform[4] - b.transform[4]);
let lineText = '';
let prevItem: TextItem | null = null;
for (const item of line) {
if (!prevItem) {
lineText = item.str;
} else {
const prevEndX = prevItem.transform[4] + (prevItem.width ?? 0);
const currentStartX = item.transform[4];
const space = currentStartX - prevEndX;
if (space > ((item.width ?? 0) * 0.3)) {
lineText += ' ' + item.str;
} else {
lineText += item.str;
}
}
prevItem = item;
}
pageText += lineText + ' ';
} }
return pageText.replace(/\s+/g, ' ').trim(); // Safely get page with error handling
let page: PDFPageProxy;
try {
page = await pdfDocument.getPage(pageNum);
} catch (err) {
console.error('Error getting PDF page:', err);
return '';
}
// If page is null or undefined, return empty string
if (!page) {
console.warn('PDF page is null or undefined');
return '';
}
const viewport = page.getViewport({ scale: 1.0 });
const pageWidth = viewport.width;
const pageHeight = viewport.height;
// Define margins based on page dimensions
const topMargin = pageHeight * margins.top;
const bottomMargin = pageHeight * margins.bottom;
const leftMargin = pageWidth * margins.left;
const rightMargin = pageWidth * margins.right;
// Extract text content from the page with a safety timeout
const textContent = await Promise.race([
page.getTextContent(),
new Promise<TextContent>((_, reject) =>
setTimeout(() => reject(new Error('PDF text extraction timeout')), 10000)
)
]) as TextContent;
if (!textContent || !textContent.items) {
console.warn('No text content found in PDF page');
return '';
}
// Process text content with margin filtering
let lastY: number | null = null;
let text = '';
// Filter items based on margins
textContent.items.forEach((item) => {
// Skip if not a TextItem
if (!isTextItem(item)) return;
const x = item.transform[4];
const y = item.transform[5];
// Skip content in margins
if (
x < leftMargin ||
x > pageWidth - rightMargin ||
y < bottomMargin ||
y > pageHeight - topMargin
) {
return;
}
// Add newlines between different y-positions
if (lastY !== null && Math.abs(y - lastY) > 5) {
text += '\n';
}
lastY = y;
text += item.str + ' ';
});
// Clean and normalize the text
text = text.replace(/\s+/g, ' ').trim();
// Add paragraph breaks where appropriate
text = text.replace(/\.\s+/g, '.\n\n');
return text;
} catch (error) { } catch (error) {
console.error('Error extracting text from PDF:', error); console.error('Error extracting text from PDF:', error);
throw new Error('Failed to extract text from PDF'); throw new Error('Failed to extract text from PDF');
} }
} }
// Highlighting functions /**
export function clearHighlights() { * Safely loads a PDF document from a URL
const textNodes = document.querySelectorAll('.react-pdf__Page__textContent span'); *
textNodes.forEach((node) => { * @param url URL of the PDF to load
const element = node as HTMLElement; * @returns Promise resolving to the PDF document proxy
element.style.backgroundColor = ''; */
element.style.opacity = '1'; export async function loadPDFDocumentSafely(url: string): Promise<PDFDocumentProxy | null> {
}); try {
} // Use the appropriate options for worker URLs in both dev and prod
const loadingTask = pdfjs.getDocument({
export function findBestTextMatch( url,
elements: Array<{ element: HTMLElement; text: string }>, cMapUrl: 'https://cdn.jsdelivr.net/npm/pdfjs-dist@3.11.174/cmaps/',
targetText: string, cMapPacked: true,
maxCombinedLength: number });
): TextMatch {
let bestMatch = { // Add proper timeout and error handling
elements: [] as HTMLElement[], const timeoutPromise = new Promise<null>((resolve) => {
rating: 0, setTimeout(() => {
text: '', loadingTask.destroy().catch(e => console.error('Error destroying PDF loading task:', e));
lengthDiff: Infinity, resolve(null);
}; }, 30000); // 30 second timeout
const SPAN_SEARCH_LIMIT = 10;
for (let i = 0; i < elements.length; i++) {
let combinedText = '';
const currentElements = [];
for (let j = i; j < Math.min(i + SPAN_SEARCH_LIMIT, elements.length); j++) {
const node = elements[j];
const newText = combinedText ? `${combinedText} ${node.text}` : node.text;
if (newText.length > maxCombinedLength) break;
combinedText = newText;
currentElements.push(node.element);
const similarity = stringSimilarity.compareTwoStrings(targetText, combinedText);
const lengthDiff = Math.abs(combinedText.length - targetText.length);
const lengthPenalty = lengthDiff / targetText.length;
const adjustedRating = similarity * (1 - lengthPenalty * 0.5);
if (adjustedRating > bestMatch.rating) {
bestMatch = {
elements: [...currentElements],
rating: adjustedRating,
text: combinedText,
lengthDiff,
};
}
}
}
return bestMatch;
}
export function highlightPattern(
text: string,
pattern: string,
containerRef: React.RefObject<HTMLDivElement>
) {
clearHighlights();
if (!pattern?.trim()) return;
const cleanPattern = pattern.trim().replace(/\s+/g, ' ');
const container = containerRef.current;
if (!container) return;
const textNodes = container.querySelectorAll('.react-pdf__Page__textContent span');
const allText = Array.from(textNodes).map((node) => ({
element: node as HTMLElement,
text: (node.textContent || '').trim(),
})).filter((node) => node.text.length > 0);
const containerRect = container.getBoundingClientRect();
const visibleTop = container.scrollTop;
const visibleBottom = visibleTop + containerRect.height;
const bufferSize = containerRect.height;
const visibleNodes = allText.filter(({ element }) => {
const rect = element.getBoundingClientRect();
const elementTop = rect.top - containerRect.top + container.scrollTop;
return elementTop >= (visibleTop - bufferSize) && elementTop <= (visibleBottom + bufferSize);
});
let bestMatch = findBestTextMatch(visibleNodes, cleanPattern, cleanPattern.length * 2);
if (bestMatch.rating < 0.3) {
bestMatch = findBestTextMatch(allText, cleanPattern, cleanPattern.length * 2);
}
const similarityThreshold = bestMatch.lengthDiff < cleanPattern.length * 0.3 ? 0.3 : 0.5;
if (bestMatch.rating >= similarityThreshold) {
bestMatch.elements.forEach((element) => {
element.style.backgroundColor = 'grey';
element.style.opacity = '0.4';
}); });
if (bestMatch.elements.length > 0) { // Race between loading and timeout
const element = bestMatch.elements[0]; const result = await Promise.race([loadingTask.promise, timeoutPromise]);
const elementRect = element.getBoundingClientRect(); return result;
const elementTop = elementRect.top - containerRect.top + container.scrollTop; } catch (error) {
console.error('Error loading PDF document:', error);
if (elementTop < visibleTop || elementTop > visibleBottom) { return null;
container.scrollTo({
top: elementTop - containerRect.height / 3,
behavior: 'smooth',
});
}
}
} }
} }
// Text Click Handler /**
* CSS class used for text highlighting
*/
const HIGHLIGHT_CLASS = 'pdf-text-highlight';
/**
* Highlights text in a PDF container based on a pattern
*
* @param text The text content to search through
* @param pattern The pattern to highlight
* @param containerRef Reference to the container element
*/
export function highlightPattern(text: string, pattern: string, containerRef: RefObject<HTMLDivElement>): void {
if (!containerRef.current || !pattern || !text) return;
clearHighlights();
// Create a RegExp for the pattern with case insensitivity
try {
const regex = new RegExp(pattern, 'gi');
// Get all text nodes in the container
const textNodes = getAllTextNodes(containerRef.current);
// Highlight matching text in each text node
textNodes.forEach(node => {
const nodeText = node.textContent || '';
const matches = [...nodeText.matchAll(regex)];
if (matches.length > 0) {
highlightTextNode(node, matches);
}
});
} catch (error) {
console.error('Error highlighting pattern:', error);
}
}
/**
* Clears all text highlights
*/
export function clearHighlights(): void {
if (typeof window === 'undefined') return;
// Remove all highlight spans
const highlights = document.querySelectorAll(`.${HIGHLIGHT_CLASS}`);
highlights.forEach(highlight => {
const parent = highlight.parentNode;
if (parent) {
// Replace highlight with its text content
parent.replaceChild(
document.createTextNode(highlight.textContent || ''),
highlight
);
// Normalize the parent to merge adjacent text nodes
parent.normalize();
}
});
}
/**
* Handles text click events to find and play from the clicked position
*
* @param event The mouse click event
* @param pdfText The full text content of the PDF
* @param containerRef Reference to the container element
* @param stopAndPlayFromIndex Callback to start playback from an index
* @param isProcessing Whether the system is currently processing a request
*/
export function handleTextClick( export function handleTextClick(
event: MouseEvent, event: MouseEvent,
pdfText: string, pdfText: string,
containerRef: React.RefObject<HTMLDivElement>, containerRef: RefObject<HTMLDivElement>,
stopAndPlayFromIndex: (index: number) => void, stopAndPlayFromIndex: (index: number) => void,
isProcessing: boolean isProcessing: boolean
) { ): void {
if (isProcessing) return; if (isProcessing || !containerRef.current || !pdfText) return;
const target = event.target as HTMLElement; // Get the click target and verify it's a text node or element
if (!target.matches('.react-pdf__Page__textContent span')) return; const target = event.target as Node;
if (!target) return;
const parentElement = target.closest('.react-pdf__Page__textContent');
if (!parentElement) return; // Find the text node that was clicked
const textNode = findTextNodeFromClick(target);
const spans = Array.from(parentElement.querySelectorAll('span')); if (!textNode || !textNode.textContent) return;
const clickedIndex = spans.indexOf(target);
const contextWindow = 3; // Get the clicked position within the text node
const startIndex = Math.max(0, clickedIndex - contextWindow); const clickedText = textNode.textContent;
const endIndex = Math.min(spans.length - 1, clickedIndex + contextWindow); const clickedTextStart = findTextNodePosition(textNode, containerRef.current);
const contextText = spans
.slice(startIndex, endIndex + 1) if (clickedTextStart !== -1) {
.map((span) => span.textContent) // Calculate the relative click position and find it in the full text
.join(' ') const clickPosition = clickedTextStart + Math.floor(clickedText.length / 2);
.trim(); const textBeforeClick = pdfText.substring(0, clickPosition);
if (!contextText?.trim()) return; // Find the sentence start position
let sentenceStart = textBeforeClick.lastIndexOf('. ');
const cleanContext = contextText.trim().replace(/\s+/g, ' '); if (sentenceStart === -1) sentenceStart = 0;
const allText = Array.from(parentElement.querySelectorAll('span')).map((node) => ({ else sentenceStart += 2; // Move past the period and space
element: node as HTMLElement,
text: (node.textContent || '').trim(), // Use the callback to start playback from this position
})).filter((node) => node.text.length > 0); stopAndPlayFromIndex(sentenceStart);
const bestMatch = findBestTextMatch(allText, cleanContext, cleanContext.length * 2);
const similarityThreshold = bestMatch.lengthDiff < cleanContext.length * 0.3 ? 0.3 : 0.5;
if (bestMatch.rating >= similarityThreshold) {
const matchText = bestMatch.text;
const sentences = nlp(pdfText).sentences().out('array') as string[];
let bestSentenceMatch = { sentence: '', rating: 0 };
for (const sentence of sentences) {
const rating = stringSimilarity.compareTwoStrings(matchText, sentence);
if (rating > bestSentenceMatch.rating) {
bestSentenceMatch = { sentence, rating };
}
}
if (bestSentenceMatch.rating >= 0.5) {
const sentenceIndex = sentences.findIndex((sentence) => sentence === bestSentenceMatch.sentence);
if (sentenceIndex !== -1) {
stopAndPlayFromIndex(sentenceIndex);
highlightPattern(pdfText, bestSentenceMatch.sentence, containerRef);
}
}
} }
} }
// Debounce for PDF viewer /**
export function debounce<T extends (...args: unknown[]) => unknown>( * Helper function to get all text nodes in an element
func: T, *
wait: number * @param element The container element
): (...args: Parameters<T>) => void { * @returns Array of text nodes
let timeout: NodeJS.Timeout; */
return (...args: Parameters<T>) => { function getAllTextNodes(element: HTMLElement): Node[] {
clearTimeout(timeout); const textNodes: Node[] = [];
timeout = setTimeout(() => func(...args), wait);
}; // Skip invisible elements
if (element.style.display === 'none' || element.style.visibility === 'hidden') {
return textNodes;
}
// Collect all text nodes using TreeWalker
const treeWalker = document.createTreeWalker(
element,
NodeFilter.SHOW_TEXT,
{
acceptNode: (node) => {
// Skip empty text nodes
return node.textContent?.trim()
? NodeFilter.FILTER_ACCEPT
: NodeFilter.FILTER_REJECT;
}
}
);
let currentNode = treeWalker.currentNode;
while (currentNode) {
if (currentNode.nodeType === Node.TEXT_NODE && currentNode.textContent?.trim()) {
textNodes.push(currentNode);
}
currentNode = treeWalker.nextNode() as Node;
}
return textNodes;
}
/**
* Highlights matches in a text node
*
* @param textNode The text node to highlight
* @param matches Array of RegExpMatchArray matches
*/
function highlightTextNode(textNode: Node, matches: RegExpMatchArray[]): void {
const parent = textNode.parentNode;
if (!parent) return;
let nodeText = textNode.textContent || '';
let lastIndex = 0;
// Create a document fragment to hold the new nodes
const fragment = document.createDocumentFragment();
// Process each match
matches.forEach(match => {
if (match.index === undefined || !match[0]) return;
// Add text before the match
if (match.index > lastIndex) {
fragment.appendChild(
document.createTextNode(nodeText.substring(lastIndex, match.index))
);
}
// Create a highlighted span for the match
const highlightSpan = document.createElement('span');
highlightSpan.textContent = match[0];
highlightSpan.className = HIGHLIGHT_CLASS;
fragment.appendChild(highlightSpan);
lastIndex = match.index + match[0].length;
});
// Add any remaining text after the last match
if (lastIndex < nodeText.length) {
fragment.appendChild(
document.createTextNode(nodeText.substring(lastIndex))
);
}
// Replace the original text node with the new fragment
parent.replaceChild(fragment, textNode);
}
/**
* Finds the clicked text node
*
* @param target The clicked element or node
* @returns The text node that was clicked, or null
*/
function findTextNodeFromClick(target: Node): Node | null {
// If target is already a text node, return it
if (target.nodeType === Node.TEXT_NODE) {
return target;
}
// If target is an element with a highlight class, get its text content
if (target.nodeType === Node.ELEMENT_NODE &&
(target as Element).classList?.contains(HIGHLIGHT_CLASS)) {
return target.firstChild;
}
// If target has child nodes, return the first text node
if (target.hasChildNodes()) {
for (let i = 0; i < target.childNodes.length; i++) {
const child = target.childNodes[i];
if (child.nodeType === Node.TEXT_NODE && child.textContent?.trim()) {
return child;
}
}
}
return null;
}
/**
* Finds the position of a text node within the full text
*
* @param textNode The text node to find
* @param container The container element
* @returns The position of the text node in the full text, or -1 if not found
*/
function findTextNodePosition(textNode: Node, container: HTMLElement): number {
const allTextNodes = getAllTextNodes(container);
let position = 0;
for (const node of allTextNodes) {
if (node === textNode) {
return position;
}
position += (node.textContent || '').length;
}
return -1;
} }

View file

@ -1,8 +1,36 @@
NEXT_PUBLIC_NODE_ENV=development # OpenReader WebUI Environment Configuration
# OpenAI API Key for Text-to-Speech functionality # Application settings
API_KEY=api_key_here_if_needed NODE_ENV=development
PORT=3003
# OpenAI API Base URL (default) # Document storage base path (optional, defaults to app directory)
# To use a local TTS model server, I suggest using https://github.com/remsky/Kokoro-FastAPI DOCUMENTS_PATH=./documents
API_BASE=https://api.openai.com/v1
# OpenAI Provider Configuration (default provider)
OPENAI_API_KEY=
OPENAI_BASE_URL=https://api.openai.com/v1
OPENAI_MODEL=tts-1
# OpenRouter Provider Configuration
OPENROUTER_API_KEY=
OPENROUTER_MODEL=openai/whisper
# Ollama Provider Configuration
OLLAMA_BASE_URL=http://localhost:11434
OLLAMA_MODEL=llama3:8b
# ElevenLabs Provider Configuration
ELEVENLABS_API_KEY=
# Optional: Default provider settings
# Valid options: openai, openrouter, ollama
DEFAULT_PROVIDER=openai
# Optional: Default voice provider settings
# Valid options: openai, elevenlabs
DEFAULT_VOICE_PROVIDER=openai
# Optional: Default voice settings
DEFAULT_VOICE=alloy
DEFAULT_VOICE_SPEED=1.0