Issue: transformers.js is an ES module package and cannot be loaded with importScripts() in classic workers. Solution: - Changed to module worker (type: 'module') - Import transformers.js from CDN as ES module - Models (42MB) still served from local server at /models/ Trade-off: - Library (900KB): Loads from cdn.jsdelivr.net once, cached - Models (42MB): Self-hosted, served from /models/ (no CDN) This is necessary because: 1. @xenova/transformers is ES module-only (package.json: "type": "module") 2. ES modules cannot use importScripts() 3. Module workers require HTTPS for imports 4. CDN is HTTPS and cacheable If CDN is blocked: - Use Web Speech API (with privacy warnings) - OR use Server Transcription (Vertex AI/AWS) Models remain self-hosted as they're 40MB+ and contain the AI.
35 lines
1.6 KiB
Docker
35 lines
1.6 KiB
Docker
FROM node:20-alpine
|
|
|
|
WORKDIR /app
|
|
|
|
# ffmpeg: audio conversion for AWS Transcribe (WebM → PCM)
|
|
# curl: download Whisper models for browser-based transcription
|
|
RUN apk add --no-cache ffmpeg curl
|
|
|
|
COPY package.json ./
|
|
RUN npm install --omit=dev
|
|
|
|
COPY . .
|
|
|
|
RUN mkdir -p /app/data/logs
|
|
|
|
# Download Browser Whisper models (self-hosted)
|
|
# Library (900KB) loads from CDN, models (40MB+) served from our server
|
|
RUN mkdir -p /app/public/models/Xenova/whisper-tiny.en/onnx && \
|
|
cd /app/public/models/Xenova/whisper-tiny.en && \
|
|
echo "Downloading Whisper model files..." && \
|
|
curl -sL -o config.json https://huggingface.co/Xenova/whisper-tiny.en/resolve/main/config.json && \
|
|
curl -sL -o tokenizer.json https://huggingface.co/Xenova/whisper-tiny.en/resolve/main/tokenizer.json && \
|
|
curl -sL -o preprocessor_config.json https://huggingface.co/Xenova/whisper-tiny.en/resolve/main/preprocessor_config.json && \
|
|
curl -sL -o generation_config.json https://huggingface.co/Xenova/whisper-tiny.en/resolve/main/generation_config.json && \
|
|
curl -sL -o onnx/encoder_model_quantized.onnx https://huggingface.co/Xenova/whisper-tiny.en/resolve/main/onnx/encoder_model_quantized.onnx && \
|
|
curl -sL -o onnx/decoder_model_merged_quantized.onnx https://huggingface.co/Xenova/whisper-tiny.en/resolve/main/onnx/decoder_model_merged_quantized.onnx && \
|
|
echo "Whisper models downloaded successfully (models: 42MB)"
|
|
|
|
EXPOSE 3000
|
|
|
|
HEALTHCHECK --interval=30s --timeout=5s --start-period=20s \
|
|
CMD wget --no-verbose --tries=1 --spider http://localhost:3000/api/health || exit 1
|
|
|
|
CMD ["node", "server.js"]
|
|
|