Enhance audio processing capabilities: add FFmpeg for audio transcoding, update database schema, and improve UI for audio file uploads

This commit is contained in:
2026-09-26 21:51:50 +02:00
parent eb83b16b54
commit 0f7d5e950c
8 changed files with 135 additions and 20 deletions
+5 -3
View File
@@ -27,21 +27,23 @@ In Safari on iPhone, open the HTTPS address and use **Share > Add to Home Screen
## Capture and offline behavior
- Recording data and its timestamp are committed to IndexedDB before upload is attempted.
- Microphone recordings use the audio format supported by the browser (commonly M4A/AAC on iPhone Safari or Opus/WebM elsewhere). You can also add an audio file from the device; both sources remain queued locally until the server confirms storage.
- The server keeps each uploaded original and creates a mono PCM WAV copy at 16-bit / 22.05 kHz for transcription. FFmpeg performs this conversion locally in the container. After a non-empty transcript is stored, both audio files are deleted; failed or empty transcriptions keep the audio for recovery.
- Uploads retry when the app opens, returns to the foreground, or the browser reports a connection. A recording is removed from the phone's queue only after the server confirms it was stored.
- Service worker caching keeps the app shell available offline. Searching the server-side archive requires a connection.
- iOS may suspend a PWA and can evict website data under storage pressure. Background uploads are not guaranteed; reopen the PWA while online to resume. Keep the phone powered and avoid clearing Safari website data for stronger practical retention.
## Local speech-to-text
The API can use a local `whisper.cpp` server when `WHISPER_SERVER_URL` is configured. Audio is sent only to that local endpoint; with the variable unset, uploads remain stored as `awaiting_transcription` and can still be searched after a transcript is edited or populated.
The API can use a local `whisper.cpp` server when `WHISPER_SERVER_URL` is configured. Audio is sent only to that local endpoint; with the variable unset, uploads remain stored as `awaiting_transcription` until a non-empty transcript is entered and saved. Saving that transcript deletes both server-side audio files.
For NVIDIA GPU transcription, run a CUDA-enabled `whisper.cpp` server with a model on the same machine, then set `WHISPER_SERVER_URL` to its local inference endpoint, for example `http://whisper:8080/inference`. Use a model that fits the RTX 3070's VRAM; start with `small` or a quantized `medium` model and measure with your audio/language. Do not expose the transcription server outside the private Docker network. See the [whisper.cpp NVIDIA instructions](https://github.com/ggml-org/whisper.cpp#nvidia-gpu-support) for CUDA builds.
This version keeps raw audio and transcript separate. A future local correction model can propose punctuation and name fixes without replacing the original transcription. Proper-name matching and offline geolocation data are not enabled yet.
The transcript remains in SQLite after the audio files are deleted. A future local correction model can propose punctuation and name fixes without replacing the stored transcription. Proper-name matching and offline geolocation data are not enabled yet.
## Data and backups
Docker persists SQLite metadata in `./data/voice-kb.sqlite3` and original audio in `./data/audio/`. Back up the whole `./data` directory while the service is stopped or use a SQLite-aware backup procedure. Model files should also be stored locally and backed up separately if needed.
Docker persists SQLite metadata in `./data/voice-kb.sqlite3` and audio awaiting transcription in `./data/audio/`. After a non-empty transcript is saved, its audio files are removed. Back up the whole `./data` directory while the service is stopped or use a SQLite-aware backup procedure. Model files should also be stored locally and backed up separately if needed.
## Development