3.6 KiB
Voice KB
A self-hosted voice-note inbox for an iPhone-installed web app. Recordings are saved in the browser first, uploaded when connectivity is available, and stored with transcripts on your own server.
Run locally
docker compose up --build
Releases
Git tags are the source of truth for release versions. From the repository root, create a semantic version tag on the commit to release, then run the publish script:
git tag 0.0.2
docker login gitea.faerro.it
./scripts/publish-image.sh
The script requires Docker and Git, and publishes both the versioned image (for example, gitea.faerro.it/andrea/faerro-kb:0.0.2) and latest. Tags may be plain versions or prefixed with v. Authenticate with a Gitea account or token that can publish packages before running the script.
Open http://localhost for a desktop smoke test. Set HTTP_PORT to publish the app on a different host port (for example, HTTP_PORT=8000 docker compose up). For iPhone use, serve the app over HTTPS through your VPN/reverse proxy. iOS only grants microphone access in a secure context. Keep the service reachable only through your VPN and firewall; this first version has no login layer.
In Safari on iPhone, open the HTTPS address and use Share > Add to Home Screen. Sign in to the VPN before opening the app if the server is only reachable there.
Capture and offline behavior
- Recording data and its timestamp are committed to IndexedDB before upload is attempted.
- Uploads retry when the app opens, returns to the foreground, or the browser reports a connection. A recording is removed from the phone's queue only after the server confirms it was stored.
- Service worker caching keeps the app shell available offline. Searching the server-side archive requires a connection.
- iOS may suspend a PWA and can evict website data under storage pressure. Background uploads are not guaranteed; reopen the PWA while online to resume. Keep the phone powered and avoid clearing Safari website data for stronger practical retention.
Local speech-to-text
The API can use a local whisper.cpp server when WHISPER_SERVER_URL is configured. Audio is sent only to that local endpoint; with the variable unset, uploads remain stored as awaiting_transcription and can still be searched after a transcript is edited or populated.
For NVIDIA GPU transcription, run a CUDA-enabled whisper.cpp server with a model on the same machine, then set WHISPER_SERVER_URL to its local inference endpoint, for example http://whisper:8080/inference. Use a model that fits the RTX 3070's VRAM; start with small or a quantized medium model and measure with your audio/language. Do not expose the transcription server outside the private Docker network. See the whisper.cpp NVIDIA instructions for CUDA builds.
This version keeps raw audio and transcript separate. A future local correction model can propose punctuation and name fixes without replacing the original transcription. Proper-name matching and offline geolocation data are not enabled yet.
Data and backups
Docker persists SQLite metadata in ./data/voice-kb.sqlite3 and original audio in ./data/audio/. Back up the whole ./data directory while the service is stopped or use a SQLite-aware backup procedure. Model files should also be stored locally and backed up separately if needed.
Development
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
uvicorn app.main:app --reload
The web app is served by FastAPI at the root URL. SQLite FTS5 indexes transcript text for plain full-text search.