2026-09-26 11:34:15 +02:00
2026-09-26 11:34:15 +02:00
2026-09-26 11:34:15 +02:00
2026-09-26 11:34:15 +02:00
2026-09-26 11:34:15 +02:00
2026-09-26 11:34:15 +02:00
2026-09-26 11:34:15 +02:00
2026-09-26 11:34:15 +02:00
2026-09-26 11:34:15 +02:00
2026-09-26 11:34:15 +02:00

Voice KB

A self-hosted voice-note inbox for an iPhone-installed web app. Recordings are saved in the browser first, uploaded when connectivity is available, and stored with transcripts on your own server.

Run locally

docker compose up --build

Open http://localhost:8000 for a desktop smoke test. For iPhone use, serve the app over HTTPS through your VPN/reverse proxy. iOS only grants microphone access in a secure context. Keep the service reachable only through your VPN and firewall; this first version has no login layer.

In Safari on iPhone, open the HTTPS address and use Share > Add to Home Screen. Sign in to the VPN before opening the app if the server is only reachable there.

Capture and offline behavior

  • Recording data and its timestamp are committed to IndexedDB before upload is attempted.
  • Uploads retry when the app opens, returns to the foreground, or the browser reports a connection. A recording is removed from the phone's queue only after the server confirms it was stored.
  • Service worker caching keeps the app shell available offline. Searching the server-side archive requires a connection.
  • iOS may suspend a PWA and can evict website data under storage pressure. Background uploads are not guaranteed; reopen the PWA while online to resume. Keep the phone powered and avoid clearing Safari website data for stronger practical retention.

Local speech-to-text

The API can use a local whisper.cpp server when WHISPER_SERVER_URL is configured. Audio is sent only to that local endpoint; with the variable unset, uploads remain stored as awaiting_transcription and can still be searched after a transcript is edited or populated.

For NVIDIA GPU transcription, run a CUDA-enabled whisper.cpp server with a model on the same machine, then set WHISPER_SERVER_URL to its local inference endpoint, for example http://whisper:8080/inference. Use a model that fits the RTX 3070's VRAM; start with small or a quantized medium model and measure with your audio/language. Do not expose the transcription server outside the private Docker network. See the whisper.cpp NVIDIA instructions for CUDA builds.

This version keeps raw audio and transcript separate. A future local correction model can propose punctuation and name fixes without replacing the original transcription. Proper-name matching and offline geolocation data are not enabled yet.

Data and backups

Docker persists SQLite metadata in ./data/voice-kb.sqlite3 and original audio in ./data/audio/. Back up the whole ./data directory while the service is stopped or use a SQLite-aware backup procedure. Model files should also be stored locally and backed up separately if needed.

Development

python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
uvicorn app.main:app --reload

The web app is served by FastAPI at the root URL. SQLite FTS5 indexes transcript text for plain full-text search.

S
Description
No description provided
Readme
138 KiB
Languages
JavaScript 43.8%
Python 23.3%
CSS 17.8%
HTML 8%
Shell 6.4%
Other 0.7%