Faerro KB
A self-hosted voice-note inbox for an iPhone-installed web app. Recordings are saved in the browser first, uploaded when connectivity is available, and stored with transcripts on your own server.
Run locally
docker compose up --build
Releases
Git tags are the source of truth for release versions. From the repository root, run the publish script to create the next patch tag and publish it:
cp .env.example .env
# Set GITEA_USERNAME and GITEA_TOKEN in .env, using a Gitea package-write token.
./scripts/publish-image.sh
Pass --minor to increment the minor version and reset patch to zero, or --major to increment the major version and reset minor and patch to zero. The options cannot be combined. The script requires at least one existing MAJOR.MINOR.PATCH Git tag, creates the next local tag on the checked-out commit, and publishes both the versioned image and latest. .env is excluded from Git; the script reads the Gitea credentials from it and passes the token to Docker through standard input. If those credentials are unset, the script uses Docker's existing login credentials.
Open http://localhost for a desktop smoke test. Set HTTP_PORT to publish the app on a different host port (for example, HTTP_PORT=8000 docker compose up). For iPhone use, serve the app over HTTPS through your VPN/reverse proxy. iOS only grants microphone access in a secure context. Keep the service reachable only through your VPN and firewall; this first version has no login layer.
In Safari on iPhone, open the HTTPS address and use Share > Add to Home Screen. Sign in to the VPN before opening the app if the server is only reachable there.
Capture and offline behavior
- Recording data and its timestamp are committed to IndexedDB before upload is attempted.
- Uploads retry when the app opens, returns to the foreground, or the browser reports a connection. A recording is removed from the phone's queue only after the server confirms it was stored.
- Service worker caching keeps the app shell available offline. Searching the server-side archive requires a connection.
- iOS may suspend a PWA and can evict website data under storage pressure. Background uploads are not guaranteed; reopen the PWA while online to resume. Keep the phone powered and avoid clearing Safari website data for stronger practical retention.
Local speech-to-text
The API can use a local whisper.cpp server when WHISPER_SERVER_URL is configured. Audio is sent only to that local endpoint; with the variable unset, uploads remain stored as awaiting_transcription and can still be searched after a transcript is edited or populated.
For NVIDIA GPU transcription, run a CUDA-enabled whisper.cpp server with a model on the same machine, then set WHISPER_SERVER_URL to its local inference endpoint, for example http://whisper:8080/inference. Use a model that fits the RTX 3070's VRAM; start with small or a quantized medium model and measure with your audio/language. Do not expose the transcription server outside the private Docker network. See the whisper.cpp NVIDIA instructions for CUDA builds.
This version keeps raw audio and transcript separate. A future local correction model can propose punctuation and name fixes without replacing the original transcription. Proper-name matching and offline geolocation data are not enabled yet.
Data and backups
Docker persists SQLite metadata in ./data/voice-kb.sqlite3 and original audio in ./data/audio/. Back up the whole ./data directory while the service is stopped or use a SQLite-aware backup procedure. Model files should also be stored locally and backed up separately if needed.
Development
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
uvicorn app.main:app --reload
The web app is served by FastAPI at the root URL. SQLite FTS5 indexes transcript text for plain full-text search.