Files
faerro-kb/README.md
T

3.8 KiB

Faerro KB

A self-hosted voice-note inbox for an iPhone-installed web app. Recordings are saved in the browser first, uploaded when connectivity is available, and stored with transcripts on your own server.

Run locally

docker compose up --build

Releases

Git tags are the source of truth for release versions. From the repository root, create a semantic version tag on the commit to release, then run the publish script:

cp .env.example .env
# Set GITEA_USERNAME and GITEA_TOKEN in .env, using a Gitea package-write token.
git tag 0.0.2
./scripts/publish-image.sh

The script requires Docker and Git, and publishes both the versioned image (for example, gitea.faerro.it/andrea/faerro-kb:0.0.2) and latest. Tags may be plain versions or prefixed with v. .env is excluded from Git; the script reads the Gitea credentials from it and passes the token to Docker through standard input. If those credentials are unset, the script uses Docker's existing login credentials.

Open http://localhost for a desktop smoke test. Set HTTP_PORT to publish the app on a different host port (for example, HTTP_PORT=8000 docker compose up). For iPhone use, serve the app over HTTPS through your VPN/reverse proxy. iOS only grants microphone access in a secure context. Keep the service reachable only through your VPN and firewall; this first version has no login layer.

In Safari on iPhone, open the HTTPS address and use Share > Add to Home Screen. Sign in to the VPN before opening the app if the server is only reachable there.

Capture and offline behavior

  • Recording data and its timestamp are committed to IndexedDB before upload is attempted.
  • Uploads retry when the app opens, returns to the foreground, or the browser reports a connection. A recording is removed from the phone's queue only after the server confirms it was stored.
  • Service worker caching keeps the app shell available offline. Searching the server-side archive requires a connection.
  • iOS may suspend a PWA and can evict website data under storage pressure. Background uploads are not guaranteed; reopen the PWA while online to resume. Keep the phone powered and avoid clearing Safari website data for stronger practical retention.

Local speech-to-text

The API can use a local whisper.cpp server when WHISPER_SERVER_URL is configured. Audio is sent only to that local endpoint; with the variable unset, uploads remain stored as awaiting_transcription and can still be searched after a transcript is edited or populated.

For NVIDIA GPU transcription, run a CUDA-enabled whisper.cpp server with a model on the same machine, then set WHISPER_SERVER_URL to its local inference endpoint, for example http://whisper:8080/inference. Use a model that fits the RTX 3070's VRAM; start with small or a quantized medium model and measure with your audio/language. Do not expose the transcription server outside the private Docker network. See the whisper.cpp NVIDIA instructions for CUDA builds.

This version keeps raw audio and transcript separate. A future local correction model can propose punctuation and name fixes without replacing the original transcription. Proper-name matching and offline geolocation data are not enabled yet.

Data and backups

Docker persists SQLite metadata in ./data/voice-kb.sqlite3 and original audio in ./data/audio/. Back up the whole ./data directory while the service is stopped or use a SQLite-aware backup procedure. Model files should also be stored locally and backed up separately if needed.

Development

python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
uvicorn app.main:app --reload

The web app is served by FastAPI at the root URL. SQLite FTS5 indexes transcript text for plain full-text search.