5.3 KiB
Faerro KB
A self-hosted voice-note inbox for an iPhone-installed web app. Recordings are saved in the browser first, uploaded when connectivity is available, and stored with transcripts on your own server.
Run locally
docker compose up --build
Releases
Git tags are the source of truth for release versions. From the repository root, run the publish script to create the next patch tag and publish it. A release requires at least one commit after the latest version tag; repeated runs without new commits are rejected:
cp .env.example .env
# Set GITEA_USERNAME and GITEA_TOKEN in .env, using a Gitea package-write token.
./scripts/publish-image.sh
Pass --minor to increment the minor version and reset patch to zero, or --major to increment the major version and reset minor and patch to zero. The options cannot be combined. Run the script from a clean worktree: staged, unstaged, and untracked changes must be committed or removed first. It creates an annotated local Git tag on the checked-out commit, then publishes both the versioned image and latest. The commit does not need to be pushed before creating the local tag, but push the commit to the Git server first when it should be shared as part of the release; the script does not push Git tags. .env is excluded from Git; the script reads the Gitea credentials from it and passes the token to Docker through standard input. If those credentials are unset, the script uses Docker's existing login credentials.
After both pushes succeed, the script removes older local Docker tags for this image while keeping the newly published version and latest. This cleanup affects the local Docker image cache, not previously published packages in the Gitea registry.
Open http://localhost for a desktop smoke test. Set HTTP_PORT to publish the app on a different host port (for example, HTTP_PORT=8000 docker compose up). For iPhone use, serve the app over HTTPS through your VPN/reverse proxy. iOS only grants microphone access in a secure context. Keep the service reachable only through your VPN and firewall; this first version has no login layer.
In Safari on iPhone, open the HTTPS address and use Share > Add to Home Screen. Sign in to the VPN before opening the app if the server is only reachable there.
Capture and offline behavior
- Recording data and its timestamp are committed to IndexedDB before upload is attempted.
- Written-note drafts and queued text notes are stored in IndexedDB on the device. A draft is restored when the app is reopened, and queued text is removed only after the server confirms storage.
- Microphone recordings use the audio format supported by the browser (commonly M4A/AAC on iPhone Safari or Opus/WebM elsewhere). You can also add an audio file from the device; both sources remain queued locally until the server confirms storage.
- The server keeps each uploaded original and creates a mono PCM WAV copy at 16-bit / 22.05 kHz for transcription. FFmpeg performs this conversion locally in the container. After a non-empty transcript is stored, both audio files are deleted; failed or empty transcriptions keep the audio for recovery.
- Uploads retry when the app opens, returns to the foreground, or the browser reports a connection. A recording is removed from the phone's queue only after the server confirms it was stored.
- Service worker caching keeps the app shell available offline. Searching the server-side archive requires a connection.
- iOS may suspend a PWA and can evict website data under storage pressure. Background uploads are not guaranteed; reopen the PWA while online to resume. Keep the phone powered and avoid clearing Safari website data for stronger practical retention.
Local speech-to-text
The API can use a local onerahmet/openai-whisper-asr-webservice server when WHISPER_SERVER_URL is configured. Set it to the service's /asr endpoint, for example http://whisper:9000/asr. Faerro KB sends the converted WAV as the audio_file multipart field with task=transcribe and output=json; audio is sent only to that local endpoint. With the variable unset, uploads remain stored as awaiting_transcription until a non-empty transcript is entered and saved. Saving that transcript deletes both server-side audio files.
Run the ASR web service with a model that fits the available hardware, then set WHISPER_SERVER_URL to its /asr endpoint. Keep the transcription server private to the Docker network and do not expose it publicly.
The transcript remains in SQLite after the audio files are deleted. A future local correction model can propose punctuation and name fixes without replacing the stored transcription. Proper-name matching and offline geolocation data are not enabled yet.
Data and backups
Docker persists SQLite metadata in ./data/voice-kb.sqlite3 and audio awaiting transcription in ./data/audio/. After a non-empty transcript is saved, its audio files are removed. Back up the whole ./data directory while the service is stopped or use a SQLite-aware backup procedure. Model files should also be stored locally and backed up separately if needed.
Development
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
uvicorn app.main:app --reload
The web app is served by FastAPI at the root URL. SQLite FTS5 indexes transcript text for plain full-text search.