60 lines
5.8 KiB
Markdown
60 lines
5.8 KiB
Markdown
# Faerro KB
|
|
|
|
A self-hosted voice-note inbox for an iPhone-installed web app. Recordings are saved in the browser first, uploaded when connectivity is available, and stored with transcripts on your own server.
|
|
|
|
## Run locally
|
|
|
|
```sh
|
|
docker compose up --build
|
|
```
|
|
|
|
## Releases
|
|
|
|
Git tags are the source of truth for release versions. From the repository root, run the publish script to create the next patch tag and publish it. A release requires at least one commit after the latest version tag; repeated runs without new commits are rejected:
|
|
|
|
```sh
|
|
cp .env.example .env
|
|
# Set GITEA_USERNAME and GITEA_TOKEN in .env, using a Gitea package-write token.
|
|
./scripts/publish-image.sh
|
|
```
|
|
|
|
Pass `--minor` to increment the minor version and reset patch to zero, or `--major` to increment the major version and reset minor and patch to zero. The options cannot be combined. Run the script from a clean worktree: staged, unstaged, and untracked changes must be committed or removed first. It creates an annotated local Git tag on the checked-out commit, then publishes both the versioned image and `latest`. The commit does not need to be pushed before creating the local tag, but push the commit to the Git server first when it should be shared as part of the release; the script does not push Git tags. `.env` is excluded from Git; the script reads the Gitea credentials from it and passes the token to Docker through standard input. If those credentials are unset, the script uses Docker's existing login credentials.
|
|
|
|
After both pushes succeed, the script removes older local Docker tags for this image while keeping the newly published version and `latest`. This cleanup affects the local Docker image cache, not previously published packages in the Gitea registry.
|
|
|
|
Open `http://localhost` for a desktop smoke test. Set `HTTP_PORT` to publish the app on a different host port (for example, `HTTP_PORT=8000 docker compose up`). For iPhone use, serve the app over HTTPS through your VPN/reverse proxy. iOS only grants microphone access in a secure context. Keep the service reachable only through your VPN and firewall; this first version has no login layer.
|
|
|
|
In Safari on iPhone, open the HTTPS address and use **Share > Add to Home Screen**. Sign in to the VPN before opening the app if the server is only reachable there.
|
|
|
|
## Capture and offline behavior
|
|
|
|
- Recording data and its timestamp are committed to IndexedDB before upload is attempted.
|
|
- Written-note drafts and queued text notes are stored in IndexedDB on the device. A draft is restored when the app is reopened, and queued text is removed only after the server confirms storage.
|
|
- Microphone recordings use the audio format supported by the browser (commonly M4A/AAC on iPhone Safari or Opus/WebM elsewhere). You can also add an audio file from the device; both sources remain queued locally until the server confirms storage.
|
|
- The server keeps each uploaded original and creates a mono PCM WAV copy at 16-bit / 22.05 kHz for transcription. FFmpeg performs this conversion locally in the container. After a non-empty transcript is stored, both audio files are deleted; failed or empty transcriptions keep the audio for recovery.
|
|
- Uploads retry when the app opens, returns to the foreground, or the browser reports a connection. A recording is removed from the phone's queue only after the server confirms it was stored.
|
|
- Service worker caching keeps the app shell available offline. Searching the server-side archive requires a connection.
|
|
- iOS may suspend a PWA and can evict website data under storage pressure. Background uploads are not guaranteed; reopen the PWA while online to resume. Keep the phone powered and avoid clearing Safari website data for stronger practical retention.
|
|
|
|
## Local speech-to-text
|
|
|
|
The API can use a local `onerahmet/openai-whisper-asr-webservice` server when `WHISPER_SERVER_URL` is configured. Set it to the service's `/asr` endpoint, including its listening port, for example `http://whisper:9000/asr` or `http://192.168.1.102:9000/asr`. Faerro KB sends the converted WAV as the `audio_file` multipart field with `task=transcribe` and `output=json`; audio is sent only to that local endpoint. At startup, the API also sends the bundled `audio test.wav` to the endpoint as a connection test before resuming queued jobs. Logs report the HTTP result, detected language, segment count, and elapsed time, but never log the transcript. With the variable unset, uploads remain stored as `awaiting_transcription` until a non-empty transcript is entered and saved. When the API starts with a successful Whisper test, it resumes awaiting, interrupted, and failed transcription jobs whose audio is still stored. Restart/recreate the API container after changing this environment variable; configuration is read at startup. Saving a non-empty transcript deletes both server-side audio files.
|
|
|
|
Run the ASR web service with a model that fits the available hardware, then set `WHISPER_SERVER_URL` to its `/asr` endpoint. Keep the transcription server private to the Docker network and do not expose it publicly.
|
|
|
|
The transcript remains in SQLite after the audio files are deleted. A future local correction model can propose punctuation and name fixes without replacing the stored transcription. Proper-name matching and offline geolocation data are not enabled yet.
|
|
|
|
## Data and backups
|
|
|
|
Docker persists SQLite metadata in `./data/voice-kb.sqlite3` and audio awaiting transcription in `./data/audio/`. After a non-empty transcript is saved, its audio files are removed. Back up the whole `./data` directory while the service is stopped or use a SQLite-aware backup procedure. Model files should also be stored locally and backed up separately if needed.
|
|
|
|
## Development
|
|
|
|
```sh
|
|
python -m venv .venv
|
|
source .venv/bin/activate
|
|
pip install -r requirements.txt
|
|
uvicorn app.main:app --reload
|
|
```
|
|
|
|
The web app is served by FastAPI at the root URL. SQLite FTS5 indexes transcript text for plain full-text search. |