Files
faerro-kb/README.md
T

55 lines
3.6 KiB
Markdown

# Voice KB
A self-hosted voice-note inbox for an iPhone-installed web app. Recordings are saved in the browser first, uploaded when connectivity is available, and stored with transcripts on your own server.
## Run locally
```sh
docker compose up --build
```
## Releases
Git tags are the source of truth for release versions. From the repository root, create a semantic version tag on the commit to release, then run the publish script:
```sh
git tag 0.0.2
docker login gitea.faerro.it
./scripts/publish-image.sh
```
The script requires Docker and Git, and publishes both the versioned image (for example, `gitea.faerro.it/andrea/faerro-kb:0.0.2`) and `latest`. Tags may be plain versions or prefixed with `v`. Authenticate with a Gitea account or token that can publish packages before running the script.
Open `http://localhost` for a desktop smoke test. Set `HTTP_PORT` to publish the app on a different host port (for example, `HTTP_PORT=8000 docker compose up`). For iPhone use, serve the app over HTTPS through your VPN/reverse proxy. iOS only grants microphone access in a secure context. Keep the service reachable only through your VPN and firewall; this first version has no login layer.
In Safari on iPhone, open the HTTPS address and use **Share > Add to Home Screen**. Sign in to the VPN before opening the app if the server is only reachable there.
## Capture and offline behavior
- Recording data and its timestamp are committed to IndexedDB before upload is attempted.
- Uploads retry when the app opens, returns to the foreground, or the browser reports a connection. A recording is removed from the phone's queue only after the server confirms it was stored.
- Service worker caching keeps the app shell available offline. Searching the server-side archive requires a connection.
- iOS may suspend a PWA and can evict website data under storage pressure. Background uploads are not guaranteed; reopen the PWA while online to resume. Keep the phone powered and avoid clearing Safari website data for stronger practical retention.
## Local speech-to-text
The API can use a local `whisper.cpp` server when `WHISPER_SERVER_URL` is configured. Audio is sent only to that local endpoint; with the variable unset, uploads remain stored as `awaiting_transcription` and can still be searched after a transcript is edited or populated.
For NVIDIA GPU transcription, run a CUDA-enabled `whisper.cpp` server with a model on the same machine, then set `WHISPER_SERVER_URL` to its local inference endpoint, for example `http://whisper:8080/inference`. Use a model that fits the RTX 3070's VRAM; start with `small` or a quantized `medium` model and measure with your audio/language. Do not expose the transcription server outside the private Docker network. See the [whisper.cpp NVIDIA instructions](https://github.com/ggml-org/whisper.cpp#nvidia-gpu-support) for CUDA builds.
This version keeps raw audio and transcript separate. A future local correction model can propose punctuation and name fixes without replacing the original transcription. Proper-name matching and offline geolocation data are not enabled yet.
## Data and backups
Docker persists SQLite metadata in `./data/voice-kb.sqlite3` and original audio in `./data/audio/`. Back up the whole `./data` directory while the service is stopped or use a SQLite-aware backup procedure. Model files should also be stored locally and backed up separately if needed.
## Development
```sh
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
uvicorn app.main:app --reload
```
The web app is served by FastAPI at the root URL. SQLite FTS5 indexes transcript text for plain full-text search.