Update README and API integration: specify local ASR server details and adjust audio upload parameters

This commit is contained in:
2026-09-28 14:05:39 +02:00
parent 8d4f46a7d1
commit 0fb66f2cee
3 changed files with 5 additions and 5 deletions
+2 -2
View File
@@ -38,9 +38,9 @@ In Safari on iPhone, open the HTTPS address and use **Share > Add to Home Screen
## Local speech-to-text
The API can use a local `whisper.cpp` server when `WHISPER_SERVER_URL` is configured. Audio is sent only to that local endpoint; with the variable unset, uploads remain stored as `awaiting_transcription` until a non-empty transcript is entered and saved. Saving that transcript deletes both server-side audio files.
The API can use a local `onerahmet/openai-whisper-asr-webservice` server when `WHISPER_SERVER_URL` is configured. Set it to the service's `/asr` endpoint, for example `http://whisper:9000/asr`. Faerro KB sends the converted WAV as the `audio_file` multipart field with `task=transcribe` and `output=json`; audio is sent only to that local endpoint. With the variable unset, uploads remain stored as `awaiting_transcription` until a non-empty transcript is entered and saved. Saving that transcript deletes both server-side audio files.
For NVIDIA GPU transcription, run a CUDA-enabled `whisper.cpp` server with a model on the same machine, then set `WHISPER_SERVER_URL` to its local inference endpoint, for example `http://whisper:8080/inference`. Use a model that fits the RTX 3070's VRAM; start with `small` or a quantized `medium` model and measure with your audio/language. Do not expose the transcription server outside the private Docker network. See the [whisper.cpp NVIDIA instructions](https://github.com/ggml-org/whisper.cpp#nvidia-gpu-support) for CUDA builds.
Run the ASR web service with a model that fits the available hardware, then set `WHISPER_SERVER_URL` to its `/asr` endpoint. Keep the transcription server private to the Docker network and do not expose it publicly.
The transcript remains in SQLite after the audio files are deleted. A future local correction model can propose punctuation and name fixes without replacing the stored transcription. Proper-name matching and offline geolocation data are not enabled yet.