178 lines
8.0 KiB
Markdown
178 lines
8.0 KiB
Markdown
# STT Runner
|
|
|
|
Speech-to-Text transcription using sherpa-onnx + Qwen3-ASR, Text-to-Speech with ZipVoice (zero-shot voice cloning), LLM/embedding inference with llama.cpp (Granite-4.2 + nomic-embed-text-v1.5), and text classification with Laya typed decisions on the CPU-only laya.cpp runtime.
|
|
|
|
> **Panduan lengkap classification runner (bahasa Indonesia): [CLASSIFY.md](CLASSIFY.md)**
|
|
|
|
## Installation
|
|
|
|
```bash
|
|
python3 -m venv .venv
|
|
.venv/bin/pip install -r requirements.txt
|
|
```
|
|
|
|
## Usage
|
|
|
|
### Speech-to-Text
|
|
|
|
```bash
|
|
python stt_runner.py [--language=Indonesian] audio1.wav audio2.wav ...
|
|
```
|
|
|
|
### Text-to-Speech
|
|
|
|
```bash
|
|
python tts_runner.py [--output=out.wav] [--ref-audio=ref.wav] [--ref-text="..."] "text to speak"
|
|
```
|
|
|
|
Output defaults to `output.wav`. The reference audio/text (voice to clone) is set in `config/tts.py` and can be overridden per-run with `--ref-audio` / `--ref-text` (the text must match the audio exactly).
|
|
|
|
### LLM Chat (llama.cpp)
|
|
|
|
```bash
|
|
python llm_runner.py "What is the capital of France?" # single prompt
|
|
python llm_runner.py # interactive chat
|
|
```
|
|
|
|
### Embedding (llama.cpp)
|
|
|
|
```bash
|
|
python embed_runner.py "text to embed" "another text"
|
|
python embed_runner.py "What is TSNE?" --prefix "search_query: "
|
|
```
|
|
|
|
### Text Classification (laya.cpp, CPU)
|
|
|
|
```bash
|
|
./build-laya.sh # once, no sudo needed
|
|
./usage-classify.sh "Saya dikenakan biaya dua kali, tolong refund."
|
|
./usage-classify.sh --state-file ticket.txt --preset triage
|
|
./usage-classify.sh --questions my_questions.json "text" --json
|
|
```
|
|
|
|
Laya is a decision model, not a single-label classifier: you supply **typed questions** and each
|
|
one declares its own answer space, so nothing has to be retrained. Three types are supported —
|
|
`choice` (picks a labelled option), `score` (returns a level from an ordered list) and `noul`
|
|
(returns a probability that a statement holds). The default `triage` preset in `config/classify.py`
|
|
asks three questions; `--questions FILE` replaces it with your own. The file may hold a bare
|
|
questions object or a full laya.cpp request object. A state starting with `{` or `[` is passed
|
|
through as a structure, so `{"subject": ..., "body": ...}` works too.
|
|
|
|
The checkpoint is the 322M multilingual model, so non-English input is expected to work:
|
|
|
|
```
|
|
$ ./usage-classify.sh "二重に請求されました。返金をお願いします。"
|
|
refund : 0.9976
|
|
```
|
|
|
|
Each run spawns `laya-cli` and reloads the 644MB weights, which dominates the wall clock
|
|
(~5.5s total, of which ~0.6s is inference). For repeated calls keep the checkpoint resident and
|
|
point the runner at the HTTP server — same answers, ~8x faster end to end:
|
|
|
|
```bash
|
|
third_party/laya.cpp/build-cpu/bin/laya-cli --server --port 8080 \
|
|
--model models/convaiinnovations/laya --variant multilingual --cpu &
|
|
LAYA_URL=http://127.0.0.1:8080 ./usage-classify.sh "..."
|
|
```
|
|
|
|
`LAYA_URL` is read from the environment, `--url` overrides it, and `LAYA_URL` in `config/classify.py`
|
|
is the fallback. If the server cannot be reached the runner says so and falls back to the CLI.
|
|
|
|
Two caveats from the upstream model card, worth knowing before you trust the numbers: the
|
|
probabilities ship **uncalibrated** and systematically over-confident, and `noul` can under-report
|
|
a clear `true`. Fit your own thresholds on held-out data before treating a probability as a decision.
|
|
|
|
## Configuration
|
|
|
|
Model paths and inference parameters are hardcoded in `config/`:
|
|
|
|
- `config/model.py` — model paths (conv_frontend, encoder, decoder, tokenizer under `models/`)
|
|
- `config/asr.py` — inference params: `LANGUAGE`, `HOTWORDS`, `NUM_THREADS`, `PROVIDER`, `SAMPLE_RATE`, `FEATURE_DIM`, `MAX_TOTAL_LEN`, `MAX_NEW_TOKENS`
|
|
- `config/tts.py` — TTS model paths, `REFERENCE_AUDIO`, `REFERENCE_TEXT`, `OUTPUT_FILE`, `NUM_THREADS`, `PROVIDER`, `NUM_STEPS`
|
|
- `config/llm.py` — Granite-4.2 model path, `N_CTX`, `N_THREADS`, `N_GPU_LAYERS`, `MAX_TOKENS`, `TEMPERATURE`, `TOP_P`, `TOP_K`, `SYSTEM_PROMPT`, `CHAT_TEMPLATE`
|
|
- `config/embed.py` — nomic-embed-text-v1.5 model path, `N_CTX`, `N_THREADS`, `PREFIX_QUERY`, `PREFIX_DOCUMENT`, `DEFAULT_PREFIX`
|
|
- `config/classify.py` — laya.cpp binary path, `MODEL_ROOT`/`VARIANT`, `BACKEND`, `LAYA_URL`, `PRESETS`, `DEFAULT_PRESET`
|
|
|
|
`LANGUAGE` defaults to `""` (all languages / auto-detect). Passing `--language` on the CLI overrides it.
|
|
|
|
## Build laya.cpp (CPU-only)
|
|
|
|
`classify_runner.py` drives the native `laya-cli` binary, so there is no Python inference
|
|
dependency — `requirements.txt` is unchanged, and in particular PyTorch is not needed.
|
|
|
|
```bash
|
|
./build-laya.sh
|
|
```
|
|
|
|
The script needs no `sudo`: `cmake` and `ninja` come from pip wheels, `nlohmann/json` is unpacked
|
|
into `third_party/nlohmann-install`, and the only system requirement is `libicu-dev`
|
|
(`apt install libicu-dev` if it is missing). Re-running it skips finished steps. The manual
|
|
equivalent:
|
|
|
|
```bash
|
|
git clone --recursive --shallow-submodules https://github.com/lkarlslund/laya.cpp third_party/laya.cpp
|
|
.venv/bin/pip install "cmake==3.31.*" "ninja<2"
|
|
export PATH="$PWD/.venv/bin:$PATH"
|
|
cmake -S third_party/laya.cpp -B third_party/laya.cpp/build-cpu -G Ninja \
|
|
-DCMAKE_BUILD_TYPE=Release \
|
|
-DLAYA_CUDA=OFF -DLAYA_VULKAN=OFF -DLAYA_COREML=OFF -DBUILD_TESTING=OFF \
|
|
-DCMAKE_PREFIX_PATH="$PWD/third_party/nlohmann-install"
|
|
cmake --build third_party/laya.cpp/build-cpu --parallel "$(nproc)"
|
|
```
|
|
|
|
Three flags matter:
|
|
|
|
- **`-DLAYA_CUDA=OFF` is mandatory** — `LAYA_CUDA` defaults to `ON` and configuration fails looking for `nvcc`.
|
|
- **`cmake` must stay on 3.x** — laya.cpp's pinned ggml declares `cmake_minimum_required(VERSION 3.14...3.28)` and does not configure on CMake 4.
|
|
- **`--cpu` is always passed at runtime** — `laya-cli` defaults to the CUDA backend, so a CPU-only build fails at startup without it.
|
|
|
|
Upstream ships no CPU-only prebuilt binary (Linux releases are CUDA 12, CUDA 13 and Vulkan only),
|
|
which is why this is built from source. On an NVIDIA box you can swap `BACKEND` in
|
|
`config/classify.py` to `cuda` and download a prebuilt `laya` instead, keeping the build step out of
|
|
the loop entirely.
|
|
|
|
`models/convaiinnovations/laya/multilingual/` holds the checkpoint in the layout the CLI expects.
|
|
Note that `laya-cli` appends the variant to `--model` for every non-English variant, so the
|
|
multilingual checkpoint must be at `<MODEL_ROOT>/multilingual/`. `config/classify.py` sets
|
|
`MODEL_ROOT` to the parent directory. A separate copy of the same weights sits at
|
|
`models/convaiinnovations/laya-multilingual/`, which this layout does **not** address — the
|
|
standalone HuggingFace repo directory is for the Python reference implementation.
|
|
|
|
## Download Model (Qwen3-ASR 1.7B int8)
|
|
|
|
```bash
|
|
BASE="https://modelscope.cn/models/zengshuishui/Qwen3-ASR-onnx/resolve/master"
|
|
mkdir -p models/model_1.7B models/tokenizer
|
|
wget -O models/model_1.7B/conv_frontend.onnx "$BASE/model_1.7B/conv_frontend.onnx"
|
|
wget -O models/model_1.7B/encoder.int8.onnx "$BASE/model_1.7B/encoder.int8.onnx"
|
|
wget -O models/model_1.7B/decoder.int8.onnx "$BASE/model_1.7B/decoder.int8.onnx"
|
|
for f in vocab.json merges.txt tokenizer_config.json preprocessor_config.json config.json chat_template.json; do
|
|
wget -O "models/tokenizer/$f" "$BASE/tokenizer/$f"
|
|
done
|
|
```
|
|
|
|
## Download Model (ZipVoice TTS)
|
|
|
|
```bash
|
|
mkdir -p models/zipvoice
|
|
wget -qO- https://github.com/k2-fsa/sherpa-onnx/releases/download/tts-models/sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2 \
|
|
| tar xjf - -C models/zipvoice --strip-components=1
|
|
wget -O models/zipvoice/vocos_24khz.onnx \
|
|
https://github.com/k2-fsa/sherpa-onnx/releases/download/vocoder-models/vocos_24khz.onnx
|
|
```
|
|
|
|
## Download Model (Granite-4.2 8B Q4_K_M)
|
|
|
|
```bash
|
|
mkdir -p models/granite-4.2
|
|
wget -O models/granite-4.2/granite-4.2-8b-Q4_K_M.gguf \
|
|
"https://huggingface.co/ibm-granite/granite-4.2-8b-GGUF/resolve/main/granite-4.2-8b-Q4_K_M.gguf"
|
|
```
|
|
|
|
## Download Model (nomic-embed-text-v1.5)
|
|
|
|
```bash
|
|
mkdir -p models/nomic-embed-text-v1.5
|
|
wget -O models/nomic-embed-text-v1.5/nomic-embed-text-v1.5.Q4_K_M.gguf \
|
|
"https://huggingface.co/nomic-ai/nomic-embed-text-v1.5-GGUF/resolve/main/nomic-embed-text-v1.5.Q4_K_M.gguf"
|
|
``` |