# STT Runner Speech-to-Text transcription using sherpa-onnx + Qwen3-ASR, Text-to-Speech with ZipVoice (zero-shot voice cloning), LLM/embedding inference with llama.cpp (Granite-4.2 + nomic-embed-text-v1.5), and text classification with Laya typed decisions on the CPU-only laya.cpp runtime. > **Panduan lengkap classification runner (bahasa Indonesia): [CLASSIFY.md](CLASSIFY.md)** ## Installation ```bash python3 -m venv .venv .venv/bin/pip install -r requirements.txt ``` ## Usage ### Speech-to-Text ```bash python stt_runner.py [--language=Indonesian] audio1.wav audio2.wav ... ``` ### Text-to-Speech ```bash python tts_runner.py [--output=out.wav] [--ref-audio=ref.wav] [--ref-text="..."] "text to speak" ``` Output defaults to `output.wav`. The reference audio/text (voice to clone) is set in `config/tts.py` and can be overridden per-run with `--ref-audio` / `--ref-text` (the text must match the audio exactly). ### LLM Chat (llama.cpp) ```bash python llm_runner.py "What is the capital of France?" # single prompt python llm_runner.py # interactive chat ``` ### Embedding (llama.cpp) ```bash python embed_runner.py "text to embed" "another text" python embed_runner.py "What is TSNE?" --prefix "search_query: " ``` ### Text Classification (laya.cpp, CPU) ```bash ./build-laya.sh # once, no sudo needed ./usage-classify.sh "Saya dikenakan biaya dua kali, tolong refund." ./usage-classify.sh --state-file ticket.txt --preset triage ./usage-classify.sh --questions my_questions.json "text" --json ``` Laya is a decision model, not a single-label classifier: you supply **typed questions** and each one declares its own answer space, so nothing has to be retrained. Three types are supported — `choice` (picks a labelled option), `score` (returns a level from an ordered list) and `noul` (returns a probability that a statement holds). The default `triage` preset in `config/classify.py` asks three questions; `--questions FILE` replaces it with your own. The file may hold a bare questions object or a full laya.cpp request object. A state starting with `{` or `[` is passed through as a structure, so `{"subject": ..., "body": ...}` works too. The checkpoint is the 322M multilingual model, so non-English input is expected to work: ``` $ ./usage-classify.sh "二重に請求されました。返金をお願いします。" refund : 0.9976 ``` Each run spawns `laya-cli` and reloads the 644MB weights, which dominates the wall clock (~5.5s total, of which ~0.6s is inference). For repeated calls keep the checkpoint resident and point the runner at the HTTP server — same answers, ~8x faster end to end: ```bash third_party/laya.cpp/build-cpu/bin/laya-cli --server --port 8080 \ --model models/convaiinnovations/laya --variant multilingual --cpu & LAYA_URL=http://127.0.0.1:8080 ./usage-classify.sh "..." ``` `LAYA_URL` is read from the environment, `--url` overrides it, and `LAYA_URL` in `config/classify.py` is the fallback. If the server cannot be reached the runner says so and falls back to the CLI. Two caveats from the upstream model card, worth knowing before you trust the numbers: the probabilities ship **uncalibrated** and systematically over-confident, and `noul` can under-report a clear `true`. Fit your own thresholds on held-out data before treating a probability as a decision. ## Configuration Model paths and inference parameters are hardcoded in `config/`: - `config/model.py` — model paths (conv_frontend, encoder, decoder, tokenizer under `models/`) - `config/asr.py` — inference params: `LANGUAGE`, `HOTWORDS`, `NUM_THREADS`, `PROVIDER`, `SAMPLE_RATE`, `FEATURE_DIM`, `MAX_TOTAL_LEN`, `MAX_NEW_TOKENS` - `config/tts.py` — TTS model paths, `REFERENCE_AUDIO`, `REFERENCE_TEXT`, `OUTPUT_FILE`, `NUM_THREADS`, `PROVIDER`, `NUM_STEPS` - `config/llm.py` — Granite-4.2 model path, `N_CTX`, `N_THREADS`, `N_GPU_LAYERS`, `MAX_TOKENS`, `TEMPERATURE`, `TOP_P`, `TOP_K`, `SYSTEM_PROMPT`, `CHAT_TEMPLATE` - `config/embed.py` — nomic-embed-text-v1.5 model path, `N_CTX`, `N_THREADS`, `PREFIX_QUERY`, `PREFIX_DOCUMENT`, `DEFAULT_PREFIX` - `config/classify.py` — laya.cpp binary path, `MODEL_ROOT`/`VARIANT`, `BACKEND`, `LAYA_URL`, `PRESETS`, `DEFAULT_PRESET` `LANGUAGE` defaults to `""` (all languages / auto-detect). Passing `--language` on the CLI overrides it. ## Build laya.cpp (CPU-only) `classify_runner.py` drives the native `laya-cli` binary, so there is no Python inference dependency — `requirements.txt` is unchanged, and in particular PyTorch is not needed. ```bash ./build-laya.sh ``` The script needs no `sudo`: `cmake` and `ninja` come from pip wheels, `nlohmann/json` is unpacked into `third_party/nlohmann-install`, and the only system requirement is `libicu-dev` (`apt install libicu-dev` if it is missing). Re-running it skips finished steps. The manual equivalent: ```bash git clone --recursive --shallow-submodules https://github.com/lkarlslund/laya.cpp third_party/laya.cpp .venv/bin/pip install "cmake==3.31.*" "ninja<2" export PATH="$PWD/.venv/bin:$PATH" cmake -S third_party/laya.cpp -B third_party/laya.cpp/build-cpu -G Ninja \ -DCMAKE_BUILD_TYPE=Release \ -DLAYA_CUDA=OFF -DLAYA_VULKAN=OFF -DLAYA_COREML=OFF -DBUILD_TESTING=OFF \ -DCMAKE_PREFIX_PATH="$PWD/third_party/nlohmann-install" cmake --build third_party/laya.cpp/build-cpu --parallel "$(nproc)" ``` Three flags matter: - **`-DLAYA_CUDA=OFF` is mandatory** — `LAYA_CUDA` defaults to `ON` and configuration fails looking for `nvcc`. - **`cmake` must stay on 3.x** — laya.cpp's pinned ggml declares `cmake_minimum_required(VERSION 3.14...3.28)` and does not configure on CMake 4. - **`--cpu` is always passed at runtime** — `laya-cli` defaults to the CUDA backend, so a CPU-only build fails at startup without it. Upstream ships no CPU-only prebuilt binary (Linux releases are CUDA 12, CUDA 13 and Vulkan only), which is why this is built from source. On an NVIDIA box you can swap `BACKEND` in `config/classify.py` to `cuda` and download a prebuilt `laya` instead, keeping the build step out of the loop entirely. `models/convaiinnovations/laya/multilingual/` holds the checkpoint in the layout the CLI expects. Note that `laya-cli` appends the variant to `--model` for every non-English variant, so the multilingual checkpoint must be at `/multilingual/`. `config/classify.py` sets `MODEL_ROOT` to the parent directory. A separate copy of the same weights sits at `models/convaiinnovations/laya-multilingual/`, which this layout does **not** address — the standalone HuggingFace repo directory is for the Python reference implementation. ## Download Model (Qwen3-ASR 1.7B int8) ```bash BASE="https://modelscope.cn/models/zengshuishui/Qwen3-ASR-onnx/resolve/master" mkdir -p models/model_1.7B models/tokenizer wget -O models/model_1.7B/conv_frontend.onnx "$BASE/model_1.7B/conv_frontend.onnx" wget -O models/model_1.7B/encoder.int8.onnx "$BASE/model_1.7B/encoder.int8.onnx" wget -O models/model_1.7B/decoder.int8.onnx "$BASE/model_1.7B/decoder.int8.onnx" for f in vocab.json merges.txt tokenizer_config.json preprocessor_config.json config.json chat_template.json; do wget -O "models/tokenizer/$f" "$BASE/tokenizer/$f" done ``` ## Download Model (ZipVoice TTS) ```bash mkdir -p models/zipvoice wget -qO- https://github.com/k2-fsa/sherpa-onnx/releases/download/tts-models/sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2 \ | tar xjf - -C models/zipvoice --strip-components=1 wget -O models/zipvoice/vocos_24khz.onnx \ https://github.com/k2-fsa/sherpa-onnx/releases/download/vocoder-models/vocos_24khz.onnx ``` ## Download Model (Granite-4.2 8B Q4_K_M) ```bash mkdir -p models/granite-4.2 wget -O models/granite-4.2/granite-4.2-8b-Q4_K_M.gguf \ "https://huggingface.co/ibm-granite/granite-4.2-8b-GGUF/resolve/main/granite-4.2-8b-Q4_K_M.gguf" ``` ## Download Model (nomic-embed-text-v1.5) ```bash mkdir -p models/nomic-embed-text-v1.5 wget -O models/nomic-embed-text-v1.5/nomic-embed-text-v1.5.Q4_K_M.gguf \ "https://huggingface.co/nomic-ai/nomic-embed-text-v1.5-GGUF/resolve/main/nomic-embed-text-v1.5.Q4_K_M.gguf" ```