8.0 KiB
STT Runner
Speech-to-Text transcription using sherpa-onnx + Qwen3-ASR, Text-to-Speech with ZipVoice (zero-shot voice cloning), LLM/embedding inference with llama.cpp (Granite-4.2 + nomic-embed-text-v1.5), and text classification with Laya typed decisions on the CPU-only laya.cpp runtime.
Panduan lengkap classification runner (bahasa Indonesia): CLASSIFY.md
Installation
python3 -m venv .venv
.venv/bin/pip install -r requirements.txt
Usage
Speech-to-Text
python stt_runner.py [--language=Indonesian] audio1.wav audio2.wav ...
Text-to-Speech
python tts_runner.py [--output=out.wav] [--ref-audio=ref.wav] [--ref-text="..."] "text to speak"
Output defaults to output.wav. The reference audio/text (voice to clone) is set in config/tts.py and can be overridden per-run with --ref-audio / --ref-text (the text must match the audio exactly).
LLM Chat (llama.cpp)
python llm_runner.py "What is the capital of France?" # single prompt
python llm_runner.py # interactive chat
Embedding (llama.cpp)
python embed_runner.py "text to embed" "another text"
python embed_runner.py "What is TSNE?" --prefix "search_query: "
Text Classification (laya.cpp, CPU)
./build-laya.sh # once, no sudo needed
./usage-classify.sh "Saya dikenakan biaya dua kali, tolong refund."
./usage-classify.sh --state-file ticket.txt --preset triage
./usage-classify.sh --questions my_questions.json "text" --json
Laya is a decision model, not a single-label classifier: you supply typed questions and each
one declares its own answer space, so nothing has to be retrained. Three types are supported —
choice (picks a labelled option), score (returns a level from an ordered list) and noul
(returns a probability that a statement holds). The default triage preset in config/classify.py
asks three questions; --questions FILE replaces it with your own. The file may hold a bare
questions object or a full laya.cpp request object. A state starting with { or [ is passed
through as a structure, so {"subject": ..., "body": ...} works too.
The checkpoint is the 322M multilingual model, so non-English input is expected to work:
$ ./usage-classify.sh "二重に請求されました。返金をお願いします。"
refund : 0.9976
Each run spawns laya-cli and reloads the 644MB weights, which dominates the wall clock
(~5.5s total, of which ~0.6s is inference). For repeated calls keep the checkpoint resident and
point the runner at the HTTP server — same answers, ~8x faster end to end:
third_party/laya.cpp/build-cpu/bin/laya-cli --server --port 8080 \
--model models/convaiinnovations/laya --variant multilingual --cpu &
LAYA_URL=http://127.0.0.1:8080 ./usage-classify.sh "..."
LAYA_URL is read from the environment, --url overrides it, and LAYA_URL in config/classify.py
is the fallback. If the server cannot be reached the runner says so and falls back to the CLI.
Two caveats from the upstream model card, worth knowing before you trust the numbers: the
probabilities ship uncalibrated and systematically over-confident, and noul can under-report
a clear true. Fit your own thresholds on held-out data before treating a probability as a decision.
Configuration
Model paths and inference parameters are hardcoded in config/:
config/model.py— model paths (conv_frontend, encoder, decoder, tokenizer undermodels/)config/asr.py— inference params:LANGUAGE,HOTWORDS,NUM_THREADS,PROVIDER,SAMPLE_RATE,FEATURE_DIM,MAX_TOTAL_LEN,MAX_NEW_TOKENSconfig/tts.py— TTS model paths,REFERENCE_AUDIO,REFERENCE_TEXT,OUTPUT_FILE,NUM_THREADS,PROVIDER,NUM_STEPSconfig/llm.py— Granite-4.2 model path,N_CTX,N_THREADS,N_GPU_LAYERS,MAX_TOKENS,TEMPERATURE,TOP_P,TOP_K,SYSTEM_PROMPT,CHAT_TEMPLATEconfig/embed.py— nomic-embed-text-v1.5 model path,N_CTX,N_THREADS,PREFIX_QUERY,PREFIX_DOCUMENT,DEFAULT_PREFIXconfig/classify.py— laya.cpp binary path,MODEL_ROOT/VARIANT,BACKEND,LAYA_URL,PRESETS,DEFAULT_PRESET
LANGUAGE defaults to "" (all languages / auto-detect). Passing --language on the CLI overrides it.
Build laya.cpp (CPU-only)
classify_runner.py drives the native laya-cli binary, so there is no Python inference
dependency — requirements.txt is unchanged, and in particular PyTorch is not needed.
./build-laya.sh
The script needs no sudo: cmake and ninja come from pip wheels, nlohmann/json is unpacked
into third_party/nlohmann-install, and the only system requirement is libicu-dev
(apt install libicu-dev if it is missing). Re-running it skips finished steps. The manual
equivalent:
git clone --recursive --shallow-submodules https://github.com/lkarlslund/laya.cpp third_party/laya.cpp
.venv/bin/pip install "cmake==3.31.*" "ninja<2"
export PATH="$PWD/.venv/bin:$PATH"
cmake -S third_party/laya.cpp -B third_party/laya.cpp/build-cpu -G Ninja \
-DCMAKE_BUILD_TYPE=Release \
-DLAYA_CUDA=OFF -DLAYA_VULKAN=OFF -DLAYA_COREML=OFF -DBUILD_TESTING=OFF \
-DCMAKE_PREFIX_PATH="$PWD/third_party/nlohmann-install"
cmake --build third_party/laya.cpp/build-cpu --parallel "$(nproc)"
Three flags matter:
-DLAYA_CUDA=OFFis mandatory —LAYA_CUDAdefaults toONand configuration fails looking fornvcc.cmakemust stay on 3.x — laya.cpp's pinned ggml declarescmake_minimum_required(VERSION 3.14...3.28)and does not configure on CMake 4.--cpuis always passed at runtime —laya-clidefaults to the CUDA backend, so a CPU-only build fails at startup without it.
Upstream ships no CPU-only prebuilt binary (Linux releases are CUDA 12, CUDA 13 and Vulkan only),
which is why this is built from source. On an NVIDIA box you can swap BACKEND in
config/classify.py to cuda and download a prebuilt laya instead, keeping the build step out of
the loop entirely.
models/convaiinnovations/laya/multilingual/ holds the checkpoint in the layout the CLI expects.
Note that laya-cli appends the variant to --model for every non-English variant, so the
multilingual checkpoint must be at <MODEL_ROOT>/multilingual/. config/classify.py sets
MODEL_ROOT to the parent directory. A separate copy of the same weights sits at
models/convaiinnovations/laya-multilingual/, which this layout does not address — the
standalone HuggingFace repo directory is for the Python reference implementation.
Download Model (Qwen3-ASR 1.7B int8)
BASE="https://modelscope.cn/models/zengshuishui/Qwen3-ASR-onnx/resolve/master"
mkdir -p models/model_1.7B models/tokenizer
wget -O models/model_1.7B/conv_frontend.onnx "$BASE/model_1.7B/conv_frontend.onnx"
wget -O models/model_1.7B/encoder.int8.onnx "$BASE/model_1.7B/encoder.int8.onnx"
wget -O models/model_1.7B/decoder.int8.onnx "$BASE/model_1.7B/decoder.int8.onnx"
for f in vocab.json merges.txt tokenizer_config.json preprocessor_config.json config.json chat_template.json; do
wget -O "models/tokenizer/$f" "$BASE/tokenizer/$f"
done
Download Model (ZipVoice TTS)
mkdir -p models/zipvoice
wget -qO- https://github.com/k2-fsa/sherpa-onnx/releases/download/tts-models/sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2 \
| tar xjf - -C models/zipvoice --strip-components=1
wget -O models/zipvoice/vocos_24khz.onnx \
https://github.com/k2-fsa/sherpa-onnx/releases/download/vocoder-models/vocos_24khz.onnx
Download Model (Granite-4.2 8B Q4_K_M)
mkdir -p models/granite-4.2
wget -O models/granite-4.2/granite-4.2-8b-Q4_K_M.gguf \
"https://huggingface.co/ibm-granite/granite-4.2-8b-GGUF/resolve/main/granite-4.2-8b-Q4_K_M.gguf"
Download Model (nomic-embed-text-v1.5)
mkdir -p models/nomic-embed-text-v1.5
wget -O models/nomic-embed-text-v1.5/nomic-embed-text-v1.5.Q4_K_M.gguf \
"https://huggingface.co/nomic-ai/nomic-embed-text-v1.5-GGUF/resolve/main/nomic-embed-text-v1.5.Q4_K_M.gguf"