ai-experiment/README.md

178 lines
8.0 KiB
Markdown

# STT Runner
Speech-to-Text transcription using sherpa-onnx + Qwen3-ASR, Text-to-Speech with ZipVoice (zero-shot voice cloning), LLM/embedding inference with llama.cpp (Granite-4.2 + nomic-embed-text-v1.5), and text classification with Laya typed decisions on the CPU-only laya.cpp runtime.
> **Panduan lengkap classification runner (bahasa Indonesia): [CLASSIFY.md](CLASSIFY.md)**
## Installation
```bash
python3 -m venv .venv
.venv/bin/pip install -r requirements.txt
```
## Usage
### Speech-to-Text
```bash
python stt_runner.py [--language=Indonesian] audio1.wav audio2.wav ...
```
### Text-to-Speech
```bash
python tts_runner.py [--output=out.wav] [--ref-audio=ref.wav] [--ref-text="..."] "text to speak"
```
Output defaults to `output.wav`. The reference audio/text (voice to clone) is set in `config/tts.py` and can be overridden per-run with `--ref-audio` / `--ref-text` (the text must match the audio exactly).
### LLM Chat (llama.cpp)
```bash
python llm_runner.py "What is the capital of France?" # single prompt
python llm_runner.py # interactive chat
```
### Embedding (llama.cpp)
```bash
python embed_runner.py "text to embed" "another text"
python embed_runner.py "What is TSNE?" --prefix "search_query: "
```
### Text Classification (laya.cpp, CPU)
```bash
./build-laya.sh # once, no sudo needed
./usage-classify.sh "Saya dikenakan biaya dua kali, tolong refund."
./usage-classify.sh --state-file ticket.txt --preset triage
./usage-classify.sh --questions my_questions.json "text" --json
```
Laya is a decision model, not a single-label classifier: you supply **typed questions** and each
one declares its own answer space, so nothing has to be retrained. Three types are supported —
`choice` (picks a labelled option), `score` (returns a level from an ordered list) and `noul`
(returns a probability that a statement holds). The default `triage` preset in `config/classify.py`
asks three questions; `--questions FILE` replaces it with your own. The file may hold a bare
questions object or a full laya.cpp request object. A state starting with `{` or `[` is passed
through as a structure, so `{"subject": ..., "body": ...}` works too.
The checkpoint is the 322M multilingual model, so non-English input is expected to work:
```
$ ./usage-classify.sh "二重に請求されました。返金をお願いします。"
refund : 0.9976
```
Each run spawns `laya-cli` and reloads the 644MB weights, which dominates the wall clock
(~5.5s total, of which ~0.6s is inference). For repeated calls keep the checkpoint resident and
point the runner at the HTTP server — same answers, ~8x faster end to end:
```bash
third_party/laya.cpp/build-cpu/bin/laya-cli --server --port 8080 \
--model models/convaiinnovations/laya --variant multilingual --cpu &
LAYA_URL=http://127.0.0.1:8080 ./usage-classify.sh "..."
```
`LAYA_URL` is read from the environment, `--url` overrides it, and `LAYA_URL` in `config/classify.py`
is the fallback. If the server cannot be reached the runner says so and falls back to the CLI.
Two caveats from the upstream model card, worth knowing before you trust the numbers: the
probabilities ship **uncalibrated** and systematically over-confident, and `noul` can under-report
a clear `true`. Fit your own thresholds on held-out data before treating a probability as a decision.
## Configuration
Model paths and inference parameters are hardcoded in `config/`:
- `config/model.py` — model paths (conv_frontend, encoder, decoder, tokenizer under `models/`)
- `config/asr.py` — inference params: `LANGUAGE`, `HOTWORDS`, `NUM_THREADS`, `PROVIDER`, `SAMPLE_RATE`, `FEATURE_DIM`, `MAX_TOTAL_LEN`, `MAX_NEW_TOKENS`
- `config/tts.py` — TTS model paths, `REFERENCE_AUDIO`, `REFERENCE_TEXT`, `OUTPUT_FILE`, `NUM_THREADS`, `PROVIDER`, `NUM_STEPS`
- `config/llm.py` — Granite-4.2 model path, `N_CTX`, `N_THREADS`, `N_GPU_LAYERS`, `MAX_TOKENS`, `TEMPERATURE`, `TOP_P`, `TOP_K`, `SYSTEM_PROMPT`, `CHAT_TEMPLATE`
- `config/embed.py` — nomic-embed-text-v1.5 model path, `N_CTX`, `N_THREADS`, `PREFIX_QUERY`, `PREFIX_DOCUMENT`, `DEFAULT_PREFIX`
- `config/classify.py` — laya.cpp binary path, `MODEL_ROOT`/`VARIANT`, `BACKEND`, `LAYA_URL`, `PRESETS`, `DEFAULT_PRESET`
`LANGUAGE` defaults to `""` (all languages / auto-detect). Passing `--language` on the CLI overrides it.
## Build laya.cpp (CPU-only)
`classify_runner.py` drives the native `laya-cli` binary, so there is no Python inference
dependency — `requirements.txt` is unchanged, and in particular PyTorch is not needed.
```bash
./build-laya.sh
```
The script needs no `sudo`: `cmake` and `ninja` come from pip wheels, `nlohmann/json` is unpacked
into `third_party/nlohmann-install`, and the only system requirement is `libicu-dev`
(`apt install libicu-dev` if it is missing). Re-running it skips finished steps. The manual
equivalent:
```bash
git clone --recursive --shallow-submodules https://github.com/lkarlslund/laya.cpp third_party/laya.cpp
.venv/bin/pip install "cmake==3.31.*" "ninja<2"
export PATH="$PWD/.venv/bin:$PATH"
cmake -S third_party/laya.cpp -B third_party/laya.cpp/build-cpu -G Ninja \
-DCMAKE_BUILD_TYPE=Release \
-DLAYA_CUDA=OFF -DLAYA_VULKAN=OFF -DLAYA_COREML=OFF -DBUILD_TESTING=OFF \
-DCMAKE_PREFIX_PATH="$PWD/third_party/nlohmann-install"
cmake --build third_party/laya.cpp/build-cpu --parallel "$(nproc)"
```
Three flags matter:
- **`-DLAYA_CUDA=OFF` is mandatory** — `LAYA_CUDA` defaults to `ON` and configuration fails looking for `nvcc`.
- **`cmake` must stay on 3.x** — laya.cpp's pinned ggml declares `cmake_minimum_required(VERSION 3.14...3.28)` and does not configure on CMake 4.
- **`--cpu` is always passed at runtime** — `laya-cli` defaults to the CUDA backend, so a CPU-only build fails at startup without it.
Upstream ships no CPU-only prebuilt binary (Linux releases are CUDA 12, CUDA 13 and Vulkan only),
which is why this is built from source. On an NVIDIA box you can swap `BACKEND` in
`config/classify.py` to `cuda` and download a prebuilt `laya` instead, keeping the build step out of
the loop entirely.
`models/convaiinnovations/laya/multilingual/` holds the checkpoint in the layout the CLI expects.
Note that `laya-cli` appends the variant to `--model` for every non-English variant, so the
multilingual checkpoint must be at `<MODEL_ROOT>/multilingual/`. `config/classify.py` sets
`MODEL_ROOT` to the parent directory. A separate copy of the same weights sits at
`models/convaiinnovations/laya-multilingual/`, which this layout does **not** address — the
standalone HuggingFace repo directory is for the Python reference implementation.
## Download Model (Qwen3-ASR 1.7B int8)
```bash
BASE="https://modelscope.cn/models/zengshuishui/Qwen3-ASR-onnx/resolve/master"
mkdir -p models/model_1.7B models/tokenizer
wget -O models/model_1.7B/conv_frontend.onnx "$BASE/model_1.7B/conv_frontend.onnx"
wget -O models/model_1.7B/encoder.int8.onnx "$BASE/model_1.7B/encoder.int8.onnx"
wget -O models/model_1.7B/decoder.int8.onnx "$BASE/model_1.7B/decoder.int8.onnx"
for f in vocab.json merges.txt tokenizer_config.json preprocessor_config.json config.json chat_template.json; do
wget -O "models/tokenizer/$f" "$BASE/tokenizer/$f"
done
```
## Download Model (ZipVoice TTS)
```bash
mkdir -p models/zipvoice
wget -qO- https://github.com/k2-fsa/sherpa-onnx/releases/download/tts-models/sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2 \
| tar xjf - -C models/zipvoice --strip-components=1
wget -O models/zipvoice/vocos_24khz.onnx \
https://github.com/k2-fsa/sherpa-onnx/releases/download/vocoder-models/vocos_24khz.onnx
```
## Download Model (Granite-4.2 8B Q4_K_M)
```bash
mkdir -p models/granite-4.2
wget -O models/granite-4.2/granite-4.2-8b-Q4_K_M.gguf \
"https://huggingface.co/ibm-granite/granite-4.2-8b-GGUF/resolve/main/granite-4.2-8b-Q4_K_M.gguf"
```
## Download Model (nomic-embed-text-v1.5)
```bash
mkdir -p models/nomic-embed-text-v1.5
wget -O models/nomic-embed-text-v1.5/nomic-embed-text-v1.5.Q4_K_M.gguf \
"https://huggingface.co/nomic-ai/nomic-embed-text-v1.5-GGUF/resolve/main/nomic-embed-text-v1.5.Q4_K_M.gguf"
```