Speech models and runtime
Kokoro and its 64 voices, Moonshine, and ONNX Runtime for the CPU or an NVIDIA card, as Debian packages that the broker brings in with it.
- Status
- Published in the archive
- Package
- amber-modelsamberlin-runtime
- Licence
- Packaging BSD-3-Clause; the weights Apache-2.0 and MIT; ONNX Runtime MIT
- Built with
- CUDAKokoroMoonshineONNX Runtime
- Part of
- Amberlin AI
- Source
- amber-modelsamberlin-runtime

The broker needs weights to speak and listen, and a runtime to run them on.
Both come from the Amber Linux archive like any other package, so a machine has
everything it needs to talk after one apt install.
The models
| Package | Holds | Download |
|---|---|---|
amber-models-tts |
Kokoro v1.0 and 64 voices, for speaking | 325 MB |
amber-models-stt |
Moonshine tiny, English, for listening | 104 MB |
amber-models |
both of the above | — |
They are separate packages because they change at different rates. Kokoro comes from hexgrad’s Kokoro-82M under Apache-2.0, Moonshine from Useful Sensors under MIT. Eight more voices that Kokoro-FastAPI carries are left out, because nothing shows they may be passed on, and one of them imitates a real person’s voice. Every file is checked against a pinned checksum when the package is built.
No language model ships here. A language model is several gigabytes, and a download that can resume and check each file suits it better than an apt package, so the one you pick is downloaded by the broker, from a signed catalogue.
The runtime
Ubuntu 24.04, which Mint 22 is built on, ships no ONNX Runtime, so
amberlin-runtime provides one, with onnxruntime-genai beside it, in a private
folder that only the broker looks in. amberlin-runtime-cuda is the same for an
NVIDIA card. It bundles no NVIDIA libraries: CUDA 13 and cuDNN come from
NVIDIA’s own apt repository, the driver from Mint’s Driver Manager, and apt will
not install the package until they are there. An RTX 5070 Ti works on the stock
build.
ONNX Runtime itself is Microsoft’s official build, passed on byte for byte. onnxruntime-genai, the library that runs the language model’s generation loop, is built from source, because no released version reads the Ministral 3 models yet.
For an NVIDIA card, install the CUDA runtime instead of the CPU one:
sudo apt install amberlin-runtime-cuda
Telemetry
Microsoft’s official ONNX Runtime builds upload usage telemetry by default. onnxruntime-genai is built here with its telemetry compiled out, and a check refuses a build in which the telemetry client turns up again. The prebuilt ONNX Runtime cannot be rebuilt the same way, so the broker switches its telemetry off before it loads. Both libraries are checked for new telemetry every month.
Install
Once the Amber Linux archive is set up (it takes three commands, on the packages page):
sudo apt install amber-models amberlin-runtime