HomeProjectsSpeech models and runtime
Packages

Speech models and runtime

Kokoro and its 64 voices, Moonshine, and ONNX Runtime for the CPU or an NVIDIA card, as Debian packages that the broker brings in with it.

Published in the archive
amber-modelsamberlin-runtime
Packaging BSD-3-Clause; the weights Apache-2.0 and MIT; ONNX Runtime MIT
CUDAKokoroMoonshineONNX Runtime
Amberlin AI
amber-modelsamberlin-runtime
Speech models and runtime: Kokoro and its 64 voices, Moonshine, and ONNX Runtime for the CPU or an NVIDIA card, as Debian packages that the broker brings in with it.

The broker needs weights to speak and listen, and a runtime to run them on. Both come from the Amber Linux archive like any other package, so a machine has everything it needs to talk after one apt install.

The models

Package Holds Download
amber-models-tts Kokoro v1.0 and 64 voices, for speaking 325 MB
amber-models-stt Moonshine tiny, English, for listening 104 MB
amber-models both of the above —

They are separate packages because they change at different rates. Kokoro comes from hexgrad’s Kokoro-82M under Apache-2.0, Moonshine from Useful Sensors under MIT. Eight more voices that Kokoro-FastAPI carries are left out, because nothing shows they may be passed on, and one of them imitates a real person’s voice. Every file is checked against a pinned checksum when the package is built.

No language model ships here. A language model is several gigabytes, and a download that can resume and check each file suits it better than an apt package, so the one you pick is downloaded by the broker, from a signed catalogue.

The runtime

Ubuntu 24.04, which Mint 22 is built on, ships no ONNX Runtime, so amberlin-runtime provides one, with onnxruntime-genai beside it, in a private folder that only the broker looks in. amberlin-runtime-cuda is the same for an NVIDIA card. It bundles no NVIDIA libraries: CUDA 13 and cuDNN come from NVIDIA’s own apt repository, the driver from Mint’s Driver Manager, and apt will not install the package until they are there. An RTX 5070 Ti works on the stock build.

ONNX Runtime itself is Microsoft’s official build, passed on byte for byte. onnxruntime-genai, the library that runs the language model’s generation loop, is built from source, because no released version reads the Ministral 3 models yet.

For an NVIDIA card, install the CUDA runtime instead of the CPU one:

sudo apt install amberlin-runtime-cuda

Telemetry

Microsoft’s official ONNX Runtime builds upload usage telemetry by default. onnxruntime-genai is built here with its telemetry compiled out, and a check refuses a build in which the telemetry client turns up again. The prebuilt ONNX Runtime cannot be rebuilt the same way, so the broker switches its telemetry off before it loads. Both libraries are checked for new telemetry every month.

Install

Once the Amber Linux archive is set up (it takes three commands, on the packages page):

bash
sudo apt install amber-models amberlin-runtime