amberlin-broker
The service at the centre of Amberlin: it holds the speech and language models once, for every application that wants an assistant, and decides who answers.
- Status
- Working; experimental until 1.0
- Package
- amberlin-brokeramber-modelsamberlin-runtime
- Licence
- Business Source License, GPL-3.0 from 1 May 2033
- Built with
- D-BusOdinONNX Runtimeonnxruntime-genai
- Part of
- Amberlin AI
- Source
- amberlin-broker

The broker is a small per-user service on the session bus,
org.amberlinux.Amberlin.Broker, and it holds the models so the applications
do not have to. Ask it a question and the answer streams back; ask it to speak
and audio comes back; hand it a recording and you get the text. Three programs
ask it questions today, Amberlin, Katai
in kat800 and elektron, and Amberlin Settings
manages its models. Any other program can call it the same way. The broker has no persona: each program brings its own. Large data never crosses the bus: every stream comes back through a
file descriptor.
Listening and speaking
Speech runs on your machine, always. Moonshine turns speech into text. Kokoro reads the answer aloud in any of 64 voices, and starts speaking on the first sentence. Both use the graphics card when there is one and the CPU otherwise.
Before Kokoro reads a sentence in an English voice, the broker writes it out the way a person would say it: “$5.30” becomes “five dollars and thirty cents”, “Dr.” becomes “Doctor”, and a phone number is read in groups. Sentences are split by the Unicode rules, so “e.g.” and “3.14” do not end one, and a sentence too long for the model is cut at its commas rather than in the middle of a phrase. That text handling is a port to Odin of the open-source Kokoro-FastAPI server’s own.
Thinking
The language model runs locally by default, from a signed catalogue:
| Model | Runs on | Download | Measured |
|---|---|---|---|
| Ministral 3 3B | CPU | 2.3 GB | 33 tokens/s on an i7-13700K |
| Ministral 3 3B | NVIDIA, 4 GB free | 2.0 GB | 323 tokens/s on an RTX 5070 Ti |
| Ministral 3 8B | NVIDIA, 7 GB free | 5.6 GB | 149 tokens/s on an RTX 5070 Ti |
A model is too big for an apt package, so the broker downloads the one you pick in Amberlin Settings, file by file, each checked against a checksum in the signed catalogue, and resumes an interrupted download. The model can call amberlin-tools to look things up, and every finished turn is kept in Ambrosia, so a conversation survives a restart.
A remote model, if you choose one
A server on your own network, or a hosted service such as DeepSeek or OpenAI, can answer instead. That is an opt-in, and when it is chosen every question leaves the machine: the settings say so, and the broker’s status tells any application that asks. An API key lives in the desktop keyring and is refused over plain HTTP to another machine. Speech stays local either way.
Not phoning home
The models run on Microsoft’s ONNX Runtime, whose official build uploads usage telemetry by default. onnxruntime-genai, the library that runs the language model, is built from source with telemetry compiled out, and the build fails if the telemetry client turns up in it again. The broker also switches telemetry off in the runtime before either library loads. No conversation text is written to the system journal.
Kokoro, its voices and Moonshine arrive as Debian packages, and so does ONNX
Runtime; the language model is the only thing downloaded at run time. The
broker needs a runtime package: amberlin-runtime for the CPU, or
amberlin-runtime-cuda for an NVIDIA card, which also needs CUDA 13 from
NVIDIA’s own apt repository. The speech models and runtime
page has the details.
Install
Once the Amber Linux archive is set up (it takes three commands, on the packages page):
sudo apt install amberlin-broker amber-models amberlin-runtime