I got tired of cloud TTS for a home project. Every option was a subscription, a per-character API bill, or a voice that sounded like a 2012 satnav. So I set up Kokoro on my mini PC — a proper neural text-to-speech model running locally, and it sounds genuinely good.

Kokoro is an open TTS model (Apache 2.0, 82M parameters). I run the full-precision ONNX export: about 340MB of model and voice data, served through ONNX Runtime on CPU. No GPU anywhere in the picture. On an N100 it renders speech faster than real time without getting warm.

The setup is a small Python server running as a systemd user service. It takes text over HTTPS and streams back PCM audio frames. The voice is af_heart, American English, normal speed. The bundle ships a whole menu of voices; I tried a few and stuck with the first one that didn’t annoy me after a hundred sentences.

Why self-host? Three reasons. Latency: no round trip to someone’s API. Cost: zero marginal cost per sentence, which matters when a chatty home assistant speaks all day. Privacy: my house’s chatter isn’t anyone’s training data.

The tradeoff is you need an always-on box with some CPU headroom. The N100 mini PC I already had does it with room to spare — the service is capped at 1.5GB of RAM and sits well under that in practice.

One design rule: degrade gracefully. If the TTS box ever goes down, the client falls back to the device’s built-in TTS. It sounds worse, but the system still talks. A silent smart home is a broken smart home.

As an Amazon Associate, Headless Diaries earns from qualifying purchases made through links on this page.