KokoroCli 0.1.1
dotnet tool install --global KokoroCli --version 0.1.1
dotnet new tool-manifest
dotnet tool install --local KokoroCli --version 0.1.1
#tool dotnet:?package=KokoroCli&version=0.1.1
nuke :add-package KokoroCli --version 0.1.1
KokoroCli
Offline text-to-speech on the command line, powered by Kokoro-82M.
Everything runs locally. There is no API key, no network call after the first run, and no telemetry.
$ kokoro say "The quick brown fox jumps over the lazy dog."
$ kokoro say "Reading this to a file." --voice bm_george --output narration.wav
$ kokoro voices --language british-english
Install
dotnet tool install -g KokoroCli
Then kokoro is available anywhere. Update with dotnet tool update -g KokoroCli, remove with
dotnet tool uninstall -g KokoroCli.
Requirements
- .NET 10 runtime to install the tool (the SDK only if you want to build it yourself)
- An audio output device for playback. On Linux this goes through OpenAL — install
openal(Arch:sudo pacman -S openal, Debian/Ubuntu:sudo apt install libopenal1). Writing to a file with--outputneeds no audio device at all.
You do not need Python, PyTorch, or a system espeak-ng install. Grapheme-to-phoneme conversion
runs in managed code via MisakiSharp, a C# port of
Kokoro's own G2P engine.
Build and run
dotnet build
dotnet run --project src/KokoroCli -- say "Hello from Kokoro."
To install it as a real kokoro command:
dotnet pack -c Release
dotnet tool install -g --add-source src/KokoroCli/bin/Release KokoroCli
Commands
kokoro say <text>
Speaks the text through your speakers, or writes a WAV file when given --output.
| Option | Description |
|---|---|
-v, --voice |
Voice name or blend. Defaults to af_heart. |
-s, --speed |
Speech rate. 1.0 is normal; useful range is about 0.5–1.3. |
-o, --output |
Write a 24 kHz mono WAV here instead of playing it. |
-m, --model |
float32 (default), float16, zh_float32, zh_float16. |
--volume |
Playback volume, 0.0–1.0. Ignored with --output. |
-q, --best-quality |
Synthesize in one pass rather than streaming segment by segment. |
By default, long text is split into segments and playback starts as soon as the first segment is
ready. --best-quality synthesizes the whole input at once instead: it takes longer to start and
caps the input at 510 tokens, but avoids any artifacts at segment boundaries.
kokoro voices
Lists the bundled voices — 60+ speakers across American and British English, Japanese, Mandarin Chinese, Spanish, French, Hindi, Italian, and Brazilian Portuguese.
| Option | Description |
|---|---|
-l, --language |
e.g. american-english, japanese, brazilian-portuguese. |
-g, --gender |
male or female. |
-s, --search |
Substring match on the voice name. |
--names-only |
Bare names, one per line, for piping. |
Voice names encode their language and gender: af_heart is american-english female,
bm_george is british-english male.
Blending voices
Kokoro represents each voice as a style vector, so voices can be averaged into new ones. Pass
several to --voice, optionally weighted with ::
kokoro say "Two speakers, evenly mixed." --voice af_heart,af_bella
kokoro say "Mostly Heart, a little Bella." --voice af_heart:3,af_bella:1
Weights are normalized, so 3:1 and 0.75:0.25 mean the same thing. The blend inherits its
language from the first voice listed — worth remembering when mixing across languages, since that
choice decides how the text gets phonemized.
Downloaded assets
Nothing large ships in the package. On first use the CLI fetches, into
${XDG_CACHE_HOME:-~/.cache}/kokoro-cli/:
| Asset | Size | When |
|---|---|---|
| Voice style vectors (157) | ~63 MB | First command that touches a voice |
| ONNX model weights | ~325 MB | First synthesis |
Both come from the KokoroSharpBinaries releases and are cached permanently. Delete the directory to force a re-download. After that first fetch the tool is fully offline — no network, no API keys, no telemetry.
Keeping these out of the package matters: bundled, they made the NuGet artifact 248 MB against nuget.org's 250 MB ceiling. Downloading them on demand, plus dropping the iOS and Android ONNX binaries that a CLI can never load, brought it to 116 MB.
Running on the GPU
The default build uses the CPU execution provider, which is comfortably faster than real time for an
82M-parameter model. To use CUDA instead, swap the runtime package in
src/KokoroCli/KokoroCli.csproj:
<PackageReference Include="KokoroSharp.GPU.Linux" Version="0.8.4" />
That requires a matching CUDA and cuDNN install for ONNX Runtime 1.22. See KokoroSharp's GPU notes.
Audio playback on Linux and macOS
The CLI substitutes its own OpenAL player
(OpenAlAudioPlayer.cs) for KokoroSharp's built-in one on
non-Windows platforms, via that library's CrossPlatformHelper.CustomAudioPlayer hook.
KokoroSharp's LinuxAudioPlayer (which MacOSAudioPlayer also derives from) generates a fixed 256
OpenAL buffers, fills only as many as the segment needs, then queues all 256. OpenAL rejects the
batch with AL_INVALID_OPERATION because the unfilled buffers carry no data, so nothing is queued —
the source reports Stopped immediately, and playback is silent while the library still reports the
speech as spoken. Reproduced against OpenAL Soft 1.24: one filled buffer plays correctly, the
256-buffer batch returns IllegalCommand and plays nothing.
Credits and licensing
This CLI is a thin wrapper around work done by others:
- Kokoro-82M by hexgrad — Apache 2.0
- KokoroSharp — MIT, the C# inference engine and voices
- MisakiSharp — Apache 2.0, the G2P engine
| Product | Versions Compatible and additional computed target framework versions. |
|---|---|
| .NET | net10.0 is compatible. net10.0-android was computed. net10.0-browser was computed. net10.0-ios was computed. net10.0-maccatalyst was computed. net10.0-macos was computed. net10.0-tvos was computed. net10.0-windows was computed. |
This package has no dependencies.
| Version | Downloads | Last Updated |
|---|---|---|
| 0.1.1 | 132 | 8/17/2026 |