KokoroCli 0.1.1

dotnet tool install --global KokoroCli --version 0.1.1
                    
This package contains a .NET tool you can call from the shell/command line.
dotnet new tool-manifest
                    
if you are setting up this repo
dotnet tool install --local KokoroCli --version 0.1.1
                    
This package contains a .NET tool you can call from the shell/command line.
#tool dotnet:?package=KokoroCli&version=0.1.1
                    
nuke :add-package KokoroCli --version 0.1.1
                    

KokoroCli

Offline text-to-speech on the command line, powered by Kokoro-82M.

Everything runs locally. There is no API key, no network call after the first run, and no telemetry.

$ kokoro say "The quick brown fox jumps over the lazy dog."
$ kokoro say "Reading this to a file." --voice bm_george --output narration.wav
$ kokoro voices --language british-english

Install

dotnet tool install -g KokoroCli

Then kokoro is available anywhere. Update with dotnet tool update -g KokoroCli, remove with dotnet tool uninstall -g KokoroCli.

Requirements

  • .NET 10 runtime to install the tool (the SDK only if you want to build it yourself)
  • An audio output device for playback. On Linux this goes through OpenAL — install openal (Arch: sudo pacman -S openal, Debian/Ubuntu: sudo apt install libopenal1). Writing to a file with --output needs no audio device at all.

You do not need Python, PyTorch, or a system espeak-ng install. Grapheme-to-phoneme conversion runs in managed code via MisakiSharp, a C# port of Kokoro's own G2P engine.

Build and run

dotnet build
dotnet run --project src/KokoroCli -- say "Hello from Kokoro."

To install it as a real kokoro command:

dotnet pack -c Release
dotnet tool install -g --add-source src/KokoroCli/bin/Release KokoroCli

Commands

kokoro say <text>

Speaks the text through your speakers, or writes a WAV file when given --output.

Option Description
-v, --voice Voice name or blend. Defaults to af_heart.
-s, --speed Speech rate. 1.0 is normal; useful range is about 0.5–1.3.
-o, --output Write a 24 kHz mono WAV here instead of playing it.
-m, --model float32 (default), float16, zh_float32, zh_float16.
--volume Playback volume, 0.0–1.0. Ignored with --output.
-q, --best-quality Synthesize in one pass rather than streaming segment by segment.

By default, long text is split into segments and playback starts as soon as the first segment is ready. --best-quality synthesizes the whole input at once instead: it takes longer to start and caps the input at 510 tokens, but avoids any artifacts at segment boundaries.

kokoro voices

Lists the bundled voices — 60+ speakers across American and British English, Japanese, Mandarin Chinese, Spanish, French, Hindi, Italian, and Brazilian Portuguese.

Option Description
-l, --language e.g. american-english, japanese, brazilian-portuguese.
-g, --gender male or female.
-s, --search Substring match on the voice name.
--names-only Bare names, one per line, for piping.

Voice names encode their language and gender: af_heart is american-english female, bm_george is british-english male.

Blending voices

Kokoro represents each voice as a style vector, so voices can be averaged into new ones. Pass several to --voice, optionally weighted with ::

kokoro say "Two speakers, evenly mixed."   --voice af_heart,af_bella
kokoro say "Mostly Heart, a little Bella." --voice af_heart:3,af_bella:1

Weights are normalized, so 3:1 and 0.75:0.25 mean the same thing. The blend inherits its language from the first voice listed — worth remembering when mixing across languages, since that choice decides how the text gets phonemized.

Downloaded assets

Nothing large ships in the package. On first use the CLI fetches, into ${XDG_CACHE_HOME:-~/.cache}/kokoro-cli/:

Asset Size When
Voice style vectors (157) ~63 MB First command that touches a voice
ONNX model weights ~325 MB First synthesis

Both come from the KokoroSharpBinaries releases and are cached permanently. Delete the directory to force a re-download. After that first fetch the tool is fully offline — no network, no API keys, no telemetry.

Keeping these out of the package matters: bundled, they made the NuGet artifact 248 MB against nuget.org's 250 MB ceiling. Downloading them on demand, plus dropping the iOS and Android ONNX binaries that a CLI can never load, brought it to 116 MB.

Running on the GPU

The default build uses the CPU execution provider, which is comfortably faster than real time for an 82M-parameter model. To use CUDA instead, swap the runtime package in src/KokoroCli/KokoroCli.csproj:

<PackageReference Include="KokoroSharp.GPU.Linux" Version="0.8.4" />

That requires a matching CUDA and cuDNN install for ONNX Runtime 1.22. See KokoroSharp's GPU notes.

Audio playback on Linux and macOS

The CLI substitutes its own OpenAL player (OpenAlAudioPlayer.cs) for KokoroSharp's built-in one on non-Windows platforms, via that library's CrossPlatformHelper.CustomAudioPlayer hook.

KokoroSharp's LinuxAudioPlayer (which MacOSAudioPlayer also derives from) generates a fixed 256 OpenAL buffers, fills only as many as the segment needs, then queues all 256. OpenAL rejects the batch with AL_INVALID_OPERATION because the unfilled buffers carry no data, so nothing is queued — the source reports Stopped immediately, and playback is silent while the library still reports the speech as spoken. Reproduced against OpenAL Soft 1.24: one filled buffer plays correctly, the 256-buffer batch returns IllegalCommand and plays nothing.

Credits and licensing

This CLI is a thin wrapper around work done by others:

Product Compatible and additional computed target framework versions.
.NET net10.0 is compatible.  net10.0-android was computed.  net10.0-browser was computed.  net10.0-ios was computed.  net10.0-maccatalyst was computed.  net10.0-macos was computed.  net10.0-tvos was computed.  net10.0-windows was computed. 
Compatible target framework(s)
Included target framework(s) (in package)
Learn more about Target Frameworks and .NET Standard.

This package has no dependencies.

Version Downloads Last Updated
0.1.1 132 8/17/2026