Graphene.AIOffice.VoiceAgent.Win
1.26.9.3
dotnet add package Graphene.AIOffice.VoiceAgent.Win --version 1.26.9.3
NuGet\Install-Package Graphene.AIOffice.VoiceAgent.Win -Version 1.26.9.3
<PackageReference Include="Graphene.AIOffice.VoiceAgent.Win" Version="1.26.9.3" />
<PackageVersion Include="Graphene.AIOffice.VoiceAgent.Win" Version="1.26.9.3" />
<PackageReference Include="Graphene.AIOffice.VoiceAgent.Win" />
paket add Graphene.AIOffice.VoiceAgent.Win --version 1.26.9.3
#r "nuget: Graphene.AIOffice.VoiceAgent.Win, 1.26.9.3"
#:package Graphene.AIOffice.VoiceAgent.Win@1.26.9.3
#addin nuget:?package=Graphene.AIOffice.VoiceAgent.Win&version=1.26.9.3
#tool nuget:?package=Graphene.AIOffice.VoiceAgent.Win&version=1.26.9.3
AIOffice.VoiceAgent.Win
Windows-only executable providing offline speech recognition (WinRT OneCore dictation) and neural text-to-speech (KokoroSharp via ONNX runtime, with Windows SAPI fallback). Communicates with the host process via JSON Lines over stdin/stdout.
Architecture
VoiceAgentWin is a thin subclass of VoiceAgentBase (in AIOffice.VoiceAgent, referenced
as a project): the JSON-Lines protocol loop, logging, the unified speak logic and the render
path are INHERITED — this project contributes only the Windows pieces:
WinRtRecognizer(IAgentRecognizer) — WinRT offline dictation on a dedicated STA threadWindowsAudioSink(IAudioSink) — continuous NAudio buffered output for streaming TTS- the SAPI fallback synthesizer
There is no duplicated protocol/log code between the two agents (the old local Log.cs and
the duplicated main loop were removed by the refactor).
Protocol
stdin (host → agent)
| Command | Description |
|---|---|
{"cmd":"start"} or {"cmd":"start","lang":"it"} |
Begin speech recognition for the given language (default: system language) |
{"cmd":"stop"} |
Stop recognition and exit |
{"cmd":"speak","text":"...","lang":"..."} |
Speak text, pause recognition, resume when done |
The optional streaming flag keeps recognition paused between consecutive calls and routes
the text to the continuous audio sink (NAudio) — first sound as soon as the first sentence
is synthesized, no gaps between sentences:
{"cmd":"speak","text":"first","lang":"it","streaming":true}— speak (streamed), stay paused{"cmd":"speak","text":"last","lang":"it","streaming":false}— speak, resume recognition{"cmd":"speak","text":"","lang":"it"}— end-of-turn: drains the TTS tail, resumes recognition
stdout (agent → host)
| Type | Description |
|---|---|
{"type":"ready","tts":"kokoro"} |
Agent initialized (tts field shows engine) |
{"type":"transcript","text":"..."} |
User speech recognized |
{"type":"done"} |
Previous speak command finished |
{"type":"error","text":"..."} |
Error occurred |
Language support
The lang field (two-letter ISO code) controls both speech recognition and TTS voice:
- Recognition: the WinRT
SpeechRecognizeris initialized with the matching language tag (e.g.it→it-IT). When no language is provided, the system default is used. - TTS: the corresponding Kokoro voice is selected for text-to-speech. When a
speakcommand omitslang, it falls back to the language set by the previousstartcommand, then to the system default.
Language fallback chain: speak lang → recognition lang (from start) → system default
Two-letter ISO code → WinRT recognition tag / Kokoro voice mapping:
| Code | Language | Recognition | TTS voice |
|---|---|---|---|
it |
Italian | it-IT |
if_sara (female) |
en |
English | en-US |
af_heart (highest quality) |
fr |
French | fr-FR |
ff_siwis |
es |
Spanish | es-ES |
ef_dora |
de |
German | de-DE |
SAPI fallback |
ja |
Japanese | ja-JP |
jf_alpha |
zh |
Chinese | zh-CN |
zf_xiaobei |
hi |
Hindi | — | hf_alpha |
pt |
Portuguese | pt-BR |
pf_dora |
ru |
Russian | ru-RU |
SAPI fallback |
| other | Unsupported | System default | Windows SAPI |
Note: Recognition requires the corresponding Windows language pack to be installed (
Settings → Time & Language → Language & region → Add a language). TheSupportedGrammarLanguageslog entry shows which languages are available on the current system.
TTS engine selection
- Kokoro neural TTS (~320MB ONNX model) — loaded on startup, auto-downloaded from GitHub Releases on first run
- Windows SAPI — fallback when Kokoro model is unavailable (offline, no cache) or the language is unsupported
The ready message includes a tts field indicating which engine loaded.
TTS method selection
| Flag | Method | Behaviour |
|---|---|---|
--tts-method=fast (default) |
SpeakFast |
Internal punctuation-based segmentation. First segment plays immediately while rest is inferred in background. Best for responsiveness. |
--tts-method=full |
Speak |
No segmentation. Full text inferred before playback starts. Best for quality on very short phrases. |
Debugging & logging
| Flag | Behaviour |
|---|---|
--debug |
Calls Debugger.Launch() to attach a Visual Studio JIT debugger |
In Debug builds, step-level logging is auto-enabled. Each run writes a log file to {AppBase}/logs/{ProcessId}.txt with the format:
[elapsed_seconds] [calling_method] message
[0,36] [MoveNext] DispatcherQueue STA thread created
[2,95] [MoveNext] Kokoro TTS loaded successfully
[2,96] [MoveNext] Compile result: Success
The log covers startup (encoding, privacy policy, microphone permission), recognizer setup (compilation, timeouts, language), every RecognizeAsync result, TTS events, and shutdown.
Streaming TTS buffering (client-side)
The client (Voice.cs) accumulates LLM tokens and flushes only when sentence-ending punctuation (., !, ?) is found with text after it. No timeout flush — waiting for punctuation produces natural phrasing.
"Ciao! Oggi è una bella giornata. Non ti sembra? Cosa fai"
→ "Ciao!" (primo flush)
→ "Oggi è una bella giornata. Non ti sembra?" (secondo flush)
→ "Cosa fai" (flush finale)
Integration
Built automatically by AIOffice's BuildVoiceAgentPlugin MSBuild target and copied to the output directory along with required voices/ and espeak/ folders.
Build
dotnet build -p:Configuration=Debug
The compiled executable targets net10.0-windows10.0.19041.0 and requires the Windows 10 SDK (19041+).
| Product | Versions Compatible and additional computed target framework versions. |
|---|---|
| .NET | net10.0-windows10.0.19041 is compatible. |
-
net10.0-windows10.0.19041
- Graphene.AIOffice.VoiceAgent (>= 1.26.9.3)
- KokoroSharp (>= 0.8.4)
- System.Speech (>= 9.0.0)
NuGet packages
This package is not used by any NuGet packages.
GitHub repositories
This package is not used by any popular GitHub repositories.
| Version | Downloads | Last Updated |
|---|---|---|
| 1.26.9.3 | 74 | 9/3/2026 |