Voxa.Services.SpeechToSpeech
0.7.2-alpha
dotnet add package Voxa.Services.SpeechToSpeech --version 0.7.2-alpha
NuGet\Install-Package Voxa.Services.SpeechToSpeech -Version 0.7.2-alpha
<PackageReference Include="Voxa.Services.SpeechToSpeech" Version="0.7.2-alpha" />
<PackageVersion Include="Voxa.Services.SpeechToSpeech" Version="0.7.2-alpha" />
<PackageReference Include="Voxa.Services.SpeechToSpeech" />
paket add Voxa.Services.SpeechToSpeech --version 0.7.2-alpha
#r "nuget: Voxa.Services.SpeechToSpeech, 0.7.2-alpha"
#:package Voxa.Services.SpeechToSpeech@0.7.2-alpha
#addin nuget:?package=Voxa.Services.SpeechToSpeech&version=0.7.2-alpha&prerelease
#tool nuget:?package=Voxa.Services.SpeechToSpeech&version=0.7.2-alpha&prerelease
Voxa.Services.SpeechToSpeech
A full-duplex speech-to-speech composite for the Voxa pipeline (VRT-005) — the local-model peer of the cloud
realtime composites (OpenAIRealtimeProcessor / AzureVoiceLiveProcessor).
SpeechToSpeechProcessor slots in exactly where those do — one processor that is the whole voice loop
(model-owned VAD/turn-taking, full-duplex, native interruption) — but it is driven by an in-process (or sidecar)
ISpeechToSpeechSession instead of a wire transport. It emits the same frame vocabulary the cloud
composites do, so Studio's Talk view, the diagnostics hub, and any sink can't tell a third-party realtime API
from a local model:
| Direction | Frames |
|---|---|
| In | AudioRawFrame (user audio → the session), InterruptionFrame (barge-in → session.CancelAsync) |
| Out | AudioRawFrame (agent audio), LlmTextChunkFrame, BotStartedSpeakingFrame/BotStoppedSpeakingFrame, UserStartedSpeakingFrame/UserStoppedSpeakingFrame, InterruptionFrame, upstream ErrorFrame |
The seam
ISpeechToSpeechSession (in Voxa.Speech.Abstractions) models speech-core's FullDuplexSpeechInterface:
AppendUserAudioAsync, RespondAsync (a stream of SpeechToSpeechChunk carrying agent audio + text + the
model's own speaking-edge events), SetVoiceAsync, SetSystemPromptAsync, ResetSessionAsync, CancelAsync.
It is shaped deliberately parallel to what the cloud realtime processors do internally, so the composite is a
true third member of that family rather than a parallel universe.
What ships vs. what's deferred
- Ships: the seam + the composite processor, with frame parity tested against a fake session (no model).
- Deferred (VRT-005 WS3 — spike-gated): a concrete
ISpeechToSpeechSession. A real local speech-to-speech model (PersonaPlex / Moshi-class) is GB-scale and often GPU-only, so it lands behind a quality/feasibility spike, likely on a GPU or out-of-process host (reusing the VLS-002 / sidecar patterns). A cloud S2S provider could validate the seam first. Function-calling (ToolCallRequestFrame/ToolCallResultFrame) is a documented optional extension once a model that supports it lands. - Deferred (small follow-up): config-driven selection (a
Voxa:Mode = SpeechToSpeechkey, or a registered provider name, soUseDefaults()short-circuits to the composite). Today the composite is reached by constructing it directly /MapVoxaVoice(...).UseProcessor(...), exactly as a custom processor is.
| Product | Versions Compatible and additional computed target framework versions. |
|---|---|
| .NET | net10.0 is compatible. net10.0-android was computed. net10.0-browser was computed. net10.0-ios was computed. net10.0-maccatalyst was computed. net10.0-macos was computed. net10.0-tvos was computed. net10.0-windows was computed. |
-
net10.0
- Microsoft.Extensions.Logging.Abstractions (>= 10.0.7)
- Voxa.Core (>= 0.7.2-alpha)
- Voxa.Speech.Abstractions (>= 0.7.2-alpha)
NuGet packages
This package is not used by any NuGet packages.
GitHub repositories
This package is not used by any popular GitHub repositories.
| Version | Downloads | Last Updated |
|---|---|---|
| 0.7.2-alpha | 81 | 7/10/2026 |
| 0.7.1-alpha | 63 | 7/10/2026 |
| 0.7.0-alpha | 74 | 7/10/2026 |
| 0.6.0-alpha | 81 | 6/22/2026 |