GnOuGo.AI.Local
0.20.0
dotnet add package GnOuGo.AI.Local --version 0.20.0
NuGet\Install-Package GnOuGo.AI.Local -Version 0.20.0
<PackageReference Include="GnOuGo.AI.Local" Version="0.20.0" />
<PackageVersion Include="GnOuGo.AI.Local" Version="0.20.0" />
<PackageReference Include="GnOuGo.AI.Local" />
paket add GnOuGo.AI.Local --version 0.20.0
#r "nuget: GnOuGo.AI.Local, 0.20.0"
#:package GnOuGo.AI.Local@0.20.0
#addin nuget:?package=GnOuGo.AI.Local&version=0.20.0
#tool nuget:?package=GnOuGo.AI.Local&version=0.20.0
GnOuGo.AI.Local
Embedded, publishable local LLM runtime for GnOuGo. It uses LLamaSharp/llama.cpp in-process and never requires an HTTP model server, API key, Python installation, Docker container, or child process.
The initial catalog contains qwen3:0.6b, the Apache-2.0 Qwen3 0.6B Q4_0 GGUF. The 428,970,080-byte model is downloaded separately, from an immutable revision, and accepted only after SHA-256 verification.
- Revision:
a41486f827d17edd055fe6b3b0ba3f8d427c0519 - SHA-256:
da2572f16c06133561ce56accaa822216f2391ef4d37fba427801cd6736417d4 - Source:
https://huggingface.co/ggml-org/Qwen3-0.6B-GGUF/resolve/a41486f827d17edd055fe6b3b0ba3f8d427c0519/Qwen3-0.6B-Q4_0.gguf
LocalModelManager downloads only catalog URLs, resumes .partial files, streams
progress, supports cancellation, checks path containment, verifies exact size and
SHA-256, and atomically promotes a completed file. Removal unloads the runtime
before deleting the model. Models are shared host assets under the workspace
.GnOuGo/models/ directory; tenant-specific defaults remain in Agent MCP storage.
Build and test
dotnet build src/GnOuGo.AI.Local/GnOuGo.AI.Local.csproj
dotnet test tests/GnOuGo.AI.Local.Tests/GnOuGo.AI.Local.Tests.csproj
The standard tests are offline. Set GNOUOGO_LOCAL_MODEL_SMOKE=1 and GNOUOGO_LOCAL_MODEL_PATH to run the real-model inference smoke test.
Runtime defaults
- Context: 8192 tokens
- CPU threads: automatic
- Windows/Linux: CPU
- macOS ARM64: Metal through the LLamaSharp CPU backend
- Default output limit: 1024 tokens
The embedded GGUF chat template is used for prompting. JSON responses and
Qwen/Hermes tool-call envelopes are mapped into the common Json and ToolCalls
response fields. Thinking is disabled by default; when requested, reasoning content
is stripped and is never logged or returned.
CUDA and Vulkan acceleration are intentionally deferred.
| Product | Versions Compatible and additional computed target framework versions. |
|---|---|
| .NET | net10.0 is compatible. net10.0-android was computed. net10.0-browser was computed. net10.0-ios was computed. net10.0-maccatalyst was computed. net10.0-macos was computed. net10.0-tvos was computed. net10.0-windows was computed. |
-
net10.0
- GnOuGo.AI.Core (>= 0.20.0)
- GnOuGo.Workspace (>= 0.20.0)
- LLamaSharp (>= 0.27.0)
- LLamaSharp.Backend.Cpu (>= 0.27.0)
- Microsoft.Extensions.Logging.Abstractions (>= 10.0.12)
- Microsoft.Extensions.Options (>= 10.0.12)
NuGet packages
This package is not used by any NuGet packages.
GitHub repositories
This package is not used by any popular GitHub repositories.
| Version | Downloads | Last Updated |
|---|---|---|
| 0.20.0 | 43 | 10/2/2026 |
| 0.20.0-dev.847 | 44 | 9/30/2026 |
| 0.20.0-dev.837 | 44 | 9/30/2026 |
| 0.20.0-dev.786 | 48 | 9/25/2026 |
| 0.20.0-dev.784 | 47 | 9/25/2026 |
| 0.19.4-dev.756 | 46 | 9/24/2026 |
| 0.19.3 | 85 | 9/24/2026 |
| 0.19.3-dev.730 | 54 | 9/23/2026 |
| 0.19.3-dev.728 | 62 | 9/22/2026 |
| 0.19.2 | 82 | 9/22/2026 |
| 0.19.2-dev.715 | 53 | 9/22/2026 |
| 0.19.1 | 118 | 8/16/2026 |
| 0.19.0 | 114 | 8/16/2026 |