LiteRtLmSharp.SemanticKernel
1.2.0
Prefix Reserved
dotnet add package LiteRtLmSharp.SemanticKernel --version 1.2.0
NuGet\Install-Package LiteRtLmSharp.SemanticKernel -Version 1.2.0
<PackageReference Include="LiteRtLmSharp.SemanticKernel" Version="1.2.0" />
<PackageVersion Include="LiteRtLmSharp.SemanticKernel" Version="1.2.0" />
<PackageReference Include="LiteRtLmSharp.SemanticKernel" />
paket add LiteRtLmSharp.SemanticKernel --version 1.2.0
#r "nuget: LiteRtLmSharp.SemanticKernel, 1.2.0"
#:package LiteRtLmSharp.SemanticKernel@1.2.0
#addin nuget:?package=LiteRtLmSharp.SemanticKernel&version=1.2.0
#tool nuget:?package=LiteRtLmSharp.SemanticKernel&version=1.2.0
LiteRtLmSharp
.NET 10 bindings for Google's LiteRT-LM — on-device LLM inference (e.g. Gemma) for any .NET app, including MAUI. No server, no cloud: the model runs locally on CPU or GPU.
Install
Install the managed package plus the native runtime package for your platform, always with the same version number:
<PackageReference Include="LiteRtLmSharp" Version="1.2.0" />
<PackageReference Include="LiteRtLmSharp.runtime.win-x64" Version="1.2.0" />
Optional integrations: LiteRtLmSharp.Extensions.AI (a Microsoft.Extensions.AI.IChatClient —
works with the Microsoft Agent Framework, Semantic Kernel and plain MEAI) and
LiteRtLmSharp.SemanticKernel (an IChatCompletionService).
First tokens
using LiteRtLmSharp;
using var engine = LiteRtEngine.Load(new LiteRtEngineOptions
{
ModelPath = "gemma-4-E2B-it.litertlm", // from huggingface.co/litert-community
Backend = LiteRtBackend.Cpu, // or .Gpu
MaxNumTokens = 4096, // total context window
});
using var chat = engine.CreateConversation();
await foreach (var chunk in chat.SendStreamingAsync("Tell me a joke"))
Console.Write(chunk.Text);
Features
- Chat: blocking, awaitable + cancellable, and streaming sends
- Function calling with constrained decoding for reliable JSON arguments
- Reasoning mode (Gemma "thinking"), surfaced separately from the answer
- Multimodal input: image and audio attachments
- Conversation restore & clone, token counting, speculative decoding, benchmarking
- AOT- and trim-compatible (source-generated P/Invoke)
Documentation
- Docs & guides: https://orihuelaconde.github.io/LiteRtLmSharp/
- API reference: https://orihuelaconde.github.io/LiteRtLmSharp/api/LiteRtLmSharp.html
- Repository & samples: https://github.com/OrihuelaConde/LiteRtLmSharp
- Changelog: https://github.com/OrihuelaConde/LiteRtLmSharp/blob/master/CHANGELOG.md
License and trademarks
Apache-2.0. This is an unofficial, community-maintained project — not affiliated with, sponsored, or endorsed by Google. LiteRT, LiteRT-LM and Gemma are trademarks of Google LLC. The native binaries are built from LiteRT-LM source (Apache-2.0) at pinned release tags.
| Product | Versions Compatible and additional computed target framework versions. |
|---|---|
| .NET | net10.0 is compatible. net10.0-android was computed. net10.0-browser was computed. net10.0-ios was computed. net10.0-maccatalyst was computed. net10.0-macos was computed. net10.0-tvos was computed. net10.0-windows was computed. |
-
net10.0
- LiteRtLmSharp (>= 1.2.0)
- LiteRtLmSharp.Extensions.AI (>= 1.2.0)
- Microsoft.Extensions.AI (>= 10.7.0)
- Microsoft.SemanticKernel.Abstractions (>= 1.77.0)
NuGet packages
This package is not used by any NuGet packages.
GitHub repositories
This package is not used by any popular GitHub repositories.
1.2.0: native binaries move to Google's official LiteRT-LM v0.16.0 C API prebuilts (one monolithic library per platform: linux-x64 now needs the system Vulkan loader (libvulkan1), win-x64 no longer needs the VC++ Redistributable, Android keeps GPU sampling with the embedded sampler, iOS ships the official CLiteRTLM.xcframework). Constrained decoding works on linux-x64 (guard removed); text LoRA is applied on LoRA-enabled bundles; new EnableYnnpack engine option; NoRepeatNgramSize / SuppressTokens on the MEAI and Semantic Kernel options. Also carries the v0.15.0 cycle: per-send decoding controls (penalties, no-repeat-ngram, suppress tokens, thinking budget, regex/JSON-schema constraints), the LlGuidance constraint provider, the KV overflow guard recalibrated to the new prefill planning, and Clone() now requiring an advanced conversation. The managed assembly requires the same-version runtime packages (LiteRT-LM v0.16.0). Full notes: https://github.com/OrihuelaConde/LiteRtLmSharp/blob/master/CHANGELOG.md