LiteRtLmSharp 1.2.0
Prefix Reserveddotnet add package LiteRtLmSharp --version 1.2.0
NuGet\Install-Package LiteRtLmSharp -Version 1.2.0
<PackageReference Include="LiteRtLmSharp" Version="1.2.0" />
<PackageVersion Include="LiteRtLmSharp" Version="1.2.0" />
<PackageReference Include="LiteRtLmSharp" />
paket add LiteRtLmSharp --version 1.2.0
#r "nuget: LiteRtLmSharp, 1.2.0"
#:package LiteRtLmSharp@1.2.0
#addin nuget:?package=LiteRtLmSharp&version=1.2.0
#tool nuget:?package=LiteRtLmSharp&version=1.2.0
LiteRtLmSharp
.NET 10 bindings for Google's LiteRT-LM — on-device LLM inference (e.g. Gemma) for any .NET app, including MAUI. No server, no cloud: the model runs locally on CPU or GPU.
Install
Install the managed package plus the native runtime package for your platform, always with the same version number:
<PackageReference Include="LiteRtLmSharp" Version="1.2.0" />
<PackageReference Include="LiteRtLmSharp.runtime.win-x64" Version="1.2.0" />
Optional integrations: LiteRtLmSharp.Extensions.AI (a Microsoft.Extensions.AI.IChatClient —
works with the Microsoft Agent Framework, Semantic Kernel and plain MEAI) and
LiteRtLmSharp.SemanticKernel (an IChatCompletionService).
First tokens
using LiteRtLmSharp;
using var engine = LiteRtEngine.Load(new LiteRtEngineOptions
{
ModelPath = "gemma-4-E2B-it.litertlm", // from huggingface.co/litert-community
Backend = LiteRtBackend.Cpu, // or .Gpu
MaxNumTokens = 4096, // total context window
});
using var chat = engine.CreateConversation();
await foreach (var chunk in chat.SendStreamingAsync("Tell me a joke"))
Console.Write(chunk.Text);
Features
- Chat: blocking, awaitable + cancellable, and streaming sends
- Function calling with constrained decoding for reliable JSON arguments
- Reasoning mode (Gemma "thinking"), surfaced separately from the answer
- Multimodal input: image and audio attachments
- Conversation restore & clone, token counting, speculative decoding, benchmarking
- AOT- and trim-compatible (source-generated P/Invoke)
Documentation
- Docs & guides: https://orihuelaconde.github.io/LiteRtLmSharp/
- API reference: https://orihuelaconde.github.io/LiteRtLmSharp/api/LiteRtLmSharp.html
- Repository & samples: https://github.com/OrihuelaConde/LiteRtLmSharp
- Changelog: https://github.com/OrihuelaConde/LiteRtLmSharp/blob/master/CHANGELOG.md
License and trademarks
Apache-2.0. This is an unofficial, community-maintained project — not affiliated with, sponsored, or endorsed by Google. LiteRT, LiteRT-LM and Gemma are trademarks of Google LLC. The native binaries are built from LiteRT-LM source (Apache-2.0) at pinned release tags.
| Product | Versions Compatible and additional computed target framework versions. |
|---|---|
| .NET | net10.0 is compatible. net10.0-android was computed. net10.0-browser was computed. net10.0-ios was computed. net10.0-maccatalyst was computed. net10.0-macos was computed. net10.0-tvos was computed. net10.0-windows was computed. |
-
net10.0
- No dependencies.
NuGet packages (2)
Showing the top 2 NuGet packages that depend on LiteRtLmSharp:
| Package | Downloads |
|---|---|
|
LiteRtLmSharp.Extensions.AI
Microsoft.Extensions.AI (IChatClient) integration for LiteRtLmSharp on-device LLM inference. Works with Microsoft Agent Framework, Semantic Kernel, and any IChatClient consumer. Companion to the LiteRtLmSharp package. Unofficial community bindings, not affiliated with or endorsed by Google. LiteRT is a trademark of Google LLC. |
|
|
LiteRtLmSharp.SemanticKernel
Microsoft Semantic Kernel connector (IChatCompletionService) for LiteRtLmSharp on-device LLM inference — a thin layer over the LiteRtLmSharp.Extensions.AI IChatClient. Companion to the LiteRtLmSharp package. Unofficial community bindings, not affiliated with or endorsed by Google. LiteRT is a trademark of Google LLC. |
GitHub repositories
This package is not used by any popular GitHub repositories.
| Version | Downloads | Last Updated |
|---|---|---|
| 1.2.0 | 187 | 9/5/2026 |
| 1.1.1 | 859 | 7/18/2026 |
| 1.1.0 | 148 | 7/17/2026 |
| 1.0.0 | 256 | 7/2/2026 |
| 0.1.0-preview.3 | 81 | 6/20/2026 |
| 0.1.0-preview.2 | 91 | 6/17/2026 |
| 0.1.0-preview.1 | 75 | 6/11/2026 |
1.2.0: native binaries move to Google's official LiteRT-LM v0.16.0 C API prebuilts (one monolithic library per platform: linux-x64 now needs the system Vulkan loader (libvulkan1), win-x64 no longer needs the VC++ Redistributable, Android keeps GPU sampling with the embedded sampler, iOS ships the official CLiteRTLM.xcframework). Constrained decoding works on linux-x64 (guard removed); text LoRA is applied on LoRA-enabled bundles; new EnableYnnpack engine option; NoRepeatNgramSize / SuppressTokens on the MEAI and Semantic Kernel options. Also carries the v0.15.0 cycle: per-send decoding controls (penalties, no-repeat-ngram, suppress tokens, thinking budget, regex/JSON-schema constraints), the LlGuidance constraint provider, the KV overflow guard recalibrated to the new prefill planning, and Clone() now requiring an advanced conversation. The managed assembly requires the same-version runtime packages (LiteRT-LM v0.16.0). Full notes: https://github.com/OrihuelaConde/LiteRtLmSharp/blob/master/CHANGELOG.md