LiteRtLmSharp.runtime.android-arm64
1.2.0
Prefix Reserved
dotnet add package LiteRtLmSharp.runtime.android-arm64 --version 1.2.0
NuGet\Install-Package LiteRtLmSharp.runtime.android-arm64 -Version 1.2.0
<PackageReference Include="LiteRtLmSharp.runtime.android-arm64" Version="1.2.0" />
<PackageVersion Include="LiteRtLmSharp.runtime.android-arm64" Version="1.2.0" />
<PackageReference Include="LiteRtLmSharp.runtime.android-arm64" />
paket add LiteRtLmSharp.runtime.android-arm64 --version 1.2.0
#r "nuget: LiteRtLmSharp.runtime.android-arm64, 1.2.0"
#:package LiteRtLmSharp.runtime.android-arm64@1.2.0
#addin nuget:?package=LiteRtLmSharp.runtime.android-arm64&version=1.2.0
#tool nuget:?package=LiteRtLmSharp.runtime.android-arm64&version=1.2.0
LiteRtLmSharp
.NET 10 bindings for Google's LiteRT-LM — on-device LLM inference (e.g. Gemma) for any .NET app, including MAUI. No server, no cloud: the model runs locally on CPU or GPU.
Install
Install the managed package plus the native runtime package for your platform, always with the same version number:
<PackageReference Include="LiteRtLmSharp" Version="1.2.0" />
<PackageReference Include="LiteRtLmSharp.runtime.win-x64" Version="1.2.0" />
Optional integrations: LiteRtLmSharp.Extensions.AI (a Microsoft.Extensions.AI.IChatClient —
works with the Microsoft Agent Framework, Semantic Kernel and plain MEAI) and
LiteRtLmSharp.SemanticKernel (an IChatCompletionService).
First tokens
using LiteRtLmSharp;
using var engine = LiteRtEngine.Load(new LiteRtEngineOptions
{
ModelPath = "gemma-4-E2B-it.litertlm", // from huggingface.co/litert-community
Backend = LiteRtBackend.Cpu, // or .Gpu
MaxNumTokens = 4096, // total context window
});
using var chat = engine.CreateConversation();
await foreach (var chunk in chat.SendStreamingAsync("Tell me a joke"))
Console.Write(chunk.Text);
Features
- Chat: blocking, awaitable + cancellable, and streaming sends
- Function calling with constrained decoding for reliable JSON arguments
- Reasoning mode (Gemma "thinking"), surfaced separately from the answer
- Multimodal input: image and audio attachments
- Conversation restore & clone, token counting, speculative decoding, benchmarking
- AOT- and trim-compatible (source-generated P/Invoke)
Documentation
- Docs & guides: https://orihuelaconde.github.io/LiteRtLmSharp/
- API reference: https://orihuelaconde.github.io/LiteRtLmSharp/api/LiteRtLmSharp.html
- Repository & samples: https://github.com/OrihuelaConde/LiteRtLmSharp
- Changelog: https://github.com/OrihuelaConde/LiteRtLmSharp/blob/master/CHANGELOG.md
License and trademarks
Apache-2.0. This is an unofficial, community-maintained project — not affiliated with, sponsored, or endorsed by Google. LiteRT, LiteRT-LM and Gemma are trademarks of Google LLC. The native binaries are built from LiteRT-LM source (Apache-2.0) at pinned release tags.
Learn more about Target Frameworks and .NET Standard.
This package has no dependencies.
NuGet packages
This package is not used by any NuGet packages.
GitHub repositories
This package is not used by any popular GitHub repositories.
| Version | Downloads | Last Updated |
|---|---|---|
| 1.2.0 | 192 | 9/5/2026 |
| 1.1.1 | 352 | 7/18/2026 |
| 1.1.0 | 188 | 7/17/2026 |
| 1.0.0 | 332 | 7/2/2026 |
| 0.1.0-preview.3 | 184 | 6/20/2026 |
| 0.1.0-preview.2 | 199 | 6/17/2026 |
| 0.1.0-preview.1 | 185 | 6/11/2026 |
1.2.0: native binaries move to Google's official LiteRT-LM v0.16.0 C API prebuilts (one monolithic library per platform: linux-x64 now needs the system Vulkan loader (libvulkan1), win-x64 no longer needs the VC++ Redistributable, Android keeps GPU sampling with the embedded sampler, iOS ships the official CLiteRTLM.xcframework). Constrained decoding works on linux-x64 (guard removed); text LoRA is applied on LoRA-enabled bundles; new EnableYnnpack engine option; NoRepeatNgramSize / SuppressTokens on the MEAI and Semantic Kernel options. Also carries the v0.15.0 cycle: per-send decoding controls (penalties, no-repeat-ngram, suppress tokens, thinking budget, regex/JSON-schema constraints), the LlGuidance constraint provider, the KV overflow guard recalibrated to the new prefill planning, and Clone() now requiring an advanced conversation. The managed assembly requires the same-version runtime packages (LiteRT-LM v0.16.0). Full notes: https://github.com/OrihuelaConde/LiteRtLmSharp/blob/master/CHANGELOG.md