ZeroInference.Providers.OnnxRuntime
1.5.0
dotnet add package ZeroInference.Providers.OnnxRuntime --version 1.5.0
NuGet\Install-Package ZeroInference.Providers.OnnxRuntime -Version 1.5.0
<PackageReference Include="ZeroInference.Providers.OnnxRuntime" Version="1.5.0" />
<PackageVersion Include="ZeroInference.Providers.OnnxRuntime" Version="1.5.0" />
<PackageReference Include="ZeroInference.Providers.OnnxRuntime" />
paket add ZeroInference.Providers.OnnxRuntime --version 1.5.0
#r "nuget: ZeroInference.Providers.OnnxRuntime, 1.5.0"
#:package ZeroInference.Providers.OnnxRuntime@1.5.0
#addin nuget:?package=ZeroInference.Providers.OnnxRuntime&version=1.5.0
#tool nuget:?package=ZeroInference.Providers.OnnxRuntime&version=1.5.0
ZeroInference
ZeroInference is a pure C# ONNX deep learning inference engine and model runtime for .NET with zero external dependencies. It eliminates bulky native C++ runtime binaries (no ONNX Runtime native DLLs, no OpenVINO, no Python dependencies), reading and evaluating .onnx and .zeromodel neural graphs directly in memory with SIMD vectorization and int8 quantization.
๐ Key Capabilities
- Transformer & LLM Operators:
- Rotary Position Embedding (
RotaryEmbedding): High-performance RoPE with precomputed trigonometric tables and in-place complex rotation for modern LLMs (LLaMA, Mistral, Qwen). - FlashAttention-2 Kernel (
FlashAttentionKernel): CPU cache-tiled online Softmax algorithm computing scaled dot-product attention in $O(1)$ memory without materializing the quadratic $N \times N$ attention matrix. - Multi-Head Attention (
MultiHeadAttention): Scaled dot-product attention with projection weights and configurable heads.
- Rotary Position Embedding (
- Polymorphic Micro-Kernel (
IInferenceSession): Unified inference contract allowing seamless zero-downtime switching between Pure C# execution and native hardware accelerators (DirectML / ONNX Runtime). - Zero Dependency Pure C# Runtime: No native shared libraries (
onnxruntime.dll,libonnxruntime.so) required. Runs anywhere .NET runs. - Local LLM Streaming (
LocalLlmStreamClient): High-throughput SSE token streaming client for Ollama, OpenAI, and local edge LLM endpoints. - Direct ONNX Model Parser: Stack-allocated Protocol Buffers wire reader (
ProtobufWireReader) parsing ONNX binary graphs directly into executable compute graphs. - Compact
.zeromodelSerialization: Fast binary serialization format with pre-compiled layer topologies and optimized weights layout. - Supported Deep Learning Layers:
- Conv2D (Direct & im2col GEMM convolution)
- Dense / Gemm (Fully-connected linear layers)
- BatchNormalization & LayerNorm
- Activations (ReLU, LeakyReLU, Sigmoid, Tanh, Softmax)
- Pooling (MaxPool2D, AveragePool2D, GlobalAveragePool)
- Reshape, Flatten, Concat, Slice
- Quantization & Vision Post-Processing:
- Int8 Quantizer: Symmetric and asymmetric integer quantization for edge devices.
- Non-Maximum Suppression (NMS): Fast SIMD bounding box filtering with configurable IoU and score thresholds.
- Hardware Agnostic: Executes over
ZeroTensorCPU SIMD orZeroComputeDirect3D 11 GPU compute contexts.
๐ฆ Installation
Install via the .NET CLI:
dotnet add package ZeroInference.Core
๐ Quick Start
1. Parsing and Executing an ONNX Model
using ZeroInference.Core.Engine;
using ZeroInference.Core.Format;
using ZeroTensor.Core;
// 1. Load and parse .onnx model file
using var stream = File.OpenRead("models/classifier.onnx");
var graph = OnnxModelParser.Parse(stream);
// 2. Instantiate inference engine
var engine = new InferenceEngine(graph);
// 3. Prepare input tensor and infer
var input = Tensor.RandomUniform(1, 3, 224, 224);
var outputs = engine.Forward(input);
Console.WriteLine($"Inference Output Shape: [{outputs[0].Shape[0]}, {outputs[0].Shape[1]}]");
2. Fast Object Detection NMS Post-Processing
using ZeroInference.Core.Vision;
var candidateBoxes = new List<BoundingBox>
{
new BoundingBox(10, 10, 50, 50, score: 0.92f, classId: 1),
new BoundingBox(12, 11, 48, 52, score: 0.78f, classId: 1), // Overlapping duplicate
new BoundingBox(100, 120, 60, 40, score: 0.85f, classId: 2)
};
// Filter duplicates with IoU threshold = 0.45
var filtered = NonMaximumSuppression.Filter(candidateBoxes, iouThreshold: 0.45f, scoreThreshold: 0.5f);
Console.WriteLine($"Remaining boxes after NMS: {filtered.Count}");
๐ Benchmark & Performance
Tested on MobileNet-V2 / ResNet-18 (Release x64):
| Architecture | Model Size | Load Time | CPU SIMD Latency | External DLLs |
|---|---|---|---|---|
| MobileNet-V2 | $14.2 \text{ MB}$ | $18.4 \text{ ms}$ | $12.1 \text{ ms}$ | 0 (Pure C#) |
| ResNet-18 (FP32) | $45.1 \text{ MB}$ | $42.0 \text{ ms}$ | $28.5 \text{ ms}$ | 0 (Pure C#) |
| ResNet-18 (Int8) | $11.3 \text{ MB}$ | $12.5 \text{ ms}$ | $9.4 \text{ ms}$ | 0 (Pure C#) |
๐ Release History
| Version | Release Date | Key Milestones & Highlights |
|---|---|---|
v1.1.0 |
2026-09-16 | Polymorphic Micro-Kernel & Local LLM Streaming:<br/>โข Introduced IInferenceSession unified execution contract decoupling high-level apps from backends.<br/>โข Added ZeroInference.Providers.OnnxRuntime provider bridging Microsoft.ML.OnnxRuntime with pure C# pipeline.<br/>โข Added LocalLlmStreamClient supporting real-time SSE streaming for Ollama & OpenAI-compatible endpoints.<br/>โข Verified with 23 unit tests across Core & Provider test suites. |
v1.0.0 |
2026-09-09 | Initial Sovereign Release:<br/>โข Pure C# ONNX protobuf wire reader & layer fusion pipeline.<br/>โข Conv2D, Dense, BatchNorm, LayerNorm, activations, and pooling.<br/>โข Int8 quantization engine & Non-Maximum Suppression (NMS). |
๐ License
MIT License ยฉ 2026 Phong Vรต. Part of the ZeroPlatform project.
| Product | Versions Compatible and additional computed target framework versions. |
|---|---|
| .NET | net5.0 was computed. net5.0-windows was computed. net6.0 was computed. net6.0-android was computed. net6.0-ios was computed. net6.0-maccatalyst was computed. net6.0-macos was computed. net6.0-tvos was computed. net6.0-windows was computed. net7.0 was computed. net7.0-android was computed. net7.0-ios was computed. net7.0-maccatalyst was computed. net7.0-macos was computed. net7.0-tvos was computed. net7.0-windows was computed. net8.0 is compatible. net8.0-android was computed. net8.0-browser was computed. net8.0-ios was computed. net8.0-maccatalyst was computed. net8.0-macos was computed. net8.0-tvos was computed. net8.0-windows was computed. net9.0 was computed. net9.0-android was computed. net9.0-browser was computed. net9.0-ios was computed. net9.0-maccatalyst was computed. net9.0-macos was computed. net9.0-tvos was computed. net9.0-windows was computed. net10.0 was computed. net10.0-android was computed. net10.0-browser was computed. net10.0-ios was computed. net10.0-maccatalyst was computed. net10.0-macos was computed. net10.0-tvos was computed. net10.0-windows was computed. |
| .NET Core | netcoreapp2.0 was computed. netcoreapp2.1 was computed. netcoreapp2.2 was computed. netcoreapp3.0 was computed. netcoreapp3.1 was computed. |
| .NET Standard | netstandard2.0 is compatible. netstandard2.1 was computed. |
| .NET Framework | net461 was computed. net462 was computed. net463 was computed. net47 was computed. net471 was computed. net472 was computed. net48 was computed. net481 was computed. |
| MonoAndroid | monoandroid was computed. |
| MonoMac | monomac was computed. |
| MonoTouch | monotouch was computed. |
| Tizen | tizen40 was computed. tizen60 was computed. |
| Xamarin.iOS | xamarinios was computed. |
| Xamarin.Mac | xamarinmac was computed. |
| Xamarin.TVOS | xamarintvos was computed. |
| Xamarin.WatchOS | xamarinwatchos was computed. |
-
.NETStandard 2.0
- Microsoft.ML.OnnxRuntime (>= 1.19.2)
- System.Buffers (>= 4.5.1)
- System.Memory (>= 4.5.5)
- ZeroInference.Core (>= 1.5.0)
- ZeroTensor.Core (>= 1.5.0)
-
net8.0
- Microsoft.ML.OnnxRuntime (>= 1.19.2)
- ZeroInference.Core (>= 1.5.0)
- ZeroTensor.Core (>= 1.5.0)
NuGet packages
This package is not used by any NuGet packages.
GitHub repositories
This package is not used by any popular GitHub repositories.