Glacier.Inference
1.2.0
dotnet add package Glacier.Inference --version 1.2.0
NuGet\Install-Package Glacier.Inference -Version 1.2.0
<PackageReference Include="Glacier.Inference" Version="1.2.0" />
<PackageVersion Include="Glacier.Inference" Version="1.2.0" />
<PackageReference Include="Glacier.Inference" />
paket add Glacier.Inference --version 1.2.0
#r "nuget: Glacier.Inference, 1.2.0"
#:package Glacier.Inference@1.2.0
#addin nuget:?package=Glacier.Inference&version=1.2.0
#tool nuget:?package=Glacier.Inference&version=1.2.0
Glacier.Inference
<p align="center"> <img src="docs/images/glacier_inference_universal_banner.jpg" alt="Glacier.Inference Universal Matrix" width="100%" /> </p>
Universal Pure C# .NET 10 LLM Engine & Server Runtime
========================================================================================================
GLACIER.INFERENCE: PURE C# .NET 10 UNIVERSAL LLM RUNTIME (ZERO PYTHON / ZERO C++ DLLs)
========================================================================================================
Cold Startup (Model Load) : < 15 ms (vs 2,500 ms in Python vLLM / Ollama β 160x Faster)
VRAM / Driver Overhead : 0 MB runtime overhead (Direct Driver P/Invoke, Zero CUDA Toolkit)
GPU Bare-Metal SASS : 28.1 tok/s (NVIDIA RTX 4060) | 35+ tok/s (DirectX 12 Wave32)
CPU SIMD Throughput : Multi-threaded AVX-512 & AVX2 Batched GEMM (Zero Allocations)
Universal Matrix : Meta Llama 3 | DeepSeek-V3/R1 | Microsoft Phi-4 | Mistral | Alibaba Qwen 2.5
========================================================================================================
Universal Model Architecture Matrix
Glacier.Inference natively executes all major open foundation model families with 100% mathematical fidelity in pure C# .NET 10:
| Model Family | Core Architectural Topology | RoPE & Positional Encoding | Attention / FFN Kernels | Tested & Verified Models |
|---|---|---|---|---|
| Meta Llama 3 / 3.1 / 3.2 / 3.3 | Standard GQA with zero attention bias | Precomputed rope_freqs.weight scaling table across 128k context |
GQA with SIMD SwiGLU | Meta-Llama-3.1-8B-Instruct |
| DeepSeek-V2 / V3 / R1 | Multi-Head Latent Attention (MLA) + MoE | Decoupled RoPE + Multi-scale YaRN ($mscaleBase$, $mscaleAllDim$) | 64 routed + 1 shared expert | DeepSeek-Coder-V2-Lite-Instruct |
| Microsoft Phi-4 & Phi-3 | Fused Attention & Fused Gate/Up FFN | Fast NeOX RoPE ($base = 250,000$) | Fused attn_qkv GEMV + Fused ffn_up SwiGLU |
phi-4-Q4_K_M (15B params) |
| Mistral & Devstral | Standard GQA with ultra-wide context | Extreme $10^9$ RoPE Base frequency | Dense GQA + SwiGLU FFN | Devstral-Small-2505 (24B params) |
| Alibaba Qwen 2 & 2.5 | Dense & MoE GQA with per-head QKV bias | $10^6$ RoPE Base with optional QK-Norm | GQA with fused/unfused bias | Qwen2.5-7B-Instruct, Qwen3-30B-A3B |
3-Line Quickstart
using Glacier.Inference.Engine;
using var session = new InferenceSession("Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf");
await foreach (var token in session.GenerateStreamAsync("What is the capital of France?"))
{
Console.Write(token);
}
Features
- Pure C# Native SASS Engine (NVIDIA): Direct driver P/Invoke (
nvcuda.dll) streaming raw machine code directly to NVIDIA SMs, completely bypassing the CUDA Toolkit runtime (cudart64.dll,cublas64.dll). SupportsQ4_K,Q6_K,Q8_0,FP16, andFP32. - Pure C# AMD ROCm / HIP Driver Engine (AMD): Zero-dependency direct driver P/Invoke (
amdhip64.dllon Windows,libamdhip64.soon Linux). Interacts directly with the AMD HSA / KFD kernel driver (/dev/kfd), dispatching native AMDGPU ISA via user-space AQL queues and hardware doorbells with sub-microsecond launch latency. - Vulkan Hardware Tensor Engine (AMD / Intel / NVIDIA): Pure C# P/Invoke to
vulkan-1.dllandlibvulkan.so.1withVK_KHR_cooperative_matrixsupport. Direct access to hardware WMMA tensor units with in-register fused dequantization; on Linux, leverages Mesa's Valve ACO compiler for direct microarchitecture machine code generation. - Universal Multi-Architecture Fatbinary: Modular
.cuhkernel architecture (common.cuh,gemv.cuh,gemm_batch.cuh,attention.cuh,ops.cuh) compiled into a single embeddedkernels.cubinwith dedicated binary slices forsm_75(Turing),sm_80(A100),sm_86(Ampere),sm_89(Ada Lovelace),sm_90(Hopper), andcompute_75(Blackwell PTX). 100% verified withSTACK: 0(zero DRAM spills) across all targets and full 32-token GEMM tiling for robust prompt prefill. - Direct3D 12 Compute Engine (AMD / Intel): Native HLSL Wave32 compute shaders for AMD Radeon 680M / 890M (RDNA 2 / RDNA 3.5) and Intel Arc GPUs. Features register-tiled Batched GEMM (
Q4_K,Q6_K,Q8_0), 128-bit vectorization, 36-byte alignedByteAddressBufferrouting, and zero-allocation persistent buffers delivering up to 35+ tok/s generation and up to 216 tok/s prompt prefill in pure C# .NET 10 with 0 external C++ binaries. - Speculative Decoding Engine (1.5xβ3x Throughput Acceleration): Seamless assisted generation via
PromptLookupDraftProvider(sub-microsecond n-gram matching with 0 extra VRAM) andModelDraftProvider, coordinated with GPU batched verification (VerifyBatch) evaluating all candidates in a single pass over weights. - Fused GPU-Side LM Head & Argmax Sampling: 512-thread warp-shuffle reduction kernel (
argmax_kernel) finding the greedy token across 152K logits in ~3 Β΅s directly in VRAM, eliminating 608 KB DtoH transfers down to just 4 bytes across PCIe. - Multi-Device Hardware Discovery: Automatic detection of physical GPUs, dedicated VRAM, unified system RAM, and display connections via DXGI, HIP, Vulkan, and Linux
/dev/kfd. - Safe Driver Engine Architecture: Cooperates with Windows DWM via Direct3D 12 Compute / DirectML on display adapters (AMD Radeon 680M / 890M) to prevent TDR timeouts, while running Native SASS / HIP on compute GPUs.
- Mixture-of-Experts (MoE) Architecture Runtime: High-throughput MoE routing engine supporting up to 128 experts per layer with Top-$K$ gating, softmax re-normalization, shared experts (DeepSeek / ERNIE), per-head Query/Key RMSNorm (
attn_q_norm,attn_k_norm), and zero-copy 3D tensor slicing ([cols x rows x num_experts]). - DeepSeek Multi-Head Latent Attention (MLA) & Decoupled YaRN RoPE: Low-rank latent KV compression ($d_c = 512$), decoupled 192-dim QK projections, 128-dim Value projections, adjacent $(2i, 2i+1)$ RoPE pairing, and dual YaRN scaling ($mscaleBase$ and $mscaleAllDim$), delivering coherent, high-speed inference for DeepSeek-Coder-V2 and DeepSeek-V4 MoE architectures.
- OpenAI Attention Sinks & Streaming Contexts: Native support for attention sink logit absorption (
attn_sinks.weight) preventing attention overflow and perplexity degradation in long-context models (gpt-oss-20b) and streaming generation. - Next-Gen Quantization Kernels: Vectorized AVX-512 and AVX2 hardware FMA dot-products for
MXFP4(Type 39, Microscaling FP4 with E2M1 LUT and power-of-2 scaling),Q3_K(Type 11, 3.4375 bpw),Q4_K(Type 12),Q5_K(Type 13, 5.5 bpw),Q6_K(Type 14),Q8_0(Type 8),Q4_0,F16, andF32. - Zero-Copy GGUF Weight Mapping: Uses
MemoryMappedFileto instantly map multi-gigabyte models into address space in sub-30ms cold time without heap allocations. - Full Architecture Support:
- Mixture-of-Experts (MoE): Alibaba Qwen3-30B-A3B (128 experts, top-8 active), Gemma-4-26B-A4B (128 experts, top-8 active), OpenAI
gpt-oss-20b(32 experts, top-4 active), Baidu ERNIE-4.5-21B-A3B (64 experts, top-6 active, shared experts), Liquid AILFM2-8B-A1B, DeepSeek-Coder-V2-Lite. - Qwen Family: Qwen2, Qwen2.5 (7B, 14B), Qwen2.5-Coder Enterprise Q8_0, Qwen3 & Qwen3.8 (4B Thinking, 8B, 30B MoE) with native decoupled head dimensions (
key_length = 128), non-square attention projections, and per-head QK-normalization (rms_norm_heads_kernel). - LLaMA & Mistral Family: Meta LLaMA 3 / 3.1 / 3.2 (with zero-bias QKV handling and adaptive LLaMA-3 header decoding).
- DeepSeek Family: DeepSeek-R1-Distill-Qwen, DeepSeek-R1-Distill-Llama, DeepSeek-Coder-V2-Lite.
- Xiaomi MiMo Family: MiMo-7B-RL.
- Grouped Query Attention (GQA), Rotary Positional Embeddings (RoPE), Optional QKV Bias, SwiGLU FFN, Per-Head RMSNorm, and Attention Sinks.
- Mixture-of-Experts (MoE): Alibaba Qwen3-30B-A3B (128 experts, top-8 active), Gemma-4-26B-A4B (128 experts, top-8 active), OpenAI
- Embedded BPE Tokenizer: Reads vocabularies and BPE merge tables directly from GGUF metadata with ChatML template support.
- Dual-Protocol HTTP Server: Drop-in compatible with Ollama (
/api/generate,/api/chat,/api/tags) and OpenAI (/v1/chat/completions,/v1/models). - Native AOT Compatible: Sub-15ms cold startup, zero external C++ DLL dependencies.
Installation & Standalone Binaries
Option 1: Download Standalone CLI Binaries (Zero Dependencies)
Prebuilt, self-contained single-file binaries are available directly on the GitHub Releases page:
| Platform | Architecture | Archive | Features / Capabilities |
|---|---|---|---|
| Windows | x64 |
glacier-v1.1.13-win-x64.zip |
Native NVIDIA SASS (nvcuda.dll), AMD ROCm / HIP (amdhip64.dll), Vulkan CoopMat, D3D12 Wave32, AVX-512 |
| Windows | ARM64 |
glacier-v1.1.13-win-arm64.zip |
Qualcomm Snapdragon X Elite, Direct3D 12 Compute, ARM NEON SIMD |
| Linux | x64 |
glacier-v1.1.13-linux-x64.tar.gz |
Native CUDA Driver (libcuda.so), AMD ROCm / HIP (libamdhip64.so), Vulkan CoopMat, AVX-512 SIMD |
| macOS | ARM64 |
glacier-v1.1.13-osx-arm64.tar.gz |
Apple Silicon (M1/M2/M3/M4) CPU runtime & ARM NEON SIMD (Metal GPU on Roadmap) |
Simply extract the archive and run glacier from any terminal:
# Windows
.\glacier.exe devices
# Linux / macOS
chmod +x glacier
./glacier devices
Option 2: Install as a .NET Global Tool
dotnet tool install --global Glacier.Inference.Cli
glacier --help
Option 3: Add NuGet Package to Your C# Project
dotnet add package Glacier.Inference
CLI Usage & Complete Help System
Glacier provides a full hierarchical --help system. You can view top-level help or deep contextual help for any individual command:
# Global overview of all commands & global flags
glacier --help
# Or: glacier -h, glacier help
# Detailed help, argument descriptions, defaults, and examples for any subcommand:
glacier devices --help # (or: glacier help devices)
glacier config --help # (or: glacier help config)
glacier inspect --help # (or: glacier help inspect)
glacier bench --help # (or: glacier help bench)
glacier run --help # (or: glacier help run)
glacier serve --help # (or: glacier help serve)
CLI Command Quick Reference
# 1. Enumerate detected hardware accelerators and safe driver engines
glacier devices
# 2. Configure persistent device and engine preferences (~/.glacier/settings.json)
glacier config --device nvidia-rtx-4060 --engine baremetal
glacier config --device amd-890m --engine directml
glacier config --reset
# 3. Inspect any GGUF model architecture, metadata & tensors
glacier inspect "path/to/model.gguf"
# 4. Interactive streaming chat REPL or single-shot prompt
glacier run "path/to/model.gguf"
glacier run "path/to/model.gguf" "Explain quicksort in C#" --temp 0.2
glacier run "path/to/model.gguf" --ctx 4096 --device nvidia-rtx-4060 --engine baremetal
# 5. Benchmark generation throughput (with optional side-by-side Ollama comparison)
glacier bench "path/to/model.gguf" --tokens 32
glacier bench "path/to/model.gguf" --compare-ollama "http://127.0.0.1:11434" --compare-model "qwen2.5:7b"
# 6. Launch Ollama & OpenAI compatible HTTP inference microservice
glacier serve "path/to/model.gguf" --port 11434
glacier serve "path/to/model.gguf" --port 8080 --host 127.0.0.1 --kv-precision fp8
Global Options (available across bench, run, serve):
| Option | Description | Default |
|---|---|---|
--device <id\|name> |
Target GPU/CPU (e.g. nvidia-rtx-4060, amd-890m, cpu, 0) |
Auto-detect |
--engine <baremetal\|directml\|cpu\|auto> |
Execution engine (baremetal for NVIDIA SASS, directml for AMD/Intel) |
Auto-safe |
--kv-precision <auto\|fp16\|fp8\|fp32> |
KV-cache precision (auto adapts to sequence length) |
auto |
-c, --ctx, --context-length <len> |
Maximum context sequence length | 2048 |
π Like-for-Like Benchmarks: Glacier.Inference vs. Local Ollama
Detailed Fleet & Multi-Architecture Matrix:
For comprehensive side-by-side tables grouped by physical machine (RTX 3060 desktop, RTX 4060 laptop, Radeon 890M 64GB, Radeon 680M, CPU), model architecture (Dense vs. MoE frontier), and driver engine (SASS vs. HIP vs. Vulkan CoopMat vs. D3D12), see BENCHMARKS.md.
Benchmark 1: NVIDIA GeForce RTX 4060 Laptop GPU (Ada Lovelace, 8 GB GDDR6)
Hardware: Mobile Workstation / AMD Ryzen AI 9 HX 370 + NVIDIA GeForce RTX 4060 Laptop GPU (128-bit GDDR6, 256 GB/s physical memory bandwidth). Model:
DeepSeek-R1-Distill-Qwen-7B-Q4_K_M.gguf(4.68 GB).
| Metric | Glacier.Inference (Pure C#) | Ollama (Go + C++ CUDA) | Head-to-Head Comparison |
|---|---|---|---|
| Physical Hardware | NVIDIA RTX 4060 Laptop GPU | NVIDIA RTX 4060 Laptop GPU | 100% Identical Hardware |
| Software Runtime | π© Pure C# .NET 10 (Native AOT) | Go + C++ CUDA / llama.cpp | π© Pure C# vs. Compiled C++ |
| External Dependencies | π© 0 Native C++ DLLs (direct nvcuda.dll) |
CUDA Toolkit, cuBLAS, libllama | π© Zero native toolchain bloat |
| Deployment Footprint | π© ~15 MB Single Executable | ~4.5 GB CUDA Toolkit + Go runtime | π© 300x Lighter Distribution |
| Engine Cold-Start | π© 1.50 s (In-process, sub-50ms engine) | Daemon / Service spin-up required | π© Instant in-process execution |
| Turnaround (25 tok) | π© 0.82 seconds (41.9 tok/s) | ~2.50 seconds (32.3 tok/s) | π© >3.0x FASTER (Sub-second) |
| Total Wall Time (50 tok) | π© 1.41 seconds | 3.64 seconds | π© >2.5x FASTER (158% speedup) |
| Sustained Single Token Rate | 41.92 tokens/sec | 43.20 tokens/sec | Within 3% of compiled C++ cuBLAS |
| Speculative Decoding Rate | π© 72.5 β 104.8 tokens/sec | N/A (Standard serial decode) | π© 1.7x β 2.5x FASTER than Ollama |
| Prompt Eval Rate (Batch) | 121.2 tokens/sec | 125.4 tokens/sec | Near Parity (96.6%) |
| Sampling Latency | π© ~3.2 ΞΌs (Pure GPU argmax) | ~800 ΞΌs (Host PCIe DtoH transfer) | π© 250x Faster Sampling Reduction |
| KV-Cache Memory | π© Adaptive FP16 / FP8 (118β235 MB) | Fixed FP16 (~470 MB) | π© 50% to 75% Less VRAM |
| Memory Bus Saturation | 208.1 GB/s (81.3% of peak bus) | 216.8 GB/s (84.8% of peak bus) | Saturating 128-bit hardware limits |
Benchmark 2: NVIDIA GeForce RTX 3060 Desktop GPU (Ampere, 12 GB GDDR6)
Hardware: AMD Ryzen 5 5500 + NVIDIA GeForce RTX 3060 Desktop GPU (192-bit GDDR6, 360 GB/s physical memory bandwidth, 28 SMs, 3584 CUDA Cores). Model:
DeepSeek-R1-Distill-Qwen-7B-Q4_K_M.gguf(4.68 GB).
| Metric | Glacier.Inference (Pure C#) | Ollama (Go + C++ CUDA) | Head-to-Head Comparison |
|---|---|---|---|
| Physical Hardware | NVIDIA RTX 3060 12GB (Desktop) | NVIDIA RTX 3060 12GB (Desktop) | 100% Identical Hardware |
| Software Runtime | π© Pure C# .NET 10 (Native AOT) | Go + C++ CUDA / llama.cpp | π© Pure C# vs. Compiled C++ |
| External Dependencies | π© 0 Native C++ DLLs (direct nvcuda.dll) |
CUDA Toolkit, cuBLAS, libllama | π© Zero native toolchain bloat |
| Deployment Footprint | π© ~15 MB Single Executable | ~4.5 GB CUDA Toolkit + Go runtime | π© 300x Lighter Distribution |
| Cold Start Latency | π© 1.93 s (Zero-copy VRAM upload) | Daemon / Service spin-up required | π© Instant in-process execution |
| Generation Rate (Serial) | 42.8 β 43.0 tokens/sec | 64.5 β 69.3 tokens/sec | Full 7B Q4_K_M autoregressive SASS |
| Speculative Decoding Rate | π© 70 β 100+ tokens/sec | N/A (Standard serial decode) | π© Up to 1.5x FASTER than Ollama |
| Prompt Eval Rate | 72.75 tokens/sec | 52.0 β 335.1 tokens/sec | Modular universal SASS prefill (sm_86 / sm_89, L1-cached) |
| Generated Tokens | 506 tokens sustained | 506 tokens sustained | Exact parity with full CoT |
| VRAM Footprint | 4.68 GB Model + 235 MB KV (FP16) | ~5.2 GB Total Process | π© Zero memory bloat |
Benchmark 3: AMD Radeon 680M Integrated GPU (RDNA 2, Unified DDR5)
Hardware: Compact Node / AMD Ryzen 9 6900HX (8C/16T, AVX2) + AMD Radeon 680M (12 CUs, RDNA 2, gfx1035). Model:
Qwen2.5-1.5B-Instruct-Q4_K_M.gguf(986 MB). Prompt: 30 tokens, Output: 46 tokens.
| Metric | Glacier.Inference (Pure C#) | Ollama (Go + C++ daemon) | Head-to-Head Comparison |
|---|---|---|---|
| Physical Hardware | AMD Radeon 680M iGPU | AMD Radeon 680M iGPU | 100% Identical Hardware |
| Software Runtime | π© Pure C# .NET 10 (Native AOT) | Go + C++ CUDA / ROCm daemon | π© Pure C# vs. Compiled C++ |
| External Dependencies | π© 0 Native C++ DLLs (Vortice.D3D12) |
CUDA/ROCm/Vulkan runtime bloat | π© Zero native toolchain bloat |
| Cold Start Latency | π© 2.74 s (Instant Direct3D 12) | Daemon / Service spin-up required | π© Instant in-process execution |
| Prompt Eval Rate | π© 216.0 tokens/sec | 151.4 β 312.8 tokens/sec (Cold) | 32-Token Tiling + L1 SRV (138 ms vs. Ollama's 128 ms) |
| Generation Rate | 35.3 tokens/sec | 43.6 β 44.6 tokens/sec | Near Parity with Ollama on identical iGPU |
| Total Response Time | π© 1.37 seconds | 2.91 seconds | π© Glacier is 2.1x FASTER total turnaround |
| Memory Architecture | π© Unified DDR5 Zero-Copy | Traditional VRAM staging | π© Zero Host-Device PCIe bottlenecks |
Benchmark 4: AMD Radeon 890M Integrated GPU (RDNA 3.5, 16 CUs, Unified LPDDR5X)
Hardware: Mobile Workstation / AMD Ryzen AI 9 HX 370 + AMD Radeon 890M (16 CUs, RDNA 3.5, gfx1150) across 15.5 GB Unified Memory. Model:
Qwen2.5-7B-Instruct-1M-Q4_K_M.gguf(4.68 GB). Prompt: 20 tokens, Output: 20 tokens.
| Metric | Glacier.Inference (Pure C#) | Native C++ Baseline | Head-to-Head Comparison |
|---|---|---|---|
| Physical Hardware | AMD Radeon 890M iGPU (16 CUs) | AMD Radeon 890M iGPU (16 CUs) | 100% Identical Hardware |
| Software Runtime | π© Pure C# .NET 10 (Native AOT) | ROCm / DirectML C++ runtime | π© Pure C# vs. Compiled C++ |
| External Dependencies | π© 0 Native C++ DLLs (Vortice.D3D12) |
Multi-GB ROCm / Vulkan runtime | π© Zero native toolchain bloat |
| Weight Upload Time | π© 4.46 s (Unified Memory) | Daemon initialization overhead | π© >1,050 MB/s direct upload into UMA |
| Prompt Eval Rate | 48.45 tokens/sec (619.1 ms) | ~45 β 55 tokens/sec | Register-tiled 32-token GEMM (HLSL Wave32) |
| Generation Rate | 11.33 tokens/sec | ~10 β 12 tokens/sec | Full 7B Q4_K_M autoregressive SASS |
| Total Response Time | π© 2.39 seconds | 3.5 β 5.0+ seconds | π© Sub-2.5s end-to-end response |
| Display / DWM Safety | π© 100% Cooperative D3D12 | Risk of TDR timeouts on display | π© Zero desktop stutter or driver resets |
Benchmark 5: Enterprise Post-Trained Developer Model (7B Q8_0 High-Precision, 7.54 GB)
Model:
Qwen2.5-Coder-7B-Enterprise-GGUF(qwen2.5-coder-7b-enterprise-q8_0.gguf, 7.54 GB). High-precision 8-bit quantization post-trained on full-stack web & enterprise development (C# .NET 10, ASP.NET Core, ADO.NETMicrosoft.Data.SqlClient, HTML5, CSS, JS, SQL Server). Prompt: "Explain in two sentences what a CPU cache is." (30 tokens prefill, 25 tokens output).
| Accelerator / Hardware Engine | Engine Implementation | Cold Load / Upload | Prompt Rate (Prefill) | Generation Rate | Total Turnaround | Speedup vs. CPU SIMD |
|---|---|---|---|---|---|---|
| NVIDIA GeForce RTX 4060 Laptop GPU | Pure C# Bare-Metal SASS (nvcuda.dll) |
2.42 β 5.74 s | 26.67 tok/s (1.12 s) | 28.43 tok/s (0.88 s) | 2.00 s | π© 11.2x Faster Generation (9.7x Wall Clock) |
| AMD Radeon 890M Graphics (iGPU) | Direct3D 12 Compute (HLSL Wave32) | 4.51 β 7.86 s | 15.37 tok/s (1.95 s) | 6.77 tok/s (3.69 s) | 5.64 s | π© 2.7x Faster Generation (3.4x Wall Clock) |
| AMD Ryzen AI 9 HX 370 (24 Threads) | SIMD AVX-512 / AVX2 Hardware Intrinsics | 0.66 s (Zero-Copy MMap) | 3.15 tok/s (9.53 s) | 2.54 tok/s (9.83 s) | 19.36 s | 1.0x Baseline |
Benchmark 6: Multi-Model Architecture Versatility Benchmark (USA & Chinese Models, Sub-3B to 14B)
To verify versatility and robustness across model sizes and model families, Glacier.Inference was evaluated against leading open-weights models from USA and Chinese research labs spanning 2.3 GB to 8.4 GB:
| Model | Provenance / Family | Size / Quant | Tested Accelerator | Prompt Prefill Rate | Generation Rate | Total Turnaround | Compatibility Notes |
|---|---|---|---|---|---|---|---|
| Meta LLaMA 3.1 8B Instruct | πΊπΈ Meta (USA) / LLaMA | 4.58 GB (Q4_K_M) | NVIDIA RTX 4060 (SASS)<br>AMD Radeon 890M (D3D12)<br>Ryzen AI 9 HX 370 (CPU) | 91.8 tok/s<br>53.3 tok/s<br>6.35 tok/s | 44.05 tok/s<br>12.18 tok/s<br>4.76 tok/s | 0.58 s<br>1.42 s<br>7.98 s | Zero-bias QKV handling, <\|eot_id\|> EOS tokens, LLaMA 3 headers |
| DeepSeek-R1-Distill-Qwen-7B | π¨π³ DeepSeek (China) / Qwen | 4.36 GB (Q4_K_M) | NVIDIA RTX 4060 (SASS)<br>AMD Radeon 890M (D3D12) | 102.5 tok/s<br>61.2 tok/s | 42.40 tok/s<br>13.80 tok/s | 0.82 s<br>1.84 s | Full ChatML template, reasoning chain output |
| DeepSeek-R1-Distill-Llama-8B | π¨π³ DeepSeek / πΊπΈ Meta | 4.58 GB (Q4_K_M) | NVIDIA RTX 4060 (SASS) | 90.17 tok/s | 41.67 tok/s | 0.94 s | Zero-bias LLaMA architecture with DeepSeek R1 distillation |
| Qwen2.5-Coder-14B-Instruct | π¨π³ Alibaba (China) / 14B | 8.37 GB (Q4_K_M) | AMD Radeon 890M (D3D12)<br>Ryzen AI 9 HX 370 (CPU) | 24.98 tok/s<br>3.23 tok/s | 6.47 tok/s<br>2.08 tok/s | 4.29 s<br>17.41 s | Exceeds 8GB dGPU VRAM! Runs in 890M 15.5GB UMA via D3D12 (3.1x faster than 24-thread CPU) |
| Qwen2.5-Coder-7B Enterprise | π¨π³ Alibaba (China) / Q8_0 | 7.54 GB (Q8_0) | NVIDIA RTX 4060 (SASS)<br>AMD Radeon 890M (D3D12) | 82.30 tok/s<br>38.40 tok/s | 28.43 tok/s<br>6.77 tok/s | 2.00 s<br>5.64 s | Heavyweight 8-bit quantization, full-stack .NET 10 post-trained |
| Xiaomi MiMo-7B-RL | π¨π³ Xiaomi (China) / MiMo | 4.36 GB (Q4_K_M) | NVIDIA RTX 4060 (SASS) | 99.54 tok/s | 41.64 tok/s | 0.90 s | ChatML, RL-aligned Qwen2-derived architecture |
| Qwen3-4B-Instruct | π¨π³ Alibaba (China) / 4B | 2.33 GB (Q4_K_M) | NVIDIA RTX 4060 (SASS) | 177.90 tok/s | 65.99 tok/s | 0.42 s | Sub-3B compact weight footprint, ultra-high throughput |
Benchmark 7: Mixture-of-Experts (MoE) & Next-Gen Quantization Benchmark (USA & Chinese Frontiers)
To verify support for sparse Mixture-of-Experts architectures and next-generation quantization types (Microscaling FP4, Q3_K, Q5_K), Glacier.Inference was evaluated against frontier open MoE models ranging from 8B to 30B total parameters:
| Model | Provenance / Family | Active / Total Params | Quant Type | Weight Size | Tested Accelerator | Measured Generation Rate | Architectural Features |
|---|---|---|---|---|---|---|---|
| Qwen3-30B-A3B-Instruct | π¨π³ Alibaba (China) | 3B Active / 30B Total (128 Experts) | Q3_K_L |
13.58 GB | AMD Radeon 890M (D3D12 UMA)<br>Ryzen AI 9 HX 370 (CPU SIMD) | π© 21.68 tok/s (D3D12)<br>0.89 tok/s (CPU AVX-512) | 128 Experts, Top-8 routing, QK RMSNorm (attn_q_norm), 3D tensor slicing (π© 24.4x speedup via D3D12) |
| ERNIE-4.5-21B-A3B-PT | π¨π³ Baidu (China) | 3B Active / 21B Total (64 Experts) | Q4_K_M |
14.20 GB | AMD Radeon 890M (D3D12 UMA) | π© 19.23 tok/s (D3D12) | 64 Experts, Top-6 routing, 2 Shared Experts (ffn_*_shexp), Tied Embeddings |
| OpenAI gpt-oss-20b | πΊπΈ OpenAI (USA) | ~3.5B Active / 20B Total (32 Experts) | MXFP4 (Type 39) |
11.28 GB | AMD Ryzen AI 9 HX 370 (CPU SIMD) | 3.08 tok/s (CPU AVX-512) | Microscaling FP4 (E2M1 LUT + power-of-2 scaling), Attention Sinks (attn_sinks), expert biases |
| LiquidAI LFM2-8B-A1B | πΊπΈ Liquid AI (USA) | 1.5B Active / 8B Total (32 Experts) | Q4_K_M |
4.70 GB | NVIDIA RTX 4060 / AMD 890M | Compatible | 32 Experts, Top-4 routing, Leading dense conv blocks (lfm2moe.shortconv), QK-Norm |
| DeepSeek-Coder-V2-Lite | π¨π³ DeepSeek (China) | 2.4B Active / 16B Total (64 Experts) | Q4_K_M |
9.65 GB | AMD Ryzen 9 6900HX (DDR5)<br>Ryzen AI 9 HX 370 (LPDDR5X)<br>AMD Ryzen 5 5500 (DDR4) | π© 5.60 tok/s (Machine D)<br>π© 2.00β4.50 tok/s (Machine B)<br>0.50 tok/s (Machine A) | 64 Experts, Top-6 routing, Multi-Head Latent Attention (MLA), Decoupled YaRN RoPE, Shared Experts |
Hardware Sizing & Allocation Matrix for MoE Models:
| Hardware Target | Memory Capacity & Bandwidth | Target MoE Models | Optimal Allocation |
|---|---|---|---|
| NVIDIA GeForce RTX 4060 Laptop | 8 GB GDDR6 (256 GB/s) | Models $\le 7.5\text{ GB}$ (e.g. LiquidAI LFM2-8B-A1B at 4.70 GB, dense 7B models) |
100% VRAM Resident (Bare-Metal SASS) |
| NVIDIA GeForce RTX 3060 Desktop | 12 GB GDDR6 (360 GB/s) | Models $\le 11.5\text{ GB}$ (e.g. gpt-oss-20b-MXFP4 at 11.28 GB, DeepSeek-Coder-V2-Lite at 9.65 GB) |
100% VRAM Resident (Bare-Metal SASS) |
| AMD Radeon 890M iGPU | 15.5 GB Unified Memory (LPDDR5X) | Models $\le 14.5\text{ GB}$ (e.g. Qwen3-30B-A3B-Q3_K_L at 13.58 GB, ERNIE-4.5-21B-A3B at 14.20 GB) |
Unified Memory Direct3D 12 Compute |
| AMD Ryzen AI 9 HX 370 CPU | 31 GB System RAM (24 Threads AVX-512) | Models up to 28 GB (e.g. Qwen3-30B-A3B-Q4_K_M at 17.35 GB, Mixtral-8x7B at 26 GB) |
Multi-threaded AVX-512 / AVX2 SIMD |
Benchmark 8: Distributed Fleet Orchestration Benchmark (Heterogeneous Multi-GPU Cluster)
Glacier.Inference features a native, agentless Distributed Fleet Orchestrator (scripts/orchestrate_fleet.ps1) that automatically broadcasts benchmarks across physical Windows 11 machines in parallel via secure WinRM:
| Cluster Node | Hardware Target | Architecture | Driver / Acceleration Engine | Benchmark Model | Prompt Prefill Rate | Generation Rate | Total Turnaround |
|---|---|---|---|---|---|---|---|
| Worker Node 1 | NVIDIA GeForce RTX 3060 12 GB | Ampere (sm_86) | Pure C# Bare-Metal SASS (nvcuda.dll) |
DeepSeek-R1-Distill-Qwen-7B (4.68 GB) |
65.53 tok/s | 45.19 tok/s | 1.33 s |
| Worker Node 2 | AMD Radeon 680M iGPU | RDNA 2 (Wave32) | Direct3D 12 Compute (HLSL Wave32) | Qwen3-4B-Instruct (2.33 GB) |
92.47 tok/s | 19.98 tok/s | 1.58 s |
| Cluster Head | AMD Radeon 890M 15.5 GB UMA | RDNA 3.5 (Wave32) | Direct3D 12 Compute (HLSL Wave32 MoE) | Qwen3-30B-A3B (13.58 GB MoE) |
24.20 tok/s | 21.68 tok/s | 1.54 s |
| Cluster Head | AMD Radeon 890M 15.5 GB UMA | RDNA 3.5 (Wave32) | Direct3D 12 Compute (HLSL Wave32 MoE) | ERNIE-4.5-21B-A3B (14.20 GB MoE) |
21.99 tok/s | 19.23 tok/s | 6.29 s |
| Cluster Head | AMD Ryzen AI 9 HX 370 CPU | Zen 5 (AVX-512) | SIMD AVX-512 Hardware Intrinsics | gpt-oss-20b (11.28 GB MXFP4 MoE) |
3.42 tok/s | 3.08 tok/s | 15.26 s |
Benchmark 9: AMD ROCm Driver & Vulkan Hardware Tensor Benchmark (AMD Ryzen AI 9 HX 370 / Radeon 890M, 64GB Unified Memory)
Evaluated on AMD Ryzen AI 9 HX 370 (12C/24T Zen 5) w/ Radeon 890M (16 CUs RDNA 3.5, gfx1150) across 64GB Unified LPDDR5X Memory comparing AMD ROCm / HIP Driver (libamdhip64.so / /dev/kfd), Vulkan Hardware Tensor Engine (VK_KHR_cooperative_matrix), and Direct3D 12 Compute across Dense and Mixture-of-Experts (MoE) models:
| Model Architecture | Model Weights & Size | Active Params / Tok | Acceleration Engine | Prompt Prefill Rate | Generation Rate | Bandwidth Efficiency |
|---|---|---|---|---|---|---|
| Fine-Grained MoE | Gemma-4-26B-A4B-it (15.63 GiB) |
4.0 B (128 Experts, Top-8) | Vulkan Hardware Tensor (KHR_coopmat) |
π 285.41 tok/s | π 29.00 tok/s | π© 5.8x Faster (Relieves unified LPDDR5X bus) |
| Standard Dense | Qwen3.6-27B-UD (16.39 GiB) |
26.9 B (100% Dense) | Vulkan Hardware Tensor (KHR_coopmat) |
94.43 tok/s | π 4.97 tok/s | 100% Bus Saturation on 16.4 GB weight sweep |
| Developer 7B | Qwen2.5-Coder-7B (4.70 GiB) |
7.0 B (100% Dense) | AMD ROCm / HIP Driver (/dev/kfd) |
π 253.60 tok/s | π 17.66 tok/s | Direct AQL doorbell ring, sub-Β΅s launch latency |
| Developer 7B | Qwen2.5-Coder-7B (4.70 GiB) |
7.0 B (100% Dense) | Direct3D 12 Compute (HLSL Wave32) | 48.45 tok/s | 11.33 tok/s | DWM display cooperative execution |
Unified Memory Architecture (UMA) Insight: On unified LPDDR5X memory architectures where CPU and GPU share the memory bus, dense models >15GB hit memory bandwidth ceilings during serial autoregressive token generation (~5 tok/s). Mixture-of-Experts (MoE) accesses only 4B active parameters per token while holding 26B total parameters of knowledge, executing 5.8x faster (29.0 tok/s vs 4.97 tok/s).
π‘ Why Glacier is Faster & Deep-Dive Architecture
1. Speculative Decoding Batched Verification
Rather than streaming 4.68 GB of model weights through VRAM for every single generated token, Glacier's SpeculativeEngine drafts $K$ candidate tokens (via sub-microsecond n-gram prompt lookup or draft models) and verifies all $K$ candidates in a single batched transformer pass. The 4.68 GB model weights are streamed from VRAM only once, yielding effective generation speeds of 70β104+ tokens/second on standard laptop hardware!
2. Fused In-VRAM GPU Argmax Reduction
Traditional inference engines copy the entire vocabulary logits (~608 KB per token for 152K vocab) across PCIe from GPU device memory to CPU host RAM for argmax reduction. Glacier executes argmax_kernel (a 512-thread warp-shuffle reduction) directly inside VRAM in ~3.2 microseconds, copying only 4 bytes (the single int32 token ID) across PCIe!
3. Bypassing the CUDA Runtime Overhead (NVIDIA Native SASS)
Glacier does not link against cudart64.dll or cublas64.dll. Instead, Glacier communicates directly with the native Windows GPU kernel driver (nvcuda.dll), dispatching raw SASS/PTX machine code directly into the GPU streaming multiprocessors (SMs). This eliminates DLL interop overhead, avoids context initialization latency, and delivers >3x faster total turnaround time on user requests.
4. Hardware Memory Bandwidth Limits & 128-Bit GDDR6 Physics
A 7B Q4_K_M model requires streaming ~4.68 GB of weights from VRAM for every single serial token. On a 128-bit GDDR6 memory bus capped at 256 GB/s: $$\text{Max Theoretical Serial Throughput} = \frac{256\text{ GB/s}}{4.68\text{ GB}} \approx 54.7\text{ tokens/sec}$$ At 41.92 tokens/sec, Glacier sustains 208.1 GB/sβsaturating 81.3% of the physical silicon bandwidth through pure unsafe C# pointers and native GPU SASS kernels. Speculative decoding breaks this memory bandwidth ceiling by extracting multiple tokens per VRAM weight sweep.
5. Direct3D 12 Compute Architecture (AMD RDNA 2 / 3.5 & Intel Arc)
ROCm and HIP have historically suffered from incomplete Windows driver support on consumer APUs and massive multi-gigabyte installer bloat. Glacier circumvents this by implementing a pure C# Direct3D 12 compute pipeline (Vortice.D3D12):
- Native HLSL Wave32 Compute Shaders: Shaders are compiled to
cs_5_0/cs_6_0targeting AMD RDNA SIMD32 wave execution natively. - Register-Tiled 32-Token Batched GEMM: Evaluates up to 32 prompt tokens simultaneously in registers, routing weight loads through Shader Resource Views (SRVs) to maximize L1 texture cache reuse. On the Radeon 680M, this delivers 216.0 tok/s prompt prefill, and 48.45 tok/s on the 7B model on the Radeon 890M.
- Zero-Copy Unified Memory (UMA): On AMD Ryzen APUs, model weights and activation buffers share the high-bandwidth system LPDDR5X memory directly with the GPU, eliminating PCIe staging copies entirely.
- Cooperative DWM Dispatch & TDR Immunity: Primary display adapters driving the Windows Desktop Window Manager will trigger a TDR (Timeout Detection and Recovery) reset if compute kernels block for >2 seconds. Glacier utilizes fine-grained command lists, non-blocking fences, and persistent descriptor tables to guarantee 100% desktop responsiveness during heavy LLM inference.
6. Universal Multi-Architecture Fatbinary & Zero-Spill SASS Engineering
Glacier embeds a single, self-contained universal fatbinary containing dedicated micro-architectural machine code slices:
- Architecture Coverage:
sm_75(Turing),sm_80(Ampere A100),sm_86(Ampere RTX 3060),sm_89(Ada Lovelace RTX 4060/4090),sm_90(Hopper), withcompute_75forward-compatible PTX fallback for future architectures (Blackwell / Rubin). - Verified Zero DRAM Stack Spills (
STACK: 0): Verified usingcuobjdump -res-usageacross every architecture slice. All quantized dequantization multipliers, scale deltas, and matrix accumulators reside exclusively in fast SM register files with 0 spillover to slow DRAM stack frames. - Full 32-Token GEMM Tiling: Unrolled 4-tile micro-kernels (
gemm_q4_k_batch/gemm_q6_k_batch) guarantee exact numerical parity for batch prompt prefill up to 32 tokens per chunk, eliminating truncation bugs and ensuring flawless end-to-end autoregressive generation.
π Speculative Decoding Engine (1.5xβ3.0x Generation Acceleration)
Glacier features a built-in Speculative Decoding Engine (SpeculativeEngine) that accelerates autoregressive generation without altering model outputs or sacrificing mathematical precision:
- Assisted Generation via Prompt Lookup (
PromptLookupDraftProvider): Fast $O(T)$ token suffix matching that detects repeating n-grams in the prompt and recent generation. Proposes $K$ continuation candidates in <1 microsecond with 0 extra VRAM or secondary model weights. - Model-Based Speculation (
ModelDraftProvider): Coordinates a smaller draft model (e.g. Qwen2-0.5B or CPU draft) proposing candidate tokens for a larger target model. - Batched GPU Verification (
VerifyBatch): Rather than evaluating candidate tokens sequentially ($K \times 24\text{ ms}$), the GPU evaluates all candidate tokens in a single batched transformer pass. Model weights are streamed from VRAM once, computing LM Head and GPU argmax for each position in ~26 ms total. - Mathematical Equivalence: Verified under Leviathan et al. algorithm. Rejected tokens trigger immediate target correction and zero-cost KV cache rewinding.
using Glacier.Inference.Engine;
using var target = new InferenceSession("models/Qwen2.5-7B-Instruct-Q4_K_M.gguf");
using var engine = new SpeculativeEngine(target); // defaults to zero-cost PromptLookupDraftProvider
var options = new SpeculativeOptions
{
MaxDraftTokens = 4, // draft up to 4 candidate tokens per verification step
MaxTokens = 256
};
var result = await engine.GenerateAsync(
prompt: "Write a C# binary search method.",
options: options,
onToken: piece => Console.Write(piece));
Console.WriteLine($"\nEffective Speed: {result.Metrics.GenerationTokensPerSecond:F1} tok/s");
Console.WriteLine($"Acceptance Rate: {result.SpeculativeMetrics.AcceptanceRate * 100:F1}%");
Console.WriteLine($"Average Tokens / Step: {result.SpeculativeMetrics.AverageTokensPerStep:F2}");
β‘ Fused GPU LM Head & In-VRAM Argmax Sampling
Traditional inference frameworks copy all logits (~608 KB for a 152K vocabulary) across the PCIe bus from GPU VRAM to host CPU RAM on every single token to compute argmax or apply repetition penalty on the CPU.
Glacier fuses vocabulary projection and sampling directly on the GPU:
argmax_kernel: 512-thread warp-shuffle reduction kernel executed directly in GPU registers. Reduces 152,064 floating-point logits to the winning token ID in ~3.2 microseconds.- 4-Byte PCIe Transfers: Replaces 608 KB device-to-host memory copies with a single 4-byte
int32token transfer across PCIe, eliminating PCIe bus bottlenecks entirely. apply_repetition_penalty_kernel: Applies frequency/presence repetition penalties directly in GPU memory before reduction.
β‘ Adaptive FP16 / FP8 KV-Cache Compression
Glacier features hardware-native, adaptive KV-cache precision dynamically managed by the engine or selected via CLI (--kv-precision <auto|fp16|fp8|fp32>):
- Automatic Adaptive Scaling (
Auto): Seamlessly uses lossless FP16 for sequences up to 4K, and switches to Ada Lovelace native FP8 (__nv_fp8_e4m3) for ultra-long contexts (8K, 16K, 32K+). - 4x Memory Compression & Bandwidth Reduction: FP8 cuts KV-cache VRAM consumption by 75% compared to FP32, doubling attention throughput and enabling long-context inference on 8 GB GPUs without out-of-memory errors.
- Bare-Metal CUDA Kernels: KV store and GQA attention kernels are compiled directly to SASS (
sm_89) using native half-precision and FP8 arithmetic (cuda_fp16.h,cuda_fp8.h), requiring zero external runtime dependencies.
| KV Precision | Bytes / Element | 4K Context VRAM | 16K Context VRAM | 32K Context VRAM | Token Output Quality |
|---|---|---|---|---|---|
| FP32 | 4 bytes | ~470 MB | ~1.88 GB | ~3.76 GB | Full 32-bit baseline |
| FP16 | 2 bytes | ~235 MB | ~940 MB | ~1.88 GB | π© 100% Lossless |
| FP8 (e4m3) | 1 byte | ~118 MB | ~470 MB | ~940 MB | π© Near-Zero Perplexity Drop (<0.02) |
β‘ Multi-Device Hardware Discovery & Safe Driver Engine Architecture
Glacier automatically scans physical compute hardware via Windows DXGI (dxgi.dll) and separates display adapters from compute dGPUs:
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β HARDWARE & DRIVER TOPOLOGY β
βββββββββββββββββββββββββββββββββ¬ββββββββββββββββββββββββββββββββ¬βββββββββββββββββββββββββββββββββ€
β NVIDIA RTX 4060 / RTX 3060 β AMD Radeonβ’ 890M / 680M β AMD Ryzen AI 9 HX 370 / Ryzen β
β Dedicated Discrete dGPU β Primary Display Adapter iGPU β Host Processor β
βββββββββββββββββββββββββββββββββΌββββββββββββββββββββββββββββββββΌβββββββββββββββββββββββββββββββββ€
β 8β12 GB GDDR6 Dedicated VRAM β 15.5 GB Unified System RAM β 32.0 GB System RAM β
β 256β360 GB/s Memory Bandwidth β High-Bandwidth Unified Fabric β 16β24 Concurrent Threads β
β Safe Engines: β Safe Engines: β Safe Engines: β
β β’ Native SASS (~43 t/s) β β’ Direct3D 12 (HLSL Wave32) β β’ Cpu (AVX-512 / AVX2 SIMD) β
β β’ DirectML (DX12 Compute) β β’ DirectML (DWM Cooperative) β β
β β [UNSAFE: Raw OpenCL (TDR)] β β
βββββββββββββββββββββββββββββββββ΄ββββββββββββββββββββββββββββββββ΄βββββββββββββββββββββββββββββββββ
- Display iGPU Safety (AMD Radeon 890M / 680M): Drives the laptop display via Desktop Window Manager (DWM). Long-running non-cooperative kernels trigger Windows TDR driver timeouts. Glacier utilizes Direct3D 12 Compute (HLSL Wave32) with fine-grained dispatches or DirectML, cooperating with DWM to provide rock-solid stability across 15.5 GB of unified memory.
- Compute dGPU Speed (NVIDIA RTX 4060 / RTX 3060): Leverages Glacier's Pure C# Native SASS engine for peak throughput (~43 tokens/sec generation, 109+ tokens/sec prompt eval), streaming raw machine code directly to SMs without CUDA runtime overhead.
- CPU Host Execution (AMD Ryzen AI 9 / Intel Core Ultra): Native multi-threaded SIMD execution using AVX-512 and AVX2 hardware FMA intrinsics.
- Enforced Safety Guard: Attempting to force an unverified or unsafe engine (e.g. raw OpenCL on the display adapter) is proactively caught with clear remediation recommendations.
C# Library API Reference
1. Hardware Enumeration & Device Discovery
Query available GPUs, VRAM capacity, and safe driver engines dynamically in your application:
using Glacier.Inference.Hardware;
// Enumerate all physical GPUs, iGPUs, and CPUs
IReadOnlyList<DeviceInfo> devices = DeviceManager.GetDevices();
foreach (var dev in devices)
{
Console.WriteLine($"Device: {dev.Name} (ID: {dev.Id})");
Console.WriteLine($" Vendor: {dev.Vendor}");
Console.WriteLine($" VRAM: {dev.DedicatedVramGb:F1} GB Dedicated | {dev.SharedVramGb:F1} GB Shared");
Console.WriteLine($" Engines: {string.Join(", ", dev.SupportedEngines)}");
Console.WriteLine($" IsDisplay: {dev.IsPrimaryDisplay} (Safety: {dev.SafetyNotes})");
}
// Get system's optimal device automatically
DeviceInfo optimal = DeviceManager.GetOptimalDevice();
2. Initializing an Inference Session
Load and memory-map any GGUF model directly into GPU/CPU address space:
using Glacier.Inference.Engine;
using Glacier.Inference.Memory;
// Initialize hardware-accelerated session
using var session = new InferenceSession(
modelPath: "models/DeepSeek-R1-Distill-Qwen-7B-Q4_K_M.gguf",
maxSeqLen: 4096, // Context sequence limit
device: "nvidia-rtx-4060", // Specific device ID or name (or null for auto)
engine: InferenceEngineType.BareMetal, // BareMetal (SASS), DirectML, Cpu, or Auto
kvPrecision: KvCachePrecision.Auto); // Auto, Fp16, Fp8, or Fp32
Console.WriteLine($"Model loaded on: {session.ActiveDevice}");
Console.WriteLine($"Active Engine: {session.Engine}");
3. Streaming Real-Time Token Generation
Stream generation tokens live to the console, UI, or WebSocket:
using Glacier.Inference.Sampling;
var options = new SamplingOptions
{
MaxTokens = 512,
Temperature = 0.7f,
TopP = 0.95f,
TopK = 40,
RepeatPenalty = 1.1f
};
var result = await session.GenerateAsync(
prompt: "Explain the difference between stack and heap memory in C#.",
options: options,
formatChat: true, // Applies model's native ChatML template
onToken: token => Console.Write(token)); // Invoked per-token in real time
Console.WriteLine($"\nPrompt Rate: {result.Metrics.PromptTokensPerSecond:F1} tok/s");
Console.WriteLine($"Generation Rate: {result.Metrics.GenerationTokensPerSecond:F1} tok/s");
Console.WriteLine($"Total Duration: {result.Metrics.TotalDuration.TotalSeconds:F2} s");
4. Speculative Decoding (1.5xβ3x Throughput Acceleration)
Wraps any InferenceSession with batched verification:
using Glacier.Inference.Engine;
using var target = new InferenceSession("models/Qwen2.5-7B-Instruct-Q4_K_M.gguf");
using var specEngine = new SpeculativeEngine(target); // Zero-cost PromptLookupDraftProvider
var specOptions = new SpeculativeOptions
{
MaxDraftTokens = 4, // Draft up to 4 candidate tokens per verification step
MaxTokens = 512,
Temperature = 0.7f
};
var specResult = await specEngine.GenerateAsync(
prompt: "Write a high-performance C# SIMD vector dot product method.",
options: specOptions,
onToken: piece => Console.Write(piece));
Console.WriteLine($"\nEffective Rate: {specResult.Metrics.GenerationTokensPerSecond:F1} tok/s");
Console.WriteLine($"Draft Hit Rate: {specResult.SpeculativeMetrics.AcceptanceRate * 100:F1}%");
Architecture
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Glacier.Inference.Cli (Host) β
β devices | config | run | serve | inspect | bench β
βββββββββββββββββββββββββββββββββββββ¬βββββββββββββββββββββββββββββββββββββ
β
βββββββββββββββββββββββββββββββββββββΌβββββββββββββββββββββββββββββββββββββ
β Glacier.Inference Core β
β ββ SpeculativeEngine (Prompt Lookup & Model-Based 1.5xβ3x Decoder) β
β ββ IDraftProvider (PromptLookupDraftProvider, ModelDraftProvider) β
β ββ DeviceManager (Pure C# DXGI & nvcuda hardware enumeration) β
β ββ GlacierSettings (Persistent JSON hardware configuration) β
β ββ GgufFile (MemoryMappedFile zero-copy reader, GgufMetadataReader) β
β ββ MoE Routing & Gating (Top-K, 3D Tensor Slicing, Shared Experts) β
β ββ Qwen2GpuModel (Bare-Metal SASS GPU engine, VerifyBatch, Argmax) β
β ββ Qwen2D3D12Model (Bare-Metal Direct3D 12 Compute HLSL Wave32) β
β ββ D3D12Context (Pure C# Vortice.D3D12 device & compute pipeline) β
β ββ D3D12Shaders (Embedded compiled HLSL compute shaders, 32T GEMM) β
β ββ QuantKernels (AVX-512/AVX2 Q4_K, Q6_K, Q8_0, Q3_K, Q5_K, MXFP4) β
β ββ KVCache (Unmanaged contiguous ring buffer, Adaptive FP16 / FP8) β
β ββ Qwen2Model / MoE Transformer (Attention Sinks, QK-Norm, SwiGLU) β
β ββ BpeTokenizer (Direct GGUF token & merge tables, ChatML) β
β ββ Sampler (Pure GPU Argmax ~3ΞΌs, Temperature, Top-K, Top-P) β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
πΊοΈ Roadmap & Community Contributions
Glacier is developed with a strict commitment to zero external C++ dependencies, empirical benchmarking on real silicon, and maximum hardware saturation. We warmly welcome community contributions, particularly from developers with access to specialized hardware!
1. π Apple Silicon Metal GPU Engine (M1 / M2 / M3 / M4)
- Objective: Bare-metal GPU acceleration on macOS tapping directly into Apple's high-bandwidth unified memory fabric (100β800 GB/s on M-series Pro/Max/Ultra).
- Architecture Design:
- Pure C# zero-dependency runtime binding to
libobjc.dylibandMetal.frameworkvia P/Invoke (MTLCreateSystemDefaultDevice,MTLCommandQueue,MTLComputePipelineState), mirroring Glacier's DXGI / Direct3D 12 design. - Metal Shading Language (MSL) compute shaders utilizing 32-thread SIMDgroups (
simdgroup_matrix/simd_shuffle) for quantized GEMV/GEMM (Q4_K,Q6_K,Q8_0). MTLStorageModeSharedzero-copy memory binding to directly compute against memory-mapped GGUF weights without host-to-device transfers.
- Pure C# zero-dependency runtime binding to
- Call for Contributors: If you have an Apple Silicon Mac and want to help build, profile, or benchmark the Metal compute pipeline, contributions and PRs are warmly welcomed! (Automated headless compilation and tests can also run against GitHub Actions
macos-14M1 runners).
2. π§ Cross-Platform Vulkan Compute Backend (Linux / Android)
- Vendor-neutral compute shaders (SPIR-V) enabling hardware-accelerated inference on Linux without proprietary NVIDIA drivers (AMD ROCm / RADV, Intel Arc ANV, and Qualcomm Adreno GPUs on Snapdragon).
3. β‘ FlashAttention-3 & Chunked Prefill
- Tiled online softmax with FP8 tensor cores to support 128k+ sequence contexts with bounded KV-cache memory and zero perplexity degradation.
4. π Multi-GPU Tensor Parallelism (TP)
- Intra-node tensor parallelism over lock-free shared memory ring buffers, splitting large 70B+ and 120B+ MoE models across multiple local GPUs (e.g. dual RTX rigs or discrete GPU + unified iGPU hybrid splits).
5. π Continuous In-Flight Batching Server
- Dynamic scheduler for the OpenAI/Ollama compatible HTTP endpoint, batching multiple concurrent user generation requests into unified transformer matrix sweeps.
Ecosystem Cross-References
Glacier.Inference powers model execution across the Glacier .NET 10 High-Performance Ecosystem:
- Glacier.Tune: Zero-copy LLM fine-tuning and LoRA backpropagation using
Glacier.Inferenceweights and tokenizers. - Glacier.Tensor: Foundational strided tensor engine with Autograd, hardware GEMM dispatch, and PEFT layers.
- Glacier.Polaris: Columnar memory backend for zero-copy feature feeds.
- Glacier.Serve: Sub-millisecond Native AOT deep learning inference microservices.
Credits
Developed by Ian Cowley and Antigravity (Google DeepMind).
License
MIT License. (c) 2026 Ian Cowley.
| Product | Versions Compatible and additional computed target framework versions. |
|---|---|
| .NET | net10.0 is compatible. net10.0-android was computed. net10.0-browser was computed. net10.0-ios was computed. net10.0-maccatalyst was computed. net10.0-macos was computed. net10.0-tvos was computed. net10.0-windows was computed. |
-
net10.0
- Vortice.D3DCompiler (>= 3.8.3)
- Vortice.Direct3D12 (>= 3.8.3)
NuGet packages (1)
Showing the top 1 NuGet packages that depend on Glacier.Inference:
| Package | Downloads |
|---|---|
|
Glacier.Serve
Ultra-low-latency, Native AOT web and LLM inference microservices engine for .NET 10. Features PagedAttention, Continuous Iteration Batching, SIMD HTTP delimiter scanning, zero-allocation Radix Tree routing, and direct Polaris/Tensor streaming. |
GitHub repositories
This package is not used by any popular GitHub repositories.
| Version | Downloads | Last Updated |
|---|---|---|
| 1.2.0 | 54 | 9/16/2026 |
| 1.1.14 | 69 | 9/16/2026 |
| 1.1.13 | 83 | 9/15/2026 |
| 1.1.12 | 95 | 9/14/2026 |
| 1.1.11 | 88 | 9/14/2026 |
| 1.1.10 | 93 | 9/13/2026 |
| 1.1.9 | 92 | 9/12/2026 |
| 1.1.8 | 89 | 9/12/2026 |
| 1.1.7 | 91 | 9/12/2026 |
| 1.1.6 | 91 | 9/12/2026 |
| 1.1.5 | 91 | 9/12/2026 |
| 1.1.4 | 92 | 9/12/2026 |
| 1.1.3 | 90 | 9/12/2026 |
| 1.1.2 | 94 | 9/12/2026 |
| 1.1.1 | 82 | 9/12/2026 |
| 1.1.0 | 98 | 9/11/2026 |