ElBruno.LocalLLMs.BlazorComponents 0.22.0

dotnet add package ElBruno.LocalLLMs.BlazorComponents --version 0.22.0
                    
NuGet\Install-Package ElBruno.LocalLLMs.BlazorComponents -Version 0.22.0
                    
This command is intended to be used within the Package Manager Console in Visual Studio, as it uses the NuGet module's version of Install-Package.
<PackageReference Include="ElBruno.LocalLLMs.BlazorComponents" Version="0.22.0" />
                    
For projects that support PackageReference, copy this XML node into the project file to reference the package.
<PackageVersion Include="ElBruno.LocalLLMs.BlazorComponents" Version="0.22.0" />
                    
Directory.Packages.props
<PackageReference Include="ElBruno.LocalLLMs.BlazorComponents" />
                    
Project file
For projects that support Central Package Management (CPM), copy this XML node into the solution Directory.Packages.props file to version the package.
paket add ElBruno.LocalLLMs.BlazorComponents --version 0.22.0
                    
#r "nuget: ElBruno.LocalLLMs.BlazorComponents, 0.22.0"
                    
#r directive can be used in F# Interactive and Polyglot Notebooks. Copy this into the interactive tool or source code of the script to reference the package.
#:package ElBruno.LocalLLMs.BlazorComponents@0.22.0
                    
#:package directive can be used in C# file-based apps starting in .NET 10 preview 4. Copy this into a .cs file before any lines of code to reference the package.
#addin nuget:?package=ElBruno.LocalLLMs.BlazorComponents&version=0.22.0
                    
Install as a Cake Addin
#tool nuget:?package=ElBruno.LocalLLMs.BlazorComponents&version=0.22.0
                    
Install as a Cake Tool

ElBruno.LocalLLMs

NuGet NuGet Downloads Build Status License: MIT HuggingFace .NET GitHub stars Twitter Follow

Run local LLMs in .NET through IChatClient ๐Ÿง 

Run local LLMs in .NET through IChatClient โ€” the same interface you'd use for Azure OpenAI, Ollama, or any other provider. Powered by ONNX Runtime GenAI and BitNet.

What's New

The last 5 notable additions to the library. Updated with each NuGet release.

  • ๐ŸŽฏ ElBruno.LocalLLMs.Decisions โ€” new package for local System One decision models. Typed choice, score and yes/no answers with full probability distributions in a single forward pass (~600 ms on CPU), running Laya fully in-process on ONNX Runtime โ€” no server, no Python. Built for routing, triage, moderation and intent detection, where a chat model is slow and overkill. Call services.AddLocalDecisions() and inject IDecisionClient. See the Local Decisions Guide and the LocalDecisions sample.
  • โฌ†๏ธ .NET 10 only โ€” every project now single-targets net10.0. .NET 8 reaches end of support in November 2026, so multi-targeting was dropped ahead of it. Breaking: consumers still on .NET 8 must stay on v0.21.0 or upgrade to .NET 10.
  • ๐Ÿงฉ ElBruno.LocalLLMs.BlazorComponents โ€” new Razor Class Library with 7 ready-to-use Blazor components: ModelStatusCard (download progress bar + actions), ModelGallery (filterable grid), ModelSelector (two-way-bindable dropdown), ChatBox (streaming token display), EnvironmentDashboard (CPU/CUDA/DirectML badges), LocalLLMHealthBadge (nav-bar status dot), and RagPlayground. Call services.AddLocalLLMsBlazorComponents() to register. See the Blazor Components Guide and the BlazorDemo sample.
  • ๐Ÿง  GPT-OSS 20B support โ€” OpenAI's open-weight MoE model (Apache-2.0) now runs locally via the official onnxruntime/gpt-oss-20b-onnx artifacts. Adds the Harmony prompt format, channel-aware output filtering (chain-of-thought is stripped, never shown to users), Harmony tool calling, and a ReasoningEffort option. Two model IDs: gpt-oss-20b (CPU INT4) and gpt-oss-20b-cuda. See the GptOssChat sample. Also fixes a token-duplication bug that repeated the final token of every generation.
  • ๐Ÿš€ v0.20.12 โ€” Corrects sibling-package assembly versions, hardens vision token probing against model context limits, and verifies Fara smart image resizing for screenshot workflows.

Features

  • ๐Ÿงฉ Blazor components โ€” ModelStatusCard, ChatBox, ModelGallery, ModelSelector, EnvironmentDashboard, LocalLLMHealthBadge, RagPlayground via ElBruno.LocalLLMs.BlazorComponents (guide)
  • ๐ŸŽฏ Local decision models โ€” typed choice/score/yes-no answers with probabilities in one forward pass via ElBruno.LocalLLMs.Decisions (guide)
  • ๐Ÿ”Œ IChatClient implementation โ€” seamless integration with Microsoft.Extensions.AI
  • ๐Ÿ“ฆ Automatic model download โ€” models are fetched from HuggingFace on first use
  • ๐Ÿš€ Zero friction โ€” works out of the box with sensible defaults (Phi-3.5 mini)
  • ๐Ÿ–ฅ๏ธ Multi-hardware โ€” CPU, CUDA, and DirectML execution providers
  • ๐Ÿ’‰ DI-friendly โ€” register with AddLocalLLMs() or AddBitNetChatClient() in ASP.NET Core
  • ๐Ÿ”„ Streaming โ€” token-by-token streaming via GetStreamingResponseAsync
  • ๐Ÿ“Š Multi-model โ€” switch between Phi-3.5, Phi-4, Qwen2.5, Qwen3, Llama 3.2, MagenticBrain, and more
  • ๐Ÿ‘๏ธ Vision-language models โ€” run Fara 1.5-9B image+text models via LocalVisionChatClient
  • ๐Ÿค– Agentic models โ€” Qwen3 / MagenticBrain support for multi-agent orchestration loops
  • ๐ŸŽฏ Fine-tuned models โ€” pre-trained Qwen2.5 variants for tool calling and RAG (guide)
  • โšก BitNet support โ€” run 1.58-bit ternary models via bitnet.cpp with extreme efficiency (guide)
  • ๐Ÿ“ˆ OpenTelemetry diagnostics โ€” lifecycle activities and metrics for queued, first-token, completion, cancellation, and failure (guide)

Packages

Package NuGet Downloads Description
ElBruno.LocalLLMs NuGet Downloads Core library โ€” ONNX Runtime GenAI models via IChatClient
ElBruno.LocalLLMs.Rag NuGet Downloads RAG pipeline โ€” document chunking, indexing, retrieval
ElBruno.LocalLLMs.BitNet NuGet Downloads BitNet 1.58-bit models via bitnet.cpp + IChatClient
ElBruno.LocalLLMs.BlazorComponents NuGet Downloads Blazor components โ€” ModelStatusCard, ChatBox, ModelGallery, and more
ElBruno.LocalLLMs.Decisions NuGet Downloads System One decision models โ€” typed choice, score and yes/no answers

Installation

dotnet add package ElBruno.LocalLLMs

For CPU scenarios, no extra package is required โ€” the transitive buildTransitive shim copies onnxruntime-genai.dll automatically on Windows.

Add a runtime package only when you want a specific GPU provider:

# ๐ŸŸข NVIDIA GPU (CUDA):
dotnet add package Microsoft.ML.OnnxRuntimeGenAI.Cuda

# ๐Ÿ”ต Any Windows GPU โ€” AMD, Intel, NVIDIA (DirectML):
dotnet add package Microsoft.ML.OnnxRuntimeGenAI.DirectML

โš ๏ธ Add at most one GPU runtime package. Do not reference both Microsoft.ML.OnnxRuntimeGenAI.Cuda and Microsoft.ML.OnnxRuntimeGenAI.DirectML simultaneously.

If you use a GPU runtime package and want to disable the transitive CPU copy shim, set: <ElBrunoLocalLLMsDisableCpuNativeCopy>true</ElBrunoLocalLLMsDisableCpuNativeCopy> in your application .csproj.

๐Ÿš€ The library defaults to ExecutionProvider.Auto โ€” on Windows it tries DirectML โ†’ CUDA โ†’ CPU, and on Linux it tries CUDA โ†’ CPU. No code changes needed.

Quick Start

using ElBruno.LocalLLMs;
using Microsoft.Extensions.AI;

// Create a local chat client (downloads Phi-3.5 mini on first run)
using var client = await LocalChatClient.CreateAsync();

var response = await client.GetResponseAsync([
    new(ChatRole.User, "What is the capital of France?")
]);

Console.WriteLine(response.Text);

First Run

The first time you create a LocalChatClient, the model is downloaded from HuggingFace to your local cache directory (~2-4 GB). This typically takes 30-60 seconds depending on your internet connection.

Track download progress:

using var client = await LocalChatClient.CreateAsync(
    new LocalLLMsOptions { Model = KnownModels.Phi35MiniInstruct },
    progress: new Progress<ModelDownloadProgress>(p =>
    {
        var percent = (p.BytesDownloaded * 100) / p.TotalBytes;
        Console.WriteLine($"{p.FileName}: {percent:F1}%");
    })
);

Subsequent runs load instantly from cache (%LOCALAPPDATA%/ElBruno/LocalLLMs/models).

Skip auto-download if using a pre-downloaded model:

var options = new LocalLLMsOptions
{
    Model = KnownModels.Phi35MiniInstruct,
    ModelPath = "/path/to/local/model",
    EnsureModelDownloaded = false
};
using var client = await LocalChatClient.CreateAsync(options);

Streaming

using ElBruno.LocalLLMs;
using Microsoft.Extensions.AI;

using var client = await LocalChatClient.CreateAsync(new LocalLLMsOptions
{
    Model = KnownModels.Phi35MiniInstruct
});

await foreach (var update in client.GetStreamingResponseAsync([
    new(ChatRole.System, "You are a helpful assistant."),
    new(ChatRole.User, "Explain quantum computing in simple terms.")
]))
{
    Console.Write(update.Text);
}

GPU Acceleration

By default, ExecutionProvider.Auto tries GPU first and falls back to CPU automatically:

// Use explicit GPU provider (fails if CUDA not installed; use Auto to fallback to CPU)
var options = new LocalLLMsOptions
{
    ExecutionProvider = ExecutionProvider.Cuda
};

// Multi-GPU systems: select device ID
var options2 = new LocalLLMsOptions
{
    ExecutionProvider = ExecutionProvider.Cuda,
    GpuDeviceId = 1  // Use second GPU
};

Auto fallback behavior:

  • Windows + DirectML available โ†’ uses a Windows GPU through DirectML
  • Windows + DirectML unavailable, CUDA available โ†’ uses NVIDIA GPU
  • Linux + CUDA available โ†’ uses NVIDIA GPU
  • GPU unavailable โ†’ falls back to CPU (no errors, just slower)

โš ๏ธ CUDA note: ONNX Runtime GenAI 0.15.x expects CUDA 13.*, cuDNN 9.*, and the latest Microsoft Visual C++ 2015-2022 runtime. When those native libraries are missing, provider diagnostics now surface the exact DLL mismatch or missing dependency instead of entering the failing native path.

See Troubleshooting: GPU Setup for debugging GPU issues.

Model Metadata

Inspect model capabilities at runtime โ€” context window size, model name, and vocabulary:

using var client = await LocalChatClient.CreateAsync();

var metadata = client.ModelInfo;
Console.WriteLine($"Model:          {metadata?.ModelName}");
Console.WriteLine($"Context window: {metadata?.MaxSequenceLength}");
Console.WriteLine($"Vocab size:     {metadata?.VocabSize}");

This is useful for prompt-length validation, adaptive chunking, and model selection logic.

Dependency Injection

builder.Services.AddLocalLLMs(options =>
{
    options.Model = KnownModels.Phi35MiniInstruct;
    options.ExecutionProvider = ExecutionProvider.DirectML;
});

// Inject IChatClient anywhere
public class MyService(IChatClient chatClient) { ... }

Error Handling

The library provides structured exception types for graceful error handling:

using ElBruno.LocalLLMs;
using Microsoft.Extensions.AI;

try
{
    using var client = await LocalChatClient.CreateAsync();
    var response = await client.GetResponseAsync([
        new(ChatRole.User, "Your question here")
    ]);
}
catch (ExecutionProviderException ex)
{
    // GPU/provider-specific error (no CUDA, DirectML not available, etc.)
    Console.WriteLine($"Provider error: {ex.Message}");
}
catch (ModelCapacityExceededException ex)
{
    // Prompt/response too long for model's context window
    Console.WriteLine($"Capacity error: {ex.Message}");
    // Solution: use a larger model or truncate the prompt
}
catch (InvalidOperationException ex)
{
    // General operation error (model not found, download failed, etc.)
    Console.WriteLine($"Operation error: {ex.Message}");
}

Observability

LocalChatClient emits generation lifecycle diagnostics through ActivitySource and Meter, both named ElBruno.LocalLLMs.

using ElBruno.LocalLLMs.Diagnostics;

builder.Services.AddOpenTelemetry()
    .WithTracing(tracing => tracing.AddSource(LocalLLMsInstrumentation.ActivitySourceName))
    .WithMetrics(metrics => metrics.AddMeter(LocalLLMsInstrumentation.MeterName));

By default, telemetry excludes prompt and completion text. Opt in only when you want content attached:

var options = new LocalLLMsOptions
{
    CaptureTelemetryContent = true
};

See docs/observability.md for the lifecycle event contract, metric names, and Aspire wiring notes, and docs/cancellation.md for voice barge-in cancellation behavior.

Local Decisions

Sometimes you don't need prose โ€” you need a decision. ElBruno.LocalLLMs.Decisions runs a System One model that returns a typed answer with its full probability distribution in a single forward pass, fast enough to call on every request.

dotnet add package ElBruno.LocalLLMs.Decisions

The model runs in-process on ONNX Runtime โ€” no server to start and no Python. The weights are downloaded from elbruno/laya-onnx on first use (about 800 MB) and cached afterwards.

Ask several questions at once โ€” they share one forward pass, so four cost about what one costs:

using var client = new LayaOnnxDecisionClient();

DecisionResult result = await client.EvaluateAsync(
    new DecisionRequest("My invoice charged me twice and I want my money back.")
        .Choose("team", new Dictionary<string, string?>
        {
            ["billing"]   = "Payment, invoice and refund problems",
            ["technical"] = "Bugs, outages and API errors",
            ["sales"]     = "Pricing, upgrades and new purchases"
        })
        .Score("urgency", new[] { "no rush", "normal", "urgent", "critical" })
        .Ask("refund", "The customer is asking for a refund."));

ChoiceResult team = result.Choice("team");

// Gate on the probability so ambiguous tickets reach a human
// instead of being confidently misrouted.
string route = team.ChoiceOrNull(0.6) ?? "human-review";

Console.WriteLine(route);                                  // billing
Console.WriteLine(result.Score("urgency").Score);          // 1.67 of 3
Console.WriteLine(result.Probability("refund").IsTrue);    // True

โš ๏ธ Laya's public checkpoints report uncalibrated confidence. Fit DecisionThreshold against your own labelled examples before relying on a boolean verdict. Answers whose temperature bucket the checkpoint got wrong are flagged with CalibrationClamped โ€” see the Local Decisions Guide and the DecisionCalibration sample.

Cache Management

Inspect and manage the local model cache programmatically:

// Remove a model from the cache (no-op if not cached)
await LocalChatClient.DeleteModelFromCacheAsync(KnownModels.Phi35MiniInstruct);

// Or use a custom cache directory
await LocalChatClient.DeleteModelFromCacheAsync(
    KnownModels.Phi35MiniInstruct,
    cacheDirectory: @"D:\my-models");

// Get cached size in bytes for one model (0 if not downloaded)
long bytes = LocalChatClient.GetModelCacheSize(KnownModels.Phi35MiniInstruct);
Console.WriteLine($"Cached: {bytes / 1024 / 1024:N0} MB");

// List all cached models with size and last-modified date
var cached = LocalChatClient.ListCachedModels();
foreach (var repo in cached)
    Console.WriteLine($"{repo.LocalDirectory}  {repo.TotalSizeBytes / 1024 / 1024:N0} MB  {repo.LastModified:yyyy-MM-dd}");

// Same APIs available on LocalVisionChatClient for vision models
await LocalVisionChatClient.DeleteModelFromCacheAsync(KnownModels.Fara15_9B);
long visionBytes = LocalVisionChatClient.GetModelCacheSize(KnownModels.Fara15_9B);

The default cache directory is %LOCALAPPDATA%/ElBruno/LocalLLMs/models (Windows) or ~/.local/share/ElBruno/LocalLLMs/models (Linux/macOS).

These operations delegate to ElBruno.HuggingFace.Downloader which provides the underlying DeleteCachedFilesAsync, GetCachedSize, and ListCachedRepos implementation.

Troubleshooting

GPU not working? Use ExecutionProvider.Cpu explicitly. See GPU Setup Validation.

Out of memory? Try a smaller model:

var options = new LocalLLMsOptions
{
    Model = KnownModels.Qwen25_05BInstruct  // 0.5B instead of 3.8B
};

Model download fails?

  • Check your internet connection
  • For private HuggingFace models, set the HF_TOKEN environment variable

For detailed troubleshooting, see docs/troubleshooting-guide.md.

Supported Models

Tier Model Parameters ONNX ID
โšช Tiny TinyLlama-1.1B-Chat 1.1B โœ… Native tinyllama-1.1b-chat
โšช Tiny SmolLM2-1.7B-Instruct 1.7B โœ… Native smollm2-1.7b-instruct
โšช Tiny Qwen2.5-0.5B-Instruct 0.5B โœ… Native qwen2.5-0.5b-instruct
โšช Tiny Qwen2.5-1.5B-Instruct 1.5B โœ… Native qwen2.5-1.5b-instruct
โšช Tiny Gemma-2B-IT 2B โœ… Native gemma-2b-it
โšช Tiny Gemma-4-E2B-IT 5.1B (2B active) ๐Ÿ”„ Convert gemma-4-e2b-it
โšช Tiny StableLM-2-1.6B-Chat 1.6B โœ… Native stablelm-2-1.6b-chat
๐ŸŸข Small Phi-3.5 mini instruct 3.8B โœ… Native phi-3.5-mini-instruct
๐ŸŸข Small Qwen2.5-3B-Instruct 3B โœ… Native qwen2.5-3b-instruct
๐ŸŸข Small Llama-3.2-3B-Instruct 3B โœ… Native llama-3.2-3b-instruct
๐ŸŸข Small Gemma-2-2B-IT 2B โœ… Native gemma-2-2b-it
๐ŸŸข Small Gemma-4-E4B-IT 8B (4B active) ๐Ÿ”„ Convert gemma-4-e4b-it
๐ŸŸก Medium Qwen2.5-7B-Instruct 7B โœ… Native qwen2.5-7b-instruct
๐ŸŸก Medium Qwen2.5-Coder-7B-Instruct 7B โœ… Native qwen2.5-coder-7b-instruct
๐ŸŸก Medium Llama-3.1-8B-Instruct 8B โœ… Native llama-3.1-8b-instruct
๐ŸŸก Medium Mistral-7B-Instruct-v0.3 7B โœ… Native mistral-7b-instruct-v0.3
๐ŸŸก Medium Gemma-2-9B-IT 9B โœ… Native gemma-2-9b-it
๐ŸŸก Medium Gemma-4-12B-IT 12B ๐Ÿ”„ Convert gemma-4-12b-it
๐ŸŸก Medium Phi-4 14B โœ… Native phi-4
๐ŸŸก Medium DeepSeek-R1-Distill-Qwen-14B 14B โœ… Native deepseek-r1-distill-qwen-14b
๐ŸŸก Medium Mistral-Small-24B-Instruct 24B โœ… Native mistral-small-24b-instruct
๐Ÿ”ด Large Qwen2.5-14B-Instruct 14B โœ… Native qwen2.5-14b-instruct
๐Ÿ”ด Large Qwen2.5-32B-Instruct 32B โœ… Native qwen2.5-32b-instruct
๐Ÿ”ด Large Llama-3.3-70B-Instruct 70B โœ… ONNX llama-3.3-70b-instruct
๐Ÿ”ด Large Mixtral-8x7B-Instruct-v0.1 8x7B โœ… Native mixtral-8x7b-instruct-v0.1
๐Ÿ”ด Large DeepSeek-R1-Distill-Llama-70B 70B โœ… Native deepseek-r1-distill-llama-70b
๐Ÿ”ด Large Command-R (35B) 35B โœ… Native command-r-35b
๐Ÿ”ด Large Gemma-4-26B-A4B-IT 25.2B (3.8B active) ๐Ÿ”„ Convert gemma-4-26b-a4b-it
๐Ÿ”ด Large Gemma-4-31B-IT 30.7B ๐Ÿ”„ Convert gemma-4-31b-it
๐ŸŸฃ Next-Gen Qwen3-14B-Instruct 14.77B โœ… Native qwen3-14b-instruct
๐Ÿง  GPT-OSS GPT-OSS 20B (CPU INT4) 21B (3.6B active, MoE) โœ… Native gpt-oss-20b
๐Ÿง  GPT-OSS GPT-OSS 20B (CUDA INT4) 21B (3.6B active, MoE) โœ… Native gpt-oss-20b-cuda
๐Ÿค– Agentic MagenticBrain ~14.77B โœ… Native magentic-brain
๐Ÿ‘๏ธ VLM Fara 1.5-9B ~9.4B โœ… Native fara-1.5-9b

๐Ÿ”„ Convert = Use the conversion scripts in scripts/ to export ONNX locally before running the model.

ยน MagenticBrain ONNX: Native ONNX hosted at elbruno/MagenticBrain-onnx (INT4 quantized). Auto-downloads when EnsureModelDownloaded=true.

ยฒ Fara 1.5-9B ONNX: elbruno/Fara1.5-9B-onnx now includes the validated multimodal package (qwen3vl-vision.onnx, qwen3vl-embedding.onnx, patched genai_config.json, and ORT-compatible processor_config.json). See ONNX Conversion โ€” Fara.

ยณ GPT-OSS 20B: Apache-2.0, from the official onnxruntime/gpt-oss-20b-onnx repository. The CPU INT4 variant is a ~12 GB download, and because GPT-OSS is a mixture-of-experts model, CPU inference is slow โ€” prefer gpt-oss-20b-cuda with Microsoft.ML.OnnxRuntimeGenAI.Cuda when a GPU is available. GPT-OSS reasons before answering; that chain-of-thought is filtered out and never surfaced, per the model card. Reasoning depth is controlled by LocalLLMsOptions.ReasoningEffort.

Fine-Tuned Models

Pre-trained variants optimized for specific tasks. A fine-tuned 0.5B model often matches or exceeds a base 1.5B on its specialized task.

Model Size Task HuggingFace ID
Qwen2.5-0.5B-ToolCalling ~1 GB Tool/function calling elbruno/Qwen2.5-0.5B-LocalLLMs-ToolCalling
Qwen2.5-0.5B-RAG ~1 GB RAG with citations elbruno/Qwen2.5-0.5B-LocalLLMs-RAG
Qwen2.5-0.5B-Instruct ~1 GB General-purpose elbruno/Qwen2.5-0.5B-LocalLLMs-Instruct

See the Supported Models Guide for detailed model cards, performance benchmarks, and selection guidance.

Samples

Sample Description
HelloChat Minimal console chat
StreamingChat Token-by-token streaming
MultiModelChat Switch models at runtime
DependencyInjection ASP.NET Core DI registration
ToolCallingAgent Function calling and tool use
FineTunedToolCalling Fine-tuned model for improved tool calling
RagChatbot RAG pipeline with document retrieval
ZeroCloudRag Zero-cloud RAG pipeline with real local embeddings and LLM inference
BitNetChat BitNet 1.58-bit model chat completion
BitNetPerformance Performance benchmark: BitNet vs ONNX models
MagenticBrainAgent Multi-agent orchestration loop using Qwen3/MagenticBrain
FaraVisionAgent Vision-language model (Fara 1.5-9B) image+text inference
GptOssChat GPT-OSS 20B chat, streaming, reasoning effort, and tool calling
MagenticUIServer ASP.NET Core + SignalR multi-agent server (FileSurfer, WebFetcher, Coder)
ConsoleAppDemo Interactive console application
LocalDecisions Support-ticket triage with a local System One decision model
DecisionCalibration Calibration clamping and fp16 batch variance, made visible

๐ŸŒ Reference App: ElBruno.MagenticUI โ€” full Blazor Server port of microsoft/magentic-ui running locally with this library.

Requirements

  • .NET 10.0
  • CPU (default), NVIDIA GPU (CUDA), or Windows GPU (DirectML)
  • ~2-8 GB disk space per model (depending on size and quantization)

Building from Source

git clone https://github.com/elbruno/ElBruno.LocalLLMs.git
cd ElBruno.LocalLLMs
dotnet restore ElBruno.LocalLLMs.slnx
dotnet build ElBruno.LocalLLMs.slnx
dotnet test ElBruno.LocalLLMs.slnx

Run integration tests (downloads real models โ€” requires internet):

RUN_INTEGRATION_TESTS=true dotnet test ElBruno.LocalLLMs.slnx

Integration tests validate the full lifecycle (download โ†’ infer โ†’ cache hit โ†’ delete) for all 35 supported models. See docs/tests/README.md for details.

Documentation

๐Ÿค Contributing

Contributions are welcome! Please:

  1. Fork the repository
  2. Create a feature branch (git checkout -b feature/amazing-feature)
  3. Commit your changes (git commit -m 'Add amazing feature')
  4. Push to the branch (git push origin feature/amazing-feature)
  5. Open a Pull Request

๐Ÿ“„ License

This project is licensed under the MIT License โ€” see the LICENSE file for details.

๐Ÿ‘‹ About the Author

Hi! I'm ElBruno ๐Ÿงก, a passionate developer and content creator exploring AI, .NET, and modern development practices.

Made with โค๏ธ by ElBruno

If you like this project, consider following my work across platforms:

  • ๐Ÿ“ป Podcast: No Tienen Nombre โ€” Spanish-language episodes on AI, development, and tech culture
  • ๐Ÿ’ป Blog: ElBruno.com โ€” Deep dives on embeddings, RAG, .NET, and local AI
  • ๐Ÿ“บ YouTube: youtube.com/elbruno โ€” Demos, tutorials, and live coding
  • ๐Ÿ”— LinkedIn: @elbruno โ€” Professional updates and insights
  • ๐• Twitter: @elbruno โ€” Quick tips, releases, and tech news

๐Ÿ™ Acknowledgments

Product Compatible and additional computed target framework versions.
.NET net10.0 is compatible.  net10.0-android was computed.  net10.0-browser was computed.  net10.0-ios was computed.  net10.0-maccatalyst was computed.  net10.0-macos was computed.  net10.0-tvos was computed.  net10.0-windows was computed. 
Compatible target framework(s)
Included target framework(s) (in package)
Learn more about Target Frameworks and .NET Standard.

NuGet packages

This package is not used by any NuGet packages.

GitHub repositories

This package is not used by any popular GitHub repositories.

Version Downloads Last Updated
0.22.0 9 9/30/2026
0.21.0 131 8/11/2026
0.20.12 99 8/11/2026
0.20.11 105 8/11/2026
0.20.10 112 8/10/2026
0.20.9 119 8/2/2026
0.20.8 123 8/1/2026
0.20.7 116 8/1/2026