Mythosia.AI.Serving.Ollama 1.0.0

dotnet add package Mythosia.AI.Serving.Ollama --version 1.0.0
                    
NuGet\Install-Package Mythosia.AI.Serving.Ollama -Version 1.0.0
                    
This command is intended to be used within the Package Manager Console in Visual Studio, as it uses the NuGet module's version of Install-Package.
<PackageReference Include="Mythosia.AI.Serving.Ollama" Version="1.0.0" />
                    
For projects that support PackageReference, copy this XML node into the project file to reference the package.
<PackageVersion Include="Mythosia.AI.Serving.Ollama" Version="1.0.0" />
                    
Directory.Packages.props
<PackageReference Include="Mythosia.AI.Serving.Ollama" />
                    
Project file
For projects that support Central Package Management (CPM), copy this XML node into the solution Directory.Packages.props file to version the package.
paket add Mythosia.AI.Serving.Ollama --version 1.0.0
                    
#r "nuget: Mythosia.AI.Serving.Ollama, 1.0.0"
                    
#r directive can be used in F# Interactive and Polyglot Notebooks. Copy this into the interactive tool or source code of the script to reference the package.
#:package Mythosia.AI.Serving.Ollama@1.0.0
                    
#:package directive can be used in C# file-based apps starting in .NET 10 preview 4. Copy this into a .cs file before any lines of code to reference the package.
#addin nuget:?package=Mythosia.AI.Serving.Ollama&version=1.0.0
                    
Install as a Cake Addin
#tool nuget:?package=Mythosia.AI.Serving.Ollama&version=1.0.0
                    
Install as a Cake Tool

Mythosia.AI.Serving.Ollama

Before sending a prompt, an application often needs to know whether its Ollama server is reachable, which models are registered, and which are already loaded. OllamaServer supplies that management layer and lets an operator explicitly download, preload or unload a model.

The client connects to an existing server. It does not start Ollama, install the runtime, run chat requests or generate embeddings. Chat and embedding APIs remain separate; OllamaEmbeddingProvider lives in Mythosia.AI.Rag.

Version and installation

This initial release implements the shared IModelServer, IModelLifecycle and IModelDownloader contracts from Mythosia.AI.Serving.Abstractions. It targets .NET Standard 2.1 and references Newtonsoft.Json 13.0.4 without a Mythosia.AI core dependency. See the release notes.

dotnet add package Mythosia.AI.Serving.Ollama --version 1.0.0

Discover the server and its models

using System;
using System.Net.Http;
using System.Threading;
using Mythosia.AI.Serving;
using Mythosia.AI.Serving.Ollama;

using var http = new HttpClient { Timeout = TimeSpan.FromMinutes(10) };
using var cancellation = new CancellationTokenSource();
var server = new OllamaServer("http://localhost:11434", http);
IModelServer management = server;

var health = await management.GetHealthAsync(cancellation.Token);
var info = await management.GetInfoAsync(cancellation.Token);
var capabilities = await management.GetCapabilitiesAsync(cancellation.Token);
var models = await management.GetModelsAsync(cancellation.Token);

Console.WriteLine($"Ollama {info.Version}: {health.Status}");
Console.WriteLine($"Download support: {capabilities.ModelDownloading}");
foreach (var model in models)
    Console.WriteLine($"{model.Id}: registered={model.InstallationState}, loaded={model.LoadState}");

The example uses C# top-level statements. Keep the supplied HttpClient alive for the client's lifetime; the serving client neither changes nor disposes it.

Input Configuration
endpoint HTTP(S) root, or a trailing /api or /v1. Reverse-proxy path prefixes are retained. Embedded URL credentials, query strings and fragments are rejected.
httpClient Caller-owned client; configure timeout and HTTP handlers before use. Its timeout covers each HTTP request, including streamed response bodies. Use a caller cancellation deadline to bound a workflow containing several requests.
apiKey Optional third argument; sends a Bearer credential per request without changing default headers.
cancellationToken Optional on every asynchronous operation. Use a token with a deadline when a bounded wait is needed.
Operation Native endpoint Limit
Health / version GET /api/version Management health does not prove model readiness.
Inventory GET /api/tags and GET /api/ps Separate observations, not an atomic snapshot.
Preload / unload POST /api/generate with an empty prompt Requires a model compatible with that endpoint.
Download POST /api/pull Completes on terminal success followed by end of stream.
Metrics No metrics interface Unsupported by this client.

Health checks read /api/version; a healthy management API does not establish that a particular model can perform inference. Discovery and capability checks never preload a model or issue a generation request. Capability discovery probes the version and both inventory endpoints. A valid Ollama protocol observation establishes management support; it does not verify mutation permissions or compatibility with every model. Failed authentication, unreachable endpoints or malformed responses leave affected capabilities Unknown. After version discovery succeeds, a missing inventory endpoint reports model listing as Unsupported.

Request a model download or change its residency

These operations are explicit because they can consume network bandwidth, disk space or device memory. Continue the previous example with:

IModelDownloader downloader = server;
var progress = new Progress<ModelDownloadProgress>(update =>
    Console.WriteLine($"{update.Stage}: {update.CompletedBytes}/{update.TotalBytes} ({update.Artifact})"));

const string modelId = "qwen2.5:0.5b";
await downloader.DownloadModelAsync(modelId, progress, cancellation.Token);

IModelLifecycle lifecycle = server;
await lifecycle.LoadModelAsync(modelId, cancellation.Token);
// Use a separate inference client while the model is resident.
// When finished, explicitly request unload:
await lifecycle.UnloadModelAsync(modelId, cancellation.Token);

Preloading uses Ollama's documented empty /api/generate request, with no user prompt. It applies the server's default keep-alive duration. This operation requires a model compatible with that endpoint; embedding-only models are not automatically redirected to another API, and a server rejection is returned to the caller. Unloading sends the same empty request with keep_alive: 0 and does not delete model files. A completed call acknowledges the server operation; it does not guarantee indefinite residency or local execution of a remote model.

Downloads use /api/pull and report progress per artifact/layer. Missing byte counters remain null; they are not zero and are not a whole-model percentage. The task completes only after a terminal success message and the end of the response. HTTP errors, streaming error messages, invalid counters, malformed JSON and a stream ending before success all fail the task. Unknown additional response fields are allowed. Canceling the HTTP operation does not delete cached layers or guarantee that server-side work was rolled back; Ollama can reuse partial downloads.

Interpret model state accurately

Observation Meaning
ModelInstallationState.Installed Listed by /api/tags: registered in Ollama. A remote-model manifest can be registered without local model weights.
ModelLoadState.Loaded Listed by /api/ps during the observation.
ModelLoadState.Unloaded Registered model absent from /api/ps, with no explicit remote metadata.
Remote registration without a runner IsRemote == true and LoadState.Unknown; remote execution cannot be inferred from the local process list.
Runner without a matching registration InstallationState.Unknown; the lists are separate requests and server state may change between them.
IsRemote == null No explicit remote metadata was reported. A name suffix is not used to guess locality.

SizeBytes comes from the registered entry, MemoryBytes from /api/ps's size_vram, and ContextLength from the running entry. MemoryBytes therefore describes reported VRAM usage, not total system memory. Model identifiers, required response shapes and known numeric fields are validated; duplicate identifiers within either list are rejected. Unused response fields are accepted for forward compatibility.

Failures and cancellation

All methods accept a cancellation token and pass it through sending the request and reading the body. Caller cancellation is propagated as OperationCanceledException. Other request failures use ServingException, with StatusCode when available and a FailureKind distinguishing transport, timeout and invalid response failures. Exception messages exclude server response bodies, API keys and server-provided error text.

GetHealthAsync converts network/server failures into ServerHealth (Unreachable, Unauthorized, NotReady or Unexpected) but still propagates caller cancellation. GetVersionAsync and IsHealthyAsync are convenience methods on the concrete client.

Validation and scope

The offline suite exercises installed/loaded/remote state separation, shared-client authentication, malformed/duplicate responses, lifecycle acknowledgement, streaming completion/errors, timeout and cancellation during HTTP body reads. Run it with:

dotnet test --project tests/Mythosia.AI.Serving.Ollama.Tests/Mythosia.AI.Serving.Ollama.Tests.csproj --configuration Release

Separate live checks exercised Ollama 0.34.4 with qwen2.5:0.5b on one NVIDIA A40: health/version/inventory, download completion including a cold download, preload/unload observations, sanitized errors and pre-cancellation. Cancellation after partial download progress was verified separately using a larger model. A small prompt sent through native HTTP also received an inference response; that check did not test a Mythosia chat adapter.

An initial pull failed with a server-declared error whose underlying cause was not established; subsequent library retry and cold-download checks passed. The result does not establish a minimum supported version or compatibility with every model, server configuration or future release. Cancellation proves that the local call stops; it does not prove remote rollback.

Model creation/copy/deletion, publishing, process hosting and metrics are outside this release. See the serving guide for the cross-runtime support and validation matrix.

Protocol references: registered models, running models, model preloading and keep-alive, stream errors, and Ollama's API specification.

Product Compatible and additional computed target framework versions.
.NET net5.0 was computed.  net5.0-windows was computed.  net6.0 was computed.  net6.0-android was computed.  net6.0-ios was computed.  net6.0-maccatalyst was computed.  net6.0-macos was computed.  net6.0-tvos was computed.  net6.0-windows was computed.  net7.0 was computed.  net7.0-android was computed.  net7.0-ios was computed.  net7.0-maccatalyst was computed.  net7.0-macos was computed.  net7.0-tvos was computed.  net7.0-windows was computed.  net8.0 was computed.  net8.0-android was computed.  net8.0-browser was computed.  net8.0-ios was computed.  net8.0-maccatalyst was computed.  net8.0-macos was computed.  net8.0-tvos was computed.  net8.0-windows was computed.  net9.0 was computed.  net9.0-android was computed.  net9.0-browser was computed.  net9.0-ios was computed.  net9.0-maccatalyst was computed.  net9.0-macos was computed.  net9.0-tvos was computed.  net9.0-windows was computed.  net10.0 was computed.  net10.0-android was computed.  net10.0-browser was computed.  net10.0-ios was computed.  net10.0-maccatalyst was computed.  net10.0-macos was computed.  net10.0-tvos was computed.  net10.0-windows was computed. 
.NET Core netcoreapp3.0 was computed.  netcoreapp3.1 was computed. 
.NET Standard netstandard2.1 is compatible. 
MonoAndroid monoandroid was computed. 
MonoMac monomac was computed. 
MonoTouch monotouch was computed. 
Tizen tizen60 was computed. 
Xamarin.iOS xamarinios was computed. 
Xamarin.Mac xamarinmac was computed. 
Xamarin.TVOS xamarintvos was computed. 
Xamarin.WatchOS xamarinwatchos was computed. 
Compatible target framework(s)
Included target framework(s) (in package)
Learn more about Target Frameworks and .NET Standard.

NuGet packages

This package is not used by any NuGet packages.

GitHub repositories

This package is not used by any popular GitHub repositories.

Version Downloads Last Updated
1.0.0 101 9/28/2026

v1.0.0 introduces OllamaServer with health/version discovery, registered/running inventory, explicit preload/unload and NDJSON downloads with terminal-success validation. Shared Serving contracts, caller-owned HTTP, request-local authentication, sanitized errors and cancellation are included. Live checks used Ollama 0.34.4 and qwen2.5:0.5b; see full notes for coverage, the initial pull failure and compatibility limits. Cancellation does not guarantee server-side rollback. Full notes: https://github.com/AJ-comp/Mythosia.AI/blob/main/src/serving/Mythosia.AI.Serving.Ollama/RELEASE_NOTES.md#v100