Lodestar.Gpu 0.2.0

Prefix Reserved
dotnet add package Lodestar.Gpu --version 0.2.0
                    
NuGet\Install-Package Lodestar.Gpu -Version 0.2.0
                    
This command is intended to be used within the Package Manager Console in Visual Studio, as it uses the NuGet module's version of Install-Package.
<PackageReference Include="Lodestar.Gpu" Version="0.2.0" />
                    
For projects that support PackageReference, copy this XML node into the project file to reference the package.
<PackageVersion Include="Lodestar.Gpu" Version="0.2.0" />
                    
Directory.Packages.props
<PackageReference Include="Lodestar.Gpu" />
                    
Project file
For projects that support Central Package Management (CPM), copy this XML node into the solution Directory.Packages.props file to version the package.
paket add Lodestar.Gpu --version 0.2.0
                    
#r "nuget: Lodestar.Gpu, 0.2.0"
                    
#r directive can be used in F# Interactive and Polyglot Notebooks. Copy this into the interactive tool or source code of the script to reference the package.
#:package Lodestar.Gpu@0.2.0
                    
#:package directive can be used in C# file-based apps starting in .NET 10 preview 4. Copy this into a .cs file before any lines of code to reference the package.
#addin nuget:?package=Lodestar.Gpu&version=0.2.0
                    
Install as a Cake Addin
#tool nuget:?package=Lodestar.Gpu&version=0.2.0
                    
Install as a Cake Tool

Lodestar

Quality Gate Status

The data-science pieces .NET has no maintained library for, at the parity of the Python library you already trust — with no Python at runtime, on .NET 10 and .NET Standard 2.0 from one package. Where .NET already has the answer, this project does not rewrite it: it tells you which library to use.

What it replaces

You arrived with an alternative in mind. Find it below: the row gives the number that settles the choice and where to go next. A ratio means this library is that many times faster; each number's machine, window and incumbent version are in the performance guide.

you would reach for to do what settles it go to
Fastenshtein, Quickenshtein edit distance Levenshtein 2.8× to 4.1× Quickenshtein, the closest, 2.9× to 33.8× Fastenshtein; allocates nothing from rapidfuzz
FuzzySharp (the Raffinert fork) fuzz.*, process.extract all four ratios 1.78× to 12.75×, allocating less on each from rapidfuzz
ML.NET FeaturizeText TF-IDF, count, hashing vectors a sparse matrix at scikit-learn semantics instead of a dense vector inside an IDataView; 5.7× to 11×, not like-for-like, since ML.NET adds character n-grams vectorization
ML.NET's binary evaluator classification metrics one call, one double: accuracy alone 87× to 322×, ML.NET's whole bundle 1.57× to 4.84× metrics
Microsoft.ML.Tokenizers a Hugging Face tokenizer reads the tokenizer.json it has no loader for; on identical ids, 1.13× to 2.57× embeddings
Accord.Statistics (archived; last release 2017), Meta.Numerics scipy.stats tests 1.18× to 5.94× Meta.Numerics on all eight shared families at n = 10,000, exact p-values where it is asymptotic hypothesis testing
Accord.Statistics, Math.NET regression inference the whole statsmodels table, VIF included, which Accord does not export and Math.NET stops short of regression inference
Cortex.TimeSeries ADF, KPSS, decomposition ADF 1.79× to 2.26×, decomposition 1.85×, KPSS level; a MacKinnon p-value where Cortex clamps at 0.01 time-series diagnostics
NumFlat, Meta.Numerics k-means Lloyd's iterations 1.52× to 3.66× NumFlat from the same centres, and runs on netstandard2.0, where NumFlat does not install scikit-learn
MAPIE or lifelines, through CSnakes or Python.NET conformal intervals, survival no C# implementation of either exists; this is one, with no Python runtime to ship conformal, survival

It is not for you if what you need is dense linear algebra (Math.NET Numerics), training a model (ML.NET, TorchSharp), or a model that exists only as a Python package (CSnakes). docs/migration/ names the .NET library for every need this project does not write, marks the ones no longer maintained, and says when calling Python is still the right answer.

Getting started

dotnet add package Lodestar.Text
using Lodestar.Text.Distances;

Levenshtein.Distance("kitten", "sitting");             // 3
Levenshtein.NormalizedSimilarity("kitten", "sitting"); // 0.5714…

Full guide: docs/guides/quickstart.md. The guides linked in the table above each start from the Python call you know. Function by function, the reference pages under docs/reference/ say what each member is for, when to prefer it to its neighbour and what the trap is; the same pages are published to the wiki, where each package's channel follows main and every release is archived under its own version.

Why not just call Python?

CSnakes and Python.NET both work, both are maintained, and for a model that only exists as a Python package they are the right answer. What they cost is a Python runtime to deploy and version alongside the application, no ahead-of-time compilation to a single artifact, and the GIL between your threads and theirs. Where a .NET library will do, that is a poor trade, and this project exists to make it an avoidable one.

Why the gap is where it is

Measured package by package, what .NET lacks is almost never the computation and almost always the apparatus around it: the tokenizer loader and not its encoder, the regression's inference table and not its coefficients, the time-series diagnostics and not the forecast, sparse decomposition and not dense. Decision 0004 decides each case, and its reading is what the table above rests on:

  • Ordinary least squares is everywhere; its inference is not. Meta.Numerics (MS-PL) reports the standard error, the interval and the F test and stops there; the commercial Numerics.NET exports the whole table.
  • Split conformal prediction — an interval instead of a point, a set instead of a class, with a finite-sample coverage guarantee — had no C# implementation at all in the survey behind #441. The guarantee assumes exchangeable calibration and test data, which the guide leads with.
  • Right-censored survival was the largest void #442 surveyed. scikit-survival is the nearest reference in any language and is refused on its licence, not its capability (decision 0002).
  • Distances, embeddings and fuzzy matching are here for pipeline coherence rather than because .NET is empty — it is not, and the first table says by how much.

The netstandard2.0 build reaches .NET Framework 4.6.1+, Mono, Xamarin and Unity with the same public API (decision 0001).

Measured against the .NET incumbents

Every package with an in-process incumbent is benchmarked against the .NET library a reader would reach for, and both sides are checked to return the same answers before either is timed — bench/README.md has the harness and the agreement checks. The rows the first table does not carry:

package incumbent how it reads
Lodestar.Text, Bm25Index LuceneSharp.Core Same ranking. The query is 1.6× faster at 1,000 documents and 1.4× slower at 20,000; from raw text Lucene is ahead, 2.1× to 2.7× (performance)
Lodestar.Decomposition ML.NET ProjectToPrincipalComponents; NumFlat; Meta.Numerics Not like-for-like against ML.NET, whose PCA is dense, centred and reports no eigenvalue. The explained variance is 0.83× to 1.36× NumFlat (net8.0 only) and 36.7× to 809× Meta.Numerics, which refuses a matrix wider than it is tall (performance)
Lodestar.Cluster NumFlat, Dbscan, Aglomera DBSCAN 2.47× to 10.92×, agglomerative clustering 24× to 590× (performance)
Lodestar.Preprocessing ML.NET's splitters, normalizers and encoders Ahead on every row where ML.NET's lazy result is read back, 1.46× to 191×; the lazy call alone is cheaper on the larger splits and the 20,000-row one-hot fit (performance)
Lodestar.Stats.Regression Accord.Statistics; Math.NET Numerics The GLM is level with Accord at 200 rows and 1.4× behind at 2,000; weighted and generalized least squares are level or ahead of Math.NET while computing the whole table (performance)
Lodestar.Onnx, Lodestar.Extensions.AI, Lodestar.Extensions.MathNet — Nothing to beat. Each calls or adapts another library, so what it could be slower than is its own conversion
Lodestar.Extensions.VectorData the Microsoft.Extensions.VectorData connectors Not measured. Of the connectors surveyed, the ones implementing hybrid search are clients of a server, which an in-process store does not race
Lodestar.Gpu — Measured against this repository's own CPU paths, and each kernel ships only where it passed that gate (performance)

Parity with the Python reference

Conformance is proven, not assumed. Every algorithm replays reference values frozen from the canonical Python library — rapidfuzz, jellyfish, textdistance, difflib, scikit-learn, scipy, statsmodels, lifelines, MAPIE, nltk, HuggingFace tokenizers, sentencepiece, numpy, ONNX Runtime — into tests/oracles/*.json, compared at 1e-9 for floats and exactly for strings. Python is a development dependency only. docs/equivalence.md maps each Python call to its C# counterpart, and every deliberate divergence is a record in docs/decisions/.

Developing

dotnet build Lodestar.slnx -c Release   # both target frameworks; warnings are errors
dotnet test Lodestar.slnx -c Release    # replays the oracles, on both

The project follows GitHub flow: main is always releasable, and every change arrives through a short-lived branch and a pull request. Branch conventions, the definition of done, the oracle-validation procedure and the analyzer policy are in CONTRIBUTING.md; release history is in CHANGELOG.md.

A runnable sample, consuming the packages exactly as you would:

for p in src/Lodestar.Abstractions src/Lodestar.Text src/Lodestar.Embeddings \
        src/Lodestar.Fuzzy src/Lodestar.Metrics src/Lodestar.Conformal \
        src/Lodestar.Decomposition src/Lodestar.Onnx src/Lodestar.Extensions.AI \
        src/Lodestar.Extensions.MathNet src/Lodestar.Extensions.VectorData \
        src/Lodestar.Cluster src/Lodestar.Preprocessing \
        src/Lodestar.Stats src/Lodestar.Stats.Regression src/Lodestar.Stats.TimeSeries \
        src/Lodestar.Survival \
        src/Lodestar.Gpu; do
  dotnet pack "$p" -c Release -o ./artifacts
done
NUGET_PACKAGES=$(mktemp -d) dotnet run -c Release --project samples/Lodestar.Sample

The isolated NUGET_PACKAGES is not decoration: the global packages folder is consulted ahead of any source, so a machine that has ever restored a published Lodestar.* at one of these versions runs the sample against that rather than against what pack just produced — see CONTRIBUTING.md's Definition of done. On PowerShell the same isolation is two lines, $env:NUGET_PACKAGES = (New-Item -ItemType Directory -Path (Join-Path $env:TEMP (New-Guid))).FullName before the dotnet run, and Remove-Item Env:NUGET_PACKAGES after it.

Structure

What each package holds, and which it depends on, is CLAUDE.md's architecture table; which document carries which fact is its Where a fact belongs.

Lodestar.slnx
├── src/Lodestar.Abstractions/              CsrMatrix and SparseNorm — the sparse primitive the others share (no dependencies)
├── src/Lodestar.Text/                      distances, similarity, tokenizers, vectorizers, stemmers
├── src/Lodestar.Embeddings/                sub-word tokenizers, pooling, SIMD kNN (no dependencies)
├── src/Lodestar.Fuzzy/                     fuzz.*, process.extract, deduplication
├── src/Lodestar.Metrics/                   confusion matrix, precision/recall/F1, report, ROC-AUC
├── src/Lodestar.Conformal/                 split conformal intervals and prediction sets (no dependencies)
├── src/Lodestar.Decomposition/             truncated SVD, NMF, the Householder QR, and PCA explained variance
├── src/Lodestar.Cluster/                   k-means by Lloyd's algorithm over a row-major span
├── src/Lodestar.Preprocessing/             feature scaling fitted on arrays and applied to spans
├── src/Lodestar.Stats/                     classical hypothesis tests, at scipy.stats parity (no dependencies)
├── src/Lodestar.Stats.Regression/          ordinary, weighted and generalized least squares with the inference table
├── src/Lodestar.Stats.TimeSeries/          autocorrelation, Ljung-Box, ADF, KPSS and seasonal decomposition
├── src/Lodestar.Survival/                  Kaplan-Meier, Nelson-Aalen and the log-rank test, right-censored
├── src/Lodestar.Onnx/                      ONNX inference — satellite, carries Microsoft.ML.OnnxRuntime (decision 0003)
├── src/Lodestar.Gpu/                       ILGPU kernels — satellite, the one package on net10.0;netstandard2.1
├── src/Lodestar.Extensions.AI/             interop: the ONNX embedding path behind IEmbeddingGenerator
├── src/Lodestar.Extensions.MathNet/        interop: CsrMatrix to and from Math.NET's sparse matrix
├── src/Lodestar.Extensions.VectorData/     interop: an in-process Microsoft.Extensions.VectorData store with hybrid search
├── tests/                                  xUnit: two projects per package — net10.0, and a mirror linking the same sources against netstandard2.0
├── tests/oracles/                          frozen JSON corpora (generated from Python) + a synthetic ONNX model
├── bench/Lodestar.Text.Benchmarks/         BenchmarkDotNet: every non-netstandard benchmark, whatever package it measures
├── bench/Lodestar.NetStandard.Benchmarks/  the netstandard2.0 assemblies, measured on the same host
├── tools/generate_oracles.py               reference generation
├── Directory.Build.props                   (root); src|tests/Directory.Packages.props (central package management)
├── src/*/Version.props                     one version per publishable package (decision 0001)
├── docs/                                   guides, equivalence table, decision log
├── docs/reference/<package>/               one reference entry per exported type and public method
└── docs/wiki-map.json                      which page ships with which package, and which namespaces the reference gate enforces

Publishing

Eighteen NuGet packages are produced: Lodestar.Abstractions, Lodestar.Text, Lodestar.Embeddings, Lodestar.Fuzzy, Lodestar.Metrics, Lodestar.Conformal, Lodestar.Decomposition, Lodestar.Cluster, Lodestar.Preprocessing, Lodestar.Stats, Lodestar.Stats.Regression, Lodestar.Stats.TimeSeries, Lodestar.Survival, Lodestar.Onnx, Lodestar.Gpu, Lodestar.Extensions.AI, Lodestar.Extensions.MathNet and Lodestar.Extensions.VectorData. Thirteen are core tier and carry no external dependency — decisions/0003. Lodestar.Onnx and Lodestar.Gpu are the two satellites, each carrying the one dependency that is its whole reason to be a package; the three Lodestar.Extensions.* are the interop tier, which decisions/0003 allows a dependency a core package refused, because converting to a foreign type is not computing with it. Each versions and releases on its own: shared metadata (license, README, repository) lives in Directory.Build.props, while the version is declared per project in src/<Package>/Version.props. Lodestar.Fuzzy depends on Lodestar.Text as a published package, not as a project reference — see docs/decisions/0001.

main carries the next revision rather than the published one. A package released at 0.2.0 reads 0.2.1 in its Version.props, so every branch packs and every sample restores a number nuget.org does not hold — a version on the feed is immutable, and a collision would make two different assemblies answer to one identity. A feature pull request therefore never touches Version.props: it lands on a number already ahead of the feed.

To cut a release, set that file to the version being cut — the number main already carries when the release is a revision, a larger one when the change earns a minor or a major — and land it on main. Add the entry under the package's heading in CHANGELOG.md, in the shape CONTRIBUTING.md's item 7 sets. Then tag. Afterwards, close the release issue by bumping the revision again, which puts main back ahead of the feed.

GitHub Packages (no nuget.org account needed — uses GitHub's automatic token). Bump the version, then tag it with the package name. The release workflow packs and publishes that package alone:

# 1. src/Lodestar.Fuzzy/Version.props declares the version being cut — already true
#    for a revision; edit, commit and merge to main for a minor or a major
# 2. tag the released version — <PackageId>/v<Version>
git tag Lodestar.Fuzzy/v0.3.0
git push origin Lodestar.Fuzzy/v0.3.0

The tag does not set the version; it names which declared version to release. The workflow refuses the job if the tag and Version.props disagree. Repository-wide v* tags are retired — there is no single version left for one to designate.

Step 1 is what decides the number. A revision needs no edit — main already carries the next one — but a minor or a major does, and tagging before that edit gives a tag the workflow refuses, because it disagrees with the version Version.props declares. Re-tagging a version the feed already holds is rejected rather than absorbed: the workflows do not pass --skip-duplicate, which used to report that case as a successful release that shipped nothing. That a declared version is still off the feed is checked directly in CI by tools/check_version_floor.py.

To consume them, add a source pointing at the owner's feed (with a GitHub token that has read:packages):

dotnet nuget add source "https://nuget.pkg.github.com/CyrilB1531/index.json" \
  --name github --username CyrilB1531 --password <GITHUB_TOKEN>
dotnet add package Lodestar.Text

nuget.org uses Trusted Publishing (OIDC, no stored key): run the Publish to nuget.org workflow from the Actions tab, choosing the package and confirming its version. By hand, with an API key, one package at a time:

dotnet pack src/Lodestar.Text -c Release -o artifacts
dotnet nuget push "artifacts/Lodestar.Text.*.nupkg" \
  --source https://api.nuget.org/v3/index.json --api-key <KEY>

License

Apache-2.0. See NOTICE and THIRD-PARTY-NOTICES.md for attributions. The license choice and the code-provenance rule are documented in docs/decisions/0002-provenance-and-the-allowed-references.md.

This repository is not legal advice.

Product Compatible and additional computed target framework versions.
.NET net5.0 was computed.  net5.0-windows was computed.  net6.0 was computed.  net6.0-android was computed.  net6.0-ios was computed.  net6.0-maccatalyst was computed.  net6.0-macos was computed.  net6.0-tvos was computed.  net6.0-windows was computed.  net7.0 was computed.  net7.0-android was computed.  net7.0-ios was computed.  net7.0-maccatalyst was computed.  net7.0-macos was computed.  net7.0-tvos was computed.  net7.0-windows was computed.  net8.0 was computed.  net8.0-android was computed.  net8.0-browser was computed.  net8.0-ios was computed.  net8.0-maccatalyst was computed.  net8.0-macos was computed.  net8.0-tvos was computed.  net8.0-windows was computed.  net9.0 was computed.  net9.0-android was computed.  net9.0-browser was computed.  net9.0-ios was computed.  net9.0-maccatalyst was computed.  net9.0-macos was computed.  net9.0-tvos was computed.  net9.0-windows was computed.  net10.0 is compatible.  net10.0-android was computed.  net10.0-browser was computed.  net10.0-ios was computed.  net10.0-maccatalyst was computed.  net10.0-macos was computed.  net10.0-tvos was computed.  net10.0-windows was computed. 
.NET Core netcoreapp3.0 was computed.  netcoreapp3.1 was computed. 
.NET Standard netstandard2.1 is compatible. 
MonoAndroid monoandroid was computed. 
MonoMac monomac was computed. 
MonoTouch monotouch was computed. 
Tizen tizen60 was computed. 
Xamarin.iOS xamarinios was computed. 
Xamarin.Mac xamarinmac was computed. 
Xamarin.TVOS xamarintvos was computed. 
Xamarin.WatchOS xamarinwatchos was computed. 
Compatible target framework(s)
Included target framework(s) (in package)
Learn more about Target Frameworks and .NET Standard.

NuGet packages

This package is not used by any NuGet packages.

GitHub repositories

This package is not used by any popular GitHub repositories.

Version Downloads Last Updated
0.2.0 85 9/24/2026
0.1.0 112 9/10/2026