Gravicode.FaissNet.Gpu 1.0.0

dotnet add package Gravicode.FaissNet.Gpu --version 1.0.0
                    
NuGet\Install-Package Gravicode.FaissNet.Gpu -Version 1.0.0
                    
This command is intended to be used within the Package Manager Console in Visual Studio, as it uses the NuGet module's version of Install-Package.
<PackageReference Include="Gravicode.FaissNet.Gpu" Version="1.0.0" />
                    
For projects that support PackageReference, copy this XML node into the project file to reference the package.
<PackageVersion Include="Gravicode.FaissNet.Gpu" Version="1.0.0" />
                    
Directory.Packages.props
<PackageReference Include="Gravicode.FaissNet.Gpu" />
                    
Project file
For projects that support Central Package Management (CPM), copy this XML node into the solution Directory.Packages.props file to version the package.
paket add Gravicode.FaissNet.Gpu --version 1.0.0
                    
#r "nuget: Gravicode.FaissNet.Gpu, 1.0.0"
                    
#r directive can be used in F# Interactive and Polyglot Notebooks. Copy this into the interactive tool or source code of the script to reference the package.
#:package Gravicode.FaissNet.Gpu@1.0.0
                    
#:package directive can be used in C# file-based apps starting in .NET 10 preview 4. Copy this into a .cs file before any lines of code to reference the package.
#addin nuget:?package=Gravicode.FaissNet.Gpu&version=1.0.0
                    
Install as a Cake Addin
#tool nuget:?package=Gravicode.FaissNet.Gpu&version=1.0.0
                    
Install as a Cake Tool

FAISS.Net.Gpu

GPU acceleration for FAISS.Net, via ILGPU. Drop-in replacements for the flat indexes that run on CUDA or OpenCL.

using Faiss.Net.Gpu;

using var index = new IndexFlatL2Gpu(dimension: 128);
index.Add(database);
var results = index.Search(queries, k: 10);   // identical API, identical results

Why brute force belongs on a GPU

Exhaustive search is the ideal GPU workload: every candidate is independent, the arithmetic is a pure multiply-add chain, and the access pattern is a straight sequential read. It is also memory-bandwidth-bound, which is precisely where a GPU has an order-of-magnitude advantage — so the speedup is largest exactly where the CPU path hurts most.

Two kernels run per query chunk. The first fills a chunk × ntotal distance matrix, one thread per (query, vector) pair. The second selects the top k per query on the device, so only chunk × k results cross the bus instead of the whole matrix — the transfer, not the arithmetic, is what would otherwise dominate.

Query batches are chunked automatically so the distance matrix stays inside a configurable device-memory budget, which lets a database far larger than device memory still be searched in one call.

It runs without a GPU

With no CUDA or OpenCL device present, ILGPU falls back to a CPU accelerator and the same kernels run. Code written against a GPU index keeps working on a machine without one — just without the speedup.

if (StandardGpuResources.IsGpuAvailable())
    Console.WriteLine(string.Join("\n", StandardGpuResources.EnumerateDevices()));

using var resources = new StandardGpuResources();
Console.WriteLine(resources.DeviceName);
Console.WriteLine(resources.IsHardwareAccelerated);   // false on the CPU fallback

Check IsHardwareAccelerated before drawing conclusions from a benchmark.

Moving indexes between CPU and GPU

var cpu = new IndexFlatL2(128);
cpu.Add(database);

using var gpu = GpuIndexFlat.FromCpu(cpu);   // faiss.index_cpu_to_gpu
var back = gpu.ToCpu();                      // faiss.index_gpu_to_cpu

Multi-GPU

One replica per device, with queries split across them — the pattern IndexReplicas exists for:

var replicas = new IndexReplicas(128);
foreach (var device in StandardGpuResources.ForEachGpu())
    replicas.AddReplica(new IndexFlatL2Gpu(128, device));
replicas.Add(database);

Scope and honesty

  • Flat indexes only. GPU IVF and PQ are on the roadmap, not in this release.
  • Validated against ILGPU's CPU fallback accelerator. The kernels are covered by tests that assert results identical to the CPU library, but the backend has not yet been benchmarked on real CUDA hardware — so this package makes no speed claims.
  • Vectors live in device memory. Add re-uploads the database, so build once and query many times.

Documentation


MIT licensed. Built by Gravicode Studios, led by Kang Fadhil.

Product Compatible and additional computed target framework versions.
.NET net10.0 is compatible.  net10.0-android was computed.  net10.0-browser was computed.  net10.0-ios was computed.  net10.0-maccatalyst was computed.  net10.0-macos was computed.  net10.0-tvos was computed.  net10.0-windows was computed. 
Compatible target framework(s)
Included target framework(s) (in package)
Learn more about Target Frameworks and .NET Standard.

NuGet packages

This package is not used by any NuGet packages.

GitHub repositories

This package is not used by any popular GitHub repositories.

Version Downloads Last Updated
1.0.0 107 9/1/2026

First release. GPU flat search with on-device top-k selection and automatic query chunking against a device-memory budget.