FastCompute 0.8.1

dotnet add package FastCompute --version 0.8.1
                    
NuGet\Install-Package FastCompute -Version 0.8.1
                    
This command is intended to be used within the Package Manager Console in Visual Studio, as it uses the NuGet module's version of Install-Package.
<PackageReference Include="FastCompute" Version="0.8.1" />
                    
For projects that support PackageReference, copy this XML node into the project file to reference the package.
<PackageVersion Include="FastCompute" Version="0.8.1" />
                    
Directory.Packages.props
<PackageReference Include="FastCompute" />
                    
Project file
For projects that support Central Package Management (CPM), copy this XML node into the solution Directory.Packages.props file to version the package.
paket add FastCompute --version 0.8.1
                    
#r "nuget: FastCompute, 0.8.1"
                    
#r directive can be used in F# Interactive and Polyglot Notebooks. Copy this into the interactive tool or source code of the script to reference the package.
#:package FastCompute@0.8.1
                    
#:package directive can be used in C# file-based apps starting in .NET 10 preview 4. Copy this into a .cs file before any lines of code to reference the package.
#addin nuget:?package=FastCompute&version=0.8.1
                    
Install as a Cake Addin
#tool nuget:?package=FastCompute&version=0.8.1
                    
Install as a Cake Tool

FastCompute

FastCompute is a strongly named .NET 8 library for fast array processing. It provides one API for single-threaded CPU, multi-threaded CPU, SIMD, and ILGPU execution and supports float, double, and int arrays. It has no image dependency; image processing ships in the separate FastCompute.ImageProcessing package.

Install

dotnet add package FastCompute --version 0.8.1

The consuming project must target .NET 8 or a compatible later framework.

Quick start

using FastCompute;

float[] source = [0.0f, 0.5f, 1.0f, 1.5f];

float[] result = source
    .AsCompute()
    .Select(value => value * 2.0f)
    .SelectInPlace(value => value + 1.0f)
    .Select(value => ComputeMath.Sin(value))
    .ToArray();

ComputeMath describes mathematical operations that every compatible backend can execute. Its use does not select or force GPU execution. Nothing is executed before ToArray: FastCompute fuses the three selectors into one expression, selects a backend once, and avoids intermediate managed arrays.

Lazy optimized pipelines

AsCompute creates an immutable lazy pipeline for float[], double[], or int[]:

ComputePipeline<float> pipeline = source
    .AsCompute()
    .Select(value => value * 2.0f)
    .SelectInPlace(value => value + 1.0f)
    .Select(value => ComputeMath.Clamp(value, 0.0f, 1.0f));

// Planning, optimization, backend selection, and execution happen here.
float[] result = pipeline.ToArray();

Consecutive selectors are substituted into one Map expression and execute as one Parallel/SIMD pass or one GPU Map kernel. A pipeline can record one binary Zip node; selectors before and after it fuse into one Zip operation:

float[] combined = left
    .AsCompute()
    .Select(value => value * 2.0f)
    .Zip(right, (first, second) => first + second)
    .Select(value => ComputeMath.Clamp(value, 0.0f, 1.0f))
    .ToArray();

The right array remains lazy by reference and must have the same length as the left array when the terminal operation runs. A second Zip is rejected because it would require more than two source arrays.

Reduction terminals are fused as well. A selector chain (or a binary Zip graph) followed by Sum, Min, Max, or Average transforms values while reducing them, without materializing an intermediate array on any backend, including chunked GPU execution.

SelectInPlace allows the optimizer to reuse intermediate storage but does not change the original managed array, so normal ToArray execution is safe to branch. ToArrayInPlace is the explicit mutating terminal:

float[] sameArray = source
    .AsCompute()
    .Select(value => value * 2.0f)
    .ToArrayInPlace();

Debug.Assert(ReferenceEquals(source, sameArray));

Available terminals are ToArray, ToArrayInPlace, Sum, Min, Max, and Average.

Core operations

float[] mapped = Compute.Run(
    source,
    value => ComputeMath.Clamp(value * 1.25f, 0.0f, 1.0f));

float[] zipped = Compute.Zip(
    left,
    right,
    (x, y) => x + y); // Both arrays must have the same length.

Compute.RunInPlace(source, value => value * 2.0f + 1.0f);
Compute.ZipInPlace(target: left, right: right, (x, y) => x + y);

float sum = Compute.Sum(source);
float minimum = Compute.Min(source);
float maximum = Compute.Max(source);
float average = Compute.Average(source);

int[] histogram = Compute.Histogram(
    samples,
    binCount: 256,
    minimum: 0.0f,
    maximum: 1.0f,
    new HistogramOptions
    {
        OutOfRangeMode = HistogramOutOfRangeMode.Ignore
    });

The same operations are available for double[] and int[]. Average returns the array element type, so integer average uses integer semantics.

Supported element types and expressions

Type Map/Zip In-place Reductions Histogram Resident buffer
float Yes Yes Yes Yes Yes
double Yes Yes Yes No Yes
int Yes Yes Yes No Yes

Float expressions support arithmetic, comparisons, conditional expressions, captured primitive constants, and these ComputeMath methods:

  • Abs, Min, Max, and Clamp;
  • Sqrt, Pow, Exp, Log, and Log10;
  • Sin, Cos, and Tan;
  • Floor, Ceiling, and Round.

The name is intentionally backend-neutral: ComputeMath.Sin, for example, can run on Scalar CPU, Parallel CPU, SIMD, or GPU. The older GpuMath name remains available as a compatibility alias. Double and integer expressions use arithmetic and the applicable supported System.Math overloads.

Calls through captured reference objects and arbitrary .NET methods are intentionally rejected because they cannot be translated to SIMD or GPU instructions. For unrestricted float CPU code, use Compute.RunDelegate on the Scalar or Parallel CPU backend.

Fast Fourier transform

Complex32 uses two contiguous single-precision components and is native to Scalar, Parallel CPU, AVX SIMD, and GPU backends. One- and two-dimensional radix-2 transforms support both allocating and in-place APIs:

Complex32[] spectrum = Compute.Fft(samples, options: computeOptions);
Compute.FftInPlace(spectrum, FourierDirection.Inverse, computeOptions);

Complex32[] spectrum2D = Compute.Fft2D(pixels, width, height, options: computeOptions);

Forward transforms are unnormalized. Inverse transforms divide by the complete element count, so a forward/inverse pair reconstructs the original input. Dimensions must be positive powers of two.

Signal processing, statistics, and convolution

float[] power = Compute.PowerSpectrum(spectrum);
float[] magnitudes = Compute.MagnitudeSpectrum(spectrum);
float[] phases = Compute.PhaseSpectrum(spectrum); // Scalar, Parallel CPU, or GPU
SignalPeak[] peaks = Compute.FindPeaks(values, minimumValue: 0.5f);

float[] smoothed = Compute.Convolve1D(values, kernel, ConvolutionBoundary.Clamp);
float[] windowed = Compute.ApplyHannWindow(values);

Phase spectrum uses Atan2, which has no SIMD instruction in the expression IR, so it runs on Scalar, Parallel CPU, or GPU and rejects explicit SIMD.

Statistics cover moments, covariance, correlation, and regression:

StatisticsResult moments = Compute.CalculateStatistics(values);
double covariance = Compute.Covariance(x, y);
double correlation = Compute.Correlation(x, y);
LinearRegressionResult regression = Compute.LinearRegression(x, y);
double entropy = Compute.ShannonEntropy(histogram);

Percentile/quantile/median sort a copy of the input because FastCompute does not yet expose a backend-native ordering primitive.

Thresholding and normalization are first-class as well:

float[] binary = Compute.Threshold(values, threshold: 0.5f);
MinMaxResult range = Compute.MinMax(values);
float[] unitRange = Compute.Normalize(values);
float[] safe = Compute.SafeDivide(numerator, denominator, zeroResult: 0.0f);

Composite values

Unmanaged structures can opt into FastCompute by implementing IComputeValue<T>. Descriptors validate a tightly packed homogeneous layout of float or byte components. Float-component transformations, including transformations between different structures, run on Scalar, Parallel CPU, SIMD, and GPU. Homogeneous byte-component values with one through four components have native SIMD layout load/store kernels and GPU execution.

Choosing a backend

var options = new ComputeOptions
{
    Backend = ComputeBackendKind.Auto,
    MaxDegreeOfParallelism = Environment.ProcessorCount
};
Backend Behavior
Auto Selects a compatible backend using expression complexity, array size, transfer cost, and the GPU memory budget.
Scalar Runs a conventional single-threaded CPU loop.
ParallelCpu Splits work into CPU chunks and processes them on multiple threads.
Simd Uses hardware-accelerated CPU vectors and a scalar tail.
Gpu Executes through ILGPU on the selected accelerator.

SIMD is not another form of Parallel.For: it processes several values per CPU instruction on the calling thread. Explicitly selected backends never silently fall back; use Auto when fallback is required.

GPU execution

Create a reusable context, precompile kernels, and keep multi-step pipelines resident on the accelerator:

using ComputeContext gpu = ComputeContext.Create(
    new ComputeContextOptions
    {
        AcceleratorIndex = preferredGpu.Index
    });

gpu.PrecompileAll();

using ComputeBuffer<float> input = gpu.Upload(source);
using ComputeBuffer<float> scaled =
    input.Select(value => value * 0.75f);
using ComputeBuffer<float> transformed =
    scaled.Select(value => ComputeMath.Sin(value));

float sum = transformed.Sum();
float[] output = new float[transformed.Length];
transformed.Download(output);
  • ComputeDefaults.PreferredGpuAcceleratorIndex sets a process-wide default GPU for Auto and explicit GPU operations.
  • Arrays larger than available GPU memory run in sequential chunks (GpuChunkElementCount, GpuMemoryBudgetBytes). Opt-in EnableGpuStreaming = true overlaps transfers and execution for explicit out-of-place float Map.
  • Precompile<T> and Prepare<T> move kernel compilation out of the first business operation.

Diagnostics and async

ComputeResult<float[]> result = Compute.RunWithDiagnostics(
    source,
    value => value * value,
    new ComputeOptions { Backend = ComputeBackendKind.Auto });

ComputeDiagnostics d = result.Diagnostics;
Console.WriteLine($"Backend:       {d.Backend}");
Console.WriteLine($"Device:        {d.DeviceName ?? "CPU"}");
Console.WriteLine($"Planning:      {d.PlanningTime}");
Console.WriteLine($"Execution:     {d.ExecutionTime}");
Console.WriteLine($"Upload bytes:  {d.UploadedBytes}");
Console.WriteLine($"Download bytes:{d.DownloadedBytes}");
Console.WriteLine($"Chunks:        {d.ChunkCount}");
using var cancellationSource = new CancellationTokenSource();

float[] mapped = await Compute.RunAsync(
    source,
    value => value * 2.0f,
    new ComputeOptions
    {
        CancellationToken = cancellationSource.Token
    });

ILGPU currently exposes synchronous completion primitives, so async methods return completed tasks and do not hide blocking work in Task.Run.

Common problems

  • GPU is slower than Parallel CPU. Expected when transfer and compilation costs exceed the kernel work. Reuse a context, precompile, keep intermediate data in resident buffers, or let Auto choose CPU.
  • An explicit backend throws instead of using CPU. This is the intended contract. Use Auto when fallback is required.
  • An expression cannot be translated. Use supported arithmetic and math methods, or use RunDelegate for unrestricted float CPU code.
  • The first GPU call is slow. The first call includes expression lowering and kernel compilation. Reuse the context and call PrecompileAll, Precompile<T>, or Prepare<T> during warm-up.

Image processing

Native image formats, filters, Bayer CFA handling, camera simulation, and GPU-resident image buffers are not part of this package. Install FastCompute.ImageProcessing when an application needs Image<TPixel> and friends.

Further documentation

Product Compatible and additional computed target framework versions.
.NET net8.0 is compatible.  net8.0-android was computed.  net8.0-browser was computed.  net8.0-ios was computed.  net8.0-maccatalyst was computed.  net8.0-macos was computed.  net8.0-tvos was computed.  net8.0-windows was computed.  net9.0 was computed.  net9.0-android was computed.  net9.0-browser was computed.  net9.0-ios was computed.  net9.0-maccatalyst was computed.  net9.0-macos was computed.  net9.0-tvos was computed.  net9.0-windows was computed.  net10.0 was computed.  net10.0-android was computed.  net10.0-browser was computed.  net10.0-ios was computed.  net10.0-maccatalyst was computed.  net10.0-macos was computed.  net10.0-tvos was computed.  net10.0-windows was computed. 
Compatible target framework(s)
Included target framework(s) (in package)
Learn more about Target Frameworks and .NET Standard.

NuGet packages (1)

Showing the top 1 NuGet packages that depend on FastCompute:

Package Downloads
FastCompute.ImageProcessing

Backend-neutral native image processing for FastCompute.NET.

GitHub repositories

This package is not used by any popular GitHub repositories.

Version Downloads Last Updated
0.8.1 122 8/14/2026
0.8.0 107 8/14/2026
0.7.0 106 7/30/2026
0.6.0 102 7/30/2026
0.5.0 100 7/30/2026

# FastCompute.NET Release Notes

## 0.8.1 - 2026-08-14

Documentation-only repack. Each NuGet package now ships its own README:
`FastCompute` carries the core guide and `FastCompute.ImageProcessing` carries
the image processing guide, while the repository root README remains the
project overview. No API or behavior changes.

## 0.8.0 - 2026-08-14

Native signal, statistics, and image processing primitives. The image and
forensics capabilities now live in their own assembly, and the generic numeric
primitives behind them moved into the core package.

### Added

- One- and two-dimensional radix-2 FFT for `Complex32[]` with allocating and
 in-place APIs (`Fft`, `FftInPlace`, `Fft2D`, `Fft2DInPlace`, inverse
 variants) on Scalar, Parallel CPU, AVX SIMD, and GPU backends.
- `Complex32` native composite value with `PowerSpectrum`,
 `MagnitudeSpectrum`, and `PhaseSpectrum` helpers; `FindPeaks`,
 `PeakToMedianRatio`, `MeanAbsoluteDifference`, `Percentile`, `Quantile`,
 `Median`, and Hann/Hamming/Blackman window functions in
 `Compute.Signal`.
- 1D and 2D convolution (`Convolve1D`, `Convolve2D`) with the same
 `ComputeOptions` contract.
- Statistics: `CalculateStatistics`, `Mean`, `Variance`, `StandardDeviation`,
 `Skewness`, `Kurtosis`, `SumOfSquares`, `Covariance`, `Correlation`,
 `AutoCorrelation`, `LinearRegression`, and `ShannonEntropy`.
- Threshold, `MinMax`, `Normalize`, and `SafeDivide` utilities.
- Homogeneous `byte`-component packed values with native SIMD layout
 load/store kernels and byte-composite GPU execution for one through four
 components.

### Changed

- Image and forensics functionality was moved into the new
 `FastCompute.ImageProcessing` assembly and NuGet package. The core
 `FastCompute` package has no image dependency; it ships with the generic
 primitives above.
- The negative image forensics pipeline now runs on the generic primitives
 instead of its own copies of FFT, statistics, convolution, Bayer handling,
 and camera simulation. See
 `docs/ai-image-forensics-algorithm-migration.md` for the ownership table.
- `Image<TPixel>` gained convolution-backed Gaussian/Sobel/Laplacian filters,
 residuals, local contrast and entropy, spectrum preparation, deterministic
 area resize, Bayer CFA sampling, and demosaicing on Scalar, Parallel CPU,
 SIMD, and GPU.

### Compatibility

- `FastCompute.ImageProcessing` 0.8.0 depends on `FastCompute` 0.8.0 and is
 strongly named with the same public key token `c76a60c96d65300c`.
- The original lazy pipeline, reduction fusion, `ComputeMath` (and the
 `GpuMath` alias), resident buffers, and chunked/streaming GPU execution
 remain unchanged.
- Explicit SIMD requests for local window entropy and phase spectrum are
 rejected instead of falling back to a hidden scalar loop; both operations
 have Scalar, Parallel CPU, and native GPU paths.