LLMFromScratch.SDK 0.3.2

dotnet add package LLMFromScratch.SDK --version 0.3.2
                    
NuGet\Install-Package LLMFromScratch.SDK -Version 0.3.2
                    
This command is intended to be used within the Package Manager Console in Visual Studio, as it uses the NuGet module's version of Install-Package.
<PackageReference Include="LLMFromScratch.SDK" Version="0.3.2" />
                    
For projects that support PackageReference, copy this XML node into the project file to reference the package.
<PackageVersion Include="LLMFromScratch.SDK" Version="0.3.2" />
                    
Directory.Packages.props
<PackageReference Include="LLMFromScratch.SDK" />
                    
Project file
For projects that support Central Package Management (CPM), copy this XML node into the solution Directory.Packages.props file to version the package.
paket add LLMFromScratch.SDK --version 0.3.2
                    
#r "nuget: LLMFromScratch.SDK, 0.3.2"
                    
#r directive can be used in F# Interactive and Polyglot Notebooks. Copy this into the interactive tool or source code of the script to reference the package.
#:package LLMFromScratch.SDK@0.3.2
                    
#:package directive can be used in C# file-based apps starting in .NET 10 preview 4. Copy this into a .cs file before any lines of code to reference the package.
#addin nuget:?package=LLMFromScratch.SDK&version=0.3.2
                    
Install as a Cake Addin
#tool nuget:?package=LLMFromScratch.SDK&version=0.3.2
                    
Install as a Cake Tool

LLMFromScratch SDK

LLMFromScratch.SDK is the base .NET 9/10 SDK for this repository's original in-process GPT model. It creates, trains, checkpoints, tool-enables, and exports that model as lossless F32 GGUF. It has no Qwen, Llama, Hugging Face, Python, Unsloth, PEFT, llama.cpp, or Ollama training dependency.

Install

dotnet add package LLMFromScratch.SDK --version 0.3.2

Add exactly one native TorchSharp runtime package for the target host:

Host Package
Windows CPU TorchSharp-cpu version 0.107.0
Windows NVIDIA CUDA TorchSharp-cuda-windows version 0.107.0
Linux CPU libtorch-cpu-linux-x64 version 2.10.0
Linux NVIDIA CUDA TorchSharp-cuda-linux version 0.107.0
Apple Silicon TorchSharp-cpu version 0.107.0

Do not mix CPU and CUDA runtime payloads in the same application. On Apple Silicon, target osx-arm64; MPS/Metal is supplied by macOS.

Train the native GPT model

Create Data/train.txt, optionally Data/validation.txt, and matching GPT-2 Models/vocab.json and Models/merges.txt files. Then configure the client:

using LLMFromScratch.Core.Configurations;
using LLMFromScratch.SDK;
using LLMFromScratch.Training.Export;

var options = new LlmFromScratchOptions
{
    Model = new GPTModelConfiguration
    {
        VocabularySize = 50_257,
        ContextLength = 128,
        EmbeddingDimension = 256,
        NumberOfHeads = 8,
        NumberOfLayers = 4,
        Dropout = 0.1
    },
    Training = new TrainingOptions
    {
        Epochs = 1,
        BatchSize = 1,
        LearningRate = 3e-4,
        MinimumLearningRate = 3e-4,
        WeightDecay = 0.01,
        GradientClip = 1.0,
        Optimizer = OptimizerType.AdamW,
        Scheduler = SchedulerType.Constant
    },
    ComputeDevice = new ComputeDeviceOptions
    {
        Preference = ComputeDevicePreference.Auto
    }
};

using var client = new LlmFromScratchClient(options);
client.SetDatasetPaths("Data/train.txt", "Data/validation.txt");

var training = await client.TrainAsync();
var gguf = client.ExportGguf(new GgufExportOptions
{
    VocabularyPath = "Models/vocab.json",
    MergeFilePath = "Models/merges.txt",
    OutputPath = "exports/my-gpt-f32.gguf",
    ModelName = "My GPT model"
});

Console.WriteLine($"Loss: {training.Loss:F4}");
Console.WriteLine($"GGUF: {gguf.OutputPath}");
Console.WriteLine($"Modelfile: {gguf.ModelfilePath}");

TrainAsync writes a checkpoint after each epoch. ExportGguf refuses to overwrite an existing output file. Use it only for the LLMFromScratch GPT-2-compatible model, whose tensor layout and tokenizer are not compatible with Qwen, Llama, or other Hugging Face models.

ExportGguf also writes a companion Ollama Modelfile next to the GGUF file (<output>.Modelfile by default), built from the checkpoint's own llmfromscratch.* control-token metadata so ollama create picks up a matching TEMPLATE and stop sequences automatically. Set GgufExportOptions.WriteModelfile = false to skip it, or GgufExportOptions.ModelfilePath to choose the path; the path actually written is reported on GgufExportResult.ModelfilePath. This only fixes the prompt format Ollama runs the model with — it does not change what the underlying checkpoint has learned, so a model that needs more or better training data will still answer incorrectly inside a correctly-templated prompt.

Optional third-party training extension

Install LLMFromScratch.SDK.ThirdPartyTraining only when you need Qwen/Llama LoRA or QLoRA training through Unsloth or PEFT, llama.cpp GGUF conversion, or Ollama import. That opt-in package has its own pipeline API and Python tooling; it does not change the base SDK's model or device lifecycle. See the repository docs/ThirdPartyModels.md for its prerequisites and examples.

Device and lifecycle notes

  • ComputeDevicePreference.Auto selects CUDA, then Apple Silicon MPS, then CPU.
  • ComputeDevicePreference.Cuda fails early if CUDA is unavailable.
  • Dispose LlmFromScratchClient to release model and native tensor resources.

MIT. See the repository license.

Product Compatible and additional computed target framework versions.
.NET net9.0 is compatible.  net9.0-android was computed.  net9.0-browser was computed.  net9.0-ios was computed.  net9.0-maccatalyst was computed.  net9.0-macos was computed.  net9.0-tvos was computed.  net9.0-windows was computed.  net10.0 is compatible.  net10.0-android was computed.  net10.0-browser was computed.  net10.0-ios was computed.  net10.0-maccatalyst was computed.  net10.0-macos was computed.  net10.0-tvos was computed.  net10.0-windows was computed. 
Compatible target framework(s)
Included target framework(s) (in package)
Learn more about Target Frameworks and .NET Standard.

NuGet packages (1)

Showing the top 1 NuGet packages that depend on LLMFromScratch.SDK:

Package Downloads
LLMFromScratch.SDK.ThirdPartyTraining

Optional LLMFromScratch SDK extension for fine-tuning Hugging Face Qwen and Llama models with Unsloth or PEFT, converting to GGUF, and importing into Ollama.

GitHub repositories

This package is not used by any popular GitHub repositories.

Version Downloads Last Updated
0.3.2 153 9/7/2026
0.3.1 119 9/7/2026
0.3.0 122 9/7/2026
0.2.1 115 9/2/2026
0.2.0 114 9/2/2026
0.1.20 89 9/1/2026
0.1.17 93 9/1/2026
0.1.15 107 8/25/2026
0.1.14 111 8/25/2026
0.1.13 108 8/25/2026
0.1.12 104 8/25/2026
0.1.11 106 8/25/2026
0.1.10 97 8/25/2026
0.1.9 101 8/23/2026
0.1.8 97 8/23/2026
0.1.7 90 8/23/2026
0.1.6 99 8/23/2026
0.1.5 99 8/23/2026
0.1.4 103 8/23/2026
0.1.3 103 8/23/2026
Loading failed

Lowers the auto-generated companion Modelfile's default sampling parameters (repeat_penalty 1.3 to 1.1, temperature 0.7 to 0.3): a control-token format reuses the same "<", "|", ">" subtokens in every tag, and the previous, more aggressive repeat_penalty could discourage the model from reusing them correctly right after closing the previous tag, contributing to malformed tag output such as "<|/answer|/reasoning|/answer|>". Documents a known, separate limitation: this tokenizer has no atomic/special tokens for the control tags, so a small or undertrained checkpoint can still corrupt a tag's exact multi-subtoken BPE sequence even with a correct TEMPLATE; that is a tokenizer/training-data property the exporter cannot fix on its own. 0.3.1's fix - Modelfile generation moved into GgufModelExporter.Export() itself so every caller (including LlmFromScratchClient.ExportGguf()) gets it, with GgufExportOptions.WriteModelfile/ModelfilePath and GgufExportResult.ModelfilePath - remains in place, as do 0.3.0's caller-configurable control tokens, export-time NaN/Infinity validation, and opt-in GenerationOptions.RepetitionPenalty.