Fuuga 1.0.28

dotnet add package Fuuga --version 1.0.28
                    
NuGet\Install-Package Fuuga -Version 1.0.28
                    
This command is intended to be used within the Package Manager Console in Visual Studio, as it uses the NuGet module's version of Install-Package.
<PackageReference Include="Fuuga" Version="1.0.28" />
                    
For projects that support PackageReference, copy this XML node into the project file to reference the package.
<PackageVersion Include="Fuuga" Version="1.0.28" />
                    
Directory.Packages.props
<PackageReference Include="Fuuga" />
                    
Project file
For projects that support Central Package Management (CPM), copy this XML node into the solution Directory.Packages.props file to version the package.
paket add Fuuga --version 1.0.28
                    
#r "nuget: Fuuga, 1.0.28"
                    
#r directive can be used in F# Interactive and Polyglot Notebooks. Copy this into the interactive tool or source code of the script to reference the package.
#:package Fuuga@1.0.28
                    
#:package directive can be used in C# file-based apps starting in .NET 10 preview 4. Copy this into a .cs file before any lines of code to reference the package.
#addin nuget:?package=Fuuga&version=1.0.28
                    
Install as a Cake Addin
#tool nuget:?package=Fuuga&version=1.0.28
                    
Install as a Cake Tool

Fuuga

Tired of paying tokens? Think you could train a better model? Well, now you can try.

An LLM built from scratch in F# and .NET. Fuuga implements a complete language model pipeline: tokenization, data ingestion, model training, fine-tuning, and text generation -- with no Python dependencies.

Built on TorchSharp for tensor operations and Microsoft.ML.Tokenizers for BPE, Fuuga uses idiomatic F# (discriminated unions, pipelines, immutability) throughout. Works on GPU or CPU.

Choose your path

New here? Each goal is a short sequence of commands driven by one recipe file. Scaffold a recipe, edit a few paths, then run the commands with --recipe. See docs/workflows.md for the full journeys.

I want to... Start with
Make an open model follow my data fuuga scaffold finetune — Fine-tune a donor (the common path)
Train a small model from scratch fuuga scaffold pretrain — Pre-train
Shrink/export a model for deployment Shrink & export
Score a model Evaluate
Serve or build an agent on a model Serve & integrate
Review pull requests in CI/CD with my own model Code review in CI/CD — scripts/ci-review.sh + fuuga-serve
Rent/sell my model to others via OpenRouter Sell as a service — pricing + prepaid credit ledger

Getting Fuuga

Fuuga comes in two forms, reflecting how it's used:

  • F# library — NuGet. Reference it from a script (#r "nuget: Fuuga") or a project (<PackageReference Include="Fuuga" />) and call the API directly. Best for power users composing steps the CLI doesn't expose — see the examples/ catalogue (complete-pipeline.fsx, finetuning-lora-qlora.fsx, modelops-compress-and-export.fsx, and more).

    The NuGet packages come in two backend flavors, and the package choice — not a runtime flag — decides CPU vs GPU: Fuuga carries CUDA 12.8 natives (win-x64; needs an NVIDIA GPU), Fuuga.cpu carries CPU natives and runs anywhere. The same split applies to the image module: Fuuga.Image (CUDA) and Fuuga.Image.cpu (CPU). When building from source, the TorchBackend MSBuild property selects the backend: dotnet build -p:TorchBackend=cuda for CUDA, plain dotnet build for the CPU default.

  • CLI and server — clone the repo. The fuuga CLI and the fuuga-serve server are not published as standalone tools: their native dependency (libtorch) is over a gigabyte — too large for a NuGet package or a dotnet tool. So clone the repo and run from source:

    • CLI — shown throughout the docs as fuuga <command>; from a clone, run it as dotnet run -- <command>.
    • Server — dotnet run --project Fuuga.Server -- --checkpoint <dir> --tokenizer <dir> [--port <N>] (this is what fuuga-serve refers to).

Improving the server's distribution story (a slimmer, CPU-only or download-on-first-run tool) is an open area for contribution.

Two ways to use Fuuga, then: the CLI for the full pipeline (config-driven via --recipe), or the F# API from scripts. Full command list below; full reference in docs/cli-reference.md.

Features

Core Pipeline:

  • BPE Tokenizer -- Train a byte-pair encoding tokenizer on your own corpus with configurable vocabulary size
  • Data Ingestion -- Discover and tokenize epub, markdown, Parquet, and plain text files into a binary corpus
  • Parquet I/O -- Read and write HuggingFace-compatible Parquet datasets for SFT, DPO, and document data
  • Corpus Compression -- Zstd compression/decompression for .fuge corpus files
  • GPT-2 Transformer -- Decoder-only causal transformer with rotary position embeddings (RoPE), grouped-query attention (GQA), RMSNorm, and SwiGLU activation
  • Multi-Head Latent Attention (MLA) -- DeepSeek-V2 style compressed KV cache with query/KV compression, decoupled RoPE keys, and optional weight absorption for reduced memory during inference
  • Kimi Delta Attention (KDA) -- Kimi K3 / Kimi-Linear style linear attention: delta-rule recurrence with channel-wise lower-bounded decay (fixed-size state, NoPE), plus per-layer hybrid layouts (KDA-local / MLA-global every Nth layer, K3's 3:1 pattern) (experimental)
  • Stable LatentMoE & Block Attention Residuals -- Kimi K3 architecture options: routed experts in a compact latent space with pre-up-projection RMSNorm; learned depth-wise attention over block residual summaries (experimental)
  • Sigmoid MoE routing + Quantile Balancing -- DeepSeek-V3/K3-style sigmoid router scores with a frozen selection bias (imported from donor e_score_correction_bias) and aux-loss-free load balancing via router-score quantiles
  • SiTU-GLU activation & Muon optimizer -- K3's softcapped GLU (bounded activations for low-precision training) and Muon (Newton-Schulz orthogonalized momentum, per-head variant); MTP draft heads can train on the EAGLE-3-style LK acceptance-rate loss
  • Vision Encoder -- Vision model support for multimodal inputs
  • Vision Bridge -- Q-Former cross-attention bridge that compresses vision patch tokens into learned query vectors for multimodal (image+text) inputs
  • Paged Attention -- Paged KV-cache attention for efficient memory usage during long-context generation
  • Memory Hierarchy -- Compressed memory with external retrieval for extended context
  • Multi-Resolution Attention -- Chunk pooling with global tokens for efficient long-context processing
  • FlashAttention Config -- SDPA backend selection and benchmarking for attention kernels
  • Auto Config -- Hardware-aware auto-resolution of DU configuration cases (norm, activation, precision, offloading, communication) at startup
  • Early Exit -- Adaptive depth inference for faster generation when confidence is high
  • Training -- AdamW optimizer with cosine learning rate scheduling, warmup, gradient clipping, mixed precision support, and gradient accumulation
  • Multi-Token Prediction (MTP) -- DeepSeek-V3 style auxiliary heads predicting multiple future tokens (configurable Depth); adds a weighted multi-depth loss during training for better sample efficiency and powers MTP-drafted speculative decoding for faster generation. Enable from the CLI with train --mtp-depth <N> [--mtp-loss-weight <f>], or set MtpConfig in the model-config JSON
  • INT8 Optimizer Moments -- Optional INT8 quantization of AdamW M/V moment tensors with per-row symmetric quantization, reducing optimizer memory ~4× (--moment-quant int8). Mutually exclusive with SWA/Lookahead/SAM, differential per-group LR, gradient offload, per-param flush, and NVMe/CPU optimizer offload (combining them fails fast). INT8 moments are not persisted across checkpoint resume — model weights resume normally while momentum/variance restart fresh.
  • Optimizer Variants -- Stochastic Weight Averaging (SWA) and Lookahead optimizer support with checkpointable optimizer state
  • Gradient Checkpointing -- Memory-efficient training via activation recomputation
  • GPU Offloading -- Layer-wise CPU/GPU offloading for reduced VRAM usage
  • Optimizer Offloading -- Offload optimizer states to CPU memory
  • NVMe Paging -- ZeRO-Infinity 3-tier GPU/CPU/NVMe memory management for training models larger than available VRAM
  • Per-Tensor Gradient Offloading -- Bulk-copy gradients to CPU after backward pass and restore before optimizer step, freeing GPU VRAM during the optimizer phase (--grad-offload)
  • Per-Parameter CUDA Flush -- Aggressive CUDA cache cleanup after optimizer step to reclaim transient VRAM spikes from M/V update temporaries (--flush-each-param)
  • VRAM Guard -- In-process background thread that polls GPU memory via nvidia-smi and signals the training loop to warn, skip batches, or abort when usage exceeds a configurable threshold (--vram-guard-gb <float>)
  • Memory Strategy Presets -- MemoryStrategyConfigs.none, .constrained (grad-offload + flush), and .full (all three with VRAM guard at 95%) with automatic ConfigWizard recommendations based on model-to-VRAM ratio
  • Model Parallelism -- Tensor and pipeline parallelism configuration for 70B+ parameter models, with automatic DataParallel recommendation for multi-GPU setups
  • Inference -- Greedy, top-k, top-p (nucleus), and temperature sampling with repetition penalty
  • Fill-in-the-Middle -- FIM support with prefix/suffix tokens for code completion
  • Checkpoints -- Save, load, resume training from checkpoints with full metadata; safetensors format support
  • Memory-Mapped Loading -- mmap-based model loading for fast startup
  • Streaming Inference -- Token-by-token generation with configurable stop conditions
  • Confidence Signals -- Entropy, repetition rate, hedging detection, calibrated confidence with Platt scaling, and stop reason reporting
  • Drift Detection -- Statistical drift monitoring (Kolmogorov-Smirnov, Population Stability Index) over confidence signals with ring-buffered accumulation
  • Drift Alerting -- Dual-threshold alerts with adaptive sigma-based thresholds, OpenTelemetry metrics, and retraining triggers that carry an approve-or-reject brief (evidence, proposed action, alternative, expected effect, risk)
  • Drift Monitor -- fuuga-serve --drift-state <dir> feeds every generation's signals to the monitor, captures the reference window, tests on a timer and files triggers; fuuga drift status|decide|reset-reference is the operator loop — verdicts are logged to the experience store and reported as per-signal trigger precision with tuning hints. Nothing is retrained automatically
  • ONNX Export -- Export to ONNX format with fp16/int8 quantization, validation, and benchmarking
  • ONNX Inference -- ONNX Runtime backend for optimized inference (--backend onnx)
  • Benchmark Evaluation -- Built-in benchmark runner for MMLU, HellaSwag, ARC-Challenge, WinoGrande, PIQA, TruthfulQA (log-likelihood MCQ scoring), GSM8K (generative, numeric-answer extraction), MATH-500 (generative; symbolic verifier grading with per-difficulty-level sub-scores and --maj-k self-consistency voting by mathematical equivalence), and HumanEval (generative; candidates execute against the official assert-based test suites via a local Python interpreter, with an approximate fallback and warning when Python is absent), plus cached dataset downloads and checkpoint-attached benchmark results
  • Math Verifier -- From-scratch symbolic answer verification (no CAS dependency): \boxed{} extraction, LaTeX normalization (fractions, surds, degrees, superscripts, mixed numbers), exact BigInteger-rational equivalence (0.5 ≡ 1/2 but 0.3333333333 ≢ 1/3), and random-rational-point identity testing for expressions with variables -- grades LaTeX mixed numbers correctly where SymPy-based reference pipelines do not
  • FP8 Dequantization -- FP8 format support for quantized weight loading with GPU-accelerated LUT path (256-entry cached lookup table using torch.index_select) auto-selected when CUDA is available
  • Validation Pipeline -- Input validation framework with composable validators
  • Scaling Heuristics -- Auto-scaling configuration from corpus and hardware stats
  • Config Wizard -- Corpus analysis and hardware-aware config generation using Chinchilla scaling laws, activation memory estimates, NTK-aware RoPE, multi-GPU detection (nvidia-smi), and memory strategy recommendations
  • CLI -- Subcommands for the full pipeline (tokenize, ingest, train, infer, info, experiments, sft, dpo, rl, merge, transfer, distill, merge-models, fisher, calibrate, plan-quant, distributed, export onnx, export gguf, export llama, export glm, compress, decompress, prune, eval, config, config wizard, rag, graph, decide, wordnet, serve, orchestrate, agent, image)
  • Experiment tracking -- fuuga experiments <dir> compares training runs side by side (train/val loss, benchmark scores, approx params, fine-tuning stage, corpus hash, date) from their checkpoints, sorted and in table/CSV/JSON form. Read-only over existing checkpoint.json files — no models loaded.
  • Data provenance -- ingest writes a <corpus>.manifest.json sidecar recording the source files behind the corpus (path, size, content hash) and the corpus's SHA256. That hash is the same value stamped into CheckpointMetadata.CorpusHash, so checkpoint → corpus manifest → source files is fully traceable for reproducibility.

Fine-Tuning:

  • Supervised Fine-Tuning (SFT) -- LoRA-based fine-tuning on instruction/chat JSONL data with configurable rank, alpha, and target modules
  • Prompt Tuning -- Soft-prompt / virtual-token fine-tuning with frozen base weights for lightweight PEFT workflows
  • Direct Preference Optimization (DPO) -- Preference learning from chosen/rejected pairs with LoRA
  • Reinforcement Learning (RL) -- REINFORCE++ / GRPO fine-tuning with pluggable reward functions, trajectory-faithful rollout scoring, chain-of-thought rollouts (--thinking), symbolic verifier rewards (--reward math), DAPO zero-signal group filtering, Dr. GRPO advantage normalization (--no-adv-std), and token-level length-normalized loss (--token-level-loss)
  • Reward Functions -- Composable reward functions for RL training (correctness, formatting, safety); accuracy rewards score only the answer span after thinking, so truncated reasoning is never rewarded as correct
  • Rejection Sampling (STaR/RFT) -- fuuga sample-traces samples N completions per prompt, keeps only verifier-passing traces, and writes them as SFT JSONL -- with --thinking, curated traces carry tagged reasoning spans
  • LoRA Adapter Merging -- Merge trained LoRA adapters back into the base model weights
  • QLoRA -- NF4-quantized base weights with LoRA adapters for memory-efficient fine-tuning on consumer GPUs
  • Data Validation -- JSONL format validation for SFT and DPO datasets with honesty pattern classification
  • Data Augmentation -- Synonym replacement, rule-based paraphrasing, token-level noise injection, and SFT/DPO oversampling for training data diversity

Weight Transfer and Model Merging:

  • Weight Transfer -- Transfer weights from donor models with architecture-aware mapping (Phi-3, Phi-4, LLaMA3, Mistral, Qwen 2.x/3.x, Gemma-3, Gemma-4, DeepSeek-V3 / Kimi K2 including MLA attention into an AttentionType = MLA target (latent down/up projections, decoupled RoPE key, and the two latent RMSNorms), Z.AI GLM-4.5/4.6/5 MoE (experimental), Moonshot Kimi K3 MoE with MXFP4 auto-dequantization (experimental), ModernBERT) and dimension adaptation for mismatched tensors. ModernBERT is the one encoder donor: transfer --mapping modernbert initialises a bidirectional model for the typed-decision path (decide train) rather than a generative one, so an open ModernBERT-backbone decision model can be used as the starting point instead of training the System-1 encoder from a causal checkpoint
  • Donor Export -- The reverse direction: export a Fuuga model back to HuggingFace-format safetensors. fuuga export llama produces LlamaForCausalLM (near-lossless for a faithful Llama fine-tune) loadable by Transformers / vLLM / TGI / AWQ-GPTQ; fuuga export glm produces GLM (Glm4Moe) safetensors (experimental, best-effort). Round-trips (export → re-import) are verified bit-exact in the test suite.
  • Knowledge Distillation -- Token-level, sequence-level (teacher-forced pseudo-labels, or true Kim & Rush via --mode sequence-gen), reverse-KLD, attention (distill --attention) and hidden-state (distill --hidden) distillation from a teacher model
  • Endpoint Teacher -- distill --teacher-endpoint distils from a teacher reachable only over HTTP. An API returns samples rather than distributions, so this is response-based (black-box) distillation shaped as a data step: prompts in, the teacher's completions out as SFT JSONL for fuuga sft. Reaches hosted teachers no white-box path can, at the cost of much less signal per token
  • Distillation Quality Gate -- distill --eval-corpus measures the student on held-out data before and after and reports whether held-out loss actually improved, saying plainly when a run was a regression -- a falling KD loss only means the student tracked the teacher on the training corpus. Also prints a feasibility read (teacher/student top-1 agreement and top-K overlap) before training commits
  • LoRA Distillation -- distill --lora-rank distils into adapters over a frozen student base, so distillation fits the same hardware sft does instead of needing more. Works with every distillation path (logits, attention, hidden-state, cached); adapters are merged before saving unless --keep-adapter is given
  • Offline Logit Cache -- distill --write-logit-cache runs the teacher over the corpus once and stores its top-K logits; --logit-cache then trains with no teacher loaded at all. Decouples teacher hardware from student hardware (one pass on a rented GPU, then train locally) and makes extra epochs free of teacher cost. Cache/run compatibility is validated against the header rather than assumed
  • Attention Distillation -- Layer-level distillation between any attention pair (GQA/MLA/KDA/NSA/H-Transformer) with KL-divergence + MSE loss, pairing by layer index rather than attention class. This is the only route across architectures that weight transfer cannot express -- Kimi K3's KDA (linear) and Gated MLA layers have no weight-space mapping onto GQA. --hidden matches whole-block hidden states instead, which stays valid when the two attention mechanisms are fundamentally different -- and cascades by default, running each model over its own stack from its own embedding so neither side is probed off-manifold. Depth reduction (a deeper teacher into a shallower student) samples across the teacher's full depth via --layer-mapping. The cascade is also the only layer-level path that supports MoE models, since it walks through MoE blocks rather than requiring them to be pairable submodules. Both layer-level probes train the backbone only -- the LM head gets no gradient, so they are an initialisation step ahead of a logits pass, not a complete training run. Layers that miss the quality gate can be rolled back to their pre-distillation weights with --revert-unconverged, and MoE routers keep their load-balancing signal throughout
  • N-ary Model Merging -- Merge multiple models with configurable strategies (TIES, DARE, Karcher mean, ModelSoups, ModelStock) and EWC protection
  • Fisher Information -- Compute diagonal Fisher information matrices for Elastic Weight Consolidation
  • N-ary Data Mixing -- Weighted multi-source mixing with static, curriculum, proxy-based DoReMi (Group DRO domain reweighting from a unigram proxy), and self-paced (difficulty-ramped) strategies

Inference Capabilities:

  • Chain-of-Thought -- Thinking mode with ThinkStart/ThinkEnd token handling and dimmed thinking display
  • Thinking Budgets (s1-style) -- --max-thinking-tokens force-closes the thinking span at a cap; --min-thinking-tokens suppresses premature end-of-thinking and injects a "Wait," continuation cue to extend reasoning (test-time scaling)
  • Self-Consistency Voting -- fuuga infer --consensus N samples N responses and majority-votes the extracted answers by mathematical equivalence; also available as the consensus request option on the server API
  • Constrained Decoding -- Grammar-guided JSON structured output generation
  • Self-Verification -- Draft/refine verification passes with learned verifier scoring for higher-confidence answers
  • Tool Calling -- MCP (Model Context Protocol) client for tool discovery and invocation during generation
  • Tool Policy -- Confidence-aware tool routing policy for deciding when external tools should be invoked
  • Web Search -- Web search integration for grounded generation with citations
  • Image Routing -- Generate or caption images via the Fuuga.Image MCP server, either with the fuuga image command or by configuring fuuga-image in mcp.json
  • A2A Protocol -- Agent-to-Agent protocol client for multi-agent communication
  • Tree-Structured Speculative Decoding -- Speculative decoding with tree-structured candidates for faster generation
  • Draft-and-Refine -- Multi-pass reasoning pipeline for improved output quality
  • Autonomous Agent -- Web search agent loop for autonomous information gathering
  • Advanced Reasoning -- Consensus voting, verifier-scored selection, and tree-of-thoughts for improved answer quality
  • Backend Client -- HTTP client for calling external OpenAI-compatible LLM endpoints with structured response types
  • Context Awareness -- Convention file discovery (AGENTS.md, CLAUDE.md, .cursorrules), language/framework detection, git context
  • Semantic Knowledge -- WordNet WNDB parser with token-to-synset mapping and multi-lingual support

Orchestration:

  • Multi-Model Orchestrator -- Route tasks to appropriate models based on capability. fuuga orchestrate --backends backends.json decomposes a task and routes each subtask across the configured backends; without --backends it runs every subtask against a single --endpoint.
  • Cost-Aware Routing -- Budget-tracked model routing with cost optimization. With --backends, each subtask runs through a cost cascade: the first (cheapest) backend is tried, then the next (most capable) is used when confidence/history and the loaded routing profile (--checkpoint's compression_codesign.json, consumed via its QualityFloor) warrant escalation. Per-subtask spend is tracked against each backend's dailyBudget, and outcome history accumulates between subtasks to improve later routing. The backends.json file is a JSON array of { id, name, endpoint, apiKey?, inputTokenCost, outputTokenCost, dailyBudget?, strengths[] } ordered cheapest-first.
  • Fan-Out Orchestration -- Decompose tasks into subtasks, run in parallel, and aggregate results
  • Resumable Orchestration -- Checkpoint and resume fan-out plans across sessions

Agentic Persistence:

  • Experience Store -- Append-only JSON Lines log of attempt outcomes with thread-safe managed access
  • Strategy Lessons -- Persist and load distilled lessons from past experience for self-improvement
  • Persistent Retrieval Store -- Disk-backed IRetrievalStore for cross-session document retrieval
  • Hybrid Retrieval -- BM25 keyword index fused with dense cosine ranking via reciprocal rank fusion (--rag-hybrid on infer and fuuga-serve); catches exact-term matches that embedding similarity misses. Language-agnostic: stopwords are handled statistically (document-frequency cutoff), not by a hardcoded English list, so Finnish/Swedish/any-language corpora get equal treatment
  • Reranking -- Maximal Marginal Relevance (diversity) or model-scored relevance reranking over a widened candidate pool (infer --rag-rerank mmr|llm)
  • Multi-Query Expansion -- Model-rewritten query variants retrieved independently and fused with RRF (infer --rag-multi-query <N>)
  • RAG Retrieval Evaluation -- fuuga rag eval scores an index against a labelled JSONL question set (hit@k, mean reciprocal rank) so retrieval changes are measured, not guessed
  • Knowledge Graph (Graph Engineering) -- GraphRAG-style structured memory as a complement to chunk retrieval: facts stored as subject-predicate-object triples with evidence quotes, confidence, and provenance (GraphKnowledge.fs). LLM-driven extraction (fuuga graph extract against any OpenAI-compatible endpoint) and batch entity resolution, alias-aware graph retrieval (entity lookup, k-hop neighborhood, shortest evidence chain for "why"/"how connected" questions, connected communities, --as-of temporal filtering), and a typed query DSL as the query-translation target (easier for small models than Cypher, no server dependency). Maintenance follows the never-silently-overwrite rule: duplicates skipped, conflicts on functional predicates (persisted graph schema) flagged for review via graph query --conflicts, superseded facts timestamped instead of deleted. Wired into generation as fuuga infer --rag-graph (the question becomes a graph query with entity-mention fallback; facts ground the prompt alone or fuse with --rag text chunks via RRF) and server-side as fuuga-serve --rag-graph (deterministic entity-mention facts fused into the per-request context, no extra generation cost). CLI: fuuga graph add|extract|query|stats
  • Typed Decisions (System 1) -- The Laya / Jev pattern for agentic workflows: instead of asking a model to reason in free text and parsing the answer, ask a typed question — choice (probabilities over named options), score (ordered rubric levels + expected level) or noul (P(true) of a proposition) — about a state and get a probability for every option (TypedDecisions.fs). Requests and answers use Laya's JSON shape. Zero-shot over any existing model via option-tag log-probs (DecisionScoring.fs: Fuuga checkpoint, GGUF, or an OpenAI-compatible endpoint with logprobs), temperature calibration per (question type, option count) fitted on labelled decisions, calibration metrics (accuracy, NLL, Brier, ECE, reliability bins, selective accuracy at each act-threshold), and an act-or-escalate gate that feeds the minion's escalation verdict. The native System-1 model (DecisionHeads.fs) is the design proper: any Fuuga checkpoint becomes a bidirectional encoder (new AttentionPattern.Bidirectional; weights carry over, LLM2Vec-style) with per-kind bilinear decision heads that score every option from its pooled token span in one forward pass — no generation — trained with a strictly proper scoring rule (log score or Brier) on labelled decisions via fuuga decide train. CLI: fuuga decide predict|calibrate|eval|train|compare|judge; server: POST /v1/decide (fuuga-serve --decision-model, --decide-calibration). Measuring a re-trained model is part of the same layer (JudgeArena.fs): fuuga decide compare --data scores two engines on exactly the recorded outcomes both answered and tests the difference paired, by McNemar's exact test over the decisions they split, and for free-text output fuuga decide judge A/Bs two response sets with a typed judge asked in both presentation orders, reporting the judge's own position bias, order-flip rate and (against any human labels) its accuracy and ECE next to the win rate.
  • Agent Session Management -- Session lifecycle (init → active → completed/failed), save/load state, step-level experience recording
  • Orchestration Checkpoints -- Save and resume fan-out orchestration plans with per-subtask completion tracking
  • Cost Outcome Tracking -- Persist cost-aware routing outcomes for budget optimization across sessions

Model Compression:

  • Structured Pruning -- Attention head removal and layer removal with importance scoring
  • NF4 Quantization -- 4-bit NormalFloat quantization for weight compression
  • INT8 Weight-Only QLoRA -- Per-row symmetric INT8 base weights (8.5 bpw) with trainable LoRA adapters as a sibling to NF4 QLoRA — see Int8QLoraLinear; selected by a [[QuantPlan]] entry of BwInt8
  • Activation Calibration -- Chat-templated forward-pass probe collects per-tensor max-abs / mean-abs / RMS statistics over SFT data (fuuga calibrate). Drives the dynamic quantisation planner; the chat-template-on-instruct-data step is the key lesson from Unsloth's Dynamic 2.0 GGUFs
  • Dynamic QuantPlan -- Greedy per-tensor bit-width allocator that blends activation magnitude with optional Fisher importance to pack the model under a target average-bpw budget (fuuga plan-quant --target-bpw 4.5 --fisher F.bin). Preserves embeddings / lm_head / norms at higher precision, pushes FFN down-projections to the lowest bpw — a Fuuga-native analogue of Unsloth Dynamic 2.0
  • Per-Architecture Recipes -- QuantPlanConfig.forArchitecture "phi-3" | "llama-3" | "gemma-3" | "deepseek-v3" | "kimi-k3" returns a config with tuned PreserveFirstNLayers / PreserveLastNLayers boundary protections. Boundary layers (first / last 1-3 blocks) tolerate aggressive quant worst — Unsloth's empirical observation
  • MoE-Aware Quantization -- QuantPlanConfig.MoeExpertBitWidth routes .moe.expert* tensors to a separate bit width while keeping .moe.router at preserved precision. Routed experts tolerate aggressive quant far better than dense layers; this is where Dynamic 1.0's biggest absolute size savings come from on MoE deployments (DeepSeek-V3, Kimi K2)
  • Quantization-Aware Training (STE) -- Straight-Through Estimator for NF4 weights: forward pass sees quantized values, backward pass flows gradients through identity (--ste)
  • Compression Pipeline -- Orchestrated prune → fine-tune → quantize workflow for production deployment

GGUF Interop (llama.cpp / Ollama / LM Studio):

  • GGUF v3 Writer -- Full container emission (header + KV metadata + tensor info + aligned data) with Fuuga → llama-arch tensor name mapping. fuuga export gguf --checkpoint X --output model.gguf [--plan plan.json --tokenizer tokenizer/]
  • GGUF v3 Reader -- The inverse path: ingest pre-quantized GGUFs (Unsloth's own, Bartowski's quants) as donor weights for fuuga transfer / fuuga sft / fuuga distill. WeightTransfer.loadDonorWeights routes .gguf paths through GgufImport.readGgufDonor automatically; the reverse name mapping treats llama-arch FFN as SwiGLU
  • Tensor Encodings -- F32, F16, Q8_0 (8.5 bpw), Q4_0 (4.5 bpw), Q6_K (6.5625 bpw), Q4_K (4.5 bpw), Q2_K (2.625 bpw), IQ4_NL (4.5 bpw, ARM / Apple Silicon friendly). Block layouts follow llama.cpp/ggml-quants.c; every encoder/decoder pair is round-trip-validated in tests (F16 bit-exact; the quantized formats to per-format SNR floors)
  • Tokenizer Metadata -- --tokenizer <dir> embeds the full BPE vocab + merges + special-token ids + Jinja chat template so the resulting GGUF is directly runnable in llama.cpp / Ollama; without it the file is gguf-dump-inspectable but unloadable
  • Embedding Safety Floor -- Embeddings / lm_head automatically bump above 4-bit even when the plan asks for NF4 / Q2_K / Q4_K / IQ4_NL: K-family defaults floor to Q6_K, non-K to Q8_0. Override per-tensor via QuantPlan

Server:

  • OpenAI-Compatible API -- Separate fuuga-serve project with /v1/chat/completions, /v1/completions, /v1/models, and /v1/embeddings endpoints
  • Function Calling -- Client-supplied tools / tool_choice (auto, none, required, or a named tool) on /v1/chat/completions; tool definitions render into the model's native <|tool_call|> prompt format and emitted calls come back as OpenAI tool_calls (non-streaming and streaming deltas)
  • Structured Outputs -- response_format: json_schema routes the schema into grammar-constrained decoding; json_object becomes a JSON-only system instruction
  • SSE Streaming -- Server-Sent Events for real-time token streaming
  • Continuous Batching -- Iteration-level scheduler with a paged KV cache, admission/preemption, and an async engine, wired into the server via fuuga-serve --continuous-batching (memory-aware cache sizing; opt-in). Requests beyond the in-flight batch queue instead of blocking a thread. Uses a true batched-matmul forward for GQA-dense models (prefill + decode, parity-proven); MoE, NSA/H-Transformer, MLA, KDA and per-layer hybrids are served per-sequence within the same engine (the paged pool sizes each layer's pages from that layer's attention kind, and a KDA layer's fixed-size recurrent state lives in a sequence-keyed side table instead of pages).
  • Dynamic Batching -- Batch scheduling engine with metrics and backpressure, shared by the continuous-batching serving path.
  • Bearer Token Auth -- Optional API key authentication middleware
  • Guard Rails -- Prompt injection detection, PII masking, and content filtering. Enable with fuuga-serve --guardrails (input + output checks on all chat endpoints) or fuuga infer --guardrails; the core lives in Fuuga.GuardRails with an Oxpecker middleware adapter in Fuuga.Server.GuardRails
  • MCP Tool Routing -- Server-side MCP tool integration for function calling
  • A2A Server -- Agent-to-Agent protocol server endpoint for multi-agent workflows

Distributed Training:

  • PyTorch/DeepSpeed Integration -- Export model weights for distributed training, import trained weights back, and auto-generate launch scripts

Image Generation (Fuuga.Image):

  • Text-to-Image / Image-to-Image -- Stable Diffusion generation from text prompts (samplers, steps, guidance) and image transformation with denoising strength control
  • From-Scratch Stable Diffusion -- A complete SD 1.5 implementation written in F# on TorchSharp (CLIP BPE tokenizer + text encoder, VAE, UNet with cross-attention, DDIM/DDPM samplers, classifier-free guidance) that loads standard SD v1.x safetensors directly — educational and hackable; fuuga-image scratch, the fuuga_image_scratch MCP tool, or fuuga image --scratch
  • Video -- txt2video, img2video, and vid2vid (edit/continue guided by text, VACE) when the loaded model supports video (e.g. Wan2.1/2.2 GGUF); gated on the model's own capability flag. Output is a dependency-free PNG image sequence by default, or mp4/gif/webp via optional ffmpeg. --audio muxes a provided audio track into mp4/webp (video models make no audio).
  • Captioning / Understanding -- Describe an image or video (frame-sampled), or transcribe/describe audio, via whatever multimodal model you load (e.g. Phi-3.5-vision or Phi-4-multimodal) — brief/standard/detailed modes, all on the same ONNX-Runtime-GenAI runtime.
  • MCP Server Mode -- Exposes txt2img / img2img / txt2video / img2video / vid2vid / caption as MCP tools over stdio for integration with Fuuga LLM

Local Minion (Delegation):

  • Minion delegation -- Run Fuuga as a local "minion" that a more capable master agent (e.g. Claude Code) delegates small, well-scoped errands to, keeping simple work on a fully-local model with minimal energy cost
  • MCP server (fuuga-serve mcp) -- Exposes a fuuga_delegate tool plus local file tools to a master over stdio JSON-RPC; register several with distinct --name/--role to run a fleet of specialists (e.g. F# coder, C# coder, project manager)
  • Swappable brain -- The minion runs on a Fuuga-trained checkpoint (--checkpoint), an in-process GGUF via LLamaSharp (--gguf), or any OpenAI-compatible endpoint such as a local Ollama (--endpoint)
  • Sandboxed local tools -- read_file, list_dir, grep, plus permission-gated write_file, replace_in_file, apply_edits, and run_command; confined to a workspace --root (resists .. and symlink escapes), with writes/shell off by default
  • Multi-file refactors -- apply_edits (atomic edits across files: all land or none do), lsp_rename (semantic rename via the language server), --verify-command (build/tests gate the result; failures go back to the minion to fix, then escalate as verification_failed), and transcript compaction for long sessions. See docs/minion.md
  • Extended reach -- Opt into the built-in web tools (--web) and configured MCP servers (--mcp-config) so a delegated errand can fetch pages and call other tools, not just touch the filesystem
  • Agent Skills -- Anthropic-style, Claude-Code-compatible SKILL.md skills (--skills <dir>, default ./skills); the minion sees a compact catalog and loads a skill's full instructions on demand via load_skill (progressive disclosure). One registry also backs the A2A agent card. See docs/skills.md
  • Escalation contract -- Each delegation returns structured JSON {status, output, files_changed, escalate, reason, confidence, verification}; the minion self-verifies and hands work back (escalate=true) when it is not confident, so the master only spends its own capacity when needed
  • One-shot CLI (fuuga delegate) -- Run a single errand locally and print the JSON result, without a master; --fail-on-escalate turns an escalation into exit code 4 for CI gates

Observability:

  • OpenTelemetry -- OTLP trace and metrics export with Serilog integration
  • Spectre.Console -- Rich terminal output for training progress and diagnostics

Prerequisites

  • .NET 10 SDK (v10.0.103 or later)
  • GPU is optional -- CPU works for the dev configuration (small model). CUDA-capable GPU recommended for larger models.
  • ~500 MB disk space for dependencies, plus space for training data and checkpoints

Quick Start

For most users, the fastest path to useful output is to start from donor weights, not from scratch training.

Recommended paths:

  • examples/donor-transfer-and-refine.fsx -- practical donor-first workflow for normal users, with two modes:
    • SmokeTest for limited hardware, using a very small donor just to prove the F# pipeline is real
    • Practical for a few-GB donor model that gives much better output quality
  • examples/complete-pipeline.fsx -- educational train-from-scratch pipeline
  • examples/finetuning-lora-qlora.fsx, preference-tuning-dpo-rl.fsx, modelops-compress-and-export.fsx, evaluation.fsx, rag-grounded-generation.fsx, agents-and-delegation.fsx, inference-techniques.fsx -- focused, runnable examples per workflow (see examples/README.md)

If you want to understand the full pipeline from scratch, use the CLI below:

# Build (CPU):
dotnet build
# Build (GPU, ~2 GB dependency):
dotnet build -p:TorchBackend=cuda

# Train a tokenizer, ingest a corpus, train, and generate text
dotnet run -- tokenize --input data/raw --vocab-size 8000 --output data/tokenizer
dotnet run -- ingest --input data/raw --output data/corpus.fuge --tokenizer data/tokenizer
dotnet run -- train --corpus data/corpus.fuge --tokenizer data/tokenizer --checkpoint-dir checkpoints/
dotnet run -- infer --checkpoint checkpoints/step-100 --tokenizer data/tokenizer --prompt "Once upon a time"

See the Getting Started Tutorial for a complete end-to-end walkthrough.

Important expectation setting:

  • scratch training is educational and flexible, but tiny early runs often produce weak or gibberish text
  • donor transfer is the better starting point when you want coherent output quickly
  • short refinement on your own domain data is usually much more useful than starting from random weights

Use .fuge for tokenized corpus files and .fuuga for portable model packages.

Compressing a trained checkpoint for llama.cpp / Ollama

For deploying a trained Fuuga model into the llama.cpp ecosystem, the three-step pipeline produces a Fuuga-dynamic GGUF where critical tensors stay at higher precision and the rest land at an aggressive quant of your choice:

# 1. Collect per-tensor activation statistics from chat-templated SFT data.
dotnet run -- calibrate \
    --checkpoint checkpoints/step-N --data data/sft.jsonl \
    --output artefacts/cal.fcalib --batches 64

# 2. Plan per-tensor bit widths under a target average-bpw budget.
#    Pass --fisher to blend Fisher-importance with activation magnitude.
dotnet run -- plan-quant \
    --checkpoint checkpoints/step-N --calibration artefacts/cal.fcalib \
    --output artefacts/plan.json --target-bpw 4.5 [--fisher artefacts/F.bin]

# 3. Emit a runnable GGUF — --tokenizer embeds the BPE vocab+merges so
#    llama.cpp can load it directly. --default-type picks the encoding for
#    tensors the plan doesn't preserve.
dotnet run -- export gguf \
    --checkpoint checkpoints/step-N --plan artefacts/plan.json \
    --tokenizer data/tokenizer --default-type q4_k \
    --output deploy/model.gguf

Pick --default-type q2_k for extreme compression, iq4_nl for ARM / Apple Silicon, q6_k for the lowest-loss K-quant. Embeddings and lm_head automatically bump above 4-bit even when the plan asks for NF4 — see "GGUF Export" above.

F# Script Examples

Prefer the F# API over the CLI?

Practical donor-first path:

dotnet fsi examples/donor-transfer-and-refine.fsx

This loads donor weights, runs transfer into a Fuuga model, evaluates prompt outputs, and can do a short refinement pass. It is the recommended starting point for users who want useful results on limited hardware or with a few-GB donor model.

Train-from-scratch path:

dotnet fsi examples/complete-pipeline.fsx

This trains a tokenizer, ingests data, trains a model, and generates text -- all using the Fuuga modules directly. See examples/complete-pipeline.fsx for the full source.

For advanced workflows -- weight transfer from Phi-3/LLaMA3/DeepSeek, LoRA/QLoRA and preference tuning, quantize-and-export to GGUF, benchmark + quality evaluation, RAG, agents/delegation, and advanced decoding (speculative, constrained, chain-of-thought, streaming) -- see the focused scripts in examples/ (catalogued in examples/README.md).

Some examples draw an animated picture of what they do with --svg:

<table> <tr> <td align="center" width="50%"><a href="examples/continuous-batching-and-paged-cache.fsx"><img src="examples/_images/continuous-batching-and-paged-cache.svg" alt="Continuous batching with a paged KV cache" width="100%"></a><br><sub>Requests join the running batch as seats free up, each holding only the cache pages it has filled</sub></td> <td align="center" width="50%"><a href="examples/agents-and-delegation.fsx"><img src="examples/_images/agents-and-delegation.svg" alt="Delegating an errand to a local minion" width="100%"></a><br><sub>A master agent hands a small errand to a local model with sandboxed tools</sub></td> </tr> </table>

For image generation and captioning, see examples/image-demo.fsx.

Image Generation

Fuuga.Image is a standalone CLI for image generation and captioning. See the Fuuga.Image README for full command reference and MCP server mode, or run examples/image-demo.fsx.

Project Structure

Fuuga.fsproj                    # Project file with layered compilation order
Types.fs                        # All shared types (ModelConfig, TrainingConfig, GenerationConfig, etc.)
Logging.fs                      # ActivitySource/Meter definitions, ILoggerFactory
GuardRails.fs                   # Guardrails core: prompt-injection / PII / content checks (shared by CLI + server)
Observability.fs                # OpenTelemetry providers, Spectre.Console, --observe flag
DriftDetection.fs               # Statistical drift monitoring (KS, PSI) over confidence signals
DriftAlerting.fs                # Dual-threshold alerts, adaptive thresholds, OTel metrics, retraining triggers + decision log
DriftMonitor.fs                 # Serve-side runtime: signal accumulation, timed analysis, reference/trigger persistence
Config.fs                       # JSON config loading, CLI arg parsing, MCP config, LoRA target parsing
Validation.fs                   # Input validation pipeline with composable validators
Scaling.fs                      # Scaling heuristics from corpus and hardware stats
ConfigWizard.fs                 # Corpus analysis + hardware-aware config generation (Chinchilla scaling)
Tokenizer.fs                    # BPE tokenizer training and loading
ParquetIO.fs                    # HuggingFace Parquet dataset read/write (Document, SFT, DPO)
TextCleanup.fs                  # Ingestion/preparation text cleanup
RagCleanup.fs                   # RAG (Retrieval-Augmented Generation) cleanup algorithms
Ingest.fs                       # Document discovery and binary corpus writing
CorpusCompression.fs            # Zstd compression/decompression for .fuge files
Tensor.fs                       # Device selection (CPU/CUDA), DisposeScope
MultiResolutionAttention.fs     # Chunk pooling, global tokens for long context
Model.fs                        # GPT-2 transformer with RoPE, GQA, RMSNorm, SwiGLU, MLA
AttentionConfig.fs              # FlashAttention verification, SDPA backend selection
AutoConfig.fs                   # Auto-resolution of DU Auto* config cases from hardware probing
Vision.fs                       # Vision encoder for multimodal inputs
VisionBridge.fs                 # Q-Former cross-attention bridge for vision-to-language compression
PagedAttention.fs               # Paged KV-cache attention
MemoryHierarchy.fs              # Compressed memory, external retrieval
PersistentRetrievalStore.fs     # Disk-backed IRetrievalStore for cross-session retrieval
ConfidenceHead.fs               # Calibrated confidence MLP, Platt scaling, bucket assignment
EarlyExit.fs                    # Early exit / adaptive depth inference
Optimizer.fs                    # AdamW, SWA, Lookahead, and INT8 moment-quantized optimizers
Checkpoint.fs                   # Checkpoint save/load/metadata, safetensors
MmapLoading.fs                  # Memory-mapped model loading
GradientCheckpointing.fs        # Gradient checkpointing for memory-efficient training
DistributedTraining.fs          # Distributed training (PyTorch/DeepSpeed export/import)
ModelParallelism.fs             # Tensor/pipeline parallelism config for 70B+ models
GpuOffloading.fs                # Layer-wise CPU/GPU offloading
OptimizerOffload.fs             # Optimizer state offloading
NvmePaging.fs                   # ZeRO-Infinity 3-tier GPU/CPU/NVMe memory management
OnnxExport.fs                   # ONNX export with quantization and validation
GgufImport.fs                   # GGUF v3 read (foundation: shared fp16↔fp32, kvaluesIq4nl, dequantizers); routes .gguf donors into WeightTransfer
GgufExport.fs                   # GGUF v3 write (F32/F16/Q8_0/Q4_0/Q6_K/Q4_K/Q2_K/IQ4_NL) with tokenizer + QuantPlan; builds on GgufImport's helpers
DataMixture.fs                  # N-ary weighted data source mixing
VramGuard.fs                    # In-process VRAM monitoring with nvidia-smi polling, signal-based training loop integration
Training.fs                     # Training loop with AdamW/cosine LR, gradient offloading, per-param flush
FineTuningData.fs               # SFT/DPO JSONL parsing, chat templates, tokenization, batching
FineTuning.fs                   # LoRA (LoraLinear), SFT training, DPO loss/training, adapter save/load
RewardFunctions.fs              # Composable reward functions for RL training
DataValidation.fs               # SFT/DPO JSONL validation, honesty pattern classification
Fp8Dequantization.fs            # FP8 format dequantization with GPU LUT acceleration
Nf4Quantizer.fs                 # NF4/FP4 4-bit weight quantization, STE for QAT
QLoraTraining.fs                # QLoRA training (NF4 base + LoRA adapters), Int8QLoraLinear (per-row INT8)
LogitCache.fs                   # Offline top-K teacher logit cache (write once, train with no teacher resident)
DistillEval.fs                  # Distillation quality gate: held-out before/after + teacher-agreement feasibility
EndpointTeacher.fs              # Response-based distillation from an HTTP teacher (prompts -> SFT JSONL)
WeightTransfer.fs               # Weight transfer from donor models (Phi-3, LLaMA3, Mistral, Qwen, Gemma-3/4, DeepSeek, GLM, Kimi-K3 mappings)
ModelMerge.fs                   # N-ary model merging (TIES, DARE, Karcher) with EWC protection
Calibration.fs                  # Chat-templated activation-statistics probes for QuantPlanner
QuantPlanner.fs                 # Per-tensor BitWidth planner — Fuuga-dynamic 4.5-bpw plans via Fisher + activations
AttentionDistillation.fs        # Layer-level distillation: any attention pair, or whole blocks (KL + MSE loss)
Pruning.fs                      # Structured pruning (attention heads, layers)
CompressionPipeline.fs          # Prune → finetune → quantize orchestration
ConstrainedDecoding.fs          # Grammar-guided JSON constrained decoding
ChainOfThought.fs               # ThinkStart/ThinkEnd token handling, phase tracking
MathVerifier.fs                 # Symbolic math answer verification (boxed extraction, LaTeX normalization, exact-rational equivalence)
Verifier.fs                     # Rule-based and learned verifier strategies
ToolPolicy.fs                   # Confidence-aware tool invocation policy
McpClient.fs                    # MCP client: connection, tool discovery, tool invocation
WebSearch.fs                    # Web search integration for grounded generation
ImageRouting.fs                 # Image query routing to Fuuga.Image
A2AClient.fs                    # A2A protocol client for agent-to-agent communication
TreeSpeculation.fs              # Tree-structured speculative decoding
ContinuousBatching.fs           # Iteration-level continuous batching for serving
Inference.fs                    # Text generation with sampling, tool-augmented generation, structured output, CoT
DraftAndRefine.fs               # Draft-and-refine multi-pass reasoning pipeline
AdvancedReasoning.fs            # Consensus voting, verifier-scored, tree-of-thoughts
Orchestrator.fs                 # Multi-model orchestrator with capability-based routing
BackendClient.fs                # HTTP client for external OpenAI-compatible LLM endpoints
ExperienceStore.fs              # Append-only experience log, strategy lessons, managed store
CostAwareRouting.fs             # Cost-aware routing with budget tracking
AgentSession.fs                 # Agent session lifecycle, save/load, orchestration checkpoints
FanOutOrchestration.fs          # Fan-out/fan-in task decomposition and aggregation
ContextAwareness.fs             # Convention file discovery, language/framework detection, git context
SemanticKnowledge.fs            # WordNet WNDB parser, token-to-synset mapping, multi-lingual
DataAugmentation.fs             # Synonym replacement, paraphrasing, token noise, oversampling
VerifierRuntime.fs              # Verifier loading and runtime integration
MinionTools.fs                  # Sandboxed local tools (read/list/grep/write/replace/run) for delegated errands
MinionBrain.fs                  # Swappable minion brain: Fuuga checkpoint, GGUF (LLamaSharp), or OpenAI endpoint
MinionEscalation.fs             # Escalation contract: self-verify + uncertainty/budget signals -> hand back to master
MinionToolEnv.fs                # Minion reach beyond local files: built-in web tools, MCP servers, and skills (load_skill)
Skills.fs                       # Anthropic-style SKILL.md registry (parse/discover/catalog); backs minion, MCP host, and A2A card
MinionAgent.fs                  # Delegate loop: tool-use rounds, structured DelegateResult, escalation verdict
MinionCli.fs                    # Shared minion flag parsing (brain, sandbox, reach, identity, escalation)
Eval.fs                         # Benchmark datasets, runners, result serialization
Program.fs                      # CLI entry point with subcommand routing (incl. `delegate`)

Fuuga.Server/                   # OpenAI-compatible HTTP server (separate project)
  ApiTypes.fs                   # Request/response types (OpenAI-compatible)
  McpToolRouting.fs             # Server-side MCP tool routing for function calling
  FuugaChatClient.fs            # IChatClient adapter for Microsoft AI ecosystem
  FuugaEmbeddingGenerator.fs    # IEmbeddingGenerator adapter for /v1/embeddings
  A2AServer.fs                  # A2A protocol server endpoint
  GuardRails.fs                 # Oxpecker/ASP.NET middleware adapter over the Fuuga.GuardRails core
  DynamicBatchingServer.fs      # Dynamic batching with HTTP/SSE integration
  Server.fs                     # Oxpecker HTTP server with SSE streaming, auth middleware
  MinionServer.fs               # Minion MCP server (stdio JSON-RPC): fuuga_delegate + local tools
  Program.fs                    # Server entry point (incl. `mcp` minion mode)

Fuuga.Image/                    # Standalone image generation and captioning CLI
  Types.fs                      # Domain types, error handling (ImageError DU)
  Config.fs                     # CLI argument parsing
  ImageIO.fs                    # Image load/save, format conversion, validation
  Diffusion.fs                  # Stable Diffusion model wrapper (txt2img, img2img)
  Caption.fs                    # Phi-3.5-vision captioning (ONNX Runtime GenAI)
  McpServer.fs                  # MCP JSON-RPC 2.0 server over stdio
  Program.fs                    # Entry point, subcommand routing

Fuuga.Tests/                    # Unit and integration tests (xUnit + FsUnit)
Fuuga.PropertyTests/            # FsCheck property tests over the RAG ingestion path
Fuuga.Image/Fuuga.Image.Tests/  # Image module tests
docs/                           # User-facing documentation
examples/                       # Runnable F# script examples
scripts/                        # Training data generation and validation scripts

Documentation

Architecture

Fuuga uses a layered module architecture with strict dependency ordering enforced by F#'s compilation model:

Layer 0:  Types, Logging, Observability           (foundation, no dependencies)
          DriftDetection, DriftAlerting, DriftMonitor (statistical drift monitoring, OTel alerts, serve-side loop)
Layer 1:  Config, Validation, Scaling             (configuration, validation, heuristics)
          ConfigWizard                             (hardware-aware config generation)
Layer 2:  Tokenizer, ParquetIO                    (BPE training/loading, Parquet dataset I/O)
Layer 3:  Ingest, CorpusCompression               (document discovery, corpus writing, Zstd compression)
Layer 4:  Tensor, MultiResolutionAttention        (device selection, chunk pooling + global tokens)
          Model, AttentionConfig, AutoConfig       (transformer with MLA, FlashAttention, auto-resolution)
          Vision, VisionBridge                     (vision encoder, Q-Former bridge)
          PagedAttention, MemoryHierarchy          (paged KV-cache, compressed memory)
          PersistentRetrievalStore                 (disk-backed retrieval for cross-session use)
          ConfidenceHead, EarlyExit                (calibration, adaptive depth)
Layer 5:  Optimizer, Checkpoint, MmapLoading       (optimizer variants incl. INT8 moments, save/load/metadata, mmap loading)
          GradientCheckpointing                    (activation recomputation)
          DistributedTraining, ModelParallelism    (PyTorch/DeepSpeed, tensor/pipeline parallel)
          GpuOffloading, OptimizerOffload          (CPU/GPU memory management)
          NvmePaging, OnnxExport                   (NVMe paging, ONNX export)
          GgufImport, GgufExport                   (GGUF v3 read/write; Import owns shared format helpers, Export builds on them)
          DataMixture, VramGuard                     (data mixing, in-process VRAM monitoring)
Layer 6:  Training, FineTuningData, FineTuning    (training loop, SFT/DPO data, LoRA training)
          RewardFunctions, DataValidation          (composable RL rewards, JSONL validation)
          Fp8Dequantization, Nf4Quantizer          (FP8/NF4 quantization support, GPU LUT, STE for QAT)
          QLoraTraining                            (NF4 + INT8 weight-only QLoRA, optional QuantPlan-driven dispatch)
          WeightTransfer, ModelMerge               (donor model transfer, N-ary merging)
          Calibration, QuantPlanner                (chat-templated activation probes, per-tensor BitWidth allocator)
          AttentionDistillation                    (any-attention-pair / block distillation)
          Pruning, CompressionPipeline             (structured pruning, prune→finetune→quantize)
Layer 7:  ConstrainedDecoding, ChainOfThought     (generation extensions)
          Verifier, ToolPolicy, McpClient          (verifier strategies, tool policy, tool calling)
          WebSearch, ImageRouting, A2AClient        (search, image routing, A2A protocol)
          TreeSpeculation, ContinuousBatching      (speculation, batch scheduling)
          Inference                                (text generation with sampling, tools, structured output)
          DraftAndRefine, AdvancedReasoning        (multi-pass reasoning, consensus/tree-of-thoughts)
          Orchestrator, BackendClient              (multi-model routing, external LLM client)
          ExperienceStore, CostAwareRouting        (cross-session persistence, cost optimization)
          AgentSession, FanOutOrchestration        (agent lifecycle, parallel task decomposition)
Layer 8:  ContextAwareness, SemanticKnowledge     (project context, WordNet)
          DataAugmentation                         (synonym replacement, paraphrasing, token noise)
          VerifierRuntime                          (verifier loading and runtime integration)
          Eval                                     (benchmark datasets, runners, result serialization)
Layer 9:  Program                                  (CLI entry point, subcommand routing)

More about design decisions can be read from the architecture document.

Running Tests

dotnet test Fuuga.Tests

Tests cover tokenization, ingestion, Parquet I/O, corpus compression, model architecture, MLA attention, vision encoder, vision bridge, paged attention, multi-resolution attention, attention configuration, early exit, training, optimizer variants, INT8 optimizer moments, gradient checkpointing, GPU offloading, optimizer offloading, ONNX export, GGUF v3 export (including Q6_K/Q4_K/Q2_K/IQ4_NL round-trip SNR + tokenizer metadata + embedding safety floor), inference, verifier-guided generation, speculative decoding, checkpoints, memory-mapped loading, fine-tuning, prompt tuning, QLoRA (NF4 and INT8 variants), reinforcement learning, reward functions, data validation, data augmentation, benchmark evaluation, FP8 dequantization, FP8 GPU LUT dequantization, NF4 quantization, STE quantization-aware training, activation calibration (chat-templated probes + binary report round-trip), QuantPlan synthesis (Fisher-blended importance ranking + bpw-budget allocator + JSON round-trip), structured pruning, compression pipeline, constrained decoding, chain-of-thought, MCP client, A2A client/server, web search, image routing, draft-and-refine, advanced reasoning, math verifier (symbolic equivalence, boxed extraction, equivalence-class voting), thinking budgets (s1-style budget forcing), orchestrator, backend client, cost-aware routing, fan-out orchestration, weight transfer, attention distillation, model merging, distributed training, model parallelism, scaling, validation, continuous batching, dynamic batching server, experience persistence, agent session management, persistent retrieval, drift detection, drift alerting, context awareness, semantic knowledge, observability, guard rails, memory strategy (gradient offloading, per-param flush, VRAM guard lifecycle, ConfigWizard recommendations, multi-GPU parallelism), and CLI integration. Tests use xUnit with FsUnit assertions and include both unit tests and end-to-end integration tests.

Assertions that FsUnit cannot express without collapsing the interesting values into a bare true — structural equality on DU-wrapped values, tensor closeness, quantifiers over collections, Result/Option cases — go through Fuuga.Tests/TestAssertions.fs (shouldEqualValue, shouldBeAllCloseAtol, shouldAllSatisfy, shouldContainSatisfying, shouldBeOk, …). The pass/fail decision is identical to the |> should equal true form they replace, but a failure reports the two sides of the equality, the worst element of a tensor comparison and the tolerance it needed, the first element to fail a predicate, or the Error that arrived instead of Ok. The helpers are themselves tested in TestAssertionsTests.fs. Tests whose assertions are tight enough to depend on a particular random draw take it from a private torch.Generator, never from torch.manual_seed, which seeds a process-global generator that a concurrently running collection can reseed mid-test (see TestDeterminism.fs).

Fuuga.PropertyTests is an FsCheck suite over the RAG ingestion path (dotnet test Fuuga.PropertyTests): clean prose comes back unchanged from TextCleanup and RagCleanup, code blocks behind <|code_start|> markers are never touched, the noise the passes are written for (control characters, curly quotes, ligatures, zero-width characters, page numbers, dot leaders, decorative lines, doubled words, mojibake, odd spaces) is gone afterwards without a letter lost, cleaning twice is cleaning once, the markdown walker keeps every fenced block verbatim with its normalised language tag, the JSONL reader turns exactly the chat records into documents, the FIM transform is a permutation of the text, and a corpus reads back token for token as it was written.

Technology Stack

Component Library Purpose
Tensors & GPU TorchSharp 0.106.0 Tensor operations, CUDA support
LibTorch libtorch-cpu 2.10.0 LibTorch CPU backend
Tokenization Microsoft.ML.Tokenizers 2.0.0 BPE tokenizer training
ONNX Runtime Microsoft.ML.OnnxRuntime 1.24.3 ONNX model inference
ONNX Export OnnxSharp 0.3.2 ONNX model construction and manipulation
Protobuf Google.Protobuf 3.34.0 Protobuf serialization for ONNX
Epub parsing VersOne.Epub 3.3.4 Extract text from epub files
Markdown Markdig 1.1.1 Parse markdown to plain text
Parquet Parquet.Net 5.5.0 HuggingFace-compatible dataset I/O
Compression ZstdSharp.Port 0.8.7 Zstandard corpus compression
Statistics MathNet.Numerics.FSharp 5.0.0 Drift detection (KS test, PSI)
Logging Serilog 4.3.0 Structured logging
Logging sinks Serilog.Sinks.Console 6.0.0, .File 6.0.0, .OpenTelemetry 4.2.0 Console, file, and OTel log sinks
Logging bridge Serilog.Extensions.Logging 9.0.0 Serilog/Microsoft.Extensions.Logging bridge
JSON FSharp.SystemTextJson 1.4.36 F# DU-aware serialization
Telemetry OpenTelemetry 1.11.2 Distributed tracing and metrics
Telemetry export OpenTelemetry.Exporter.OpenTelemetryProtocol 1.11.2 OTLP protocol export
Telemetry hosting OpenTelemetry.Extensions.Hosting 1.11.2 OpenTelemetry hosting integration
Terminal UI Spectre.Console 0.49.1 Rich terminal output
HuggingFace TorchSharp.PyBridge 1.4.3 HuggingFace weight loading
SIMD System.Numerics.Tensors 10.0.5 Preprocessing acceleration
AI abstractions Microsoft.Extensions.AI 10.4.0 IChatClient adapter
AI evaluation Microsoft.Extensions.AI.Evaluation 10.4.0 Benchmark evaluation framework
MCP ModelContextProtocol 1.1.0 MCP client SDK for tool calling
A2A A2A 0.3.3-preview Agent-to-Agent protocol client
Testing xUnit 2.9.3 + FsUnit.xUnit 7.1.1 Unit and integration tests
Server: HTTP Oxpecker 2.0.0 F# HTTP server framework
Server: A2A A2A.AspNetCore 0.3.3-preview A2A protocol server
Image: Diffusion StableDiffusion.NET 5.0.0 Stable Diffusion model wrapper
Image: Captioning Microsoft.ML.OnnxRuntimeGenAI 0.12.1 Phi-3.5-vision captioning
Image: Processing HPPH.SkiaSharp 1.0.0 Image load/save, format conversion

License

See LICENSE for details.

Product Compatible and additional computed target framework versions.
.NET net10.0 is compatible.  net10.0-android was computed.  net10.0-browser was computed.  net10.0-ios was computed.  net10.0-maccatalyst was computed.  net10.0-macos was computed.  net10.0-tvos was computed.  net10.0-windows was computed. 
Compatible target framework(s)
Included target framework(s) (in package)
Learn more about Target Frameworks and .NET Standard.

NuGet packages

This package is not used by any NuGet packages.

GitHub repositories

This package is not used by any popular GitHub repositories.

Version Downloads Last Updated
1.0.28 37 10/1/2026
1.0.27 46 9/29/2026
1.0.26 96 9/26/2026
1.0.25 84 9/25/2026
1.0.24 86 9/25/2026
1.0.23 98 9/22/2026
1.0.22 94 9/20/2026
1.0.21 88 9/18/2026
1.0.20 128 8/15/2026
1.0.19 130 8/1/2026
1.0.18 118 7/31/2026
1.0.17 116 7/29/2026
1.0.16 141 7/26/2026
1.0.15 128 7/12/2026
1.0.14 140 7/11/2026
1.0.13 124 7/6/2026
1.0.12 132 7/3/2026
1.0.11 131 7/2/2026
1.0.10 149 6/30/2026
1.0.9 143 6/29/2026
Loading failed

MINION MULTI-FILE REFACTORING. New apply_edits tool: several exact-text replacements across files as one change; nothing is written if any edit fails. New lsp_rename tool: semantic rename across files through the language server (needs --allow-write). fuuga delegate --verify-command (with --verify-timeout, --verify-retries) runs a build or test command after file changes, hands failures back to the minion and escalates with verification_failed if they persist; the result JSON gains "verification". --max-context-chars elides the oldest tool results from a long transcript (default 48000). --fail-on-escalate exits with code 4. LIBRARY API: MinionAgent.runDelegateWithOptions and DelegateOptions; LspClient.Rename, applyTextEdits, parseWorkspaceEdit. BREAKING for code that constructs these records: DelegateResult gains Verification, EscalationInputs gains VerificationFailed and VerificationPassed; EscalationReason gains the case VerificationFailed. DEPENDENCY UPDATES: Microsoft.Extensions.AI 10.10.0 (and the three Evaluation packages), Microsoft.ML.OnnxRuntime 1.30.0, OpenTelemetry 1.19.1 (with the OTLP exporter and Extensions.Hosting), ModelContextProtocol 2.2.0, Google.Protobuf 3.36.2, Markdig 1.4.0, Parquet.Net 6.1.0, System.Numerics.Tensors 10.0.12, TorchSharp.PyBridge 1.4.4. TESTS: four shared-state races between parallel test collections fixed. PREVIOUS (1.0.27) - DONOR IMPORTS: LATENT ATTENTION AND ENCODER DONORS. MLA ATTENTION NOW TRANSFERS (DeepSeek-V3, Kimi K2). --mapping deepseek / deepseek-moe previously skipped attention entirely, documented as "MLA, incompatible with Fuuga's GQA" - true only of a GQA TARGET. Against a target whose AttentionType is MLA the donor maps almost one-to-one: q_a_proj to wDQ, kv_a_proj_with_mqa split to wDKV + wKR, q_b_proj split to wUQ + wQR, kv_b_proj split to wUK + wUV, o_proj to wO, and q_a_layernorm / kv_a_layernorm into new latent RMSNorms. The two _b_proj tensors INTERLEAVE their halves within each head, so the split point is read from the target's MlaConfig rather than from the tensor: a contiguous split would hand each part the wrong heads' rows at identical shapes, and a row count the geometry cannot explain is reported as a shape error naming the tensor rather than reshaped into something plausible. Kimi K2 ships as model_type deepseek_v3, so one mapping covers both. MlaConfig gains NopeHeadDim, VHeadDim and LatentNormEps, all optional - omitted they reproduce the previous behaviour exactly, and configs written before they existed keep loading unchanged. They exist because the per-head dims cannot stay tied to EmbDim / NumHeads: DeepSeek-V3 runs 128 heads over a 7168-wide residual, so that ratio is 56 while its content and value head dims are both 128, which is precisely what made those donors' attention unrepresentable. MultiHeadLatentAttention decouples the content and value head dims across its projections, its absorption math and all three forward paths, and gains RMSNorm on the compressed latents; the KV cache layout is unchanged. The target must carry the donor's per-head geometry - layer count and FFN/expert widths may still be smaller. Kimi K3 is NOT covered: its KDA and gated-MLA layers have no key map. DeepSeek's YaRN RoPE scaling has no equivalent (Linear / NTKAware only) - no weights are involved, but long-context behaviour will differ. The layout is verified by construction and by tests that would catch a mis-split by value; it has not been run against a real DeepSeek or Kimi checkpoint. MODERNBERT ENCODER DONOR (--mapping modernbert): the one encoder in the list, producing a BIDIRECTIONAL model for the typed-decision path rather than a generative one - which is how an open judge model such as Laya 1 comes in, its encoder subfolder being a stock ModernBertForMaskedLM. Fused Wqkv and Wi reuse the Phi-3 splitters; layer 0's attn_norm is nn.Identity and is never saved, so block0.norm1 is filled from embeddings.norm, the norm that actually runs before layer 0's attention. The masked-LM head is not transferred, and the decision heads are Fuuga's own, trained by fuuga decide train. TRANSFER NOW WRITES checkpoint.json beside model_weights.dat, so its output is a checkpoint directory rather than loose weights and fuuga decide train, which has no --model-config, can consume a transfer result directly. It is written on every transfer including over an existing file, because model_weights.dat has just been replaced and metadata left from whatever used to live there would now describe a different model; an explicit --model-config still wins downstream. DISTILL REFUSES A TEACHER WHOSE ATTENTION WAS NEVER POPULATED: --attention and --hidden match the student against the teacher's own modules, so a mapping that fills no attention leaves the teacher at initialisation and trains the student against noise while the loss still falls convincingly. A .gguf teacher is exempt - it arrives keyed by Fuuga parameter names and bypasses the key map entirely. SMALLER: checkpoint metadata records the real assembly version instead of a hardcoded "0.1.0"; a cross-collection Console race in the CLI drift guard is fixed. EXAMPLE: examples/kimi-donor-laya-judge.fsx runs the whole scenario offline on CPU - a Kimi/DeepSeek-shaped MLA donor and a Laya/ModernBERT encoder donor imported, decision heads trained on the result, and the A/B arena run twice to show what the report says when a judge is unusable versus usable. PREVIOUS (1.0.26) - ROBUSTNESS FIXES FROM A STATIC-ANALYSIS PASS (fsharp-refactor notes, each finding verified in the source before it was changed). GUARDRAILS: one invalid regex in CustomPatterns threw out of the input check on every checked request, and every custom pattern was recompiled per input; the input check now shares the output side's compiled cache, which logs an invalid pattern once and skips it while the valid ones still apply. RESUME NO LONGER DESTROYS STATE: `fuuga agent --resume` reported an unreadable agent-session.json as "no previous session" and then overwrote it, and `fuuga orchestrate --resume` silently replaced a checkpoint it could not read while re-running every subtask; both files are now moved aside as <name>.unreadable-<UTC timestamp> with a message saying so (new AgentSession.setAsideUnreadable). SERVER (fuuga-serve): a literal JSON null request body bound to a null record and the handler answered 500; it is now a 400, and all six JSON endpoints report the serializer's own message ("Could not parse request body: ...") where five said only "Invalid JSON" even for valid JSON that did not bind. Tool-call arguments the model produced that are not a JSON object now travel as the call's error (FunctionCallContent.Exception) instead of running the tool with no arguments. The guardrail-filtered streaming path awaits the output enumerator's disposal instead of blocking a request thread on it. SMALLER: the LSP client disposes its cancellation source; the weight-streaming autograd warning no longer names {Blocks} twice in one template; the repetition penalty enumerates the token history once; single-character StartsWith/EndsWith checks (comment markers, JSON braces, trailing newlines, the BPE space marker) compare ordinally, so an invisible character such as a soft hyphen no longer changes the answer. No public API removed. PREVIOUS (1.0.25) - DRIFT MONITORING AND TRAINING SAFETY FIXES. SUSTAINED DRIFT NOW KEEPS ALERTING. The adaptive threshold (mean + 2 sigma over the last 30 windows) is a rolling one, so it absorbs any level that lasts - which is exactly what stops it alerting on a signal's own noise, and equally stopped it alerting on drift nobody had fixed. A sustained shift fired one alert in 52 windows and then went quiet; a shift that began soon after a reference capture, or in the very first analysed window, never alerted at all; and a critical plateau sitting just over the threshold never alerted under any previous version. Operators saw Critical in GET /v1/drift with no pending retraining trigger. The adaptive gate may now delay an alert but never suppress one: a new SustainedCriticalWindowsRequired (default 10) fires on consecutive critical windows whatever the baseline makes of them, because a signal's own noise crosses back below critical long before ten windows and a real shift does not. The gate also now needs AdaptiveMinimumSamples (default 5) windows before it has an opinion at all - a threshold built from two scores sits barely above the pair and the next ordinary window clears it. Measured on the same series, before and after, at the shipped defaults: sustained shift after a healthy run 1 alert then silence, now alerting every 10 windows indefinitely; shift two windows after a reference capture 0, now 4 in 42 windows; shift from the first window 0, now 4 in 40; critical plateau just over the line 1, now 4. Nothing new fires on quiet signals: isolated spikes 0 before and after, a noisy stationary signal wandering across the warning line 0 before and after, a stationary signal sitting on the critical line 0 before and 0.15 per 300 windows after - that residual is inherent, since ten critical windows in a row out of a signal that is critical half the time by chance cannot be told from a genuine plateau except by waiting longer; raise SustainedCriticalWindowsRequired to trade it away. A trigger the backstop raises says so in its brief - it reports that the adaptive threshold has absorbed the level rather than quoting it as exceeded - so an operator approving from the brief sees why it fired. An earlier attempt that kept critical scores out of the baseline was measured and rejected: it truncates the baseline to the quiet tail, so a signal that normally sits near the threshold is compared against a picture of its calmest windows and crosses constantly. TRAINING NO LONGER OVERWRITES ITS OWN SOURCE. fuuga decide train and fuuga merge both write model_weights.dat and checkpoint.json flat into --output, under the same names they read from their input, so pointing both at one directory replaced the source in place: decide train left the donor re-tagged Bidirectional, which fuuga serve then refuses as a generative model, and merge dropped BaseCheckpointPath, so a second run would load the already-merged weights and apply the adapters on top of them again. Both now exit 1 with an explanation before any work happens (Checkpoint.resolvesToSameDirectory). TESTS: every alerting test shrank the configuration to 1-2 consecutive windows, a 3-5 window baseline and 1 sigma, so the shipped numbers were never exercised; a new section runs the defaults from both directions - drift that does not resolve must keep alerting, and a stationary signal must not alert however close to the threshold it sits - with the stationary cases swept over seeds rather than resting on one.