Fuuga 1.0.28
dotnet add package Fuuga --version 1.0.28
NuGet\Install-Package Fuuga -Version 1.0.28
<PackageReference Include="Fuuga" Version="1.0.28" />
<PackageVersion Include="Fuuga" Version="1.0.28" />
<PackageReference Include="Fuuga" />
paket add Fuuga --version 1.0.28
#r "nuget: Fuuga, 1.0.28"
#:package Fuuga@1.0.28
#addin nuget:?package=Fuuga&version=1.0.28
#tool nuget:?package=Fuuga&version=1.0.28
Fuuga
Tired of paying tokens? Think you could train a better model? Well, now you can try.
An LLM built from scratch in F# and .NET. Fuuga implements a complete language model pipeline: tokenization, data ingestion, model training, fine-tuning, and text generation -- with no Python dependencies.
Built on TorchSharp for tensor operations and Microsoft.ML.Tokenizers for BPE, Fuuga uses idiomatic F# (discriminated unions, pipelines, immutability) throughout. Works on GPU or CPU.
Choose your path
New here? Each goal is a short sequence of commands driven by one recipe file. Scaffold a recipe, edit a few paths, then run the commands with --recipe. See docs/workflows.md for the full journeys.
| I want to... | Start with |
|---|---|
| Make an open model follow my data | fuuga scaffold finetune — Fine-tune a donor (the common path) |
| Train a small model from scratch | fuuga scaffold pretrain — Pre-train |
| Shrink/export a model for deployment | Shrink & export |
| Score a model | Evaluate |
| Serve or build an agent on a model | Serve & integrate |
| Review pull requests in CI/CD with my own model | Code review in CI/CD — scripts/ci-review.sh + fuuga-serve |
| Rent/sell my model to others via OpenRouter | Sell as a service — pricing + prepaid credit ledger |
Getting Fuuga
Fuuga comes in two forms, reflecting how it's used:
F# library — NuGet. Reference it from a script (
#r "nuget: Fuuga") or a project (<PackageReference Include="Fuuga" />) and call the API directly. Best for power users composing steps the CLI doesn't expose — see theexamples/catalogue (complete-pipeline.fsx,finetuning-lora-qlora.fsx,modelops-compress-and-export.fsx, and more).The NuGet packages come in two backend flavors, and the package choice — not a runtime flag — decides CPU vs GPU:
Fuugacarries CUDA 12.8 natives (win-x64; needs an NVIDIA GPU),Fuuga.cpucarries CPU natives and runs anywhere. The same split applies to the image module:Fuuga.Image(CUDA) andFuuga.Image.cpu(CPU). When building from source, theTorchBackendMSBuild property selects the backend:dotnet build -p:TorchBackend=cudafor CUDA, plaindotnet buildfor the CPU default.CLI and server — clone the repo. The
fuugaCLI and thefuuga-serveserver are not published as standalone tools: their native dependency (libtorch) is over a gigabyte — too large for a NuGet package or adotnet tool. So clone the repo and run from source:- CLI — shown throughout the docs as
fuuga <command>; from a clone, run it asdotnet run -- <command>. - Server —
dotnet run --project Fuuga.Server -- --checkpoint <dir> --tokenizer <dir> [--port <N>](this is whatfuuga-serverefers to).
- CLI — shown throughout the docs as
Improving the server's distribution story (a slimmer, CPU-only or download-on-first-run tool) is an open area for contribution.
Two ways to use Fuuga, then: the CLI for the full pipeline (config-driven via --recipe), or the F# API from scripts. Full command list below; full reference in docs/cli-reference.md.
Features
Core Pipeline:
- BPE Tokenizer -- Train a byte-pair encoding tokenizer on your own corpus with configurable vocabulary size
- Data Ingestion -- Discover and tokenize epub, markdown, Parquet, and plain text files into a binary corpus
- Parquet I/O -- Read and write HuggingFace-compatible Parquet datasets for SFT, DPO, and document data
- Corpus Compression -- Zstd compression/decompression for
.fugecorpus files - GPT-2 Transformer -- Decoder-only causal transformer with rotary position embeddings (RoPE), grouped-query attention (GQA), RMSNorm, and SwiGLU activation
- Multi-Head Latent Attention (MLA) -- DeepSeek-V2 style compressed KV cache with query/KV compression, decoupled RoPE keys, and optional weight absorption for reduced memory during inference
- Kimi Delta Attention (KDA) -- Kimi K3 / Kimi-Linear style linear attention: delta-rule recurrence with channel-wise lower-bounded decay (fixed-size state, NoPE), plus per-layer hybrid layouts (KDA-local / MLA-global every Nth layer, K3's 3:1 pattern) (experimental)
- Stable LatentMoE & Block Attention Residuals -- Kimi K3 architecture options: routed experts in a compact latent space with pre-up-projection RMSNorm; learned depth-wise attention over block residual summaries (experimental)
- Sigmoid MoE routing + Quantile Balancing -- DeepSeek-V3/K3-style sigmoid router scores with a frozen selection bias (imported from donor
e_score_correction_bias) and aux-loss-free load balancing via router-score quantiles - SiTU-GLU activation & Muon optimizer -- K3's softcapped GLU (bounded activations for low-precision training) and Muon (Newton-Schulz orthogonalized momentum, per-head variant); MTP draft heads can train on the EAGLE-3-style LK acceptance-rate loss
- Vision Encoder -- Vision model support for multimodal inputs
- Vision Bridge -- Q-Former cross-attention bridge that compresses vision patch tokens into learned query vectors for multimodal (image+text) inputs
- Paged Attention -- Paged KV-cache attention for efficient memory usage during long-context generation
- Memory Hierarchy -- Compressed memory with external retrieval for extended context
- Multi-Resolution Attention -- Chunk pooling with global tokens for efficient long-context processing
- FlashAttention Config -- SDPA backend selection and benchmarking for attention kernels
- Auto Config -- Hardware-aware auto-resolution of DU configuration cases (norm, activation, precision, offloading, communication) at startup
- Early Exit -- Adaptive depth inference for faster generation when confidence is high
- Training -- AdamW optimizer with cosine learning rate scheduling, warmup, gradient clipping, mixed precision support, and gradient accumulation
- Multi-Token Prediction (MTP) -- DeepSeek-V3 style auxiliary heads predicting multiple future tokens (configurable
Depth); adds a weighted multi-depth loss during training for better sample efficiency and powers MTP-drafted speculative decoding for faster generation. Enable from the CLI withtrain --mtp-depth <N> [--mtp-loss-weight <f>], or setMtpConfigin the model-config JSON - INT8 Optimizer Moments -- Optional INT8 quantization of AdamW M/V moment tensors with per-row symmetric quantization, reducing optimizer memory ~4× (
--moment-quant int8). Mutually exclusive with SWA/Lookahead/SAM, differential per-group LR, gradient offload, per-param flush, and NVMe/CPU optimizer offload (combining them fails fast). INT8 moments are not persisted across checkpoint resume — model weights resume normally while momentum/variance restart fresh. - Optimizer Variants -- Stochastic Weight Averaging (SWA) and Lookahead optimizer support with checkpointable optimizer state
- Gradient Checkpointing -- Memory-efficient training via activation recomputation
- GPU Offloading -- Layer-wise CPU/GPU offloading for reduced VRAM usage
- Optimizer Offloading -- Offload optimizer states to CPU memory
- NVMe Paging -- ZeRO-Infinity 3-tier GPU/CPU/NVMe memory management for training models larger than available VRAM
- Per-Tensor Gradient Offloading -- Bulk-copy gradients to CPU after backward pass and restore before optimizer step, freeing GPU VRAM during the optimizer phase (
--grad-offload) - Per-Parameter CUDA Flush -- Aggressive CUDA cache cleanup after optimizer step to reclaim transient VRAM spikes from M/V update temporaries (
--flush-each-param) - VRAM Guard -- In-process background thread that polls GPU memory via nvidia-smi and signals the training loop to warn, skip batches, or abort when usage exceeds a configurable threshold (
--vram-guard-gb <float>) - Memory Strategy Presets --
MemoryStrategyConfigs.none,.constrained(grad-offload + flush), and.full(all three with VRAM guard at 95%) with automatic ConfigWizard recommendations based on model-to-VRAM ratio - Model Parallelism -- Tensor and pipeline parallelism configuration for 70B+ parameter models, with automatic DataParallel recommendation for multi-GPU setups
- Inference -- Greedy, top-k, top-p (nucleus), and temperature sampling with repetition penalty
- Fill-in-the-Middle -- FIM support with prefix/suffix tokens for code completion
- Checkpoints -- Save, load, resume training from checkpoints with full metadata; safetensors format support
- Memory-Mapped Loading -- mmap-based model loading for fast startup
- Streaming Inference -- Token-by-token generation with configurable stop conditions
- Confidence Signals -- Entropy, repetition rate, hedging detection, calibrated confidence with Platt scaling, and stop reason reporting
- Drift Detection -- Statistical drift monitoring (Kolmogorov-Smirnov, Population Stability Index) over confidence signals with ring-buffered accumulation
- Drift Alerting -- Dual-threshold alerts with adaptive sigma-based thresholds, OpenTelemetry metrics, and retraining triggers that carry an approve-or-reject brief (evidence, proposed action, alternative, expected effect, risk)
- Drift Monitor --
fuuga-serve --drift-state <dir>feeds every generation's signals to the monitor, captures the reference window, tests on a timer and files triggers;fuuga drift status|decide|reset-referenceis the operator loop — verdicts are logged to the experience store and reported as per-signal trigger precision with tuning hints. Nothing is retrained automatically - ONNX Export -- Export to ONNX format with fp16/int8 quantization, validation, and benchmarking
- ONNX Inference -- ONNX Runtime backend for optimized inference (
--backend onnx) - Benchmark Evaluation -- Built-in benchmark runner for MMLU, HellaSwag, ARC-Challenge, WinoGrande, PIQA, TruthfulQA (log-likelihood MCQ scoring), GSM8K (generative, numeric-answer extraction), MATH-500 (generative; symbolic verifier grading with per-difficulty-level sub-scores and
--maj-kself-consistency voting by mathematical equivalence), and HumanEval (generative; candidates execute against the official assert-based test suites via a local Python interpreter, with an approximate fallback and warning when Python is absent), plus cached dataset downloads and checkpoint-attached benchmark results - Math Verifier -- From-scratch symbolic answer verification (no CAS dependency):
\boxed{}extraction, LaTeX normalization (fractions, surds, degrees, superscripts, mixed numbers), exact BigInteger-rational equivalence (0.5 ≡ 1/2but0.3333333333 ≢ 1/3), and random-rational-point identity testing for expressions with variables -- grades LaTeX mixed numbers correctly where SymPy-based reference pipelines do not - FP8 Dequantization -- FP8 format support for quantized weight loading with GPU-accelerated LUT path (256-entry cached lookup table using
torch.index_select) auto-selected when CUDA is available - Validation Pipeline -- Input validation framework with composable validators
- Scaling Heuristics -- Auto-scaling configuration from corpus and hardware stats
- Config Wizard -- Corpus analysis and hardware-aware config generation using Chinchilla scaling laws, activation memory estimates, NTK-aware RoPE, multi-GPU detection (nvidia-smi), and memory strategy recommendations
- CLI -- Subcommands for the full pipeline (
tokenize,ingest,train,infer,info,experiments,sft,dpo,rl,merge,transfer,distill,merge-models,fisher,calibrate,plan-quant,distributed,export onnx,export gguf,export llama,export glm,compress,decompress,prune,eval,config,config wizard,rag,graph,decide,wordnet,serve,orchestrate,agent,image) - Experiment tracking --
fuuga experiments <dir>compares training runs side by side (train/val loss, benchmark scores, approx params, fine-tuning stage, corpus hash, date) from their checkpoints, sorted and in table/CSV/JSON form. Read-only over existingcheckpoint.jsonfiles — no models loaded. - Data provenance --
ingestwrites a<corpus>.manifest.jsonsidecar recording the source files behind the corpus (path, size, content hash) and the corpus's SHA256. That hash is the same value stamped intoCheckpointMetadata.CorpusHash, socheckpoint → corpus manifest → source filesis fully traceable for reproducibility.
Fine-Tuning:
- Supervised Fine-Tuning (SFT) -- LoRA-based fine-tuning on instruction/chat JSONL data with configurable rank, alpha, and target modules
- Prompt Tuning -- Soft-prompt / virtual-token fine-tuning with frozen base weights for lightweight PEFT workflows
- Direct Preference Optimization (DPO) -- Preference learning from chosen/rejected pairs with LoRA
- Reinforcement Learning (RL) -- REINFORCE++ / GRPO fine-tuning with pluggable reward functions, trajectory-faithful rollout scoring, chain-of-thought rollouts (
--thinking), symbolic verifier rewards (--reward math), DAPO zero-signal group filtering, Dr. GRPO advantage normalization (--no-adv-std), and token-level length-normalized loss (--token-level-loss) - Reward Functions -- Composable reward functions for RL training (correctness, formatting, safety); accuracy rewards score only the answer span after thinking, so truncated reasoning is never rewarded as correct
- Rejection Sampling (STaR/RFT) --
fuuga sample-tracessamples N completions per prompt, keeps only verifier-passing traces, and writes them as SFT JSONL -- with--thinking, curated traces carry tagged reasoning spans - LoRA Adapter Merging -- Merge trained LoRA adapters back into the base model weights
- QLoRA -- NF4-quantized base weights with LoRA adapters for memory-efficient fine-tuning on consumer GPUs
- Data Validation -- JSONL format validation for SFT and DPO datasets with honesty pattern classification
- Data Augmentation -- Synonym replacement, rule-based paraphrasing, token-level noise injection, and SFT/DPO oversampling for training data diversity
Weight Transfer and Model Merging:
- Weight Transfer -- Transfer weights from donor models with architecture-aware mapping (Phi-3, Phi-4, LLaMA3, Mistral, Qwen 2.x/3.x, Gemma-3, Gemma-4, DeepSeek-V3 / Kimi K2 including MLA attention into an
AttentionType = MLAtarget (latent down/up projections, decoupled RoPE key, and the two latent RMSNorms), Z.AI GLM-4.5/4.6/5 MoE (experimental), Moonshot Kimi K3 MoE with MXFP4 auto-dequantization (experimental), ModernBERT) and dimension adaptation for mismatched tensors. ModernBERT is the one encoder donor:transfer --mapping modernbertinitialises a bidirectional model for the typed-decision path (decide train) rather than a generative one, so an open ModernBERT-backbone decision model can be used as the starting point instead of training the System-1 encoder from a causal checkpoint - Donor Export -- The reverse direction: export a Fuuga model back to HuggingFace-format safetensors.
fuuga export llamaproducesLlamaForCausalLM(near-lossless for a faithful Llama fine-tune) loadable by Transformers / vLLM / TGI / AWQ-GPTQ;fuuga export glmproduces GLM (Glm4Moe) safetensors (experimental, best-effort). Round-trips (export → re-import) are verified bit-exact in the test suite. - Knowledge Distillation -- Token-level, sequence-level (teacher-forced pseudo-labels, or true Kim & Rush via
--mode sequence-gen), reverse-KLD, attention (distill --attention) and hidden-state (distill --hidden) distillation from a teacher model - Endpoint Teacher --
distill --teacher-endpointdistils from a teacher reachable only over HTTP. An API returns samples rather than distributions, so this is response-based (black-box) distillation shaped as a data step: prompts in, the teacher's completions out as SFT JSONL forfuuga sft. Reaches hosted teachers no white-box path can, at the cost of much less signal per token - Distillation Quality Gate --
distill --eval-corpusmeasures the student on held-out data before and after and reports whether held-out loss actually improved, saying plainly when a run was a regression -- a falling KD loss only means the student tracked the teacher on the training corpus. Also prints a feasibility read (teacher/student top-1 agreement and top-K overlap) before training commits - LoRA Distillation --
distill --lora-rankdistils into adapters over a frozen student base, so distillation fits the same hardwaresftdoes instead of needing more. Works with every distillation path (logits, attention, hidden-state, cached); adapters are merged before saving unless--keep-adapteris given - Offline Logit Cache --
distill --write-logit-cacheruns the teacher over the corpus once and stores its top-K logits;--logit-cachethen trains with no teacher loaded at all. Decouples teacher hardware from student hardware (one pass on a rented GPU, then train locally) and makes extra epochs free of teacher cost. Cache/run compatibility is validated against the header rather than assumed - Attention Distillation -- Layer-level distillation between any attention pair (GQA/MLA/KDA/NSA/H-Transformer) with KL-divergence + MSE loss, pairing by layer index rather than attention class. This is the only route across architectures that weight transfer cannot express -- Kimi K3's KDA (linear) and Gated MLA layers have no weight-space mapping onto GQA.
--hiddenmatches whole-block hidden states instead, which stays valid when the two attention mechanisms are fundamentally different -- and cascades by default, running each model over its own stack from its own embedding so neither side is probed off-manifold. Depth reduction (a deeper teacher into a shallower student) samples across the teacher's full depth via--layer-mapping. The cascade is also the only layer-level path that supports MoE models, since it walks through MoE blocks rather than requiring them to be pairable submodules. Both layer-level probes train the backbone only -- the LM head gets no gradient, so they are an initialisation step ahead of a logits pass, not a complete training run. Layers that miss the quality gate can be rolled back to their pre-distillation weights with--revert-unconverged, and MoE routers keep their load-balancing signal throughout - N-ary Model Merging -- Merge multiple models with configurable strategies (TIES, DARE, Karcher mean, ModelSoups, ModelStock) and EWC protection
- Fisher Information -- Compute diagonal Fisher information matrices for Elastic Weight Consolidation
- N-ary Data Mixing -- Weighted multi-source mixing with static, curriculum, proxy-based DoReMi (Group DRO domain reweighting from a unigram proxy), and self-paced (difficulty-ramped) strategies
Inference Capabilities:
- Chain-of-Thought -- Thinking mode with ThinkStart/ThinkEnd token handling and dimmed thinking display
- Thinking Budgets (s1-style) --
--max-thinking-tokensforce-closes the thinking span at a cap;--min-thinking-tokenssuppresses premature end-of-thinking and injects a "Wait," continuation cue to extend reasoning (test-time scaling) - Self-Consistency Voting --
fuuga infer --consensus Nsamples N responses and majority-votes the extracted answers by mathematical equivalence; also available as theconsensusrequest option on the server API - Constrained Decoding -- Grammar-guided JSON structured output generation
- Self-Verification -- Draft/refine verification passes with learned verifier scoring for higher-confidence answers
- Tool Calling -- MCP (Model Context Protocol) client for tool discovery and invocation during generation
- Tool Policy -- Confidence-aware tool routing policy for deciding when external tools should be invoked
- Web Search -- Web search integration for grounded generation with citations
- Image Routing -- Generate or caption images via the Fuuga.Image MCP server, either with the
fuuga imagecommand or by configuring fuuga-image inmcp.json - A2A Protocol -- Agent-to-Agent protocol client for multi-agent communication
- Tree-Structured Speculative Decoding -- Speculative decoding with tree-structured candidates for faster generation
- Draft-and-Refine -- Multi-pass reasoning pipeline for improved output quality
- Autonomous Agent -- Web search agent loop for autonomous information gathering
- Advanced Reasoning -- Consensus voting, verifier-scored selection, and tree-of-thoughts for improved answer quality
- Backend Client -- HTTP client for calling external OpenAI-compatible LLM endpoints with structured response types
- Context Awareness -- Convention file discovery (AGENTS.md, CLAUDE.md, .cursorrules), language/framework detection, git context
- Semantic Knowledge -- WordNet WNDB parser with token-to-synset mapping and multi-lingual support
Orchestration:
- Multi-Model Orchestrator -- Route tasks to appropriate models based on capability.
fuuga orchestrate --backends backends.jsondecomposes a task and routes each subtask across the configured backends; without--backendsit runs every subtask against a single--endpoint. - Cost-Aware Routing -- Budget-tracked model routing with cost optimization. With
--backends, each subtask runs through a cost cascade: the first (cheapest) backend is tried, then the next (most capable) is used when confidence/history and the loaded routing profile (--checkpoint'scompression_codesign.json, consumed via itsQualityFloor) warrant escalation. Per-subtask spend is tracked against each backend'sdailyBudget, and outcome history accumulates between subtasks to improve later routing. Thebackends.jsonfile is a JSON array of{ id, name, endpoint, apiKey?, inputTokenCost, outputTokenCost, dailyBudget?, strengths[] }ordered cheapest-first. - Fan-Out Orchestration -- Decompose tasks into subtasks, run in parallel, and aggregate results
- Resumable Orchestration -- Checkpoint and resume fan-out plans across sessions
Agentic Persistence:
- Experience Store -- Append-only JSON Lines log of attempt outcomes with thread-safe managed access
- Strategy Lessons -- Persist and load distilled lessons from past experience for self-improvement
- Persistent Retrieval Store -- Disk-backed IRetrievalStore for cross-session document retrieval
- Hybrid Retrieval -- BM25 keyword index fused with dense cosine ranking via reciprocal rank fusion (
--rag-hybridoninferandfuuga-serve); catches exact-term matches that embedding similarity misses. Language-agnostic: stopwords are handled statistically (document-frequency cutoff), not by a hardcoded English list, so Finnish/Swedish/any-language corpora get equal treatment - Reranking -- Maximal Marginal Relevance (diversity) or model-scored relevance reranking over a widened candidate pool (
infer --rag-rerank mmr|llm) - Multi-Query Expansion -- Model-rewritten query variants retrieved independently and fused with RRF (
infer --rag-multi-query <N>) - RAG Retrieval Evaluation --
fuuga rag evalscores an index against a labelled JSONL question set (hit@k, mean reciprocal rank) so retrieval changes are measured, not guessed - Knowledge Graph (Graph Engineering) -- GraphRAG-style structured memory as a complement to chunk retrieval: facts stored as subject-predicate-object triples with evidence quotes, confidence, and provenance (
GraphKnowledge.fs). LLM-driven extraction (fuuga graph extractagainst any OpenAI-compatible endpoint) and batch entity resolution, alias-aware graph retrieval (entity lookup, k-hop neighborhood, shortest evidence chain for "why"/"how connected" questions, connected communities,--as-oftemporal filtering), and a typed query DSL as the query-translation target (easier for small models than Cypher, no server dependency). Maintenance follows the never-silently-overwrite rule: duplicates skipped, conflicts on functional predicates (persisted graph schema) flagged for review viagraph query --conflicts, superseded facts timestamped instead of deleted. Wired into generation asfuuga infer --rag-graph(the question becomes a graph query with entity-mention fallback; facts ground the prompt alone or fuse with--ragtext chunks via RRF) and server-side asfuuga-serve --rag-graph(deterministic entity-mention facts fused into the per-request context, no extra generation cost). CLI:fuuga graph add|extract|query|stats - Typed Decisions (System 1) -- The Laya / Jev pattern for agentic workflows: instead of asking a model to reason in free text and parsing the answer, ask a typed question —
choice(probabilities over named options),score(ordered rubric levels + expected level) ornoul(P(true) of a proposition) — about a state and get a probability for every option (TypedDecisions.fs). Requests and answers use Laya's JSON shape. Zero-shot over any existing model via option-tag log-probs (DecisionScoring.fs: Fuuga checkpoint, GGUF, or an OpenAI-compatible endpoint withlogprobs), temperature calibration per (question type, option count) fitted on labelled decisions, calibration metrics (accuracy, NLL, Brier, ECE, reliability bins, selective accuracy at each act-threshold), and an act-or-escalate gate that feeds the minion's escalation verdict. The native System-1 model (DecisionHeads.fs) is the design proper: any Fuuga checkpoint becomes a bidirectional encoder (newAttentionPattern.Bidirectional; weights carry over, LLM2Vec-style) with per-kind bilinear decision heads that score every option from its pooled token span in one forward pass — no generation — trained with a strictly proper scoring rule (log score or Brier) on labelled decisions viafuuga decide train. CLI:fuuga decide predict|calibrate|eval|train|compare|judge; server:POST /v1/decide(fuuga-serve --decision-model,--decide-calibration). Measuring a re-trained model is part of the same layer (JudgeArena.fs):fuuga decide compare --datascores two engines on exactly the recorded outcomes both answered and tests the difference paired, by McNemar's exact test over the decisions they split, and for free-text outputfuuga decide judgeA/Bs two response sets with a typed judge asked in both presentation orders, reporting the judge's own position bias, order-flip rate and (against any human labels) its accuracy and ECE next to the win rate. - Agent Session Management -- Session lifecycle (init → active → completed/failed), save/load state, step-level experience recording
- Orchestration Checkpoints -- Save and resume fan-out orchestration plans with per-subtask completion tracking
- Cost Outcome Tracking -- Persist cost-aware routing outcomes for budget optimization across sessions
Model Compression:
- Structured Pruning -- Attention head removal and layer removal with importance scoring
- NF4 Quantization -- 4-bit NormalFloat quantization for weight compression
- INT8 Weight-Only QLoRA -- Per-row symmetric INT8 base weights (8.5 bpw) with trainable LoRA adapters as a sibling to NF4 QLoRA — see
Int8QLoraLinear; selected by a [[QuantPlan]] entry ofBwInt8 - Activation Calibration -- Chat-templated forward-pass probe collects per-tensor
max-abs / mean-abs / RMSstatistics over SFT data (fuuga calibrate). Drives the dynamic quantisation planner; the chat-template-on-instruct-data step is the key lesson from Unsloth's Dynamic 2.0 GGUFs - Dynamic QuantPlan -- Greedy per-tensor bit-width allocator that blends activation magnitude with optional Fisher importance to pack the model under a target average-bpw budget (
fuuga plan-quant --target-bpw 4.5 --fisher F.bin). Preserves embeddings /lm_head/ norms at higher precision, pushes FFN down-projections to the lowest bpw — a Fuuga-native analogue of Unsloth Dynamic 2.0 - Per-Architecture Recipes --
QuantPlanConfig.forArchitecture "phi-3" | "llama-3" | "gemma-3" | "deepseek-v3" | "kimi-k3"returns a config with tunedPreserveFirstNLayers/PreserveLastNLayersboundary protections. Boundary layers (first / last 1-3 blocks) tolerate aggressive quant worst — Unsloth's empirical observation - MoE-Aware Quantization --
QuantPlanConfig.MoeExpertBitWidthroutes.moe.expert*tensors to a separate bit width while keeping.moe.routerat preserved precision. Routed experts tolerate aggressive quant far better than dense layers; this is where Dynamic 1.0's biggest absolute size savings come from on MoE deployments (DeepSeek-V3, Kimi K2) - Quantization-Aware Training (STE) -- Straight-Through Estimator for NF4 weights: forward pass sees quantized values, backward pass flows gradients through identity (
--ste) - Compression Pipeline -- Orchestrated prune → fine-tune → quantize workflow for production deployment
GGUF Interop (llama.cpp / Ollama / LM Studio):
- GGUF v3 Writer -- Full container emission (header + KV metadata + tensor info + aligned data) with Fuuga → llama-arch tensor name mapping.
fuuga export gguf --checkpoint X --output model.gguf [--plan plan.json --tokenizer tokenizer/] - GGUF v3 Reader -- The inverse path: ingest pre-quantized GGUFs (Unsloth's own, Bartowski's quants) as donor weights for
fuuga transfer/fuuga sft/fuuga distill.WeightTransfer.loadDonorWeightsroutes.ggufpaths throughGgufImport.readGgufDonorautomatically; the reverse name mapping treats llama-arch FFN as SwiGLU - Tensor Encodings --
F32,F16,Q8_0(8.5 bpw),Q4_0(4.5 bpw),Q6_K(6.5625 bpw),Q4_K(4.5 bpw),Q2_K(2.625 bpw),IQ4_NL(4.5 bpw, ARM / Apple Silicon friendly). Block layouts followllama.cpp/ggml-quants.c; every encoder/decoder pair is round-trip-validated in tests (F16 bit-exact; the quantized formats to per-format SNR floors) - Tokenizer Metadata --
--tokenizer <dir>embeds the full BPE vocab + merges + special-token ids + Jinja chat template so the resulting GGUF is directly runnable in llama.cpp / Ollama; without it the file isgguf-dump-inspectable but unloadable - Embedding Safety Floor -- Embeddings /
lm_headautomatically bump above 4-bit even when the plan asks for NF4 / Q2_K / Q4_K / IQ4_NL: K-family defaults floor to Q6_K, non-K to Q8_0. Override per-tensor via QuantPlan
Server:
- OpenAI-Compatible API -- Separate
fuuga-serveproject with/v1/chat/completions,/v1/completions,/v1/models, and/v1/embeddingsendpoints - Function Calling -- Client-supplied
tools/tool_choice(auto, none, required, or a named tool) on/v1/chat/completions; tool definitions render into the model's native<|tool_call|>prompt format and emitted calls come back as OpenAItool_calls(non-streaming and streaming deltas) - Structured Outputs --
response_format: json_schemaroutes the schema into grammar-constrained decoding;json_objectbecomes a JSON-only system instruction - SSE Streaming -- Server-Sent Events for real-time token streaming
- Continuous Batching -- Iteration-level scheduler with a paged KV cache, admission/preemption, and an async engine, wired into the server via
fuuga-serve --continuous-batching(memory-aware cache sizing; opt-in). Requests beyond the in-flight batch queue instead of blocking a thread. Uses a true batched-matmul forward for GQA-dense models (prefill + decode, parity-proven); MoE, NSA/H-Transformer, MLA, KDA and per-layer hybrids are served per-sequence within the same engine (the paged pool sizes each layer's pages from that layer's attention kind, and a KDA layer's fixed-size recurrent state lives in a sequence-keyed side table instead of pages). - Dynamic Batching -- Batch scheduling engine with metrics and backpressure, shared by the continuous-batching serving path.
- Bearer Token Auth -- Optional API key authentication middleware
- Guard Rails -- Prompt injection detection, PII masking, and content filtering. Enable with
fuuga-serve --guardrails(input + output checks on all chat endpoints) orfuuga infer --guardrails; the core lives inFuuga.GuardRailswith an Oxpecker middleware adapter inFuuga.Server.GuardRails - MCP Tool Routing -- Server-side MCP tool integration for function calling
- A2A Server -- Agent-to-Agent protocol server endpoint for multi-agent workflows
Distributed Training:
- PyTorch/DeepSpeed Integration -- Export model weights for distributed training, import trained weights back, and auto-generate launch scripts
Image Generation (Fuuga.Image):
- Text-to-Image / Image-to-Image -- Stable Diffusion generation from text prompts (samplers, steps, guidance) and image transformation with denoising strength control
- From-Scratch Stable Diffusion -- A complete SD 1.5 implementation written in F# on TorchSharp (CLIP BPE tokenizer + text encoder, VAE, UNet with cross-attention, DDIM/DDPM samplers, classifier-free guidance) that loads standard SD v1.x safetensors directly — educational and hackable;
fuuga-image scratch, thefuuga_image_scratchMCP tool, orfuuga image --scratch - Video --
txt2video,img2video, andvid2vid(edit/continue guided by text, VACE) when the loaded model supports video (e.g. Wan2.1/2.2 GGUF); gated on the model's own capability flag. Output is a dependency-free PNG image sequence by default, or mp4/gif/webp via optional ffmpeg.--audiomuxes a provided audio track into mp4/webp (video models make no audio). - Captioning / Understanding -- Describe an image or video (frame-sampled), or transcribe/describe audio, via whatever multimodal model you load (e.g. Phi-3.5-vision or Phi-4-multimodal) — brief/standard/detailed modes, all on the same ONNX-Runtime-GenAI runtime.
- MCP Server Mode -- Exposes txt2img / img2img / txt2video / img2video / vid2vid / caption as MCP tools over stdio for integration with Fuuga LLM
Local Minion (Delegation):
- Minion delegation -- Run Fuuga as a local "minion" that a more capable master agent (e.g. Claude Code) delegates small, well-scoped errands to, keeping simple work on a fully-local model with minimal energy cost
- MCP server (
fuuga-serve mcp) -- Exposes afuuga_delegatetool plus local file tools to a master over stdio JSON-RPC; register several with distinct--name/--roleto run a fleet of specialists (e.g. F# coder, C# coder, project manager) - Swappable brain -- The minion runs on a Fuuga-trained checkpoint (
--checkpoint), an in-process GGUF via LLamaSharp (--gguf), or any OpenAI-compatible endpoint such as a local Ollama (--endpoint) - Sandboxed local tools --
read_file,list_dir,grep, plus permission-gatedwrite_file,replace_in_file,apply_edits, andrun_command; confined to a workspace--root(resists..and symlink escapes), with writes/shell off by default - Multi-file refactors --
apply_edits(atomic edits across files: all land or none do),lsp_rename(semantic rename via the language server),--verify-command(build/tests gate the result; failures go back to the minion to fix, then escalate asverification_failed), and transcript compaction for long sessions. See docs/minion.md - Extended reach -- Opt into the built-in web tools (
--web) and configured MCP servers (--mcp-config) so a delegated errand can fetch pages and call other tools, not just touch the filesystem - Agent Skills -- Anthropic-style, Claude-Code-compatible
SKILL.mdskills (--skills <dir>, default./skills); the minion sees a compact catalog and loads a skill's full instructions on demand viaload_skill(progressive disclosure). One registry also backs the A2A agent card. See docs/skills.md - Escalation contract -- Each delegation returns structured JSON
{status, output, files_changed, escalate, reason, confidence, verification}; the minion self-verifies and hands work back (escalate=true) when it is not confident, so the master only spends its own capacity when needed - One-shot CLI (
fuuga delegate) -- Run a single errand locally and print the JSON result, without a master;--fail-on-escalateturns an escalation into exit code 4 for CI gates
Observability:
- OpenTelemetry -- OTLP trace and metrics export with Serilog integration
- Spectre.Console -- Rich terminal output for training progress and diagnostics
Prerequisites
- .NET 10 SDK (v10.0.103 or later)
- GPU is optional -- CPU works for the dev configuration (small model). CUDA-capable GPU recommended for larger models.
- ~500 MB disk space for dependencies, plus space for training data and checkpoints
Quick Start
For most users, the fastest path to useful output is to start from donor weights, not from scratch training.
Recommended paths:
examples/donor-transfer-and-refine.fsx-- practical donor-first workflow for normal users, with two modes:SmokeTestfor limited hardware, using a very small donor just to prove the F# pipeline is realPracticalfor a few-GB donor model that gives much better output quality
examples/complete-pipeline.fsx-- educational train-from-scratch pipelineexamples/finetuning-lora-qlora.fsx,preference-tuning-dpo-rl.fsx,modelops-compress-and-export.fsx,evaluation.fsx,rag-grounded-generation.fsx,agents-and-delegation.fsx,inference-techniques.fsx-- focused, runnable examples per workflow (seeexamples/README.md)
If you want to understand the full pipeline from scratch, use the CLI below:
# Build (CPU):
dotnet build
# Build (GPU, ~2 GB dependency):
dotnet build -p:TorchBackend=cuda
# Train a tokenizer, ingest a corpus, train, and generate text
dotnet run -- tokenize --input data/raw --vocab-size 8000 --output data/tokenizer
dotnet run -- ingest --input data/raw --output data/corpus.fuge --tokenizer data/tokenizer
dotnet run -- train --corpus data/corpus.fuge --tokenizer data/tokenizer --checkpoint-dir checkpoints/
dotnet run -- infer --checkpoint checkpoints/step-100 --tokenizer data/tokenizer --prompt "Once upon a time"
See the Getting Started Tutorial for a complete end-to-end walkthrough.
Important expectation setting:
- scratch training is educational and flexible, but tiny early runs often produce weak or gibberish text
- donor transfer is the better starting point when you want coherent output quickly
- short refinement on your own domain data is usually much more useful than starting from random weights
Use .fuge for tokenized corpus files and .fuuga for portable model packages.
Compressing a trained checkpoint for llama.cpp / Ollama
For deploying a trained Fuuga model into the llama.cpp ecosystem, the three-step pipeline produces a Fuuga-dynamic GGUF where critical tensors stay at higher precision and the rest land at an aggressive quant of your choice:
# 1. Collect per-tensor activation statistics from chat-templated SFT data.
dotnet run -- calibrate \
--checkpoint checkpoints/step-N --data data/sft.jsonl \
--output artefacts/cal.fcalib --batches 64
# 2. Plan per-tensor bit widths under a target average-bpw budget.
# Pass --fisher to blend Fisher-importance with activation magnitude.
dotnet run -- plan-quant \
--checkpoint checkpoints/step-N --calibration artefacts/cal.fcalib \
--output artefacts/plan.json --target-bpw 4.5 [--fisher artefacts/F.bin]
# 3. Emit a runnable GGUF — --tokenizer embeds the BPE vocab+merges so
# llama.cpp can load it directly. --default-type picks the encoding for
# tensors the plan doesn't preserve.
dotnet run -- export gguf \
--checkpoint checkpoints/step-N --plan artefacts/plan.json \
--tokenizer data/tokenizer --default-type q4_k \
--output deploy/model.gguf
Pick --default-type q2_k for extreme compression, iq4_nl for ARM / Apple Silicon, q6_k for the lowest-loss K-quant. Embeddings and lm_head automatically bump above 4-bit even when the plan asks for NF4 — see "GGUF Export" above.
F# Script Examples
Prefer the F# API over the CLI?
Practical donor-first path:
dotnet fsi examples/donor-transfer-and-refine.fsx
This loads donor weights, runs transfer into a Fuuga model, evaluates prompt outputs, and can do a short refinement pass. It is the recommended starting point for users who want useful results on limited hardware or with a few-GB donor model.
Train-from-scratch path:
dotnet fsi examples/complete-pipeline.fsx
This trains a tokenizer, ingests data, trains a model, and generates text -- all using the Fuuga modules directly. See examples/complete-pipeline.fsx for the full source.
For advanced workflows -- weight transfer from Phi-3/LLaMA3/DeepSeek, LoRA/QLoRA and preference tuning, quantize-and-export to GGUF, benchmark + quality evaluation, RAG, agents/delegation, and advanced decoding (speculative, constrained, chain-of-thought, streaming) -- see the focused scripts in examples/ (catalogued in examples/README.md).
Some examples draw an animated picture of what they do with --svg:
<table> <tr> <td align="center" width="50%"><a href="examples/continuous-batching-and-paged-cache.fsx"><img src="examples/_images/continuous-batching-and-paged-cache.svg" alt="Continuous batching with a paged KV cache" width="100%"></a><br><sub>Requests join the running batch as seats free up, each holding only the cache pages it has filled</sub></td> <td align="center" width="50%"><a href="examples/agents-and-delegation.fsx"><img src="examples/_images/agents-and-delegation.svg" alt="Delegating an errand to a local minion" width="100%"></a><br><sub>A master agent hands a small errand to a local model with sandboxed tools</sub></td> </tr> </table>
For image generation and captioning, see examples/image-demo.fsx.
Image Generation
Fuuga.Image is a standalone CLI for image generation and captioning. See the Fuuga.Image README for full command reference and MCP server mode, or run examples/image-demo.fsx.
Project Structure
Fuuga.fsproj # Project file with layered compilation order
Types.fs # All shared types (ModelConfig, TrainingConfig, GenerationConfig, etc.)
Logging.fs # ActivitySource/Meter definitions, ILoggerFactory
GuardRails.fs # Guardrails core: prompt-injection / PII / content checks (shared by CLI + server)
Observability.fs # OpenTelemetry providers, Spectre.Console, --observe flag
DriftDetection.fs # Statistical drift monitoring (KS, PSI) over confidence signals
DriftAlerting.fs # Dual-threshold alerts, adaptive thresholds, OTel metrics, retraining triggers + decision log
DriftMonitor.fs # Serve-side runtime: signal accumulation, timed analysis, reference/trigger persistence
Config.fs # JSON config loading, CLI arg parsing, MCP config, LoRA target parsing
Validation.fs # Input validation pipeline with composable validators
Scaling.fs # Scaling heuristics from corpus and hardware stats
ConfigWizard.fs # Corpus analysis + hardware-aware config generation (Chinchilla scaling)
Tokenizer.fs # BPE tokenizer training and loading
ParquetIO.fs # HuggingFace Parquet dataset read/write (Document, SFT, DPO)
TextCleanup.fs # Ingestion/preparation text cleanup
RagCleanup.fs # RAG (Retrieval-Augmented Generation) cleanup algorithms
Ingest.fs # Document discovery and binary corpus writing
CorpusCompression.fs # Zstd compression/decompression for .fuge files
Tensor.fs # Device selection (CPU/CUDA), DisposeScope
MultiResolutionAttention.fs # Chunk pooling, global tokens for long context
Model.fs # GPT-2 transformer with RoPE, GQA, RMSNorm, SwiGLU, MLA
AttentionConfig.fs # FlashAttention verification, SDPA backend selection
AutoConfig.fs # Auto-resolution of DU Auto* config cases from hardware probing
Vision.fs # Vision encoder for multimodal inputs
VisionBridge.fs # Q-Former cross-attention bridge for vision-to-language compression
PagedAttention.fs # Paged KV-cache attention
MemoryHierarchy.fs # Compressed memory, external retrieval
PersistentRetrievalStore.fs # Disk-backed IRetrievalStore for cross-session retrieval
ConfidenceHead.fs # Calibrated confidence MLP, Platt scaling, bucket assignment
EarlyExit.fs # Early exit / adaptive depth inference
Optimizer.fs # AdamW, SWA, Lookahead, and INT8 moment-quantized optimizers
Checkpoint.fs # Checkpoint save/load/metadata, safetensors
MmapLoading.fs # Memory-mapped model loading
GradientCheckpointing.fs # Gradient checkpointing for memory-efficient training
DistributedTraining.fs # Distributed training (PyTorch/DeepSpeed export/import)
ModelParallelism.fs # Tensor/pipeline parallelism config for 70B+ models
GpuOffloading.fs # Layer-wise CPU/GPU offloading
OptimizerOffload.fs # Optimizer state offloading
NvmePaging.fs # ZeRO-Infinity 3-tier GPU/CPU/NVMe memory management
OnnxExport.fs # ONNX export with quantization and validation
GgufImport.fs # GGUF v3 read (foundation: shared fp16↔fp32, kvaluesIq4nl, dequantizers); routes .gguf donors into WeightTransfer
GgufExport.fs # GGUF v3 write (F32/F16/Q8_0/Q4_0/Q6_K/Q4_K/Q2_K/IQ4_NL) with tokenizer + QuantPlan; builds on GgufImport's helpers
DataMixture.fs # N-ary weighted data source mixing
VramGuard.fs # In-process VRAM monitoring with nvidia-smi polling, signal-based training loop integration
Training.fs # Training loop with AdamW/cosine LR, gradient offloading, per-param flush
FineTuningData.fs # SFT/DPO JSONL parsing, chat templates, tokenization, batching
FineTuning.fs # LoRA (LoraLinear), SFT training, DPO loss/training, adapter save/load
RewardFunctions.fs # Composable reward functions for RL training
DataValidation.fs # SFT/DPO JSONL validation, honesty pattern classification
Fp8Dequantization.fs # FP8 format dequantization with GPU LUT acceleration
Nf4Quantizer.fs # NF4/FP4 4-bit weight quantization, STE for QAT
QLoraTraining.fs # QLoRA training (NF4 base + LoRA adapters), Int8QLoraLinear (per-row INT8)
LogitCache.fs # Offline top-K teacher logit cache (write once, train with no teacher resident)
DistillEval.fs # Distillation quality gate: held-out before/after + teacher-agreement feasibility
EndpointTeacher.fs # Response-based distillation from an HTTP teacher (prompts -> SFT JSONL)
WeightTransfer.fs # Weight transfer from donor models (Phi-3, LLaMA3, Mistral, Qwen, Gemma-3/4, DeepSeek, GLM, Kimi-K3 mappings)
ModelMerge.fs # N-ary model merging (TIES, DARE, Karcher) with EWC protection
Calibration.fs # Chat-templated activation-statistics probes for QuantPlanner
QuantPlanner.fs # Per-tensor BitWidth planner — Fuuga-dynamic 4.5-bpw plans via Fisher + activations
AttentionDistillation.fs # Layer-level distillation: any attention pair, or whole blocks (KL + MSE loss)
Pruning.fs # Structured pruning (attention heads, layers)
CompressionPipeline.fs # Prune → finetune → quantize orchestration
ConstrainedDecoding.fs # Grammar-guided JSON constrained decoding
ChainOfThought.fs # ThinkStart/ThinkEnd token handling, phase tracking
MathVerifier.fs # Symbolic math answer verification (boxed extraction, LaTeX normalization, exact-rational equivalence)
Verifier.fs # Rule-based and learned verifier strategies
ToolPolicy.fs # Confidence-aware tool invocation policy
McpClient.fs # MCP client: connection, tool discovery, tool invocation
WebSearch.fs # Web search integration for grounded generation
ImageRouting.fs # Image query routing to Fuuga.Image
A2AClient.fs # A2A protocol client for agent-to-agent communication
TreeSpeculation.fs # Tree-structured speculative decoding
ContinuousBatching.fs # Iteration-level continuous batching for serving
Inference.fs # Text generation with sampling, tool-augmented generation, structured output, CoT
DraftAndRefine.fs # Draft-and-refine multi-pass reasoning pipeline
AdvancedReasoning.fs # Consensus voting, verifier-scored, tree-of-thoughts
Orchestrator.fs # Multi-model orchestrator with capability-based routing
BackendClient.fs # HTTP client for external OpenAI-compatible LLM endpoints
ExperienceStore.fs # Append-only experience log, strategy lessons, managed store
CostAwareRouting.fs # Cost-aware routing with budget tracking
AgentSession.fs # Agent session lifecycle, save/load, orchestration checkpoints
FanOutOrchestration.fs # Fan-out/fan-in task decomposition and aggregation
ContextAwareness.fs # Convention file discovery, language/framework detection, git context
SemanticKnowledge.fs # WordNet WNDB parser, token-to-synset mapping, multi-lingual
DataAugmentation.fs # Synonym replacement, paraphrasing, token noise, oversampling
VerifierRuntime.fs # Verifier loading and runtime integration
MinionTools.fs # Sandboxed local tools (read/list/grep/write/replace/run) for delegated errands
MinionBrain.fs # Swappable minion brain: Fuuga checkpoint, GGUF (LLamaSharp), or OpenAI endpoint
MinionEscalation.fs # Escalation contract: self-verify + uncertainty/budget signals -> hand back to master
MinionToolEnv.fs # Minion reach beyond local files: built-in web tools, MCP servers, and skills (load_skill)
Skills.fs # Anthropic-style SKILL.md registry (parse/discover/catalog); backs minion, MCP host, and A2A card
MinionAgent.fs # Delegate loop: tool-use rounds, structured DelegateResult, escalation verdict
MinionCli.fs # Shared minion flag parsing (brain, sandbox, reach, identity, escalation)
Eval.fs # Benchmark datasets, runners, result serialization
Program.fs # CLI entry point with subcommand routing (incl. `delegate`)
Fuuga.Server/ # OpenAI-compatible HTTP server (separate project)
ApiTypes.fs # Request/response types (OpenAI-compatible)
McpToolRouting.fs # Server-side MCP tool routing for function calling
FuugaChatClient.fs # IChatClient adapter for Microsoft AI ecosystem
FuugaEmbeddingGenerator.fs # IEmbeddingGenerator adapter for /v1/embeddings
A2AServer.fs # A2A protocol server endpoint
GuardRails.fs # Oxpecker/ASP.NET middleware adapter over the Fuuga.GuardRails core
DynamicBatchingServer.fs # Dynamic batching with HTTP/SSE integration
Server.fs # Oxpecker HTTP server with SSE streaming, auth middleware
MinionServer.fs # Minion MCP server (stdio JSON-RPC): fuuga_delegate + local tools
Program.fs # Server entry point (incl. `mcp` minion mode)
Fuuga.Image/ # Standalone image generation and captioning CLI
Types.fs # Domain types, error handling (ImageError DU)
Config.fs # CLI argument parsing
ImageIO.fs # Image load/save, format conversion, validation
Diffusion.fs # Stable Diffusion model wrapper (txt2img, img2img)
Caption.fs # Phi-3.5-vision captioning (ONNX Runtime GenAI)
McpServer.fs # MCP JSON-RPC 2.0 server over stdio
Program.fs # Entry point, subcommand routing
Fuuga.Tests/ # Unit and integration tests (xUnit + FsUnit)
Fuuga.PropertyTests/ # FsCheck property tests over the RAG ingestion path
Fuuga.Image/Fuuga.Image.Tests/ # Image module tests
docs/ # User-facing documentation
examples/ # Runnable F# script examples
scripts/ # Training data generation and validation scripts
Documentation
- Getting Started Tutorial -- End-to-end walkthrough from build to text generation
- CLI Reference -- All subcommands, flags, defaults, and exit codes
- Configuration Reference -- Model, training, generation, and fine-tuning parameters explained
- Agent Skills -- Authoring and using Anthropic-style
SKILL.mdskills across the minion, MCP host, and A2A card - Ecosystem Comparison -- Fuuga vs .NET ecosystem comparison
- Language Centralization -- Adding new natural languages via the KnownNaturalLanguages registry
- Fuuga.Image README -- Image generation CLI reference and MCP server mode
- Examples catalogue -- focused, runnable F# scripts per workflow (pre-train, fine-tune/LoRA/QLoRA, DPO/RL, modelops/GGUF, evaluation, RAG, agents, advanced decoding, image)
Architecture
Fuuga uses a layered module architecture with strict dependency ordering enforced by F#'s compilation model:
Layer 0: Types, Logging, Observability (foundation, no dependencies)
DriftDetection, DriftAlerting, DriftMonitor (statistical drift monitoring, OTel alerts, serve-side loop)
Layer 1: Config, Validation, Scaling (configuration, validation, heuristics)
ConfigWizard (hardware-aware config generation)
Layer 2: Tokenizer, ParquetIO (BPE training/loading, Parquet dataset I/O)
Layer 3: Ingest, CorpusCompression (document discovery, corpus writing, Zstd compression)
Layer 4: Tensor, MultiResolutionAttention (device selection, chunk pooling + global tokens)
Model, AttentionConfig, AutoConfig (transformer with MLA, FlashAttention, auto-resolution)
Vision, VisionBridge (vision encoder, Q-Former bridge)
PagedAttention, MemoryHierarchy (paged KV-cache, compressed memory)
PersistentRetrievalStore (disk-backed retrieval for cross-session use)
ConfidenceHead, EarlyExit (calibration, adaptive depth)
Layer 5: Optimizer, Checkpoint, MmapLoading (optimizer variants incl. INT8 moments, save/load/metadata, mmap loading)
GradientCheckpointing (activation recomputation)
DistributedTraining, ModelParallelism (PyTorch/DeepSpeed, tensor/pipeline parallel)
GpuOffloading, OptimizerOffload (CPU/GPU memory management)
NvmePaging, OnnxExport (NVMe paging, ONNX export)
GgufImport, GgufExport (GGUF v3 read/write; Import owns shared format helpers, Export builds on them)
DataMixture, VramGuard (data mixing, in-process VRAM monitoring)
Layer 6: Training, FineTuningData, FineTuning (training loop, SFT/DPO data, LoRA training)
RewardFunctions, DataValidation (composable RL rewards, JSONL validation)
Fp8Dequantization, Nf4Quantizer (FP8/NF4 quantization support, GPU LUT, STE for QAT)
QLoraTraining (NF4 + INT8 weight-only QLoRA, optional QuantPlan-driven dispatch)
WeightTransfer, ModelMerge (donor model transfer, N-ary merging)
Calibration, QuantPlanner (chat-templated activation probes, per-tensor BitWidth allocator)
AttentionDistillation (any-attention-pair / block distillation)
Pruning, CompressionPipeline (structured pruning, prune→finetune→quantize)
Layer 7: ConstrainedDecoding, ChainOfThought (generation extensions)
Verifier, ToolPolicy, McpClient (verifier strategies, tool policy, tool calling)
WebSearch, ImageRouting, A2AClient (search, image routing, A2A protocol)
TreeSpeculation, ContinuousBatching (speculation, batch scheduling)
Inference (text generation with sampling, tools, structured output)
DraftAndRefine, AdvancedReasoning (multi-pass reasoning, consensus/tree-of-thoughts)
Orchestrator, BackendClient (multi-model routing, external LLM client)
ExperienceStore, CostAwareRouting (cross-session persistence, cost optimization)
AgentSession, FanOutOrchestration (agent lifecycle, parallel task decomposition)
Layer 8: ContextAwareness, SemanticKnowledge (project context, WordNet)
DataAugmentation (synonym replacement, paraphrasing, token noise)
VerifierRuntime (verifier loading and runtime integration)
Eval (benchmark datasets, runners, result serialization)
Layer 9: Program (CLI entry point, subcommand routing)
More about design decisions can be read from the architecture document.
Running Tests
dotnet test Fuuga.Tests
Tests cover tokenization, ingestion, Parquet I/O, corpus compression, model architecture, MLA attention, vision encoder, vision bridge, paged attention, multi-resolution attention, attention configuration, early exit, training, optimizer variants, INT8 optimizer moments, gradient checkpointing, GPU offloading, optimizer offloading, ONNX export, GGUF v3 export (including Q6_K/Q4_K/Q2_K/IQ4_NL round-trip SNR + tokenizer metadata + embedding safety floor), inference, verifier-guided generation, speculative decoding, checkpoints, memory-mapped loading, fine-tuning, prompt tuning, QLoRA (NF4 and INT8 variants), reinforcement learning, reward functions, data validation, data augmentation, benchmark evaluation, FP8 dequantization, FP8 GPU LUT dequantization, NF4 quantization, STE quantization-aware training, activation calibration (chat-templated probes + binary report round-trip), QuantPlan synthesis (Fisher-blended importance ranking + bpw-budget allocator + JSON round-trip), structured pruning, compression pipeline, constrained decoding, chain-of-thought, MCP client, A2A client/server, web search, image routing, draft-and-refine, advanced reasoning, math verifier (symbolic equivalence, boxed extraction, equivalence-class voting), thinking budgets (s1-style budget forcing), orchestrator, backend client, cost-aware routing, fan-out orchestration, weight transfer, attention distillation, model merging, distributed training, model parallelism, scaling, validation, continuous batching, dynamic batching server, experience persistence, agent session management, persistent retrieval, drift detection, drift alerting, context awareness, semantic knowledge, observability, guard rails, memory strategy (gradient offloading, per-param flush, VRAM guard lifecycle, ConfigWizard recommendations, multi-GPU parallelism), and CLI integration. Tests use xUnit with FsUnit assertions and include both unit tests and end-to-end integration tests.
Assertions that FsUnit cannot express without collapsing the interesting values into a bare true — structural equality on DU-wrapped values, tensor closeness, quantifiers over collections, Result/Option cases — go through Fuuga.Tests/TestAssertions.fs (shouldEqualValue, shouldBeAllCloseAtol, shouldAllSatisfy, shouldContainSatisfying, shouldBeOk, …). The pass/fail decision is identical to the |> should equal true form they replace, but a failure reports the two sides of the equality, the worst element of a tensor comparison and the tolerance it needed, the first element to fail a predicate, or the Error that arrived instead of Ok. The helpers are themselves tested in TestAssertionsTests.fs. Tests whose assertions are tight enough to depend on a particular random draw take it from a private torch.Generator, never from torch.manual_seed, which seeds a process-global generator that a concurrently running collection can reseed mid-test (see TestDeterminism.fs).
Fuuga.PropertyTests is an FsCheck suite over the RAG ingestion path (dotnet test Fuuga.PropertyTests): clean prose comes back unchanged from TextCleanup and RagCleanup, code blocks behind <|code_start|> markers are never touched, the noise the passes are written for (control characters, curly quotes, ligatures, zero-width characters, page numbers, dot leaders, decorative lines, doubled words, mojibake, odd spaces) is gone afterwards without a letter lost, cleaning twice is cleaning once, the markdown walker keeps every fenced block verbatim with its normalised language tag, the JSONL reader turns exactly the chat records into documents, the FIM transform is a permutation of the text, and a corpus reads back token for token as it was written.
Technology Stack
| Component | Library | Purpose |
|---|---|---|
| Tensors & GPU | TorchSharp 0.106.0 | Tensor operations, CUDA support |
| LibTorch | libtorch-cpu 2.10.0 | LibTorch CPU backend |
| Tokenization | Microsoft.ML.Tokenizers 2.0.0 | BPE tokenizer training |
| ONNX Runtime | Microsoft.ML.OnnxRuntime 1.24.3 | ONNX model inference |
| ONNX Export | OnnxSharp 0.3.2 | ONNX model construction and manipulation |
| Protobuf | Google.Protobuf 3.34.0 | Protobuf serialization for ONNX |
| Epub parsing | VersOne.Epub 3.3.4 | Extract text from epub files |
| Markdown | Markdig 1.1.1 | Parse markdown to plain text |
| Parquet | Parquet.Net 5.5.0 | HuggingFace-compatible dataset I/O |
| Compression | ZstdSharp.Port 0.8.7 | Zstandard corpus compression |
| Statistics | MathNet.Numerics.FSharp 5.0.0 | Drift detection (KS test, PSI) |
| Logging | Serilog 4.3.0 | Structured logging |
| Logging sinks | Serilog.Sinks.Console 6.0.0, .File 6.0.0, .OpenTelemetry 4.2.0 | Console, file, and OTel log sinks |
| Logging bridge | Serilog.Extensions.Logging 9.0.0 | Serilog/Microsoft.Extensions.Logging bridge |
| JSON | FSharp.SystemTextJson 1.4.36 | F# DU-aware serialization |
| Telemetry | OpenTelemetry 1.11.2 | Distributed tracing and metrics |
| Telemetry export | OpenTelemetry.Exporter.OpenTelemetryProtocol 1.11.2 | OTLP protocol export |
| Telemetry hosting | OpenTelemetry.Extensions.Hosting 1.11.2 | OpenTelemetry hosting integration |
| Terminal UI | Spectre.Console 0.49.1 | Rich terminal output |
| HuggingFace | TorchSharp.PyBridge 1.4.3 | HuggingFace weight loading |
| SIMD | System.Numerics.Tensors 10.0.5 | Preprocessing acceleration |
| AI abstractions | Microsoft.Extensions.AI 10.4.0 | IChatClient adapter |
| AI evaluation | Microsoft.Extensions.AI.Evaluation 10.4.0 | Benchmark evaluation framework |
| MCP | ModelContextProtocol 1.1.0 | MCP client SDK for tool calling |
| A2A | A2A 0.3.3-preview | Agent-to-Agent protocol client |
| Testing | xUnit 2.9.3 + FsUnit.xUnit 7.1.1 | Unit and integration tests |
| Server: HTTP | Oxpecker 2.0.0 | F# HTTP server framework |
| Server: A2A | A2A.AspNetCore 0.3.3-preview | A2A protocol server |
| Image: Diffusion | StableDiffusion.NET 5.0.0 | Stable Diffusion model wrapper |
| Image: Captioning | Microsoft.ML.OnnxRuntimeGenAI 0.12.1 | Phi-3.5-vision captioning |
| Image: Processing | HPPH.SkiaSharp 1.0.0 | Image load/save, format conversion |
License
See LICENSE for details.
| Product | Versions Compatible and additional computed target framework versions. |
|---|---|
| .NET | net10.0 is compatible. net10.0-android was computed. net10.0-browser was computed. net10.0-ios was computed. net10.0-maccatalyst was computed. net10.0-macos was computed. net10.0-tvos was computed. net10.0-windows was computed. |
-
net10.0
- A2A (>= 1.0.0-preview2)
- FSharp.Core (>= 10.1.302)
- FSharp.SystemTextJson (>= 1.4.36)
- Google.Protobuf (>= 3.36.2)
- libtorch-cuda-12.8-win-x64 (>= 2.10.0)
- LLamaSharp (>= 0.27.0)
- LLamaSharp.Backend.Cpu (>= 0.27.0)
- Markdig (>= 1.4.0)
- MathNet.Numerics.FSharp (>= 5.0.0)
- Microsoft.Extensions.AI (>= 10.10.0)
- Microsoft.Extensions.AI.Evaluation (>= 10.10.0)
- Microsoft.Extensions.AI.Evaluation.Quality (>= 10.10.0)
- Microsoft.Extensions.AI.Evaluation.Reporting (>= 10.10.0)
- Microsoft.ML.OnnxRuntime (>= 1.30.0)
- Microsoft.ML.Tokenizers (>= 2.0.0)
- ModelContextProtocol (>= 2.2.0)
- OnnxSharp (>= 0.3.2)
- OpenTelemetry (>= 1.19.1)
- OpenTelemetry.Exporter.OpenTelemetryProtocol (>= 1.19.1)
- OpenTelemetry.Extensions.Hosting (>= 1.19.1)
- Parquet.Net (>= 6.1.0)
- Serilog (>= 4.4.0)
- Serilog.Extensions.Logging (>= 10.0.0)
- Serilog.Sinks.Console (>= 6.1.1)
- Serilog.Sinks.File (>= 7.0.0)
- Serilog.Sinks.OpenTelemetry (>= 4.2.0)
- SkiaSharp (>= 2.88.8)
- SkiaSharp.NativeAssets.Linux (>= 2.88.8)
- Spectre.Console (>= 0.57.2)
- System.Numerics.Tensors (>= 10.0.12)
- TorchSharp (>= 0.107.0)
- TorchSharp.PyBridge (>= 1.4.4)
- TreeSitter.DotNet (>= 1.3.0)
- VersOne.Epub (>= 3.3.6)
- ZstdSharp.Port (>= 0.8.8)
NuGet packages
This package is not used by any NuGet packages.
GitHub repositories
This package is not used by any popular GitHub repositories.
| Version | Downloads | Last Updated |
|---|---|---|
| 1.0.28 | 37 | 10/1/2026 |
| 1.0.27 | 46 | 9/29/2026 |
| 1.0.26 | 96 | 9/26/2026 |
| 1.0.25 | 84 | 9/25/2026 |
| 1.0.24 | 86 | 9/25/2026 |
| 1.0.23 | 98 | 9/22/2026 |
| 1.0.22 | 94 | 9/20/2026 |
| 1.0.21 | 88 | 9/18/2026 |
| 1.0.20 | 128 | 8/15/2026 |
| 1.0.19 | 130 | 8/1/2026 |
| 1.0.18 | 118 | 7/31/2026 |
| 1.0.17 | 116 | 7/29/2026 |
| 1.0.16 | 141 | 7/26/2026 |
| 1.0.15 | 128 | 7/12/2026 |
| 1.0.14 | 140 | 7/11/2026 |
| 1.0.13 | 124 | 7/6/2026 |
| 1.0.12 | 132 | 7/3/2026 |
| 1.0.11 | 131 | 7/2/2026 |
| 1.0.10 | 149 | 6/30/2026 |
| 1.0.9 | 143 | 6/29/2026 |
MINION MULTI-FILE REFACTORING. New apply_edits tool: several exact-text replacements across files as one change; nothing is written if any edit fails. New lsp_rename tool: semantic rename across files through the language server (needs --allow-write). fuuga delegate --verify-command (with --verify-timeout, --verify-retries) runs a build or test command after file changes, hands failures back to the minion and escalates with verification_failed if they persist; the result JSON gains "verification". --max-context-chars elides the oldest tool results from a long transcript (default 48000). --fail-on-escalate exits with code 4. LIBRARY API: MinionAgent.runDelegateWithOptions and DelegateOptions; LspClient.Rename, applyTextEdits, parseWorkspaceEdit. BREAKING for code that constructs these records: DelegateResult gains Verification, EscalationInputs gains VerificationFailed and VerificationPassed; EscalationReason gains the case VerificationFailed. DEPENDENCY UPDATES: Microsoft.Extensions.AI 10.10.0 (and the three Evaluation packages), Microsoft.ML.OnnxRuntime 1.30.0, OpenTelemetry 1.19.1 (with the OTLP exporter and Extensions.Hosting), ModelContextProtocol 2.2.0, Google.Protobuf 3.36.2, Markdig 1.4.0, Parquet.Net 6.1.0, System.Numerics.Tensors 10.0.12, TorchSharp.PyBridge 1.4.4. TESTS: four shared-state races between parallel test collections fixed. PREVIOUS (1.0.27) - DONOR IMPORTS: LATENT ATTENTION AND ENCODER DONORS. MLA ATTENTION NOW TRANSFERS (DeepSeek-V3, Kimi K2). --mapping deepseek / deepseek-moe previously skipped attention entirely, documented as "MLA, incompatible with Fuuga's GQA" - true only of a GQA TARGET. Against a target whose AttentionType is MLA the donor maps almost one-to-one: q_a_proj to wDQ, kv_a_proj_with_mqa split to wDKV + wKR, q_b_proj split to wUQ + wQR, kv_b_proj split to wUK + wUV, o_proj to wO, and q_a_layernorm / kv_a_layernorm into new latent RMSNorms. The two _b_proj tensors INTERLEAVE their halves within each head, so the split point is read from the target's MlaConfig rather than from the tensor: a contiguous split would hand each part the wrong heads' rows at identical shapes, and a row count the geometry cannot explain is reported as a shape error naming the tensor rather than reshaped into something plausible. Kimi K2 ships as model_type deepseek_v3, so one mapping covers both. MlaConfig gains NopeHeadDim, VHeadDim and LatentNormEps, all optional - omitted they reproduce the previous behaviour exactly, and configs written before they existed keep loading unchanged. They exist because the per-head dims cannot stay tied to EmbDim / NumHeads: DeepSeek-V3 runs 128 heads over a 7168-wide residual, so that ratio is 56 while its content and value head dims are both 128, which is precisely what made those donors' attention unrepresentable. MultiHeadLatentAttention decouples the content and value head dims across its projections, its absorption math and all three forward paths, and gains RMSNorm on the compressed latents; the KV cache layout is unchanged. The target must carry the donor's per-head geometry - layer count and FFN/expert widths may still be smaller. Kimi K3 is NOT covered: its KDA and gated-MLA layers have no key map. DeepSeek's YaRN RoPE scaling has no equivalent (Linear / NTKAware only) - no weights are involved, but long-context behaviour will differ. The layout is verified by construction and by tests that would catch a mis-split by value; it has not been run against a real DeepSeek or Kimi checkpoint. MODERNBERT ENCODER DONOR (--mapping modernbert): the one encoder in the list, producing a BIDIRECTIONAL model for the typed-decision path rather than a generative one - which is how an open judge model such as Laya 1 comes in, its encoder subfolder being a stock ModernBertForMaskedLM. Fused Wqkv and Wi reuse the Phi-3 splitters; layer 0's attn_norm is nn.Identity and is never saved, so block0.norm1 is filled from embeddings.norm, the norm that actually runs before layer 0's attention. The masked-LM head is not transferred, and the decision heads are Fuuga's own, trained by fuuga decide train. TRANSFER NOW WRITES checkpoint.json beside model_weights.dat, so its output is a checkpoint directory rather than loose weights and fuuga decide train, which has no --model-config, can consume a transfer result directly. It is written on every transfer including over an existing file, because model_weights.dat has just been replaced and metadata left from whatever used to live there would now describe a different model; an explicit --model-config still wins downstream. DISTILL REFUSES A TEACHER WHOSE ATTENTION WAS NEVER POPULATED: --attention and --hidden match the student against the teacher's own modules, so a mapping that fills no attention leaves the teacher at initialisation and trains the student against noise while the loss still falls convincingly. A .gguf teacher is exempt - it arrives keyed by Fuuga parameter names and bypasses the key map entirely. SMALLER: checkpoint metadata records the real assembly version instead of a hardcoded "0.1.0"; a cross-collection Console race in the CLI drift guard is fixed. EXAMPLE: examples/kimi-donor-laya-judge.fsx runs the whole scenario offline on CPU - a Kimi/DeepSeek-shaped MLA donor and a Laya/ModernBERT encoder donor imported, decision heads trained on the result, and the A/B arena run twice to show what the report says when a judge is unusable versus usable. PREVIOUS (1.0.26) - ROBUSTNESS FIXES FROM A STATIC-ANALYSIS PASS (fsharp-refactor notes, each finding verified in the source before it was changed). GUARDRAILS: one invalid regex in CustomPatterns threw out of the input check on every checked request, and every custom pattern was recompiled per input; the input check now shares the output side's compiled cache, which logs an invalid pattern once and skips it while the valid ones still apply. RESUME NO LONGER DESTROYS STATE: `fuuga agent --resume` reported an unreadable agent-session.json as "no previous session" and then overwrote it, and `fuuga orchestrate --resume` silently replaced a checkpoint it could not read while re-running every subtask; both files are now moved aside as <name>.unreadable-<UTC timestamp> with a message saying so (new AgentSession.setAsideUnreadable). SERVER (fuuga-serve): a literal JSON null request body bound to a null record and the handler answered 500; it is now a 400, and all six JSON endpoints report the serializer's own message ("Could not parse request body: ...") where five said only "Invalid JSON" even for valid JSON that did not bind. Tool-call arguments the model produced that are not a JSON object now travel as the call's error (FunctionCallContent.Exception) instead of running the tool with no arguments. The guardrail-filtered streaming path awaits the output enumerator's disposal instead of blocking a request thread on it. SMALLER: the LSP client disposes its cancellation source; the weight-streaming autograd warning no longer names {Blocks} twice in one template; the repetition penalty enumerates the token history once; single-character StartsWith/EndsWith checks (comment markers, JSON braces, trailing newlines, the BPE space marker) compare ordinally, so an invisible character such as a soft hyphen no longer changes the answer. No public API removed. PREVIOUS (1.0.25) - DRIFT MONITORING AND TRAINING SAFETY FIXES. SUSTAINED DRIFT NOW KEEPS ALERTING. The adaptive threshold (mean + 2 sigma over the last 30 windows) is a rolling one, so it absorbs any level that lasts - which is exactly what stops it alerting on a signal's own noise, and equally stopped it alerting on drift nobody had fixed. A sustained shift fired one alert in 52 windows and then went quiet; a shift that began soon after a reference capture, or in the very first analysed window, never alerted at all; and a critical plateau sitting just over the threshold never alerted under any previous version. Operators saw Critical in GET /v1/drift with no pending retraining trigger. The adaptive gate may now delay an alert but never suppress one: a new SustainedCriticalWindowsRequired (default 10) fires on consecutive critical windows whatever the baseline makes of them, because a signal's own noise crosses back below critical long before ten windows and a real shift does not. The gate also now needs AdaptiveMinimumSamples (default 5) windows before it has an opinion at all - a threshold built from two scores sits barely above the pair and the next ordinary window clears it. Measured on the same series, before and after, at the shipped defaults: sustained shift after a healthy run 1 alert then silence, now alerting every 10 windows indefinitely; shift two windows after a reference capture 0, now 4 in 42 windows; shift from the first window 0, now 4 in 40; critical plateau just over the line 1, now 4. Nothing new fires on quiet signals: isolated spikes 0 before and after, a noisy stationary signal wandering across the warning line 0 before and after, a stationary signal sitting on the critical line 0 before and 0.15 per 300 windows after - that residual is inherent, since ten critical windows in a row out of a signal that is critical half the time by chance cannot be told from a genuine plateau except by waiting longer; raise SustainedCriticalWindowsRequired to trade it away. A trigger the backstop raises says so in its brief - it reports that the adaptive threshold has absorbed the level rather than quoting it as exceeded - so an operator approving from the brief sees why it fired. An earlier attempt that kept critical scores out of the baseline was measured and rejected: it truncates the baseline to the quiet tail, so a signal that normally sits near the threshold is compared against a picture of its calmest windows and crosses constantly. TRAINING NO LONGER OVERWRITES ITS OWN SOURCE. fuuga decide train and fuuga merge both write model_weights.dat and checkpoint.json flat into --output, under the same names they read from their input, so pointing both at one directory replaced the source in place: decide train left the donor re-tagged Bidirectional, which fuuga serve then refuses as a generative model, and merge dropped BaseCheckpointPath, so a second run would load the already-merged weights and apply the adapters on top of them again. Both now exit 1 with an explanation before any work happens (Checkpoint.resolvesToSameDirectory). TESTS: every alerting test shrank the configuration to 1-2 consecutive windows, a 3-5 window baseline and 1 sigma, so the shipped numbers were never exercised; a new section runs the defaults from both directions - drift that does not resolve must keep alerting, and a stationary signal must not alert however close to the threshold it sits - with the stationary cases swept over seeds rather than resting on one.