SapphireGuard.AgentEvals 0.2.0

dotnet add package SapphireGuard.AgentEvals --version 0.2.0
                    
NuGet\Install-Package SapphireGuard.AgentEvals -Version 0.2.0
                    
This command is intended to be used within the Package Manager Console in Visual Studio, as it uses the NuGet module's version of Install-Package.
<PackageReference Include="SapphireGuard.AgentEvals" Version="0.2.0" />
                    
For projects that support PackageReference, copy this XML node into the project file to reference the package.
<PackageVersion Include="SapphireGuard.AgentEvals" Version="0.2.0" />
                    
Directory.Packages.props
<PackageReference Include="SapphireGuard.AgentEvals" />
                    
Project file
For projects that support Central Package Management (CPM), copy this XML node into the solution Directory.Packages.props file to version the package.
paket add SapphireGuard.AgentEvals --version 0.2.0
                    
#r "nuget: SapphireGuard.AgentEvals, 0.2.0"
                    
#r directive can be used in F# Interactive and Polyglot Notebooks. Copy this into the interactive tool or source code of the script to reference the package.
#:package SapphireGuard.AgentEvals@0.2.0
                    
#:package directive can be used in C# file-based apps starting in .NET 10 preview 4. Copy this into a .cs file before any lines of code to reference the package.
#addin nuget:?package=SapphireGuard.AgentEvals&version=0.2.0
                    
Install as a Cake Addin
#tool nuget:?package=SapphireGuard.AgentEvals&version=0.2.0
                    
Install as a Cake Tool

SapphireGuard.AgentEvals

How do you know the AI you've shipped is actually working reliably in production? Not "it looked good in the demo" — but measurably, repeatably, with evidence you can put in front of a stakeholder?

That's the question this project is built around. SapphireGuard.AgentEvals is a framework for evaluating AI agents. It runs an agent against a library of test cases, captures everything the agent did, and scores that behaviour against clear expectations — producing a defensible report on quality, cost, and reliability before anything reaches users.

The framework is runtime-agnostic: you evaluate any agent by implementing one seam, IAgentRunner. It ships with a skill-testing harness — the reference host that wires the framework to Claude skills — as its first fully-worked consumer. That's the use case driving the framework today.

The stakes for getting this right only increase as AI systems become more autonomous. That's the motivation for building this discipline early.


Framework, and the skill harness on top

Two distinct things live in this repo, and it's worth being clear which is which:

  • The framework — the reusable, runtime-agnostic core: the case model, the evaluators, the orchestration, the scoring and statistics, and the reporting. It has no dependency on any particular agent runtime. You evaluate your agent by implementing a single interface, IAgentRunner. This is the product, published as the SapphireGuard.AgentEvals NuGet package.
  • The skill harness — a reference host that wires the framework to Claude skills (AgentSkillRunner plus the CLI). It's one consumer of the framework — the first, fully-worked one — and it's what the examples in these docs use.

Historically the skill harness came first, and the framework was factored out of it — which is why the git repo is still named eval-harness even though the package and the product are SapphireGuard.AgentEvals. The core is a framework; the skill harness is one of its hosts. Everything below describes the framework, illustrated through the skill harness.


The core idea: measure before you ship

The standard approach to AI skill development goes roughly like this: write the skill, test it manually, ship it, and find out about problems when users complain. This is the same mistake the industry made with software before automated testing — and it has the same fix.

Evaluation-driven development (EDD) inverts the order. You define what "good" looks like first — as a set of test cases with clear pass/fail expectations — then build the skill to satisfy them, then measure. Every change produces a number. Regressions are caught locally. Quality is something you can claim with evidence, not just assert with confidence.

The discipline that makes this work is the baseline. Every evaluation run has two configurations: with_skill (the agent with the skill loaded) and baseline (the same agent, same task, no skill). The number that matters isn't the absolute pass rate — it's the delta. A skill scoring 95% against a 93% baseline isn't earning its keep. A skill scoring 90% against a 30% baseline is delivering real value. Without the baseline, you're measuring noise.


What evaluation actually tests

A common instinct when thinking about agent evaluation is to reach for a running system — spin up the agent, attach a tracer, observe what it does. That model isn't wrong, but it misses something important.

If you stub all tool calls — which you should, to keep evals reproducible and free of side effects — then for the question evals answer, the runtime is largely factored out. You are not testing whether the tools work, or how the runtime behaves when they fail. You are testing whether the agent reasons correctly given what tools return. Tool implementations have their own tests; the runtime's error handling and latency have theirs. Evals are not those tests.

What you actually need from an agent to run evals is its configuration:

  • A system prompt
  • A set of tool schemas
  • A model

The framework provides stub implementations of those tools from each case, runs the model against the system prompt, captures the trace, and scores it. The agent does not need to be deployed. Its infrastructure, database connections, and external integrations play no part.

This is already how skill evaluation works. A skill is a system prompt (SKILL.md) plus tool schemas (tools/*.json). The skill harness reconstructs it entirely from those files. Agent evaluation is the same shape — the difference is that agents bring their own system prompt rather than loading one dynamically.

The implication for agent teams: you do not expose your runtime to run evals. You provide your configuration.

Two honest limits come with stubbing, and it's worth naming them. The runtime is factored out, so evals don't exercise real tool-integration failures, error recovery, or latency — those need their own tests. And a case only tests the scenario its author imagined; the canned responses encode assumptions about what the world returns. The framework measures whether the agent reasons correctly on the scenarios you thought of — not whether it will survive everything production sends. You close that gap with more cases, ideally drawn from real usage — not with a cleverer framework. CONCEPTS.md works through the full mental model, limits included.


Why a dedicated evaluation framework

Whether you control the agent runtime or not, you need a way to measure what your agent actually does — before it reaches users.

This framework is the engineer's primary reliability mechanism. A few requirements follow from that:

  • It needs to be more reliable than what it's measuring — so most checks are deterministic trace inspection, and the LLM judges it does use are calibrated against human labels before they're trusted. A flaky framework produces noise, not verdicts.
  • It runs multiple times per case to separate genuine regressions from the natural variance of LLM responses — and reports reliability (pass^k) and confidence intervals, not bare point estimates that read as certainty.
  • It tracks cost, because a skill that doubles token usage for a marginal quality improvement may not be worth shipping.
  • It must be reproducible — the same inputs, the same evaluators, the same verdict, every time.

Why not just use the evals built into each skill?

The Agent Skills spec lets authors bundle an evals/evals.json file with a skill. Those are a valuable author-time sanity check while you're writing — but they're built for a different job than a pre-ship reliability gate. The difference isn't quality, it's purpose:

In-skill evals SapphireGuard.AgentEvals
Primary use A quick check while authoring the skill A repeatable pre-ship reliability gate
Baseline No control group — no way to isolate what the skill adds Baseline runs alongside every eval to produce a defensible delta
Grading Model-graded Deterministic checks for most things; LLM-as-judge only for genuinely subjective quality
Signal depth Pass/fail on the final response Activation → Trajectory → Outcome pinpoints exactly where something broke
Reproducibility Tool responses can vary between runs Canned tool responses in each case guarantee identical replay
Attribution "It failed" "Activation passed, trajectory passed, outcome failed"
Runtime independence Tied to the runtime Runs entirely outside the agent — works even when you don't own the runtime

The distinguishing move here is deterministic-first grading. Most of what looks like a quality judgement — did the agent call the right tools? does the output contain these fields? — can be expressed as a structural check on the trace, which is faster, cheaper, and free of the variance a model grader introduces. The framework reserves LLM-as-judge for the questions that genuinely can't be answered any other way, and calibrates those judges against human labels before trusting them.


The three evaluation layers

Each test scenario is evaluated across three independent layers. This triangulation means a failure can be pinpointed precisely — not just "something went wrong" but exactly where in the agent's behaviour.

1. Activation — Did the agent correctly identify that the skill should activate for this request? Evaluated deterministically against a known activation expectation.

2. Trajectory — Did the agent consult the right data sources, avoid prohibited actions, and follow the correct sequence of steps? Evaluated by inspecting the agent's tool call trace.

3. Outcome — Did the agent produce the correct structured result? Evaluated both deterministically (checking expected fields and structure) and by independent AI judges for subjective quality.

Every layer is scored in both configurations, so the delta applies at each one — you can see exactly which layer the skill improves.


Current scope and where this is going

The skill harness is the first consumer, and it's fully operational — it runs an agent, loads a skill, and measures whether the skill is doing useful work. Claude skills are the use case driving the framework today.

But the evaluation discipline — define cases, run the agent, score the trace, report the delta — isn't specific to skills. It applies to any agent. The architecture reflects this: the framework (orchestrator, evaluators, case model, reporting) has no dependency on MAF or any particular runtime. The MAF-backed skill runner is just one implementation of a single IAgentRunner interface — and the seam is already proven by a second, independent consumer: the debtor-verification-agent project implements IAgentRunner against a different runtime (model-harness) and gets the full evaluation stack unchanged.

The framework is available as a NuGet package: SapphireGuard.AgentEvals. A team building an agent in their own solution can install it, implement IAgentRunner against their runtime, and get the full stack — deterministic evaluators, multi-run reliability, governance reporting — without taking a dependency on MAF or the skill-harness CLI. The test cases they write are portable across runtimes.

See ROADMAP.md for what's next.


  • CONCEPTS.md — the mental model behind the framework: why it's shaped this way, and its limits
  • RUNNING.md — setup, configuration, and CLI command reference
  • EXTENDING.md — how to add evaluators and test cases
  • GLOSSARY.md — definitions of terms used throughout this project
  • INDUSTRY-ALIGNMENT.md — how these choices map to current published agent-eval practice
  • ROADMAP.md — what's built, what's next, and what's out of scope
Product Compatible and additional computed target framework versions.
.NET net10.0 is compatible.  net10.0-android was computed.  net10.0-browser was computed.  net10.0-ios was computed.  net10.0-maccatalyst was computed.  net10.0-macos was computed.  net10.0-tvos was computed.  net10.0-windows was computed. 
Compatible target framework(s)
Included target framework(s) (in package)
Learn more about Target Frameworks and .NET Standard.

NuGet packages

This package is not used by any NuGet packages.

GitHub repositories

This package is not used by any popular GitHub repositories.

Version Downloads Last Updated
0.2.0 122 9/15/2026
0.1.1-alpha.0.37 47 9/15/2026