SapphireGuard.EvalHarness 0.1.0

Suggested Alternatives

SapphireGuard.AgentEvals

Additional Details

Renamed the package to be more accurate to what it is.

There is a newer prerelease version of this package available.
See the version list below for details.
dotnet add package SapphireGuard.EvalHarness --version 0.1.0
                    
NuGet\Install-Package SapphireGuard.EvalHarness -Version 0.1.0
                    
This command is intended to be used within the Package Manager Console in Visual Studio, as it uses the NuGet module's version of Install-Package.
<PackageReference Include="SapphireGuard.EvalHarness" Version="0.1.0" />
                    
For projects that support PackageReference, copy this XML node into the project file to reference the package.
<PackageVersion Include="SapphireGuard.EvalHarness" Version="0.1.0" />
                    
Directory.Packages.props
<PackageReference Include="SapphireGuard.EvalHarness" />
                    
Project file
For projects that support Central Package Management (CPM), copy this XML node into the solution Directory.Packages.props file to version the package.
paket add SapphireGuard.EvalHarness --version 0.1.0
                    
#r "nuget: SapphireGuard.EvalHarness, 0.1.0"
                    
#r directive can be used in F# Interactive and Polyglot Notebooks. Copy this into the interactive tool or source code of the script to reference the package.
#:package SapphireGuard.EvalHarness@0.1.0
                    
#:package directive can be used in C# file-based apps starting in .NET 10 preview 4. Copy this into a .cs file before any lines of code to reference the package.
#addin nuget:?package=SapphireGuard.EvalHarness&version=0.1.0
                    
Install as a Cake Addin
#tool nuget:?package=SapphireGuard.EvalHarness&version=0.1.0
                    
Install as a Cake Tool

SapphireGuard.EvalHarness

How do you know the AI you've shipped is actually working reliably in production? Not "it looked good in the demo" — but measurably, repeatably, with evidence you can put in front of a stakeholder?

That's the question this project is built around. SapphireGuard.EvalHarness runs an AI agent against a library of test cases, captures everything the agent did, and scores that behaviour against clear expectations. The result is a defensible report on quality, cost, and reliability — produced before anything reaches users.

The stakes for getting this right only increase as AI systems become more autonomous. That's the motivation for building this discipline early.


The core idea: measure before you ship

The standard approach to AI skill development goes roughly like this: write the skill, test it manually, ship it, and find out about problems when users complain. This is the same mistake the industry made with software before automated testing — and it has the same fix.

Evaluation-driven development (EDD) inverts the order. You define what "good" looks like first — as a set of test cases with clear pass/fail expectations — then build the skill to satisfy them, then measure. Every change produces a number. Regressions are caught locally. Quality is something you can claim with evidence, not just assert with confidence.

The discipline that makes this work is the baseline. Every evaluation run has two configurations: with_skill (the agent with the skill loaded) and baseline (the same agent, same task, no skill). The number that matters isn't the absolute pass rate — it's the delta. A skill scoring 95% against a 93% baseline isn't earning its keep. A skill scoring 90% against a 30% baseline is delivering real value. Without the baseline, you're measuring noise.


What evaluation actually tests

A common instinct when thinking about agent evaluation is to reach for a running system — spin up the agent, attach a tracer, observe what it does. That model isn't wrong, but it misses something important.

If you stub all tool calls — which you should, to keep evals reproducible and free of side effects — then the agent's runtime becomes irrelevant. You are not testing whether the tools work. You are testing whether the agent reasons correctly given what tools return. Tool implementations have their own tests. Evals are not those tests.

What you actually need from an agent to run evals is its configuration:

  • A system prompt
  • A set of tool schemas
  • A model

The harness provides stub implementations of those tools from each case, runs the model against the system prompt, captures the trace, and scores it. The agent does not need to be deployed. Its infrastructure, database connections, and external integrations play no part.

This is already how skill evaluation works. A skill is a system prompt (SKILL.md) plus tool schemas (tools/*.json). The harness reconstructs it entirely from those files. Agent evaluation is the same shape — the difference is that agents bring their own system prompt rather than loading one dynamically.

The implication for agent teams: you do not expose your runtime to run evals. You provide your configuration.


Why we needed a dedicated harness

Whether you control the agent runtime or not, you need a way to measure what your agent actually does — before it reaches users.

The eval harness is the engineer's primary reliability mechanism. A few requirements follow from that:

  • It needs to be more reliable than what it's measuring — if the harness itself is flaky, its verdicts are noise.
  • It runs multiple times per case to separate genuine regressions from the natural variance of LLM responses.
  • It tracks cost, because a skill that doubles token usage for a marginal quality improvement may not be worth shipping.
  • It must be reproducible — the same inputs, the same evaluators, the same verdict, every time.

Why not just use the evals built into each skill?

The Agent Skills spec allows skill authors to bundle an evals/evals.json file with their skill. These are useful during authoring — a quick sanity check while you're writing. But they're not a reliability mechanism:

In-skill evals SapphireGuard.EvalHarness
When they run Manually, during authoring only On demand throughout development
Baseline No control group — no way to measure whether the skill adds value Baseline runs alongside every eval to produce a defensible delta
Grading LLM-as-judge for every assertion Deterministic checks for most things; LLM-as-judge only for genuinely subjective quality
Signal depth Pass/fail on the final response Activation → Trajectory → Outcome pinpoints exactly where something broke
Reproducibility Live tool responses vary between runs Canned tool responses in each case guarantee identical replay
Attribution "It failed" "Activation passed, trajectory passed, outcome failed"
Runtime independence Tied to the runtime Runs entirely outside the agent — works even when you don't own the runtime

The in-skill evals also have no equivalent of deterministic evaluation. Most of what looks like a quality judgement — did the agent call the right tools? does the output contain these fields? — can be expressed as a structural check on the trace. Sending everything through an LLM judge is slower, more expensive, and introduces variance. The harness reserves LLM-as-judge for questions that genuinely can't be answered any other way.


The three evaluation layers

Each test scenario is evaluated across three independent layers. This triangulation means a failure can be pinpointed precisely — not just "something went wrong" but exactly where in the agent's behaviour.

1. Activation — Did the agent correctly identify that the skill should activate for this request? Evaluated deterministically against a known activation expectation.

2. Trajectory — Did the agent consult the right data sources, avoid prohibited actions, and follow the correct sequence of steps? Evaluated by inspecting the agent's tool call trace.

3. Outcome — Did the agent produce the correct structured result? Evaluated both deterministically (checking expected fields and structure) and by independent AI judges for subjective quality.

The baseline configuration runs the same scenarios without the skill loaded, so the model answers from general knowledge alone. The delta between baseline and skill shows the value the skill adds — a skill that doesn't meaningfully beat the baseline is not earning its place.


Current scope and where this is going

Right now this harness is focused on Claude skills — it runs an agent, loads a skill, and measures whether the skill is doing useful work. That's the first use case driving it, and it's fully operational.

But the evaluation discipline — define cases, run the agent, score the trace, report the delta — isn't specific to skills. It applies to any agent. The architecture reflects this: the core framework (orchestrator, evaluators, case model, reporting) has no dependency on MAF or any particular runtime. The MAF runner is one implementation of a single IAgentRunner interface.

The plan is to extract that framework as a package. A team building an agent in their own solution would install it, implement IAgentRunner against their runtime, and get the full evaluation stack — deterministic and inferential evaluators, multi-run variance, governance reporting — for free. The test cases they write would be portable across runtimes.

The trigger for that work is a real second consumer. Until then, the separation is maintained so extraction stays a refactor, not a redesign. See ROADMAP.md for the specifics.


  • RUNNING.md — setup, configuration, and CLI command reference
  • EXTENDING.md — how to add evaluators and test cases
  • GLOSSARY.md — definitions of terms used throughout this project
  • ROADMAP.md — what's built, what's next, and what's out of scope
Product Compatible and additional computed target framework versions.
.NET net10.0 is compatible.  net10.0-android was computed.  net10.0-browser was computed.  net10.0-ios was computed.  net10.0-maccatalyst was computed.  net10.0-macos was computed.  net10.0-tvos was computed.  net10.0-windows was computed. 
Compatible target framework(s)
Included target framework(s) (in package)
Learn more about Target Frameworks and .NET Standard.

NuGet packages

This package is not used by any NuGet packages.

GitHub repositories

This package is not used by any popular GitHub repositories.