SapphireGuard.EvalHarness
0.1.0
Renamed the package to be more accurate to what it is.
See the version list below for details.
dotnet add package SapphireGuard.EvalHarness --version 0.1.0
NuGet\Install-Package SapphireGuard.EvalHarness -Version 0.1.0
<PackageReference Include="SapphireGuard.EvalHarness" Version="0.1.0" />
<PackageVersion Include="SapphireGuard.EvalHarness" Version="0.1.0" />
<PackageReference Include="SapphireGuard.EvalHarness" />
paket add SapphireGuard.EvalHarness --version 0.1.0
#r "nuget: SapphireGuard.EvalHarness, 0.1.0"
#:package SapphireGuard.EvalHarness@0.1.0
#addin nuget:?package=SapphireGuard.EvalHarness&version=0.1.0
#tool nuget:?package=SapphireGuard.EvalHarness&version=0.1.0
SapphireGuard.EvalHarness
How do you know the AI you've shipped is actually working reliably in production? Not "it looked good in the demo" — but measurably, repeatably, with evidence you can put in front of a stakeholder?
That's the question this project is built around. SapphireGuard.EvalHarness runs an AI agent against a library of test cases, captures everything the agent did, and scores that behaviour against clear expectations. The result is a defensible report on quality, cost, and reliability — produced before anything reaches users.
The stakes for getting this right only increase as AI systems become more autonomous. That's the motivation for building this discipline early.
The core idea: measure before you ship
The standard approach to AI skill development goes roughly like this: write the skill, test it manually, ship it, and find out about problems when users complain. This is the same mistake the industry made with software before automated testing — and it has the same fix.
Evaluation-driven development (EDD) inverts the order. You define what "good" looks like first — as a set of test cases with clear pass/fail expectations — then build the skill to satisfy them, then measure. Every change produces a number. Regressions are caught locally. Quality is something you can claim with evidence, not just assert with confidence.
The discipline that makes this work is the baseline. Every evaluation run has two configurations: with_skill (the agent with the skill loaded) and baseline (the same agent, same task, no skill). The number that matters isn't the absolute pass rate — it's the delta. A skill scoring 95% against a 93% baseline isn't earning its keep. A skill scoring 90% against a 30% baseline is delivering real value. Without the baseline, you're measuring noise.
What evaluation actually tests
A common instinct when thinking about agent evaluation is to reach for a running system — spin up the agent, attach a tracer, observe what it does. That model isn't wrong, but it misses something important.
If you stub all tool calls — which you should, to keep evals reproducible and free of side effects — then the agent's runtime becomes irrelevant. You are not testing whether the tools work. You are testing whether the agent reasons correctly given what tools return. Tool implementations have their own tests. Evals are not those tests.
What you actually need from an agent to run evals is its configuration:
- A system prompt
- A set of tool schemas
- A model
The harness provides stub implementations of those tools from each case, runs the model against the system prompt, captures the trace, and scores it. The agent does not need to be deployed. Its infrastructure, database connections, and external integrations play no part.
This is already how skill evaluation works. A skill is a system prompt (SKILL.md) plus tool schemas (tools/*.json). The harness reconstructs it entirely from those files. Agent evaluation is the same shape — the difference is that agents bring their own system prompt rather than loading one dynamically.
The implication for agent teams: you do not expose your runtime to run evals. You provide your configuration.
Why we needed a dedicated harness
Whether you control the agent runtime or not, you need a way to measure what your agent actually does — before it reaches users.
The eval harness is the engineer's primary reliability mechanism. A few requirements follow from that:
- It needs to be more reliable than what it's measuring — if the harness itself is flaky, its verdicts are noise.
- It runs multiple times per case to separate genuine regressions from the natural variance of LLM responses.
- It tracks cost, because a skill that doubles token usage for a marginal quality improvement may not be worth shipping.
- It must be reproducible — the same inputs, the same evaluators, the same verdict, every time.
Why not just use the evals built into each skill?
The Agent Skills spec allows skill authors to bundle an evals/evals.json file with their skill. These are useful during authoring — a quick sanity check while you're writing. But they're not a reliability mechanism:
| In-skill evals | SapphireGuard.EvalHarness | |
|---|---|---|
| When they run | Manually, during authoring only | On demand throughout development |
| Baseline | No control group — no way to measure whether the skill adds value | Baseline runs alongside every eval to produce a defensible delta |
| Grading | LLM-as-judge for every assertion | Deterministic checks for most things; LLM-as-judge only for genuinely subjective quality |
| Signal depth | Pass/fail on the final response | Activation → Trajectory → Outcome pinpoints exactly where something broke |
| Reproducibility | Live tool responses vary between runs | Canned tool responses in each case guarantee identical replay |
| Attribution | "It failed" | "Activation passed, trajectory passed, outcome failed" |
| Runtime independence | Tied to the runtime | Runs entirely outside the agent — works even when you don't own the runtime |
The in-skill evals also have no equivalent of deterministic evaluation. Most of what looks like a quality judgement — did the agent call the right tools? does the output contain these fields? — can be expressed as a structural check on the trace. Sending everything through an LLM judge is slower, more expensive, and introduces variance. The harness reserves LLM-as-judge for questions that genuinely can't be answered any other way.
The three evaluation layers
Each test scenario is evaluated across three independent layers. This triangulation means a failure can be pinpointed precisely — not just "something went wrong" but exactly where in the agent's behaviour.
1. Activation — Did the agent correctly identify that the skill should activate for this request? Evaluated deterministically against a known activation expectation.
2. Trajectory — Did the agent consult the right data sources, avoid prohibited actions, and follow the correct sequence of steps? Evaluated by inspecting the agent's tool call trace.
3. Outcome — Did the agent produce the correct structured result? Evaluated both deterministically (checking expected fields and structure) and by independent AI judges for subjective quality.
The baseline configuration runs the same scenarios without the skill loaded, so the model answers from general knowledge alone. The delta between baseline and skill shows the value the skill adds — a skill that doesn't meaningfully beat the baseline is not earning its place.
Current scope and where this is going
Right now this harness is focused on Claude skills — it runs an agent, loads a skill, and measures whether the skill is doing useful work. That's the first use case driving it, and it's fully operational.
But the evaluation discipline — define cases, run the agent, score the trace, report the delta — isn't specific to skills. It applies to any agent. The architecture reflects this: the core framework (orchestrator, evaluators, case model, reporting) has no dependency on MAF or any particular runtime. The MAF runner is one implementation of a single IAgentRunner interface.
The plan is to extract that framework as a package. A team building an agent in their own solution would install it, implement IAgentRunner against their runtime, and get the full evaluation stack — deterministic and inferential evaluators, multi-run variance, governance reporting — for free. The test cases they write would be portable across runtimes.
The trigger for that work is a real second consumer. Until then, the separation is maintained so extraction stays a refactor, not a redesign. See ROADMAP.md for the specifics.
Links
- RUNNING.md — setup, configuration, and CLI command reference
- EXTENDING.md — how to add evaluators and test cases
- GLOSSARY.md — definitions of terms used throughout this project
- ROADMAP.md — what's built, what's next, and what's out of scope
| Product | Versions Compatible and additional computed target framework versions. |
|---|---|
| .NET | net10.0 is compatible. net10.0-android was computed. net10.0-browser was computed. net10.0-ios was computed. net10.0-maccatalyst was computed. net10.0-macos was computed. net10.0-tvos was computed. net10.0-windows was computed. |
-
net10.0
- Microsoft.Extensions.Configuration.Binder (>= 10.0.7)
- Microsoft.Extensions.DependencyInjection.Abstractions (>= 10.0.7)
NuGet packages
This package is not used by any NuGet packages.
GitHub repositories
This package is not used by any popular GitHub repositories.
| Version | Downloads | Last Updated | |
|---|---|---|---|
| 0.1.1-alpha.0.18 | 111 | 7/11/2026 | |
| 0.1.1-alpha.0.16 | 101 | 7/11/2026 | |
| 0.1.1-alpha.0.15 | 85 | 7/11/2026 | |
| 0.1.1-alpha.0.14 | 82 | 7/11/2026 | |
| 0.1.1-alpha.0.13 | 85 | 7/11/2026 | |
| 0.1.1-alpha.0.10 | 99 | 6/25/2026 | |
| 0.1.1-alpha.0.9 | 91 | 6/24/2026 | |
| 0.1.1-alpha.0.8 | 94 | 6/24/2026 | |
| 0.1.1-alpha.0.6 | 92 | 6/11/2026 | |
| 0.1.1-alpha.0.5 | 93 | 6/11/2026 | |
| 0.1.1-alpha.0.3 | 93 | 6/10/2026 | |
| 0.1.1-alpha.0.1 | 95 | 6/10/2026 | |
| 0.1.0 | 173 | 6/10/2026 |