Eyu.Rdf
0.7.0
dotnet add package Eyu.Rdf --version 0.7.0
NuGet\Install-Package Eyu.Rdf -Version 0.7.0
<PackageReference Include="Eyu.Rdf" Version="0.7.0" />
<PackageVersion Include="Eyu.Rdf" Version="0.7.0" />
<PackageReference Include="Eyu.Rdf" />
paket add Eyu.Rdf --version 0.7.0
#r "nuget: Eyu.Rdf, 0.7.0"
#:package Eyu.Rdf@0.7.0
#addin nuget:?package=Eyu.Rdf&version=0.7.0
#tool nuget:?package=Eyu.Rdf&version=0.7.0
Eyu
Eyu (이유) — Korean for "reason" or "why". Pronounced roughly "eh-yoo".
A source-agnostic ontology inference engine: turns declared structure and raw records into proposed entities, relations, and grounded claims — the reason a piece of data is shaped the way it is, made explicit and citable.
Status: contract implemented; grounding-integrity and response-parsing reliability validated
against a real model; entity-resolution accuracy measured on a labeled public benchmark (B-cubed F1
0.99–1.00 on Febrl synthetic person records) and not yet on any other data; exercised by one external
package consumer on a Korean-language document corpus, whose first measurement is what the
entity Name, per-element Rejections and stamped VocabularyOrigin answer. All five ports exist as C# types (Eyu.Core), plus a first source
adapter (Eyu.Formbase). IOntologyProposer's judgment logic (SinglePassOntologyProposer +
HttpModelClient) is implemented and has been measured against a real model (GPUStack
qwen3.8-27b) across two record domains, and over two Korean company-profile documents passed as
chunks — every cited source id existed among the records given, 0 violations every time. That check alone does not prove a citation actually backs its claim, so
the same measurement now also runs a mechanical content-overlap check
(GroundingOverlapCheck) between each claim
and the record content it cites — text in any script, with Chinese, Japanese and Korean
compared as overlapping character pairs so a Korean claim still matches the record it restates
when the particles differ — a heuristic signal to spot-check per run, not a confirmed
defect count (see the type's doc comment for why token overlap is not a semantic verifier; a
run's own counts are in its report, not restated here since the pre-filter below changes them
run to run). SinglePassOntologyProposer also takes an optional LinkageOptions
(Eyu.Core.Linkage) that runs a Fellegi-Sunter record-linkage pre-filter before the model call:
record pairs it confirms as a match are injected into the prompt as "merge them, do not
re-decide", gray-zone pairs are injected as a log-odds hint for the model to judge itself, and
confirmed non-matches get neither — so the prompt does now carry a resolution instruction, where
before it carried none. The linkage evidence then adjusts an entity's confidence only through the
records the entity claims denote it (DenotedBy), never through records that merely mention it:
a machine several work orders name is not penalized because the work orders are different orders. A live run measured that this classification is observable (per-pair
Match/GrayZone/NonMatch, EM convergence status, match prior) and that changing the thresholds
measurably changes both the resulting prompt and the grounding-overlap counts. What that run does
not establish: whether any particular threshold setting is more correct — no labeled
ground truth exists in these domains to score the pre-filter's own match/non-match calls against
(on a labeled public benchmark it has been scored — docs/linkage-benchmark.md:
B-cubed F1 0.99–1.00 on Febrl person records, 0.991–0.994 on names and addresses alone with exact
comparison); what the pre-filter does report is an
unlabeled estimate of its own error rates (LinkageAnalysis.ErrorRates: the false-match and
false-non-match rates the fitted EM mixture expects of its own calls, with the preconditions of
that estimate flagged when they did not hold — an estimate is not a measurement, and it says
whether a clerical review of a sample is worth running, not what it would find; the unit such a
review will be scored on is already fixed as B-cubed per record, not per pair,
ClusteringMetrics.BCubed, and the sampling and adjudication design for that review is
docs/clerical-review.md, whose estimator is in the library as
ClericalReviewEstimator, and the path from a reviewer's verdicts to the per-record scores it
consumes is ClericalReviewScoring — the arithmetic that turns a reviewed sample into a
population score with an interval exists end to end and has been run with a benchmark's labels as
the reviewer — which showed a pilot's interval collapsing where errors are rare, since fixed so a
stratum that shows no error keeps the uncertainty its sample allows;
what is still missing is a review by people on a corpus of the domains above) — or whether
the proposer's self-reported
confidence is informative on its own, which it was not in an earlier measurement (see
design rationale, §D); gray-zone cases now combine it with the
Fellegi-Sunter prior via a Bayesian update instead of using it alone. Field comparison inside
that pre-filter is exact-match (case-insensitive, ignoring surrounding whitespace and Unicode form —
decomposed Hangul, full-width digits) by default; LinkageOptions. UseStringSimilarityComparator opts into a Jaro-Winkler threshold instead, so two records
denoting the same entity but differing only in notation (punctuation, spacing) still register as
agreeing on that field — still no ground truth to say which mode classifies better on any given
dataset, so the option exists but the default is unchanged.
Eyu.Core and Eyu.Formbase are published on NuGet, and the package is referenced by the
external consumer named above; inside this repository the
coupling smoke test
references Formbase.Core the same way, as a package.
Install
Current release: 0.7.0. All three packages ship from this repository and move together:
dotnet add package Eyu.Core --version 0.7.0
dotnet add package Eyu.Formbase --version 0.7.0
dotnet add package Eyu.Rdf --version 0.7.0
Eyu.Core alone is enough to implement the five ports against your own source; Eyu.Formbase is
the adapter for one of them and reads Formbase.Core as a package, so it pairs with a Formbase
release — this one is built against Formbase.* 0.17.0. Eyu.Rdf writes a proposal as OWL in Turtle
(see Exporting as RDF/OWL).
Quick start
Hand Eyu a few records, get back the entities, relations and grounded claims it proposes — then, optionally, the same proposal as OWL:
using System.Net.Http.Headers;
using Eyu.Core.Declared;
using Eyu.Core.Inference.Http;
using Eyu.Core.Judgment;
using Eyu.Core.Primitives;
using Eyu.Core.Records;
using Eyu.Rdf;
// Any OpenAI-compatible server. Requests go to "chat/completions" relative to the base address,
// so it ends with the API version segment and a slash.
var http = new HttpClient
{
BaseAddress = new Uri("https://api.openai.com/v1/"),
Timeout = TimeSpan.FromMinutes(5), // a thinking model can take minutes per call
};
http.DefaultRequestHeaders.Authorization =
new AuthenticationHeaderValue("Bearer", Environment.GetEnvironmentVariable("OPENAI_API_KEY"));
var model = new HttpModelClient(http, "gpt-4o-mini");
// Records are handed in, never fetched: an id, and the fields as the source holds them.
RawRecord[] records =
[
new("emp-1", new Dictionary<string, string?> { ["name"] = "Kim Minji", ["email"] = "minji.kim@example.com" }),
new("emp-2", new Dictionary<string, string?> { ["name"] = "Minji Kim", ["email"] = "minji.kim@example.com" }),
new("wo-1", new Dictionary<string, string?> { ["title"] = "Replace bearing on press #4", ["reviewer"] = "Kim Minji" }),
];
// Declare what you already know -- here only that some records are employees. Declared always wins;
// the model proposes the rest.
DeclaredStructure[] declared = [new(SubjectRef.Create("Employee"), [], [])];
var proposal = await new SinglePassOntologyProposer(model).ProposeAsync(declared, records);
foreach (var entity in proposal.Entities)
{
Console.WriteLine($"{entity.Name} : {entity.EntityType} ({entity.Confidence:0.00}), denoted by {string.Join(", ", entity.DenotedBy)}");
}
foreach (var relation in proposal.Relations)
{
Console.WriteLine($"{relation.FromEntityId} -{relation.RelationName}-> {relation.ToEntityId}: {relation.Claim.Claim}");
}
Console.WriteLine($"{proposal.Rejections.Count} element(s) left out, each with its reason in proposal.Rejections");
Console.WriteLine(OntologyTurtle.ToTurtle(proposal, new RdfExportOptions(new Uri("https://example.org/plant#"))));
The two employee rows are one person written two ways. Whether they come back as one entity denoted
by both is the proposal's judgment, and DenotedBy shows which it made — two runs of this very
sample have answered both ways, which is why a proposal is routed rather than applied (Status, above,
says what is and is not measured). A self-hosted thinking model
can be told not to think through extraBody (see Ports); a failed call throws
HttpRequestException naming the status, the request URI and what the server answered.
Why
Every system that wants an ontology chatbot ends up rebuilding the same brain: infer entities and relations from structure, decide how confident that inference is, and refuse to answer without a citable path back to the source. That logic has nothing to do with where the data lives — whether it's a raw document store you own, a schema you declared elsewhere, or a table in someone else's database you can only read. Building it once, separately from any storage or access model, is the only way it doesn't get rebuilt every time a new consumer needs it.
This is the bet the design makes, not yet a cross-validated claim: today Eyu
has one source adapter (Eyu.Formbase, in this same repo) and one external
package consumer, on one document corpus (see Status above). Read "every
system" as the target the architecture is built toward, not as a track record.
The idea
[structure hints] ──┐
├─▶ Eyu ─▶ proposed entities / relations
[raw records] ──┘ + confidence
+ {claim, sources[], path[]}
- Input, not fetch. Eyu never pulls data. A caller hands it declared
structure (field hints, a schema, an M3L-style declaration) and/or raw
records to look at. What the caller doesn't supply, Eyu doesn't know.
That includes its own earlier answers: a caller that keeps proposals can
hand back the entities they identified as known entities (
KnownEntity— a key of the caller's choosing, the name and type, and the records that denoted it), and an entity the new records show that is one of them comes back with that key inKnownEntityKeyinstead of as a new entity. Eyu keeps nothing between calls; the known entity's records are compared, never cited. An entity earlier records only mentioned — a work order naming its machine, with no record of the machine itself — is not known: nothing said what it is. Hand it back as aMentionedEntity(key, name, type, the records that mentioned it), and when a later call meets a record of a thing that may be it, the proposal lists aMergeCandidate(the entity, the mentioned key, a grounded claim, a confidence) inMergeCandidates. A candidate is not an identity — whether to join the two is the caller's decision — so identity does not depend on which source arrives first, and no mention is ever promoted to a known entity. - Declared always wins. Where structure is explicitly declared, the
declaration is the answer. Inference only fills what nothing declared.
A call takes one declaration per subject, so a caller whose records name
several kinds of thing — chunks of a document name companies, people and
products at once — declares each kind, by name alone when that is all it
knows (
new DeclaredStructure(SubjectRef.Create("Organization"), [], [])), and the roles it needs kept apart as typed relations (PartnerOf,CustomerOf). A declaration is neither a filter nor a renaming: a type the model proposes that nothing declares still comes back, as inferred, under the name the model used. Measured on two Korean company-profile documents, three attempts each, declaringOrganization,PersonandProductwith four typed relations took the entity names proposed under more than one type (companyin one attempt,organizationin the next) from 7 of 9 to 0 of 9 in a one-chunk document, and to 0 of 19 in a seven-chunk one that drifted little without it (1 of 18), while kinds nothing declared — locations, services, dates — kept coming. Two earlier wordings of the declaration sentence failed that measurement in opposite regimes: one held document chunks to exactly the declared types, the other, allowing the model its own types only where nothing declared describes an entity, left nothing to propose beyond a form declaration. A measurement, not a promise. Enforced after the model answers, not only asked of it: every proposal carries whether its type was declared (ProposalBasis), and a relation proposed under a declared name whose ends contradict the declaration is left out and reported (RejectionReason.ContradictsDeclaration) rather than returned. One caveat from measurement: handed a partial declaration under an earlier prompt that asked the model to "infer only what nothing declares", it proposed only what was declared and inferred nothing beyond it, so a partial declaration reached fewer competency questions than no declaration at all (the declared-completeness ablation in the test project). The prompt now says a declaration is a floor, not a ceiling; whether the model treats it that way is what the ablation's "inferred reach" column measures — a measurement, not a promise. - Confidence routes, it doesn't decide. Every proposal carries a
confidence score; a caller-defined threshold routes it to auto-apply,
human review, or draft-only. Eyu proposes — it never applies anything.
The routing is code, not a convention left to the caller: a
RoutingPolicyholds the caller's thresholds (per origin — see §C),Routereturns the tier, andTracewraps the proposal in HoneAI'sITracedPrediction<T>provenance stamp so anyIHitlGate-style review flow can consume it. There is no default threshold, on purpose: self-reported confidence is not trustworthy uncalibrated — see design rationale, §D. - No claim without a reason. A response is a structure of
{claim, sources[], path[]}. A claim that can't cite its sources cannot be expressed — this is enforced by the output shape, not by a prompt — and a source must be a record the call was actually given: a response that cites an id it was never shown is refused whole, not passed through — the check can see that an id exists, never that the record says what the claim says, so one invented citation leaves no ground for trusting the others. Every other defect costs only its own element: an entity or relation with a missing field or an out-of-range confidence, every entity under a duplicated id, and a relation whose end is not a surviving entity of the same response are left out, and each appears inOntologyProposal.Rejectionswith aRejectionReason. Nothing is dropped silently — a caller that wants all-or-nothing checksRejections.Count, and a caller measuring the model counts the reasons. ("Grounding" throughout means this citation back to records — not the normalization of a mention to an ontology term id that biomedical extraction tools call ontology grounding; Eyu does no such lookup. See design rationale, §B.)
What Eyu is
- A judgment library. Given structure and records, it proposes what the
entities, relations, and their confidence are — including which records
refer to the same real-world entity (entity resolution is part of
proposing what the entities are, not a separate concern). An entity
cites every record it appears in, and says which of those are it
(
EntityProposal.DenotedBy) — the rest only mention it, the way a work order names its machine. Only the denoting records are a claim that they are one entity. That's the whole surface. - Storage-agnostic. It has no raw store, no projection target, no query engine of its own.
- Provider-agnostic. Model access is a single injected port; local or hosted inference both work unmodified.
What Eyu is not
- Not a store. It doesn't own raw data, projected tables, or a graph database. Something upstream owns storage; Eyu only judges what's in it.
- Not a permission system. It has no concept of who is allowed to see what — that's a consuming system's job, enforced before or after Eyu is called, never inside it.
- Not a federation layer. It doesn't know how to reach a remote system, retry a query, or merge live records. It receives records; it doesn't fetch them.
- Not an agent, not a UI, not a chat surface. It answers a structural question with a grounded proposal — nothing about how that proposal reaches a person is in scope here.
Ports
| Port | Responsibility |
|---|---|
IStructureSource |
What the caller has already declared — field hints, relations, version |
IRecordSample |
Raw records to infer from, when declaration alone is insufficient |
IOntologyProposer |
The core judgment: entities, relations, confidence, and entity resolution (merging records that denote the same entity) — all from the above. Each entity carries its Name as the records write it, and each proposal a VocabularyOrigin (Innate | Acquired) — see design rationale, §C |
IGroundingContract |
{claim, sources[], path[]} — the shape every answer is expressed in |
IModelClient |
Provider-neutral inference access (local or hosted) |
Each consumer implements IStructureSource/IRecordSample for its own world
— an owned raw store, a declared schema, a federated read — and gets the same
proposal logic back through IOntologyProposer.
Routing is also a value, not a port. A caller builds a
RoutingPolicy from its own
RoutingThresholds (auto-apply / review lower bounds, one pair per
VocabularyOrigin), then calls proposal.Route(policy) for the tier or
proposal.Trace(policy) for the proposal wrapped in a HoneAI
PredictionProvenance (SourceLayer = Frontier, the confidence, the claim as
rationale, RequiresReview for every tier a machine may not act on, and the
route / origin / basis / cited record ids as annotations). Those annotation keys
are constants on
ProvenanceAnnotations, and the
stamp's value type is string: route, origin and basis are enum names, while the
cited ids arrive as a JSON array of strings — record ids are caller-supplied
and may contain any delimiter, so they are serialized rather than joined. A
consumer reading the trace parses that one value; the rest are read as-is. Eyu
references only
HoneAI.Abstractions —
the zero-dependency contract package — and never implements IHitlGate: opening
a gate, awaiting the reviewer, and applying an approved proposal are the
consumer's, because Eyu never applies anything.
The bundled HttpModelClient speaks
the OpenAI-compatible chat/completions shape; the caller owns the HttpClient's base
address, auth header and timeout. A provider-specific request field the library does not
model — a self-hosted server's thinking control (chat_template_kwargs, reasoning), a
temperature — is passed through an optional extraBody of opaque JSON merged into every
request (the OpenAI SDK's extra_body convention; model and messages stay owned by the
client). When the provider reports token counts, ModelResponse.Usage carries them
(prompt / completion / total, each optional) verbatim, so a caller can budget context or
bill against the server's own numbers.
The proposer describes the JSON it expects in the prompt, but a sentence in a prompt does not
stop every model from wrapping its answer in a markdown code fence — some self-hosted
instruction models do so routinely — and a fenced answer is refused as invalid JSON (the
exception carries the text, so the cause is visible). The parser does not strip fences: a
formatting violation it absorbed would stop showing up in measurement. So the proposer also
hands the model client the response's JSON Schema (ModelRequest.ResponseSchema, strict:
every field required, nothing extra), and HttpModelClient sends it as OpenAI structured
output — response_format: {"type": "json_schema", ...} — with no configuration. Its
json_schema form is used rather than the older {"type": "json_object"} because servers do
not all honor the latter: measured against one self-hosted OpenAI-compatible server and
instruction model, json_object was accepted and ignored (three of three answers still
fenced), while the proposer's schema sent this way produced parseable JSON on five of five
calls across a one-chunk and a seven-chunk document — and the same calls with structured
output turned off came back fenced again.
A server that rejects json_schema, or handles it badly, is the caller's to override: a
response_format in extraBody replaces the mapped one, and {"type": "text"} turns
structured output off.
var extraBody = new Dictionary<string, JsonElement>
{
["response_format"] = JsonSerializer.SerializeToElement(new { type = "text" }),
};
var client = new HttpModelClient(httpClient, model, extraBody);
A caller with its own IModelClient maps ResponseSchema to its provider's structured
output the same way, or ignores it and relies on the prompt.
Entity resolution is tuned through a value, not a port: SinglePassOntologyProposer
accepts an optional LinkageOptions record
covering the record-linkage pre-filter's classification thresholds (compared against the
log-likelihood ratio, not the posterior — docs/linkage-benchmark.md), its EM iteration
limits, and whether field comparison is exact or similarity-based. Every default
reproduces the behavior of passing nothing, so a caller reaches for it only once a live
run shows the defaults classifying that caller's data badly — what each value does, and
what is still unmeasured about them, is in Status above.
Gray-zone pairs grow with the square of the batch, so they — not the records — are what fills a
prompt first: measured, 47 records produced 542 pairs, more lines than the records themselves, and
104 records produced 2,187 pairs and a prompt past a 131,072-token context.
LinkageOptions.MaxGrayZonePairsInPrompt (default 200) caps how many are put to the model; past
it the pairs with the strongest prior toward the same entity are kept, and the proposal says how
many were left out (OntologyProposal.Linkage.GrayZonePairsOmitted). A pair left out is not judged
different — the model still sees both records — it is only not singled out. Size a batch by its
records and their length plus that cap.
Every proposal from SinglePassOntologyProposer also carries what the pre-filter contributed
(OntologyProposal.Linkage): its full result for the call (Analysis) and the confirmed groups
the answer spread over more than one entity anyway (SplitClusters). The pre-filter compares every
field two records share, so documents that copy the same values from a record they refer to — the
same machine's number and location on every inspection — agree, and read as one entity; a model
that sees they are different documents will not merge them. That list is where the disagreement
shows, for a caller to settle from the records.
One default is a premise rather than a tuning value. The pre-filter assumes a record
is one mention of one entity — a row, a form submission, a directory entry — so that
two records agreeing on their fields is evidence they denote the same thing. A document
fragment is not that: a text chunk with a title and a path names many entities and
denotes none, and there the premise inverts — measured, chunks of one document agree on
their metadata, the estimator reads the agreement as identity, and the whole document is
pre-linked as one entity before the model sees it. Records of that kind still propose and
ground correctly (a chunk id is a fine source id); pass
LinkageOptions with RecordsDenoteEntities: false and no pair is compared, every
record stays its own singleton, and the prompt carries no pre-linked groups or gray-zone
pairs. With the pre-filter off there is no linkage prior, so nothing adjusts the model's
self-reported confidence up or down — it is carried through as given. The cited sources then
no longer identify an entity either (a chunk cites many), and no chunk denotes one —
EntityProposal.DenotedBy is always empty here, whatever the model answered — so
EntityProposal.Name — the entity as the records write it — is what a caller links and stores by; EntityId only ties
relations to entities inside one proposal and must never be persisted as an identity. Keep one document
per batch, and mind that every record is rendered into the
prompt in full — the batch size is bounded by the model's context, not by Eyu.
Exporting as RDF/OWL
Eyu.Rdf writes an OntologyProposal as an OWL ontology in RDF Turtle, so a proposal can be
opened in an OWL editor, loaded into a triple store or checked by a SHACL validator:
var turtle = OntologyTurtle.ToTurtle(proposal, new RdfExportOptions(new Uri("https://example.org/plant#")));
Each distinct entity type becomes an owl:Class, each distinct relation name an
owl:ObjectProperty, each entity an owl:NamedIndividual and each relation an assertion between
two of them. The claim, the cited records, which of them denote an individual (eyu:denotedBy), the
confidence and whether a type was declared travel as annotations — on a relation, as an OWL axiom annotation (owl:Axiom), the form an OWL editor attaches
to the assertion itself. Acquired terms are minted
under the namespace you pass; innate ones are Eyu's own terms, and neither is aligned to an outside
vocabulary. An acquired term's IRI comes from its name compared the way Eyu compares names — case,
separators and Unicode form ignored (Hangul decomposed or precomposed, Latin full-width or not) — with a class's first letter upper-cased and a property's lower-cased, so
Work Order and WorkOrder from two exports are one class :Workorder, and maintains and
Maintains one property :maintains; the spelling each export used is kept as the term's
rdfs:label. Names, labels and claims are written in Unicode NFC; record ids and field names exactly as
given. Rejections are not written. The package depends on nothing beyond Eyu.Core.
An individual's IRI never comes from EntityId, which the model picks afresh on every call. It is
derived from the entity's name and type, compared ignoring case, separators and Unicode form, together with the
records that denote it when any does — the records alone are not enough, since a model reading one
row that reports an event says the row denotes the aircraft, the part and the event alike. The key is
hashed under entity/ in your namespace, so two exports of the same records name one entity alike
and a triple store merging them merges its individuals. The IRI holds only as long as its inputs do:
a renamed entity, a type named differently beyond case and separators (a declared vocabulary holds
types still) or a different set of denoting records is a different IRI, and two different things with one name, type and set of
denoting records share one — within one proposal too, where a model reading a document chunk by chunk
proposes the same company once per chunk: those entities are one individual carrying every claim.
OntologyTurtle.IndividualIris returns the IRI each entity is written under, for linking your own
triples to them. An entity matched to a known entity (KnownEntityKey) is written under that entity
instead: a key that is an absolute IRI is the IRI — pass the IRI an earlier export gave the entity and
a match keeps it, however this call named it — and any other key is hashed under entity/known/.
Merge candidates are not written: a candidate is not owl:sameAs, and joining it is the caller's call.
The ontology names the rule its IRIs were minted under — <ontology> eyu:iriRule "0.5.0", the
release that introduced the rule. The value changes only in a release that changes how a class,
property or individual IRI is minted, and says so in its changelog. Individual IRIs survive such a
change and term IRIs may not, so a store that keeps exports from two rules types one individual into
both the old class and the new one. Load each export into its own named graph and eyu:iriRule
tells which graphs an older rule wrote — drop those and re-export. (Every export names the same
ontology node, so merged into one default graph the annotations merge too and no longer say which
triple came from where.) An export without it was written by 0.5.0, under the rule "0.5.0" names
but before the annotation existed, or by 0.4.0 or earlier, whose rules differ (see the changelog) —
so a graph without it is not by that alone one to drop.
Further reading
Design rationale — why judgment requires structure first, what kind of thing a proposed ontology is, and why proposals carry an innate/acquired origin tag, each with a confidence grade on how well-anchored the reasoning is.
License
| Product | Versions Compatible and additional computed target framework versions. |
|---|---|
| .NET | net10.0 is compatible. net10.0-android was computed. net10.0-browser was computed. net10.0-ios was computed. net10.0-maccatalyst was computed. net10.0-macos was computed. net10.0-tvos was computed. net10.0-windows was computed. |
-
net10.0
- Eyu.Core (>= 0.7.0)
NuGet packages
This package is not used by any NuGet packages.
GitHub repositories
This package is not used by any popular GitHub repositories.
See CHANGELOG.md