AnonymousPatternDiscovery 0.1.0
dotnet add package AnonymousPatternDiscovery --version 0.1.0
NuGet\Install-Package AnonymousPatternDiscovery -Version 0.1.0
<PackageReference Include="AnonymousPatternDiscovery" Version="0.1.0" />
<PackageVersion Include="AnonymousPatternDiscovery" Version="0.1.0" />
<PackageReference Include="AnonymousPatternDiscovery" />
paket add AnonymousPatternDiscovery --version 0.1.0
#r "nuget: AnonymousPatternDiscovery, 0.1.0"
#:package AnonymousPatternDiscovery@0.1.0
#addin nuget:?package=AnonymousPatternDiscovery&version=0.1.0
#tool nuget:?package=AnonymousPatternDiscovery&version=0.1.0
AnonymousPatternDiscovery
A zero-dependency, deterministic toolkit of classic (non-neural, CPU-only) pattern-discovery and data-mining algorithms for .NET 8. It covers the surface you'd reach to Python/Java for today — frequent itemsets, association rules, sequential & process mining, clustering, anomaly detection, graph mining, causal-structure tests, time-series motifs, survival analysis — implemented in pure C# with no external package references.
The design premise: hand the algorithms data anonymized to bare integer ids and let them surface structure without being told what to look for. A fluent front door does the integer encoding for you, so you can speak in your own labels and get named results back.
- No dependencies. A single
net8.0assembly. Nothing to vendor, nothing transitive. - Deterministic. Same input → same output, every time. Counts are cross-checked against published benchmarks and independent implementations.
- All public. Every analyzer is a small, directly-constructable class. Use the fluent API for the common cases, or call any analyzer directly for full control.
- IntelliSense included. The package ships its XML docs, so every type and method is documented in your editor.
Contents
- Install
- The 60-second version
- Core idea: anonymous integer encoding
- Path 1 — the fluent on-ramp
- Path 2 — the one-call orchestrator
- Path 3 — calling analyzers directly
- Data shaping: the six core shapes
- "I have X / I want Y" → analyzer
- Analyzer catalog
- Reading results
- Determinism, threading & performance
- Versioning
Install
dotnet add package AnonymousPatternDiscovery --version 0.1.0
<PackageReference Include="AnonymousPatternDiscovery" Version="0.1.0" />
A single net8.0 assembly with no transitive dependencies. The package ships its XML docs, so you get full
IntelliSense from the binary alone — the consuming project never needs the source.
The 60-second version
using AnonymousPatternDiscovery.Fluent;
// You have transactions of items, in your own labels. Find everything interesting.
var rows = new[]
{
("order-1", "milk"), ("order-1", "bread"), ("order-1", "eggs"),
("order-2", "milk"), ("order-2", "bread"),
("order-3", "beer"), ("order-3", "diapers"),
// ...
};
NamedDiscoveryReport report = PatternDiscovery.Baskets(
rows,
transaction: r => r.Item1, // what groups items together
item: r => r.Item2); // the thing that co-occurs
Console.WriteLine(report.ToMarkdown()); // ranked findings, ready to share
No id maps, no configuration, no statistical setup. You passed strings; you got named findings back.
Core idea: anonymous integer encoding
Every analyzer consumes integer ids only — no strings ever reach the algorithms. That's deliberate: it keeps the methods hypothesis-free (they can't be "told" what a thing means) and makes them fast and reproducible. The flip side is that the missing ingredient for discovery is structure, not names: an analyzer needs to know which thing is this, is it the same as that, and in what transaction/period/group did it occur — and that's all carried by integer ids plus the relations you supply.
You normally never see this layer — the fluent API encodes your labels on the
way in and decodes them on the way out. When you call analyzers directly, you manage the mapping yourself with
the included LabelCodec<T>:
using AnonymousPatternDiscovery.Fluent;
var items = new LabelCodec<string>(); // deterministic two-way string ↔ int map
int milk = items.Encode("milk"); // assigns 1 the first time
int bread = items.Encode("bread"); // 2
// ... run an analyzer over the ints ...
string name = items.Decode(1); // "milk" — reattach labels for display
The one hard part of using this library is shaping your data — picking what counts as a "transaction" and mapping your fields to ids. The data-shaping section is the practical reference. The algorithms take it from there.
Path 1 — the fluent on-ramp
AnonymousPatternDiscovery.Fluent.PatternDiscovery is the string-first front door. Three static entry points
cover the most common tasks; each handles the integer encoding for you and returns named results.
Baskets — "here are my transactions, find everything"
One row per (transaction, item) participation. Co-occurrence is derived by grouping on the transaction key. Supplying a timestamp unlocks the temporal analyzers; supplying a group unlocks contrast mining.
NamedDiscoveryReport report = PatternDiscovery.Baskets(
salesLines,
transaction: r => r.OrderId,
item: r => r.Sku,
timestamp: r => r.PurchasedAt, // optional → enables burst / change-point / seasonality
group: r => r.Region); // optional → enables contrast (this group vs the rest)
foreach (var h in report.Highlights)
{
Console.WriteLine($"[{h.Kind}] {string.Join(", ", h.Entities)} (score {h.Score:0.00}) — {h.Detail}");
}
string markdown = report.ToMarkdown(top: 25); // shareable table of the top findings
Edits — "a subset was changed/selected; what's the rule, and what was missed?"
Given a set of elements, which were "changed", and their named attributes, this recovers the intent(s) behind the selection, flags elements that match a rule but weren't changed (likely missed), and flags changes that fit no rule (off-pattern). Numeric attributes are left raw — the learner discovers the thresholds.
NamedEditExplanation<int> explained = PatternDiscovery.Edits(
parts,
id: p => p.Id,
changed: p => p.WasUpdated,
categorical: p => new Dictionary<string, string> { ["type"] = p.Type, ["zone"] = p.Zone },
numeric: p => new Dictionary<string, double> { ["length"] = p.Length });
Console.WriteLine(explained.ToMarkdown());
// e.g. "Intent 1: type=Frame ∧ length ≥ 40.5 covers 60 edits (precision 95%, recall 59%)"
// "Likely missed: 12, 88, 140"
// "Off-pattern: 7"
Diff — "explain what changed between two snapshots"
Same as Edits, but it derives the changed set for you by comparing before and after snapshots of the same
elements (an element counts as changed if any attribute differs, or it was added/removed).
NamedEditExplanation<string> diff = PatternDiscovery.Diff(
before, after,
id: e => e.Key,
categorical: e => e.Tags,
numeric: e => e.Metrics);
Path 2 — the one-call orchestrator
When you want a no-configuration baseline at the integer level, DiscoveryOrchestratorInt is the "throw a
dataset at it" entry point. It inspects the data (timestamps? groups? size), scales each analyzer's thresholds
to the row count, runs every applicable analyzer, and returns one ranked report.
using AnonymousPatternDiscovery;
var obs = new List<BasketObservation>();
// BasketObservation(TransactionId, EntityId, RoleId?, GroupId?, Timestamp)
obs.Add(new BasketObservation(0, milk, null, null, DateTimeOffset.MinValue));
obs.Add(new BasketObservation(0, bread, null, null, DateTimeOffset.MinValue));
// ...
OrchestratedReportInt report = new DiscoveryOrchestratorInt().Run(obs);
Console.WriteLine($"ran: {string.Join(", ", report.AnalyzersRun)}");
Console.WriteLine($"skipped: {string.Join("; ", report.AnalyzersSkipped)}");
foreach (var h in report.Highlights)
{
Console.WriteLine($"[{h.Kind}] score {h.Score:0.00}");
}
(PatternDiscovery.Baskets is a thin, string-friendly wrapper around exactly this.)
Path 3 — calling analyzers directly
Every analyzer is public and directly constructable. Build the integer input, set any options, call the one
entry method. A few representative examples:
Association rules (Apriori, multi-item antecedents):
using AnonymousPatternDiscovery;
IReadOnlyList<AssociationRuleFindingInt> rules =
new AssociationRuleMinerInt().Mine(obs, new RuleMiningOptionsInt { MaxAntecedentSize = 3 });
foreach (var r in rules)
{
// Antecedent (ids) ⇒ Consequent (id), with Confidence / Lift / Leverage / Conviction
Console.WriteLine($"{string.Join("+", r.Antecedent)} ⇒ {r.Consequent} " +
$"conf {r.Confidence:P0}, lift {r.Lift:0.0}");
}
All frequent itemsets (counts by size, optional itemset retention):
FrequentItemsetResultInt fi =
new FrequentItemsetMinerInt().Mine(obs, new FrequentItemsetMiningOptionsInt { MinSupportCount = 50 });
Console.WriteLine($"{fi.TotalFrequentItemsets} frequent itemsets; by size: {string.Join(",", fi.CountBySize)}");
Sequential patterns (order matters — PrefixSpan):
var sequences = new[]
{
new[] { search, compare, wishlist, checkout },
new[] { search, checkout },
};
IReadOnlyList<SequentialPatternInt> seqs =
new SequentialPatternMinerInt { MinSupport = 2 }.Mine(sequences);
// each: Pattern (ordered ids) + Support
Process mining — discover control flow from an event log (one ordered trace per case):
// traces: each case's ordered activity ids (a workflow run, an editing session, a customer journey…)
var traces = new[]
{
new[] { receive, review, approve, ship },
new[] { receive, review, reject },
new[] { receive, review, review, approve, ship }, // a repeated step = rework
};
// 1) the directly-follows model + Heuristics-Miner dependency backbone
ProcessModelInt model = new DirectlyFollowsMinerInt { DependencyThreshold = 0.8, MinEdgeFrequency = 2 }.Mine(traces);
foreach (var e in model.Backbone) // the de-noised control flow
{
Console.WriteLine($"{e.From} → {e.To} (count {e.Count}, dependency {e.Dependency:0.00})");
}
// model.StartActivities / model.EndActivities / model.SelfLoops (rework) are also populated
// 2) which paths actually happen, ranked (the "happy path" vs the long tail)
VariantReportInt variants = new VariantAnalysisInt().Analyze(traces);
Console.WriteLine($"{variants.DistinctVariants} distinct variants over {variants.TotalCases} cases");
// 3) (optional) a block-structured process tree
ProcessTreeNodeInt tree = new InductiveMinerInt().Discover(traces);
Deviant subgroups w.r.t. a binary target (WRAcc beam search):
// records: one set of present ids per row; the target is just another id present-or-not
IReadOnlyList<SubgroupInt> subgroups =
new SubgroupDiscoveryInt().Discover(records, target: convertedFlagId);
// each: Description (ids) + Size + TargetRate + Wracc
Recurring shapes & anomalies in a numeric series (matrix profile):
double[] series = LoadSignal();
MatrixProfileResultInt mp = new MatrixProfileInt { Window = 30 }.Analyze(series);
// mp.MotifA / mp.MotifB = the most-recurring shape; mp.Discord = the biggest anomaly
Communities in a typed graph (Louvain):
var g = new TypedGraphInt(nodeType); // nodeType: id → type id
g.AddEdge(a, b, Friend); g.AddEdge(b, a, Friend); // add both directions for "undirected"
var communities = new CommunityDetectorInt().Detect(g, Friend);
Every other analyzer follows the same shape:
new XxxInt { ...options... }.DoIt(input). Hover any type in your editor for the full signature and notes — the XML docs ship with the package.
Data shaping: the six core shapes
Picking the right input shape (and the right transaction grain) is the whole game.
| Shape | Type | One-liner |
|---|---|---|
| Basket | BasketObservation(TxnId, EntityId, RoleId?, GroupId?, Timestamp) |
one row per (transaction, item) |
| Records-as-sets | IReadOnlyCollection<int> per record |
the set of "present" ids per record; the target is just another id |
| Typed features | ElementFeaturesInt(Categorical: int→int, Numeric: int→double) |
per element: categorical attr→value + numeric attr→value |
| Sequence(s) | IReadOnlyList<int> per sequence (or one long stream) |
ordered event ids; order matters |
| Numeric series | IReadOnlyList<double> |
a real-valued signal |
| Typed graph | TypedGraphInt (node→type) + AddEdge(src, dst, relType) |
typed nodes, typed directed edges |
Three golden rules
- Everything is integer ids. Map labels to ints (use
LabelCodec<T>), reattach names after discovery. - Pick the transaction grain first — it's THE decision. A "transaction" is one unit of co-occurrence: a basket, an element, a person, a session, a time period. Too coarse → everything co-occurs (no signal); too fine → nothing co-occurs.
- Only add what you'll use. Timestamp unlocks temporal analyzers; Group unlocks contrast; quantities unlock utility. Don't pre-bin numerics — the analyzers discretize and discover thresholds themselves.
Common pitfalls: wrong transaction grain (#1 mistake) · pre-binned numerics (hides the threshold) · dirty
variant values (AB12 vs ab12) fragmenting a rule · high-cardinality items → raise the support floor ·
colliding id namespaces (encode attributes, values, and targets in disjoint ranges, e.g. attrs 1–99, values
100_000+, targets 2_000_000+) · forgetting the name map.
"I have X / I want Y" → analyzer
| You have… | You want to find… | Shape | Analyzer |
|---|---|---|---|
| transactions of items | what co-occurs / rules / bundles | Basket | DiscoveryOrchestratorInt (one call) or PatternDiscovery.Baskets |
| items with timestamps | spikes, regime shifts, cycles, A-then-B | Basket + Timestamp | burst / change-point / seasonality / sequence |
| a set someone changed/selected | the common thread + what was missed | Typed features + changed-set | MixedConditionExplainerInt / PatternDiscovery.Edits |
| two correlated variables | is it real or a confound? | Records-as-sets | ConditionalIndependenceTestInt / ConfoundingAndContextInt / PcAlgorithmInt |
| records with a binary target | the most deviant subgroup | Records-as-sets + target id | SubgroupDiscoveryInt |
| ordered events | the ordered pattern (journey) | Sequence(s) / stream | SequentialPatternMinerInt / EpisodeMinerInt |
| a numeric signal | recurring shapes / anomalies | Numeric series | MatrixProfileInt |
| entities + relations | groups / rings / motifs / connection rules | Typed graph | community / dense / motif / metapath / FSM |
Analyzer catalog
A non-exhaustive map of what's in the box (every entry is a public …Int class):
- Itemsets & association —
AssociationStrengthAnalyzerInt(lift/PMI/NPMI),AssociationRuleMinerInt(Apriori rules),FrequentItemsetMinerInt(all frequent),CondensedItemsetMinerInt(closed & maximal),RareCorrelatedMinerInt,HighUtilityMinerInt(value ≠ frequency),KrimpMinerInt(MDL/compression),InteractionMinerInt(XOR/synergy),ContrastSetMinerInt,EclatInt,QuantitativeAssociationMinerInt. - Temporal —
TemporalBurstAnalyzerInt,ChangePointAnalyzerInt,SeasonalityAnalyzerInt,SequenceMinerInt(lead–lag),SequentialPatternMinerInt(PrefixSpan),EpisodeMinerInt(WINEPI),PeriodicItemsetMinerInt. - Anomaly —
FpofAnalyzerInt,IsolationForestInt,LocalOutlierFactorInt,HbosInt,KnnOutlierInt. - Clustering —
KMeansInt,KMedoidsInt,DbscanInt,OpticsInt,GaussianMixtureInt,AgglomerativeClusteringInt,BisectingKMeansInt,FuzzyCMeansInt, plusClusterValidityInt(choose-k). - Graph mining (
TypedGraphInt) —CommunityDetectorInt(Louvain),DenseSubgraphMinerInt(k-core / densest),MotifCounterInt,MetapathMinerInt,FrequentSubgraphMinerInt,GraphCentralityInt(PageRank/betweenness/…),LinkPredictionInt,KCoreDecompositionInt,KTrussInt,LabelPropagationInt,PersonalizedPageRankInt,WeightedPathsInt,BipartiteMatchingInt,GraphAlgorithmsInt. - Causal & significance —
ConditionalIndependenceTestInt(G-test),PcAlgorithmInt,ChowLiuTreeInt,PartialCorrelationInt,ConfoundingAndContextInt,SubgroupDiscoveryInt(WRAcc),SignificantPatternMinerInt(Westfall–Young FWER),EvidenceEngineInt+TrustScore,StabilityAnalyzerInt,HoldoutValidatorInt. - Edit / change explanation —
ChangeSetExplainerInt,MixedConditionExplainerInt,RuleListExplainerInt(multi-intent),ProgressiveRuleDiscoveryInt(how-many-intents scan),ContextThresholdInt,RelationalFeatureBuilderInt,DerivedFeatureBuilderInt. - Time series —
MatrixProfileInt(motifs/discords),DynamicTimeWarpingInt,SaxPaaInt,AutocorrelationInt,SeasonalDecompositionInt. - Supervised (interpretable) —
DecisionTreeInt(CART),NaiveBayesInt,KnnClassifierInt,KnnRegressorInt,RidgeRegressionInt,LassoRegressionInt, withCrossValidationInt/ClassificationMetricsInt. - Dimensionality reduction —
PcaInt,TruncatedSvdInt,NmfInt,RandomProjectionInt. - Process mining —
AlphaMinerInt,InductiveMinerInt(+Plus),DirectlyFollowsMinerInt(Heuristics),FuzzyMinerInt,ConformanceCheckingInt,AlignmentConformanceInt,VariantAnalysisInt,TransitionSystemMinerInt,SocialNetworkMinerInt,OrganizationalRoleMinerInt,PerformanceAnalysisInt,RemainingTimePredictorInt. - Text / linkage —
StringDistance(Levenshtein/Jaro–Winkler/…),TfIdfInt,MinHashLshInt,EntityResolutionInt. - Streaming —
StreamingStatisticsInt,ChangeDetectorInt(CUSUM/Page–Hinkley/ADWIN),Sketches(Count–Min, HyperLogLog). - Survival —
KaplanMeierInt(+ log-rank). - Stats toolbox —
CorrelationInt,HypothesisTestsInt,EntropyInt,RankDistanceInt. - Preprocessing — scaling, discretization (incl. MDLP), imputation, encoders, feature selection, SMOTE, quantile transform.
If you can think of a classic data-mining method, check the catalog first — it's probably here.
Reading results
For the edit/selection explainers, the vocabulary is consistent:
- Rule / Intent = one thing that was going on (one conjunction of predicates).
- Precision = purity of a rule (of the elements it matches, how many were actually changed). A precision gap below 100% is the signal — those matched-but-unchanged elements are the likely missed list.
- Recall = that rule's share of the whole change set.
- Off-pattern = changed but fits no rule → the genuine exceptions.
For Baskets / the orchestrator, each highlight has a Kind (e.g. association, burst, outlier), a Score,
the Entities involved, and a Detail string. NamedDiscoveryReport.ToMarkdown() and
NamedEditExplanation.ToMarkdown() render a shareable summary directly.
A pattern is not a cause. Findings are statistical structure over anonymous ids; a pattern can be real,
robust, and still non-causal. The evidence helpers (EvidenceEngineInt, confidence intervals, Benjamini–
Hochberg q-values, bootstrap stability, holdout persistence) tell you "this is robust and worth
investigating," not "this caused that." Causal claims need design (interventions, matched controls, natural
experiments), not just a surviving correlation.
Determinism, threading & performance
- Deterministic by design. No wall-clock or RNG leaks into results; the same input always yields the same output. (This is also why the test suite can assert exact counts.)
- Single-threaded analyzers. Each analyzer runs synchronously on the calling thread — no hidden parallelism, no shared global state. Run independent analyzers concurrently yourself if you want; they don't contend. In a UI, run on a background thread so the UI stays responsive.
- Pure CPU. No GPU, no native interop, no model files. Complexity is algorithmic — for very large /
high-cardinality inputs, raise the support/min-count thresholds rather than expecting SIMD to save you (an
AVX2 path was measured slower than scalar
POPCNThere; the headroom is algorithmic). - Correctness is validated. Frequent-itemset counts match independent Eclat and Apriori implementations exactly on the standard FIMI benchmark datasets; every "weird" scenario plants a known truth the algorithm must recover while decoys stay non-significant.
Versioning
0.1.0 — first packaged release. The library is mature and broadly tested, but the public API may still
change before 1.0 while the surface is finalized. Pin an exact version for reproducible builds.
Semantic versioning from here: patch = fixes, minor = additive, major = breaking.
| Product | Versions Compatible and additional computed target framework versions. |
|---|---|
| .NET | net8.0 is compatible. net8.0-android was computed. net8.0-browser was computed. net8.0-ios was computed. net8.0-maccatalyst was computed. net8.0-macos was computed. net8.0-tvos was computed. net8.0-windows was computed. net9.0 was computed. net9.0-android was computed. net9.0-browser was computed. net9.0-ios was computed. net9.0-maccatalyst was computed. net9.0-macos was computed. net9.0-tvos was computed. net9.0-windows was computed. net10.0 was computed. net10.0-android was computed. net10.0-browser was computed. net10.0-ios was computed. net10.0-maccatalyst was computed. net10.0-macos was computed. net10.0-tvos was computed. net10.0-windows was computed. |
-
net8.0
- No dependencies.
NuGet packages
This package is not used by any NuGet packages.
GitHub repositories
This package is not used by any popular GitHub repositories.
| Version | Downloads | Last Updated |
|---|---|---|
| 0.1.0 | 149 | 6/14/2026 |