AnonymousPatternDiscovery 0.1.0

dotnet add package AnonymousPatternDiscovery --version 0.1.0
                    
NuGet\Install-Package AnonymousPatternDiscovery -Version 0.1.0
                    
This command is intended to be used within the Package Manager Console in Visual Studio, as it uses the NuGet module's version of Install-Package.
<PackageReference Include="AnonymousPatternDiscovery" Version="0.1.0" />
                    
For projects that support PackageReference, copy this XML node into the project file to reference the package.
<PackageVersion Include="AnonymousPatternDiscovery" Version="0.1.0" />
                    
Directory.Packages.props
<PackageReference Include="AnonymousPatternDiscovery" />
                    
Project file
For projects that support Central Package Management (CPM), copy this XML node into the solution Directory.Packages.props file to version the package.
paket add AnonymousPatternDiscovery --version 0.1.0
                    
#r "nuget: AnonymousPatternDiscovery, 0.1.0"
                    
#r directive can be used in F# Interactive and Polyglot Notebooks. Copy this into the interactive tool or source code of the script to reference the package.
#:package AnonymousPatternDiscovery@0.1.0
                    
#:package directive can be used in C# file-based apps starting in .NET 10 preview 4. Copy this into a .cs file before any lines of code to reference the package.
#addin nuget:?package=AnonymousPatternDiscovery&version=0.1.0
                    
Install as a Cake Addin
#tool nuget:?package=AnonymousPatternDiscovery&version=0.1.0
                    
Install as a Cake Tool

AnonymousPatternDiscovery

A zero-dependency, deterministic toolkit of classic (non-neural, CPU-only) pattern-discovery and data-mining algorithms for .NET 8. It covers the surface you'd reach to Python/Java for today — frequent itemsets, association rules, sequential & process mining, clustering, anomaly detection, graph mining, causal-structure tests, time-series motifs, survival analysis — implemented in pure C# with no external package references.

The design premise: hand the algorithms data anonymized to bare integer ids and let them surface structure without being told what to look for. A fluent front door does the integer encoding for you, so you can speak in your own labels and get named results back.

  • No dependencies. A single net8.0 assembly. Nothing to vendor, nothing transitive.
  • Deterministic. Same input → same output, every time. Counts are cross-checked against published benchmarks and independent implementations.
  • All public. Every analyzer is a small, directly-constructable class. Use the fluent API for the common cases, or call any analyzer directly for full control.
  • IntelliSense included. The package ships its XML docs, so every type and method is documented in your editor.

Contents

  1. Install
  2. The 60-second version
  3. Core idea: anonymous integer encoding
  4. Path 1 — the fluent on-ramp
  5. Path 2 — the one-call orchestrator
  6. Path 3 — calling analyzers directly
  7. Data shaping: the six core shapes
  8. "I have X / I want Y" → analyzer
  9. Analyzer catalog
  10. Reading results
  11. Determinism, threading & performance
  12. Versioning

Install

dotnet add package AnonymousPatternDiscovery --version 0.1.0
<PackageReference Include="AnonymousPatternDiscovery" Version="0.1.0" />

A single net8.0 assembly with no transitive dependencies. The package ships its XML docs, so you get full IntelliSense from the binary alone — the consuming project never needs the source.


The 60-second version

using AnonymousPatternDiscovery.Fluent;

// You have transactions of items, in your own labels. Find everything interesting.
var rows = new[]
{
    ("order-1", "milk"), ("order-1", "bread"), ("order-1", "eggs"),
    ("order-2", "milk"), ("order-2", "bread"),
    ("order-3", "beer"), ("order-3", "diapers"),
    // ...
};

NamedDiscoveryReport report = PatternDiscovery.Baskets(
    rows,
    transaction: r => r.Item1,   // what groups items together
    item:        r => r.Item2);  // the thing that co-occurs

Console.WriteLine(report.ToMarkdown());   // ranked findings, ready to share

No id maps, no configuration, no statistical setup. You passed strings; you got named findings back.


Core idea: anonymous integer encoding

Every analyzer consumes integer ids only — no strings ever reach the algorithms. That's deliberate: it keeps the methods hypothesis-free (they can't be "told" what a thing means) and makes them fast and reproducible. The flip side is that the missing ingredient for discovery is structure, not names: an analyzer needs to know which thing is this, is it the same as that, and in what transaction/period/group did it occur — and that's all carried by integer ids plus the relations you supply.

You normally never see this layer — the fluent API encodes your labels on the way in and decodes them on the way out. When you call analyzers directly, you manage the mapping yourself with the included LabelCodec<T>:

using AnonymousPatternDiscovery.Fluent;

var items = new LabelCodec<string>();   // deterministic two-way string ↔ int map
int milk  = items.Encode("milk");       // assigns 1 the first time
int bread = items.Encode("bread");      // 2
// ... run an analyzer over the ints ...
string name = items.Decode(1);          // "milk"  — reattach labels for display

The one hard part of using this library is shaping your data — picking what counts as a "transaction" and mapping your fields to ids. The data-shaping section is the practical reference. The algorithms take it from there.


Path 1 — the fluent on-ramp

AnonymousPatternDiscovery.Fluent.PatternDiscovery is the string-first front door. Three static entry points cover the most common tasks; each handles the integer encoding for you and returns named results.

Baskets — "here are my transactions, find everything"

One row per (transaction, item) participation. Co-occurrence is derived by grouping on the transaction key. Supplying a timestamp unlocks the temporal analyzers; supplying a group unlocks contrast mining.

NamedDiscoveryReport report = PatternDiscovery.Baskets(
    salesLines,
    transaction: r => r.OrderId,
    item:        r => r.Sku,
    timestamp:   r => r.PurchasedAt,   // optional → enables burst / change-point / seasonality
    group:       r => r.Region);       // optional → enables contrast (this group vs the rest)

foreach (var h in report.Highlights)
{
    Console.WriteLine($"[{h.Kind}] {string.Join(", ", h.Entities)}  (score {h.Score:0.00}) — {h.Detail}");
}

string markdown = report.ToMarkdown(top: 25);   // shareable table of the top findings

Edits — "a subset was changed/selected; what's the rule, and what was missed?"

Given a set of elements, which were "changed", and their named attributes, this recovers the intent(s) behind the selection, flags elements that match a rule but weren't changed (likely missed), and flags changes that fit no rule (off-pattern). Numeric attributes are left raw — the learner discovers the thresholds.

NamedEditExplanation<int> explained = PatternDiscovery.Edits(
    parts,
    id:      p => p.Id,
    changed: p => p.WasUpdated,
    categorical: p => new Dictionary<string, string> { ["type"] = p.Type, ["zone"] = p.Zone },
    numeric:     p => new Dictionary<string, double> { ["length"] = p.Length });

Console.WriteLine(explained.ToMarkdown());
// e.g. "Intent 1: type=Frame ∧ length ≥ 40.5  covers 60 edits (precision 95%, recall 59%)"
//      "Likely missed: 12, 88, 140"
//      "Off-pattern: 7"

Diff — "explain what changed between two snapshots"

Same as Edits, but it derives the changed set for you by comparing before and after snapshots of the same elements (an element counts as changed if any attribute differs, or it was added/removed).

NamedEditExplanation<string> diff = PatternDiscovery.Diff(
    before, after,
    id:          e => e.Key,
    categorical: e => e.Tags,
    numeric:     e => e.Metrics);

Path 2 — the one-call orchestrator

When you want a no-configuration baseline at the integer level, DiscoveryOrchestratorInt is the "throw a dataset at it" entry point. It inspects the data (timestamps? groups? size), scales each analyzer's thresholds to the row count, runs every applicable analyzer, and returns one ranked report.

using AnonymousPatternDiscovery;

var obs = new List<BasketObservation>();
// BasketObservation(TransactionId, EntityId, RoleId?, GroupId?, Timestamp)
obs.Add(new BasketObservation(0, milk,  null, null, DateTimeOffset.MinValue));
obs.Add(new BasketObservation(0, bread, null, null, DateTimeOffset.MinValue));
// ...

OrchestratedReportInt report = new DiscoveryOrchestratorInt().Run(obs);

Console.WriteLine($"ran: {string.Join(", ", report.AnalyzersRun)}");
Console.WriteLine($"skipped: {string.Join("; ", report.AnalyzersSkipped)}");
foreach (var h in report.Highlights)
{
    Console.WriteLine($"[{h.Kind}] score {h.Score:0.00}");
}

(PatternDiscovery.Baskets is a thin, string-friendly wrapper around exactly this.)


Path 3 — calling analyzers directly

Every analyzer is public and directly constructable. Build the integer input, set any options, call the one entry method. A few representative examples:

Association rules (Apriori, multi-item antecedents):

using AnonymousPatternDiscovery;

IReadOnlyList<AssociationRuleFindingInt> rules =
    new AssociationRuleMinerInt().Mine(obs, new RuleMiningOptionsInt { MaxAntecedentSize = 3 });

foreach (var r in rules)
{
    // Antecedent (ids) ⇒ Consequent (id), with Confidence / Lift / Leverage / Conviction
    Console.WriteLine($"{string.Join("+", r.Antecedent)} ⇒ {r.Consequent}  " +
                      $"conf {r.Confidence:P0}, lift {r.Lift:0.0}");
}

All frequent itemsets (counts by size, optional itemset retention):

FrequentItemsetResultInt fi =
    new FrequentItemsetMinerInt().Mine(obs, new FrequentItemsetMiningOptionsInt { MinSupportCount = 50 });

Console.WriteLine($"{fi.TotalFrequentItemsets} frequent itemsets; by size: {string.Join(",", fi.CountBySize)}");

Sequential patterns (order matters — PrefixSpan):

var sequences = new[]
{
    new[] { search, compare, wishlist, checkout },
    new[] { search, checkout },
};
IReadOnlyList<SequentialPatternInt> seqs =
    new SequentialPatternMinerInt { MinSupport = 2 }.Mine(sequences);
// each: Pattern (ordered ids) + Support

Process mining — discover control flow from an event log (one ordered trace per case):

// traces: each case's ordered activity ids (a workflow run, an editing session, a customer journey…)
var traces = new[]
{
    new[] { receive, review, approve, ship },
    new[] { receive, review, reject },
    new[] { receive, review, review, approve, ship },   // a repeated step = rework
};

// 1) the directly-follows model + Heuristics-Miner dependency backbone
ProcessModelInt model = new DirectlyFollowsMinerInt { DependencyThreshold = 0.8, MinEdgeFrequency = 2 }.Mine(traces);
foreach (var e in model.Backbone)                       // the de-noised control flow
{
    Console.WriteLine($"{e.From} → {e.To}  (count {e.Count}, dependency {e.Dependency:0.00})");
}
// model.StartActivities / model.EndActivities / model.SelfLoops (rework) are also populated

// 2) which paths actually happen, ranked (the "happy path" vs the long tail)
VariantReportInt variants = new VariantAnalysisInt().Analyze(traces);
Console.WriteLine($"{variants.DistinctVariants} distinct variants over {variants.TotalCases} cases");

// 3) (optional) a block-structured process tree
ProcessTreeNodeInt tree = new InductiveMinerInt().Discover(traces);

Deviant subgroups w.r.t. a binary target (WRAcc beam search):

// records: one set of present ids per row; the target is just another id present-or-not
IReadOnlyList<SubgroupInt> subgroups =
    new SubgroupDiscoveryInt().Discover(records, target: convertedFlagId);
// each: Description (ids) + Size + TargetRate + Wracc

Recurring shapes & anomalies in a numeric series (matrix profile):

double[] series = LoadSignal();
MatrixProfileResultInt mp = new MatrixProfileInt { Window = 30 }.Analyze(series);
// mp.MotifA / mp.MotifB = the most-recurring shape; mp.Discord = the biggest anomaly

Communities in a typed graph (Louvain):

var g = new TypedGraphInt(nodeType);                 // nodeType: id → type id
g.AddEdge(a, b, Friend); g.AddEdge(b, a, Friend);    // add both directions for "undirected"
var communities = new CommunityDetectorInt().Detect(g, Friend);

Every other analyzer follows the same shape: new XxxInt { ...options... }.DoIt(input). Hover any type in your editor for the full signature and notes — the XML docs ship with the package.


Data shaping: the six core shapes

Picking the right input shape (and the right transaction grain) is the whole game.

Shape Type One-liner
Basket BasketObservation(TxnId, EntityId, RoleId?, GroupId?, Timestamp) one row per (transaction, item)
Records-as-sets IReadOnlyCollection<int> per record the set of "present" ids per record; the target is just another id
Typed features ElementFeaturesInt(Categorical: int→int, Numeric: int→double) per element: categorical attr→value + numeric attr→value
Sequence(s) IReadOnlyList<int> per sequence (or one long stream) ordered event ids; order matters
Numeric series IReadOnlyList<double> a real-valued signal
Typed graph TypedGraphInt (node→type) + AddEdge(src, dst, relType) typed nodes, typed directed edges

Three golden rules

  1. Everything is integer ids. Map labels to ints (use LabelCodec<T>), reattach names after discovery.
  2. Pick the transaction grain first — it's THE decision. A "transaction" is one unit of co-occurrence: a basket, an element, a person, a session, a time period. Too coarse → everything co-occurs (no signal); too fine → nothing co-occurs.
  3. Only add what you'll use. Timestamp unlocks temporal analyzers; Group unlocks contrast; quantities unlock utility. Don't pre-bin numerics — the analyzers discretize and discover thresholds themselves.

Common pitfalls: wrong transaction grain (#1 mistake) · pre-binned numerics (hides the threshold) · dirty variant values (AB12 vs ab12) fragmenting a rule · high-cardinality items → raise the support floor · colliding id namespaces (encode attributes, values, and targets in disjoint ranges, e.g. attrs 1–99, values 100_000+, targets 2_000_000+) · forgetting the name map.


"I have X / I want Y" → analyzer

You have… You want to find… Shape Analyzer
transactions of items what co-occurs / rules / bundles Basket DiscoveryOrchestratorInt (one call) or PatternDiscovery.Baskets
items with timestamps spikes, regime shifts, cycles, A-then-B Basket + Timestamp burst / change-point / seasonality / sequence
a set someone changed/selected the common thread + what was missed Typed features + changed-set MixedConditionExplainerInt / PatternDiscovery.Edits
two correlated variables is it real or a confound? Records-as-sets ConditionalIndependenceTestInt / ConfoundingAndContextInt / PcAlgorithmInt
records with a binary target the most deviant subgroup Records-as-sets + target id SubgroupDiscoveryInt
ordered events the ordered pattern (journey) Sequence(s) / stream SequentialPatternMinerInt / EpisodeMinerInt
a numeric signal recurring shapes / anomalies Numeric series MatrixProfileInt
entities + relations groups / rings / motifs / connection rules Typed graph community / dense / motif / metapath / FSM

Analyzer catalog

A non-exhaustive map of what's in the box (every entry is a public …Int class):

  • Itemsets & association — AssociationStrengthAnalyzerInt (lift/PMI/NPMI), AssociationRuleMinerInt (Apriori rules), FrequentItemsetMinerInt (all frequent), CondensedItemsetMinerInt (closed & maximal), RareCorrelatedMinerInt, HighUtilityMinerInt (value ≠ frequency), KrimpMinerInt (MDL/compression), InteractionMinerInt (XOR/synergy), ContrastSetMinerInt, EclatInt, QuantitativeAssociationMinerInt.
  • Temporal — TemporalBurstAnalyzerInt, ChangePointAnalyzerInt, SeasonalityAnalyzerInt, SequenceMinerInt (lead–lag), SequentialPatternMinerInt (PrefixSpan), EpisodeMinerInt (WINEPI), PeriodicItemsetMinerInt.
  • Anomaly — FpofAnalyzerInt, IsolationForestInt, LocalOutlierFactorInt, HbosInt, KnnOutlierInt.
  • Clustering — KMeansInt, KMedoidsInt, DbscanInt, OpticsInt, GaussianMixtureInt, AgglomerativeClusteringInt, BisectingKMeansInt, FuzzyCMeansInt, plus ClusterValidityInt (choose-k).
  • Graph mining (TypedGraphInt) — CommunityDetectorInt (Louvain), DenseSubgraphMinerInt (k-core / densest), MotifCounterInt, MetapathMinerInt, FrequentSubgraphMinerInt, GraphCentralityInt (PageRank/betweenness/…), LinkPredictionInt, KCoreDecompositionInt, KTrussInt, LabelPropagationInt, PersonalizedPageRankInt, WeightedPathsInt, BipartiteMatchingInt, GraphAlgorithmsInt.
  • Causal & significance — ConditionalIndependenceTestInt (G-test), PcAlgorithmInt, ChowLiuTreeInt, PartialCorrelationInt, ConfoundingAndContextInt, SubgroupDiscoveryInt (WRAcc), SignificantPatternMinerInt (Westfall–Young FWER), EvidenceEngineInt + TrustScore, StabilityAnalyzerInt, HoldoutValidatorInt.
  • Edit / change explanation — ChangeSetExplainerInt, MixedConditionExplainerInt, RuleListExplainerInt (multi-intent), ProgressiveRuleDiscoveryInt (how-many-intents scan), ContextThresholdInt, RelationalFeatureBuilderInt, DerivedFeatureBuilderInt.
  • Time series — MatrixProfileInt (motifs/discords), DynamicTimeWarpingInt, SaxPaaInt, AutocorrelationInt, SeasonalDecompositionInt.
  • Supervised (interpretable) — DecisionTreeInt (CART), NaiveBayesInt, KnnClassifierInt, KnnRegressorInt, RidgeRegressionInt, LassoRegressionInt, with CrossValidationInt / ClassificationMetricsInt.
  • Dimensionality reduction — PcaInt, TruncatedSvdInt, NmfInt, RandomProjectionInt.
  • Process mining — AlphaMinerInt, InductiveMinerInt (+Plus), DirectlyFollowsMinerInt (Heuristics), FuzzyMinerInt, ConformanceCheckingInt, AlignmentConformanceInt, VariantAnalysisInt, TransitionSystemMinerInt, SocialNetworkMinerInt, OrganizationalRoleMinerInt, PerformanceAnalysisInt, RemainingTimePredictorInt.
  • Text / linkage — StringDistance (Levenshtein/Jaro–Winkler/…), TfIdfInt, MinHashLshInt, EntityResolutionInt.
  • Streaming — StreamingStatisticsInt, ChangeDetectorInt (CUSUM/Page–Hinkley/ADWIN), Sketches (Count–Min, HyperLogLog).
  • Survival — KaplanMeierInt (+ log-rank).
  • Stats toolbox — CorrelationInt, HypothesisTestsInt, EntropyInt, RankDistanceInt.
  • Preprocessing — scaling, discretization (incl. MDLP), imputation, encoders, feature selection, SMOTE, quantile transform.

If you can think of a classic data-mining method, check the catalog first — it's probably here.


Reading results

For the edit/selection explainers, the vocabulary is consistent:

  • Rule / Intent = one thing that was going on (one conjunction of predicates).
  • Precision = purity of a rule (of the elements it matches, how many were actually changed). A precision gap below 100% is the signal — those matched-but-unchanged elements are the likely missed list.
  • Recall = that rule's share of the whole change set.
  • Off-pattern = changed but fits no rule → the genuine exceptions.

For Baskets / the orchestrator, each highlight has a Kind (e.g. association, burst, outlier), a Score, the Entities involved, and a Detail string. NamedDiscoveryReport.ToMarkdown() and NamedEditExplanation.ToMarkdown() render a shareable summary directly.

A pattern is not a cause. Findings are statistical structure over anonymous ids; a pattern can be real, robust, and still non-causal. The evidence helpers (EvidenceEngineInt, confidence intervals, Benjamini– Hochberg q-values, bootstrap stability, holdout persistence) tell you "this is robust and worth investigating," not "this caused that." Causal claims need design (interventions, matched controls, natural experiments), not just a surviving correlation.


Determinism, threading & performance

  • Deterministic by design. No wall-clock or RNG leaks into results; the same input always yields the same output. (This is also why the test suite can assert exact counts.)
  • Single-threaded analyzers. Each analyzer runs synchronously on the calling thread — no hidden parallelism, no shared global state. Run independent analyzers concurrently yourself if you want; they don't contend. In a UI, run on a background thread so the UI stays responsive.
  • Pure CPU. No GPU, no native interop, no model files. Complexity is algorithmic — for very large / high-cardinality inputs, raise the support/min-count thresholds rather than expecting SIMD to save you (an AVX2 path was measured slower than scalar POPCNT here; the headroom is algorithmic).
  • Correctness is validated. Frequent-itemset counts match independent Eclat and Apriori implementations exactly on the standard FIMI benchmark datasets; every "weird" scenario plants a known truth the algorithm must recover while decoys stay non-significant.

Versioning

0.1.0 — first packaged release. The library is mature and broadly tested, but the public API may still change before 1.0 while the surface is finalized. Pin an exact version for reproducible builds.

Semantic versioning from here: patch = fixes, minor = additive, major = breaking.

Product Compatible and additional computed target framework versions.
.NET net8.0 is compatible.  net8.0-android was computed.  net8.0-browser was computed.  net8.0-ios was computed.  net8.0-maccatalyst was computed.  net8.0-macos was computed.  net8.0-tvos was computed.  net8.0-windows was computed.  net9.0 was computed.  net9.0-android was computed.  net9.0-browser was computed.  net9.0-ios was computed.  net9.0-maccatalyst was computed.  net9.0-macos was computed.  net9.0-tvos was computed.  net9.0-windows was computed.  net10.0 was computed.  net10.0-android was computed.  net10.0-browser was computed.  net10.0-ios was computed.  net10.0-maccatalyst was computed.  net10.0-macos was computed.  net10.0-tvos was computed.  net10.0-windows was computed. 
Compatible target framework(s)
Included target framework(s) (in package)
Learn more about Target Frameworks and .NET Standard.
  • net8.0

    • No dependencies.

NuGet packages

This package is not used by any NuGet packages.

GitHub repositories

This package is not used by any popular GitHub repositories.

Version Downloads Last Updated
0.1.0 149 6/14/2026