LexiSharp 0.1.0

dotnet add package LexiSharp --version 0.1.0
                    
NuGet\Install-Package LexiSharp -Version 0.1.0
                    
This command is intended to be used within the Package Manager Console in Visual Studio, as it uses the NuGet module's version of Install-Package.
<PackageReference Include="LexiSharp" Version="0.1.0" />
                    
For projects that support PackageReference, copy this XML node into the project file to reference the package.
<PackageVersion Include="LexiSharp" Version="0.1.0" />
                    
Directory.Packages.props
<PackageReference Include="LexiSharp" />
                    
Project file
For projects that support Central Package Management (CPM), copy this XML node into the solution Directory.Packages.props file to version the package.
paket add LexiSharp --version 0.1.0
                    
#r "nuget: LexiSharp, 0.1.0"
                    
#r directive can be used in F# Interactive and Polyglot Notebooks. Copy this into the interactive tool or source code of the script to reference the package.
#:package LexiSharp@0.1.0
                    
#:package directive can be used in C# file-based apps starting in .NET 10 preview 4. Copy this into a .cs file before any lines of code to reference the package.
#addin nuget:?package=LexiSharp&version=0.1.0
                    
Install as a Cake Addin
#tool nuget:?package=LexiSharp&version=0.1.0
                    
Install as a Cake Tool

LexiSharp

CI CodeQL codecov SonarCloud NuGet License: MIT

A lightweight lexical text search and classification library for .NET.

LexiSharp provides a small, dependency-free set of interfaces and implementations for indexing plain text and retrieving/ranking/classifying documents without any semantic or ML model — pure lexical statistics.

Status & provenance

  • Young, single-maintainer project (v0.1.0). No production track record and no external contributors yet; the public API may still change between minor versions. Evaluate it as such before adopting it.
  • Developed with AI assistance. Most of the code and this README were written with LLM coding agents, then reviewed and tested by the maintainer. The techniques implemented (BM25, RRF, SPLADE-style sparse retrieval, MaxSim) follow established IR literature; this repository contributes no novel research.
  • Claims vs. evidence. Behavior described in this README is covered by the xUnit suite (Postgres/ParadeDB integration tests self-skip without a live instance — see Building & testing). Comparative or performance statements are kept to a minimum; the few measured numbers live in BENCHMARKS.md and are indicative only.
  • Not every combination is exercised. Engines, scorers, rerankers and mergers are tested individually and in a few documented combinations, but the full cross-product is not: treat unusual pairings as supported by construction, not yet stress-tested.

Features

  • Pluggable architecture: an ITextIndex, ITextScorer and ITokenizer are independent contracts; algortihms can be swapped without touching the engine.
  • Four ranking strategies behind the same ITextSearchEngine:
    • Bm25Scorer — Okapi BM25, with ready-made Bm25Parameters profiles (Balanced, Aggressive, Conservative),
    • TfIdfScorer — TF-IDF,
    • QueryLikelihoodScorer — probabilistic language model (Jelinek-Mercer smoothing),
    • BooleanScorer — exact AND/OR filter.
  • In-memory inverted index (InMemoryTextIndex) with term positions, document frequencies, corpus statistics and incremental Add/Remove, plus an index statistics snapshot (GetStatistics: documents, vocabulary, tokens, average length, vocabulary richness).
  • Score boosting (BoostedTextSearchEngine): a decorator that applies signed score adjustments (multiplicative factor and/or additive offset) per result — boost a category or a priority, damp or penalize stale matches — without touching the underlying engine.
  • Second-stage reranking: an IReranker seam and the RerankedTextSearchEngine decorator (over-fetch, re-rank, guard rails) in the core; shipped rerankers include a diversity-preserving MMR, a cascade pipeline that chains any number of reranking stages with per-stage trimming, a cross-encoder reranker driven by a consumer-provided pairwise scoring model (ICrossEncoderScorer), and a ColBERT MaxSim reranker that re-scores a shortlist token-by-token with late interaction (ITokenEmbeddingProvider).
  • Sparse learned embeddings: SparseTextSearchEngine and its ISparseEmbeddingProvider seam bring SPLADE/uniCOIL-style retrieval (.NET-core only, weights learned, inverted-index scoring kept) without pulling ONNX into the library — the model lives in the consumer.
  • Metadata filters: declarative, AND-composed filters over document fields (MetadataFilterOperator: equal, not-equal, contains, numeric-or-ordinal greater/less than) in SearchOptions — applied by the stock engine before any relevance math.
  • Lexical similarity (LexiSharp.Similarity): pairwise token-set measures (Jaccard, Sørensen–Dice) over the library tokenizer, a pg_trgm-style character trigram similarity, and a rolling Levenshtein edit distance — near-duplicate detection and fuzzy matching with no index.
  • Keyword extraction (LexiSharp.Keywords): a corpus-backed TF-IDF extractor (demotes corpus-frequent words) and a graph-based TextRank extractor (weighted co-occurrence graph + PageRank), both deterministic and tokenizer-configurable.
  • Explainable scoring: Bm25Scorer implements IScoreExplainer, and RankedTextSearchEngine.Explain returns a per-term breakdown (TF, IDF, term score, length normalization, parameter values) of any ranking decision.
  • Evaluation metrics (RetrievalMetrics): Precision@k, Recall@k, F1@k, binary and graded nDCG@k (exponential gains), plus ReciprocalRank@k (→ MRR) and AveragePrecision@k (→ MAP).
  • BM25 tuning: Bm25ParameterTuner grid-searches k1/b against your own validation queries, judged by Precision@k, Recall@k, F1@k or nDCG@k.
  • Supervised classification (NaiveBayesClassifier): multinomial Naive Bayes with Laplace smoothing, exposing a dedicated ITextClassifier interface.
  • Configurable tokenizer: Unicode NFKD normalization and diacritics removal, lowercasing, optional stop-word removal, optional n-grams, and a pluggable IStemmer seam (no stemmer is bundled on purpose — bring your own, e.g. Snowball). Tokenization is SIMD-accelerated (SearchValues + IndexOfAnyExcept, with a System.Text.Ascii fast path in normalization).
  • Optional backends, shipped as separate packages:
    • LexiSharp.Postgres — PostgreSQL backends implementing the same ITextSearchEngine: a lexical engine over tsvector + GIN + unaccent, an ANN engine over pgvector (HNSW/IVFFlat) driven by an external IEmbeddingProvider, a learned-sparse engine over pgvector sparsevec (HNSW) driven by an ISparseEmbeddingProvider, and an approximate fuzzy engine over the pg_trgm trigram extension (with optional fuzzystrmatch refinement);
    • LexiSharp.ParadeDBtrue Okapi BM25 on top of the pg_search Tantivy extension (AGPL-3, requires the ParadeDB Docker image or self-hosted extension).

Quick start

using LexiSharp.Core;
using LexiSharp.Indexing;
using LexiSharp.Ranking;

ITextSearchEngine engine = new RankedTextSearchEngine(
    new InMemoryTextIndex(),
    new Bm25Scorer());

engine.Index(new[]
{
    new SearchDocument("1", "The search engine uses BM25 to rank the results"),
    new SearchDocument("2", "TF-IDF is a classic method of textual search"),
    new SearchDocument("3", "Italian cuisine is renowned in Rome"),
});

IReadOnlyList<SearchResult> results = engine.Search("textual search");

foreach (var result in results)
    Console.WriteLine($"{result.DocumentId} - {result.Score:0.###}: {result.Document.Text}");

Change the ranking algorithm without rebuilding the index:

var bm25Engine   = new RankedTextSearchEngine(index, new Bm25Scorer());
var tfIdfEngine  = new RankedTextSearchEngine(index, new TfIdfScorer());
var booleanEngine = new RankedTextSearchEngine(index, new BooleanScorer(BooleanMatch.AllTerms));

Classification

using LexiSharp.Classification;

var classifier = new NaiveBayesClassifier();
classifier.Train(trainingDocuments); // requires a non-null SearchDocument.Category

foreach (var prediction in classifier.Predict("i cannot connect to the internet"))
    Console.WriteLine($"{prediction.Category}: {prediction.Probability:P}");

Tokenizer customization

using LexiSharp.Linguistics;

var tokenizer = new Tokenizer(new TokenizerOptions
{
    RemoveStopWords = true,       // English list, or provide StopWords.Create(...)
    NGramMax = 2,                 // produce unigrams + bigrams
    Stemmer = new MyStemmer(),    // implement IStemmer (French, Snowball, ...)
});

Score boosting (BoostedTextSearchEngine)

Wrap any engine to boost or damp its ranking without changing the engine. The boost is a function of the whole result, so it can read the score, the document metadata or external data (a closure over your own store):

ITextSearchEngine boosted = new BoostedTextSearchEngine(baseEngine,
    result =>
    {
        double factor = 1.0;
        if (result.Document.Category == "priority")
            factor = 2.0;                                  // up-weighted metadata
        if (result.Document.Fields.TryGetValue("stale", out _))
            factor *= 0.5;                                 // damp old matches

        return new ScoreBoost(Multiply: factor, Add: -0.5); // factor and/or offset, signed
    });

Writes are forwarded to the inner engine; Search applies score * Multiply + Add to every candidate, then re-sorts and re-applies Limit/MinimumScore. A plain double is accepted as a multiplicative factor (result => 2.0). Positive boosts (factor > 1, positive offset) and negative ones (factor in (0, 1) damp, negative offset penalty) are equally expressible; factor 0 drops the document entirely. The decorator requests more candidates than the final limit (maxCandidates, default 50) so boosted documents can surface.

Reranking (IReranker, MMR, cascade)

Retrieve with recall, then re-rank a shortlist with precision. The core seam is IReranker; RerankedTextSearchEngine decorates any engine (over-fetches, re-ranks, applies MinimumScore/Limit on the final scores):

using LexiSharp.Core;

IReranker reranker = ...;                                    // yours, or the MMR one below
ITextSearchEngine engine = new RerankedTextSearchEngine(baseEngine, reranker, maxCandidates: 100);

LexiSharp ships two built-in rerankers. MMR (Maximal Marginal Relevance) re-orders candidates so each next pick is relevant and different from the picks before it — near-duplicate results are pushed back; candidates without a vector are never penalized:

using LexiSharp.Hybrid;

var vectors = new Dictionary<string, ReadOnlyMemory<float>>
{
    ["doc-1"] = embedding1, // pre-computed with your IEmbeddingProvider
    ["doc-2"] = embedding2,
};

IReranker mmr = new MaximalMarginalRelevanceReranker(vectors, lambda: 0.7, limit: 5);

Cascade chains any number of stages, trimming between stages so only the strongest candidates reach the expensive final ones; it is itself an IReranker, so cascades nest:

var pipeline = new CascadeRerankPipeline(
    new IReranker[] { lexicalReranker, mmr },
    new CascadeRerankOptions(StageLimit: 20, FinalLimit: 5, MinimumScore: 0.01));

Cross-encoder re-scores the shortlist with a pairwise model — a precision stage for cases where whole-corpus scoring would be too expensive (ColBERT-style late interaction, an LLM judge, ...). The model itself is a consumer-provided seam (ICrossEncoderScorer, same contract as IEmbeddingProvider: LexiSharp never runs the model):

IReranker cross = new CrossEncoderReranker(myOnnxCrossEncoder, limit: 5);

CrossEncoderReranker replaces each candidate's score with the model's, drops 0/NaN/infinity scores, and applies an optional MinimumScore and Limit.

MaxSim (ColBERT-style late interaction) re-scores the shortlist token-by-token instead of as a single embedding: every token of the query is embedded, each scores against the whole candidate's token embeddings (max similarity per query token, summed), so a query token never has to "average itself away" across the document:

using LexiSharp.Core;
using LexiSharp.Hybrid;

var tokenVectors = new Dictionary<string, IReadOnlyList<ReadOnlyMemory<float>>>
{
    ["doc-1"] = doc1TokenEmbeddings, // token embeddings pre-computed at index time
    ["doc-2"] = doc2TokenEmbeddings,
};

IReranker maxsim = new MaxSimReranker(
    myTokenEmbedder,                 // ITokenEmbeddingProvider (core seam, consumer-provided)
    tokenVectors,
    limit: 5,
    minimumScore: 0.0);

MaxSimReranker scores each candidate as Σₜ max_tok cosine(q_t, d_tok) — for each query token, the best cosine against any of the candidate's token embeddings — drops 0/NaN/infinity scores and ranks by total. It is the middle ground between whole-document cosine and the full pairwise pass of a cross-encoder.

Metadata filters

Gate the corpus with structured predicates over SearchDocument.Fields — every filter must hold (AND), and filtering happens before scoring:

var options = new SearchOptions(
    Limit: 10,
    Filters:
    [
        new MetadataFilter("kind", MetadataFilterOperator.Equal, "article"),
        new MetadataFilter("year", MetadataFilterOperator.GreaterThan, "2023"),
        new MetadataFilter("tags", MetadataFilterOperator.Contains, "nlp"),
    ]);

var results = engine.Search("vector search", options);

Comparisons are culture-invariant; greater/less-than go numeric when both sides parse as numbers, otherwise ordinal. Documents missing a field fail everything except NotEqual.

Lexical similarity and keyword extraction

Pairwise similarity for near-duplicate detection and record de-duplication — token-set measures over the library tokenizer, pg_trgm-style trigrams, and Levenshtein:

using LexiSharp.Similarity;

bool duplicate = LexicalSimilarity.Jaccard(stored, incoming) > 0.5;
double fuzzy   = LexicalSimilarity.Trigram("kubernetes cluster", "kubernetes clusters");
int edits      = LevenshteinDistance.Distance("kitten", "sitting"); // 3

Keyword extraction pulls the representative terms out of a text. TF-IDF becomes corpus-aware when built over an ITextIndex; TextRank needs no corpus at all:

using LexiSharp.Keywords;

IKeywordExtractor tags = new TfIdfKeywordExtractor(someIndex, StopWordTokenizer);
IKeywordExtractor graph = new TextRankKeywordExtractor(StopWordTokenizer); // co-occurrence + PageRank

foreach (var keyword in graph.Extract(document.Text, topN: 5))
    Console.WriteLine($"{keyword.Term}: {keyword.Score:F3}");

Index persistence (LexiSharp.MessagePack)

Save and reload an InMemoryTextIndex as compact, LZ4-compressed MessagePack binary — documents (id, text, fields, category) and tokenizer configuration:

// install once:  dotnet add package LexiSharp.MessagePack
using LexiSharp.MessagePack;

MessagePackTextIndexPersistence.Save(index, "corpus.bin");
var reloaded = MessagePackTextIndexPersistence.Load("corpus.bin"); // identical statistics, no re-indexing

A Tokenizer (stop words, n-grams, single-char terms) is reconstructed automatically. A custom ITokenizer cannot be serialized: hand the same implementation to Load — a type-name check protects against rebuilding with the wrong pipeline. Stemmed tokenizers likewise require the original tokenizer at load time (stemmers are not serializable).

The same package persists a sparse engine through MessagePackSparseIndexPersistence: the stored corpus is the documents plus their learned weights, so reloading bypasses the model — only queries need the ISparseEmbeddingProvider again:

MessagePackSparseIndexPersistence.Save(sparseEngine, "splade.bin");
var reloaded = MessagePackSparseIndexPersistence.Load("splade.bin", mySplade); // exact same search scores

Explainable scoring and BM25 tuning

Audit any ranking decision term by term, then let the corpus pick its own parameters:

var engine = new RankedTextSearchEngine(index, new Bm25Scorer());
ScoreExplanation? why = engine.Explain("doc-1", "search engine");
// why.Terms -> per-term TF, IDF and score contribution; why.LengthRatio, why.Parameters...

var tuner = new Bm25ParameterTuner(index, validationQueries: [
    new Bm25ValidationQuery("search engine", ["doc-1", "doc-7"]),
    new Bm25ValidationQuery("fuzzy matching", ["doc-3"]),
]);
Bm25TuningResult tuning = tuner.Tune(topK: 5);          // grid search over k1 x b
var tunedEngine = new RankedTextSearchEngine(index, new Bm25Scorer(tuning.Parameters));

Each validation query lists the relevant document ids; candidates are judged with Precision@k, Recall@k, F1@k (default) or nDCG@k (RetrievalMetrics), averaged over the set. The index is never mutated; tuning.Grid exposes every evaluated (k1, b) point.

PostgreSQL backend (LexiSharp.Postgres)

Persistent, shared, concurrent search on top of a classic PostgreSQL setup. Several engines, all implementing ITextSearchEngine and sharing the same documents table (so the hybrid engine can fan out and merge lexical + vector + fuzzy results with ReciprocalRankFusionMerger):

Lexical (PostgresTextSearchEngine) — full-text over tsvector:

// install once:  dotnet add package LexiSharp.Postgres
using LexiSharp.Postgres;

ITextSearchEngine engine = new PostgresTextSearchEngine(
    "Host=db;Port=5432;Username=app;Password=secret;Database=search");

engine.Add(new SearchDocument("1", "the quick brown fox jumps over the lazy dog",
    new Dictionary<string, string> { ["kind"] = "fable" }, "fable"));

The provider installs (idempotently) the unaccent extension, a documents table with a tsv tsvector column and a GIN index, then queries it with websearch_to_tsquery and ranks with ts_rank_cd. Combined with the simple config, unaccent mirrors LexiSharp's accent-insensitive normalization. Scores are PostgreSQL-native, so they are not numerically comparable to Bm25Scorer/TfIdfScorer — feed both backends into the hybrid engine below when you need one consistent ordering.

Vector (PostgresVectorSearchEngine) — ANN over pgvector (HNSW or IVFFlat), fed by your own embeddings:

ITextSearchEngine vector = new PostgresVectorSearchEngine(
    "Host=db;Port=5432;Username=app;Password=secret;Database=search",
    new MyEmbeddingProvider(),                                // your ONNX/model-server deps, never LexiSharp
    new PostgresVectorOptions { Dimension = 384, Distance = VectorDistance.Cosine });

vector.Add(new SearchDocument("1", "the quick brown fox ..."));
vector.Search("a fast fox");   // scores: cosine→1-dist, L2→1/(1+dist), inner product→-dist

The engine installs (idempotently) the vector extension, adds an embedding vector(D) column and an HNSW (or IVFFlat) index on the same documents table. IVFFlat needs rows to cluster lists, so the index is created on the first EnsureSchema() call after your first inserts. ANN results are approximate: combine with HybridTextSearchEngine + RRF to trade recall for speed — exact cosine behavior is verified in the integration suite.

By default the integration tests are skipped unless POSTGRES_TEST_CONNECTION points at a live instance (e.g. Host=localhost;Port=5432;Username=postgres;Password=postgres;Database=lexisharp). The vector and sparse tests additionally require the vector extension: use the pgvector/pgvector:pg16 image (lexical tests only need stock PostgreSQL).

Sparse (PostgresSparseSearchEngine) — learned-sparse ANN over pgvector sparsevec (HNSW only — IVFFlat is unavailable for sparsevec), fed by your own sparse model:

ITextSearchEngine sparse = new PostgresSparseSearchEngine(
    connectionString,
    mySpladeProvider,                                       // ISparseEmbeddingProvider, never LexiSharp
    new PostgresSparseOptions
    {
        Vocabulary = vocabulary,                            // term → coordinate, fixed up front
        Distance = SparseDistance.InnerProduct,             // default; dot product suits SPLADE
    });

sparse.Add(new SearchDocument("1", "the quick brown fox ..."));
sparse.Search("a fast fox");   // scores: inner product→-dist, cosine→1-dist, L2/L1→1/(1+dist)

The engine installs (idempotently) the vector extension, adds a sparse sparsevec(D) column and an HNSW index on the same shared documents table. Because sparsevec is a positional format, the vocabulary (term → coordinate) is an index-layout decision: it must be fixed once and shared between the provider at index time and the one at query time. Terms outside the vocabulary are ignored; a query with no known term returns nothing. HNSW — unlike IVFFlat — works on empty tables and supports inserts, so there is no "index after first batch" step. sparsevec caps a vector at 1000 non-zero elements. Scores are PostgreSQL-native and again depend on the chosen distance; route through ReciprocalRankFusionMerger when mixing with the lexical engine.

Fuzzy (PostgresFuzzySearchEngine) — approximate, typo-tolerant matching over pg_trgm trigrams, with optional fuzzystrmatch (edit distance + phonetics):

ITextSearchEngine fuzzy = new PostgresFuzzySearchEngine(connectionString, new PostgresFuzzyOptions
{
    SearchMode = TrgmSearchMode.Nearest,        // kNN: closest labels first (autocomplete)
    // SearchMode = TrgmSearchMode.Similarity,  // threshold: content % query (de-dup, did-you-mean)
    SimilarityThreshold = 0.3,                  // honored via set_limit() in Similarity mode
    UseLevenshteinRefinement = true,            // exact edit-distance post-filter
    IncludePhonetic = true,                     // metaphone column; phonetic matches (Similarity mode)
});
fuzzy.Index(new[]
{
    new SearchDocument("1", "katherine"),
    new SearchDocument("2", "catherine"),
});
fuzzy.Search("caterin");   // typo-tolerant: both labels come back

The engine installs (idempotently) pg_trgm (+ fuzzystrmatch when enabled) and GiST and GIN trigram indexes on the same shared documents table. Scores are trigram similarities in [0, 1] (1 identical, 0 no shared trigram → excluded, honoring the library's score-0 convention). Nearest mode orders with the GiST kNN operator (content <-> query); Similarity mode ranks by similarity() above the configured threshold. These are the classic building blocks for autocomplete, de-duplication of names/addresses and "did you mean". PostgreSQL-native scores again call for ReciprocalRankFusionMerger when mixing with other engines.

ParadeDB backend (LexiSharp.ParadeDB)

Okapi BM25 ranking computed by Tantivy inside PostgreSQL through the pg_search extension — an option to consider when ts_rank_cd ranking is not good enough and true BM25 is wanted.

// install once:  dotnet add package LexiSharp.ParadeDB
using LexiSharp.ParadeDB;

ITextSearchEngine engine = new ParadeDBTextSearchEngine(connectionString);
engine.Add(new SearchDocument("1", "the quick brown fox jumps over the lazy dog"));

The engine installs (idempotently) the pg_search extension and a ParadeDB index (USING paradedb, the renamed USING bm25) on the same shared documents table as the Postgres engines, then matches with the ||| disjunction operator and ranks with pdb.score(id) — real BM25 (Tantivy variant), unlike ts_rank_cd. The default content tokenizer (pdb.simple with ASCII folding) mirrors LexiSharp's diacritic-insensitive, lowercase normalization:

new ParadeDBTextSearchEngine(
    connectionString,
    new ParadeDBOptions { ContentTokenizer = "pdb.simple('ascii_folding=true')" });

BM25 scores are PostgreSQL-native, so they are not numerically comparable to Bm25Scorer/TfIdfScorer — use the hybrid engine's ReciprocalRankFusionMerger (or re-rank) for a single cross-engine ordering. Note the extension is AGPL-3 licensed: fine for SaaS/internal use, but review it if you distribute the stack.

By default the tests are skipped unless POSTGRES_TEST_CONNECTION points at a live instance with pg_search available (the paradedb/paradedb:pg16 image ships it, preloaded) — they self-skip when the extension is absent.

Hybrid engine

Federate a hot in-memory index and a cold persistent backend, and produce one consistent global ranking:

using LexiSharp.Core;
using LexiSharp.Hybrid;
using LexiSharp.Indexing;
using LexiSharp.Ranking;

ITextSearchEngine hybrid = new HybridTextSearchEngine(new ITextSearchEngine[]
{
    new RankedTextSearchEngine(new InMemoryTextIndex(), new Bm25Scorer()), // hot subset
    postgresEngine,                                                       // cold backend
});

IReadOnlyList<SearchResult> results = hybrid.Search("textual search");

HybridTextSearchEngine queries every engine, de-duplicates the candidates by document id, then re-ranks the whole union with a single scorer (RerankingResultMerger, default BM25). Writes fan out to every engine. Three merge strategies are available:

Merger Behavior Best for
RerankingResultMerger (default) re-scores the union with one ITextScorer comparable stats, identical score scale wanted
ReciprocalRankFusionMerger Σ 1/(k + rank) (k=60), rank-only engines with incomparable scales — lexical + vector, ts_rank_cd vs BM25 (Postgres vs ParadeDB vs in-memory)
WeightedScoreResultMerger normalized per-engine score blend native scores trusted, per-engine weights wanted

Reciprocal Rank Fusion never looks at scores, so it bridges engines whose scores are not comparable — the sparse and dense embedding backends land in the same formula without calibration.

Embeddings are an agreed seam, not a feature here: IEmbeddingProvider (core) describes how a consumer project (ONNX model, model server, ...) would produce vectors — LexiSharp never computes embeddings — and VectorSimilarity provides pure cosine math. PostgresVectorSearchEngine is the reference consumer: it turns any provider into an ANN backend that the same HybridTextSearchEngine merges exactly like a lexical engine.

Sparse learned embeddings (ISparseEmbeddingProvider, SparseTextSearchEngine)

SPLADE-style models (SPLADE, uniCOIL, ...) produce sparse learned vectors: a handful of term → weight pairs where the weights are learned instead of tf-idf/BM25 frequencies. The key insight of this family is that it does not replace the inverted-index infrastructure, only the scoring function. SparseTextSearchEngine applies exactly that: it keeps a classic term → document → weight inverted index, and a query scores each document by sparse dot product Σ_t w_q(t)·w_d(t,d) — only the terms the learned model activated are visited:

using LexiSharp.Core;
using LexiSharp.Indexing;

ISparseEmbeddingProvider splade = myOnnxSplade; // consumer-provided, incl. vocab mapping
ITextSearchEngine engine = new SparseTextSearchEngine(splade);

engine.Add(new SearchDocument("doc-1", "sparse retrievers beat dense on exact terms"));

var results = engine.Search("learned sparse retrieval");

Like IEmbeddingProvider, ISparseEmbeddingProvider is a pure seam in the core: the ONNX model, tokenizer and vocabulary live in the consumer. A walkthrough of writing a SPLADE provider (ONNX + HuggingFace tokenizer + vocabulary mapping) is in docs/SPLADE.md — an outline, not a tested reference implementation. Weights are expected non-negative (ReLU-like); non-positive values are treated as "term absent". The engine implements ITextSearchEngine, so it drops straight into HybridTextSearchEngine where it merges with BM25 and dense engines via ReciprocalRankFusionMerger — RRF keeps sparse-only hits (matching terms the lexical scorer and the dense cosine disagree on) that a BM25 re-scoring merge would drop.

The engine never re-embeds on reload: Export() / Import() decouple inference from persistence, MessagePackSparseIndexPersistence serializes the stored weights directly, and PostgresSparseSearchEngine is the same model over pgvector sparsevec. In every backend, only queries keep needing the provider after the corpus is loaded.

Reranking stage

The reranking stage composes with all of it: wrap the hybrid in a RerankedTextSearchEngine (core decorator) and pass a CascadeRerankPipeline, a MaximalMarginalRelevanceReranker, a CrossEncoderReranker or a MaxSimReranker to add a precision or diversity pass on top of the fused ranking — the same two-stage retrieve-then-rerank shape, one line of composition.

Architecture

Package            Responsibilities
─────────────────────────────────────────────────────────────────────────────
LexiSharp         records + interfaces + in-memory index + scorers + tokenizer + IEmbeddingProvider + ISparseEmbeddingProvider + sparse engine (export/import) + boost/rerank decorators + filters + similarity + keywords + metrics + hybrid federation (RRF, weighted, cascade, cross-encoder, MMR, MaxSim)
LexiSharp.Postgres  PostgreSQL providers: tsvector+unaccent (lexical), pgvector ANN (vector), pgvector sparsevec (sparse), pg_trgm+fuzzystrmatch (fuzzy)
LexiSharp.ParadeDB   true BM25 provider on the pg_search (Tantivy) extension
LexiSharp.MessagePack   MessagePack (binary) persistence for the in-memory index and the sparse engine

Within the core package, separation of concerns mirrors the recommendations the library was designed from:

  • the index owns corpus statistics (tf, df, document length, positions, vocabulary);
  • the scorer is a pure strategy reading from the index;
  • the engine orchestrates query tokenization, scoring, filtering and ranking.

Scoring conventions

  • A score of exactly 0 means not a match and the document is excluded from results (all built-in scorers honor this).
  • Every sub-system is culture-agnostic; text is normalized to lowercase without accents so that "Résumé" and "resume" match.

Building & testing

dotnet build LexiSharp.slnx
dotnet test  tests/LexiSharp.Tests                # xUnit suite (Postgres tests need POSTGRES_TEST_CONNECTION)
dotnet run  --project bench/LexiSharp.Benchmarks  # BenchmarkDotNet suite (published numbers: BENCHMARKS.md)

Postgres/ParadeDB integration tests run against whatever POSTGRES_TEST_CONNECTION points to: pgvector/pgvector:pg16 covers the lexical + vector + sparse + fuzzy suites, paradedb/paradedb:pg16 covers the lexical + ParadeDB (BM25) + fuzzy suites. The sparse tests self-skip when the vector extension is unavailable. The fuzzy tests self-skip when pg_trgm (and fuzzystrmatch, when exercised) are unavailable.

License

MIT

Product Compatible and additional computed target framework versions.
.NET net10.0 is compatible.  net10.0-android was computed.  net10.0-browser was computed.  net10.0-ios was computed.  net10.0-maccatalyst was computed.  net10.0-macos was computed.  net10.0-tvos was computed.  net10.0-windows was computed. 
Compatible target framework(s)
Included target framework(s) (in package)
Learn more about Target Frameworks and .NET Standard.
  • net10.0

    • No dependencies.

NuGet packages

This package is not used by any NuGet packages.

GitHub repositories

This package is not used by any popular GitHub repositories.

Version Downloads Last Updated
0.1.0 43 9/19/2026