Brainograph.Utilities.Search 2.4.0

The owner has unlisted this package. This could mean that the package is deprecated, has security vulnerabilities or shouldn't be used anymore.
dotnet add package Brainograph.Utilities.Search --version 2.4.0
                    
NuGet\Install-Package Brainograph.Utilities.Search -Version 2.4.0
                    
This command is intended to be used within the Package Manager Console in Visual Studio, as it uses the NuGet module's version of Install-Package.
<PackageReference Include="Brainograph.Utilities.Search" Version="2.4.0" />
                    
For projects that support PackageReference, copy this XML node into the project file to reference the package.
<PackageVersion Include="Brainograph.Utilities.Search" Version="2.4.0" />
                    
Directory.Packages.props
<PackageReference Include="Brainograph.Utilities.Search" />
                    
Project file
For projects that support Central Package Management (CPM), copy this XML node into the solution Directory.Packages.props file to version the package.
paket add Brainograph.Utilities.Search --version 2.4.0
                    
#r "nuget: Brainograph.Utilities.Search, 2.4.0"
                    
#r directive can be used in F# Interactive and Polyglot Notebooks. Copy this into the interactive tool or source code of the script to reference the package.
#:package Brainograph.Utilities.Search@2.4.0
                    
#:package directive can be used in C# file-based apps starting in .NET 10 preview 4. Copy this into a .cs file before any lines of code to reference the package.
#addin nuget:?package=Brainograph.Utilities.Search&version=2.4.0
                    
Install as a Cake Addin
#tool nuget:?package=Brainograph.Utilities.Search&version=2.4.0
                    
Install as a Cake Tool

Brainograph.Utilities.Search

PostgreSQL full-text search (tsvector + GIN) blended with pg_trgm fuzzy/prefix/substring matching, packaged as a reusable kernel for services that own their own searchable tables. No Elasticsearch, no dedicated indexing service, no cross-service federation — a service searches its own PostgreSQL database directly. See dev-kit/plans/2026-07-20-search-design.md for the full design and docs/adr/0001-utilities-search-package.md for why this ships as its own package.

Consumed today by the in-service search endpoints of bg-content, bg-gis and bg-usercontent.

Add search to a service

Every searchable entity gets the same three pieces: two maintained columns, a migration that wires up the PostgreSQL plumbing, and a query built from RankedSearch + ToSearchPageAsync.

1. Migration — columns, extensions, trigger, indexes

Add the search_vector/search_text columns in an EF Core migration, then run the SearchSql builders via migrationBuilder.Sql(...). SearchSql only returns SQL strings — nothing here talks to the database itself, so it composes cleanly inside a normal migration Up:

protected override void Up(MigrationBuilder migrationBuilder)
{
    migrationBuilder.AddColumn<NpgsqlTsVector>(
        name: "search_vector",
        table: "articles",
        type: "tsvector",
        nullable: true);

    migrationBuilder.AddColumn<string>(
        name: "search_text",
        table: "articles",
        type: "text",
        nullable: true);

    // pg_trgm + unaccent — CREATE EXTENSION IF NOT EXISTS for both.
    migrationBuilder.Sql(SearchSql.EnsureExtensions());

    // Weight order = priority order: title (A) outranks intro (B) outranks body (C).
    (string Column, char Weight)[] fields =
    [
        ("main_title", 'A'),
        ("intro_text", 'B'),
        ("formatted_text", 'C'),
    ];

    migrationBuilder.Sql(SearchSql.TriggerFunction(
        fnName: "articles_search_fn",
        vectorCol: "search_vector",
        textCol: "search_text",
        config: "simple",
        fields: fields));

    migrationBuilder.Sql(SearchSql.CreateTrigger(
        triggerName: "articles_search_trg",
        fnName: "articles_search_fn",
        table: "articles"));

    migrationBuilder.Sql(SearchSql.GinIndexes(
        table: "articles",
        vectorCol: "search_vector",
        textCol: "search_text"));

    // Stage-B: backfill pre-existing rows. The trigger only fires on INSERT/UPDATE, so rows that
    // already existed before this migration keep NULL search_vector/search_text until they are
    // next written or explicitly backfilled, e.g.:
    //   migrationBuilder.Sql("UPDATE articles SET main_title = main_title;");
    // (a no-op write that re-runs the trigger) or an equivalent bounded UPDATE ... SET
    // search_vector = <TsVectorExpression>, search_text = ... sized per table at cutover.
}

SearchSql's DDL is idempotent-friendly (IF NOT EXISTS / CREATE OR REPLACE / drop-then-create for the trigger, since PostgreSQL has no CREATE TRIGGER IF NOT EXISTS), so re-running the same migration statements is safe.

2. Entity mapping

Map the two columns as NpgsqlTsVector/string. The trigger — not EF — owns their values, so mark them [DatabaseGenerated(DatabaseGeneratedOption.Computed)]: EF reads them back after a save but never sends them on INSERT/UPDATE.

public class Article
{
    public int Id { get; set; }

    public string MainTitle { get; set; } = string.Empty;

    public string IntroText { get; set; } = string.Empty;

    public string FormattedText { get; set; } = string.Empty;

    [DatabaseGenerated(DatabaseGeneratedOption.Computed)]
    public NpgsqlTsVector SearchVector { get; set; } = null!;

    [DatabaseGenerated(DatabaseGeneratedOption.Computed)]
    public string SearchText { get; set; } = string.Empty;
}

If the entity's columns aren't already snake_case-mapped to match the migration's identifiers, add the usual HasColumnName(...) calls in OnModelCreating.

3. Query — RankedSearch + ToSearchPageAsync

public Task<PagedList<ArticleSearchDto>> SearchAsync(
    string term, int pageNumber, int pageSize, CancellationToken ct)
{
    var options = new SearchOptions { Config = "simple" };

    return _db.Articles
        .Where(a => a.Status == Status.Published)   // service-owned filters go here
        .RankedSearch(term, e => e.SearchVector, e => e.SearchText, options)
        .ToSearchPageAsync(
            r => new ArticleSearchDto(r.Item.Id, r.Item.MainTitle, r.Score),
            r => r.Item.Id,   // stable tiebreak: ordering is Score desc, then this key asc
            pageNumber,
            pageSize,
            ct);
}

RankedSearch filters and scores but does not order or page. The filter is @@ websearch_to_tsquery(...) OR a to_tsquery(..., '<token>:*') prefix disjunct OR a pg_trgm word_similarity() fuzzy disjunct (length-independent, so a short query still matches a word buried in a long body). The score is a three-tier CASE, not a flat blend:

tier filter score band
exact @@ websearch_to_tsquery 2.0 + ts_rank_cd(…, 32) [2.0, 3.0)
prefix @@ to_tsquery('tok:*') 1.0 + ts_rank_cd(…, 32) [1.0, 2.0)
fuzzy word_similarity >= threshold word_similarity [0.0, 1.0)

ts_rank_cd's normalization argument is DivideByItselfPlusOne (32) — rank/(rank+1), bounded below 1.0 — so the bands cannot overlap however many times a term matches one row. That transform is strictly increasing, so ordering within a tier is unchanged. The exact arm is tested first because an exact hit is also a prefix hit; swapped, every exact match would collapse into the prefix band.

Why the prefix tier exists (bg-gis#16): the structured half matches whole lexemes, so a reader typing three characters of a seven-character place name matched nothing — websearch_to_tsquery on վան does not match the lexeme վանք. The trigram disjunct could cover it, but it is gated by FuzzyMinTermLength and is not index-accelerated in its word_similarity(...) form. The prefix disjunct is index-backed by the same GIN index the structured half uses, and it is precise: napo prefix-matches napoleon and not naples, where the trigram path returns both.

So exact / phrase hits rank above prefix hits, and both above fuzzy hits, with field weighting intact (a plain GREATEST blend would let an exact word saturate word_similarity to 1.0 and flatten that ordering). Apply any service-specific filters before calling it, on the base IQueryable<T>. ToSearchPageAsync orders by Score descending then the caller's tiebreak key (r => r.Item.Id above — required so tied scores page deterministically), clamps pageNumber/pageSize, and projects to the platform PagedList<TDto> shape. An empty or below-MinTermLength term returns an empty page, not an error.

Notes

  • unaccent fallback. Both search-support columns are accent-folded via the unaccent extension: TsVectorExpression/TriggerFunction wrap each field in unaccent(coalesce(...)) before to_tsvector, and the trigger does the same before lowercasing into search_text. If a target environment cannot install unaccent (e.g. a managed Postgres without superuser access), drop the CREATE EXTENSION IF NOT EXISTS unaccent; line and neutralize both unaccent(...) calls in the generated trigger body (the one inside the search_vector assignment and the one inside the search_text assignment) before running it — the trigger then falls back to plain lower()/no accent-folding on both columns. SearchTerm.Normalize still strips diacritics from the query term in C#, so unaccented data keeps matching correctly either way; without unaccent in the DB, accented data only matches an unaccented query if the stored value happens to already be unaccented.
  • Maintenance MUST be a trigger, not a GENERATED ... STORED column. unaccent() is STABLE, not IMMUTABLE, and PostgreSQL only allows IMMUTABLE functions in a GENERATED ALWAYS AS (...) STORED column expression (or a plain index expression). A BEFORE INSERT OR UPDATE trigger has no such restriction, which is why SearchSql.TriggerFunction builds a plpgsql trigger body instead of a generated column — do not "simplify" the migration into a generated column, it will fail to create.

Tuning (SearchOptions)

Property Default Meaning
Config "simple" tsvector/tsquery text-search configuration (regconfig). The platform default is simple (no language stemmer) — multilingual content with no built-in Armenian stemmer.
MinTermLength 1 Minimum normalized-term length before a search runs at all.
EnableFuzzy true Whether the pg_trgm word_similarity disjunct is OR-ed into the filter. Not a fallback gated on the structured match missing; its word_similarity is the score for fuzzy-only rows (structured @@ matches score in the higher 1.0 + ts_rank_cd tier).
FuzzyMinTermLength 3 Minimum normalized-term length before the fuzzy disjunct is added (e.g. GIS raises this to 7 to mirror the old fuzzy gate). Raising it does not cost in-word matching any more — that is the prefix disjunct's job, with its own floor.
FuzzyThreshold 0.3 Minimum pg_trgm word_similarity() score for a fuzzy match to count as a hit. Because word_similarity scores the best-matching word extent (not the whole string), 0.3 is a usable recall default even against long bodies.
EnablePrefix true Whether the to_tsquery(config, '<token>:*') prefix disjunct is OR-ed into the filter. This is what makes a term match inside a word. Skipped automatically for a term using websearch operator syntax — a quoted phrase, a -negated word, a bare or — because rebuilding those from sanitised tokens would change what the caller asked for (napoleon -naples would turn an exclusion into a required term).
PrefixMinTermLength 2 Minimum normalized-term length before the prefix disjunct is added. Its own floor, independent of FuzzyMinTermLength: a one-character prefix matches most of a corpus, whereas the length at which guessing at a typo becomes worthwhile is a different, higher number.

Choosing how widely to match — SearchMatchTier

EnablePrefix/EnableFuzzy are the mechanism; SearchMatchTier is the vocabulary a service should expose and reason in. It names the widest tier a caller will accept, and the same three words label what a row actually matched — so a request and a response can never describe matching differently.

var options = new SearchOptions { FuzzyMinTermLength = 7 };
options.ApplyQuality(quality);          // Exact | Prefix | Fuzzy  -> the two toggles
...
MatchType = r.Score >= SearchMatchTiers.ExactFloor  ? SearchMatchTier.Exact
          : r.Score >= SearchMatchTiers.PrefixFloor ? SearchMatchTier.Prefix
                                                    : SearchMatchTier.Fuzzy

Name the constants in the projection rather than writing 1.0/2.0: a const folds into the expression tree and translates to a SQL CASE, while a method call over the projected row does not — and a band restated in three service repositories drifts the first time the scoring is retuned.

The enum is ordered widest-last, so default(SearchMatchTier) is the strictest value: a forgotten quality errs toward too few results rather than opening the search up.

Widening only ever adds rows. A row that matched exactly keeps its Exact label when fuzzy is admitted alongside — the tiers are bands on one score, not alternative queries.

Build / test / pack

dotnet build Bg.Shared.slnx -c Release
dotnet test tests/Brainograph.Utilities.Search.Tests
dotnet pack src/Brainograph.Utilities.Search -c Release
Product Compatible and additional computed target framework versions.
.NET net10.0 is compatible.  net10.0-android was computed.  net10.0-browser was computed.  net10.0-ios was computed.  net10.0-maccatalyst was computed.  net10.0-macos was computed.  net10.0-tvos was computed.  net10.0-windows was computed. 
Compatible target framework(s)
Included target framework(s) (in package)
Learn more about Target Frameworks and .NET Standard.

NuGet packages

This package is not used by any NuGet packages.

GitHub repositories

This package is not used by any popular GitHub repositories.

Version Downloads Last Updated