Brainograph.Utilities.Search
2.4.0
dotnet add package Brainograph.Utilities.Search --version 2.4.0
NuGet\Install-Package Brainograph.Utilities.Search -Version 2.4.0
<PackageReference Include="Brainograph.Utilities.Search" Version="2.4.0" />
<PackageVersion Include="Brainograph.Utilities.Search" Version="2.4.0" />
<PackageReference Include="Brainograph.Utilities.Search" />
paket add Brainograph.Utilities.Search --version 2.4.0
#r "nuget: Brainograph.Utilities.Search, 2.4.0"
#:package Brainograph.Utilities.Search@2.4.0
#addin nuget:?package=Brainograph.Utilities.Search&version=2.4.0
#tool nuget:?package=Brainograph.Utilities.Search&version=2.4.0
Brainograph.Utilities.Search
PostgreSQL full-text search (tsvector + GIN) blended with pg_trgm fuzzy/prefix/substring
matching, packaged as a reusable kernel for services that own their own searchable tables. No
Elasticsearch, no dedicated indexing service, no cross-service federation — a service searches its
own PostgreSQL database directly. See dev-kit/plans/2026-07-20-search-design.md for the full
design and docs/adr/0001-utilities-search-package.md for why this ships as its own package.
Consumed today by the in-service search endpoints of bg-content, bg-gis and bg-usercontent.
Add search to a service
Every searchable entity gets the same three pieces: two maintained columns, a migration that wires
up the PostgreSQL plumbing, and a query built from RankedSearch + ToSearchPageAsync.
1. Migration — columns, extensions, trigger, indexes
Add the search_vector/search_text columns in an EF Core migration, then run the SearchSql
builders via migrationBuilder.Sql(...). SearchSql only returns SQL strings — nothing here talks
to the database itself, so it composes cleanly inside a normal migration Up:
protected override void Up(MigrationBuilder migrationBuilder)
{
migrationBuilder.AddColumn<NpgsqlTsVector>(
name: "search_vector",
table: "articles",
type: "tsvector",
nullable: true);
migrationBuilder.AddColumn<string>(
name: "search_text",
table: "articles",
type: "text",
nullable: true);
// pg_trgm + unaccent — CREATE EXTENSION IF NOT EXISTS for both.
migrationBuilder.Sql(SearchSql.EnsureExtensions());
// Weight order = priority order: title (A) outranks intro (B) outranks body (C).
(string Column, char Weight)[] fields =
[
("main_title", 'A'),
("intro_text", 'B'),
("formatted_text", 'C'),
];
migrationBuilder.Sql(SearchSql.TriggerFunction(
fnName: "articles_search_fn",
vectorCol: "search_vector",
textCol: "search_text",
config: "simple",
fields: fields));
migrationBuilder.Sql(SearchSql.CreateTrigger(
triggerName: "articles_search_trg",
fnName: "articles_search_fn",
table: "articles"));
migrationBuilder.Sql(SearchSql.GinIndexes(
table: "articles",
vectorCol: "search_vector",
textCol: "search_text"));
// Stage-B: backfill pre-existing rows. The trigger only fires on INSERT/UPDATE, so rows that
// already existed before this migration keep NULL search_vector/search_text until they are
// next written or explicitly backfilled, e.g.:
// migrationBuilder.Sql("UPDATE articles SET main_title = main_title;");
// (a no-op write that re-runs the trigger) or an equivalent bounded UPDATE ... SET
// search_vector = <TsVectorExpression>, search_text = ... sized per table at cutover.
}
SearchSql's DDL is idempotent-friendly (IF NOT EXISTS / CREATE OR REPLACE / drop-then-create
for the trigger, since PostgreSQL has no CREATE TRIGGER IF NOT EXISTS), so re-running the same
migration statements is safe.
2. Entity mapping
Map the two columns as NpgsqlTsVector/string. The trigger — not EF — owns their values, so mark
them [DatabaseGenerated(DatabaseGeneratedOption.Computed)]: EF reads them back after a save but
never sends them on INSERT/UPDATE.
public class Article
{
public int Id { get; set; }
public string MainTitle { get; set; } = string.Empty;
public string IntroText { get; set; } = string.Empty;
public string FormattedText { get; set; } = string.Empty;
[DatabaseGenerated(DatabaseGeneratedOption.Computed)]
public NpgsqlTsVector SearchVector { get; set; } = null!;
[DatabaseGenerated(DatabaseGeneratedOption.Computed)]
public string SearchText { get; set; } = string.Empty;
}
If the entity's columns aren't already snake_case-mapped to match the migration's identifiers, add
the usual HasColumnName(...) calls in OnModelCreating.
3. Query — RankedSearch + ToSearchPageAsync
public Task<PagedList<ArticleSearchDto>> SearchAsync(
string term, int pageNumber, int pageSize, CancellationToken ct)
{
var options = new SearchOptions { Config = "simple" };
return _db.Articles
.Where(a => a.Status == Status.Published) // service-owned filters go here
.RankedSearch(term, e => e.SearchVector, e => e.SearchText, options)
.ToSearchPageAsync(
r => new ArticleSearchDto(r.Item.Id, r.Item.MainTitle, r.Score),
r => r.Item.Id, // stable tiebreak: ordering is Score desc, then this key asc
pageNumber,
pageSize,
ct);
}
RankedSearch filters and scores but does not order or page. The filter is
@@ websearch_to_tsquery(...) OR a to_tsquery(..., '<token>:*') prefix disjunct OR a pg_trgm
word_similarity() fuzzy disjunct (length-independent, so a short query still matches a word buried in
a long body). The score is a three-tier CASE, not a flat blend:
| tier | filter | score | band |
|---|---|---|---|
| exact | @@ websearch_to_tsquery |
2.0 + ts_rank_cd(…, 32) |
[2.0, 3.0) |
| prefix | @@ to_tsquery('tok:*') |
1.0 + ts_rank_cd(…, 32) |
[1.0, 2.0) |
| fuzzy | word_similarity >= threshold |
word_similarity |
[0.0, 1.0) |
ts_rank_cd's normalization argument is DivideByItselfPlusOne (32) — rank/(rank+1), bounded below
1.0 — so the bands cannot overlap however many times a term matches one row. That transform is
strictly increasing, so ordering within a tier is unchanged. The exact arm is tested first because an
exact hit is also a prefix hit; swapped, every exact match would collapse into the prefix band.
Why the prefix tier exists (bg-gis#16): the structured half matches whole lexemes, so a reader
typing three characters of a seven-character place name matched nothing — websearch_to_tsquery on
վան does not match the lexeme վանք. The trigram disjunct could cover it, but it is gated by
FuzzyMinTermLength and is not index-accelerated in its word_similarity(...) form. The prefix
disjunct is index-backed by the same GIN index the structured half uses, and it is precise: napo
prefix-matches napoleon and not naples, where the trigram path returns both.
So exact / phrase hits rank above prefix hits, and both above fuzzy hits, with field weighting intact
(a plain GREATEST blend would
let an exact word saturate word_similarity to 1.0 and flatten that ordering). Apply any
service-specific filters before calling it, on the base IQueryable<T>. ToSearchPageAsync orders by
Score descending then the caller's tiebreak key (r => r.Item.Id above — required so tied scores
page deterministically), clamps pageNumber/pageSize, and projects to the platform PagedList<TDto>
shape. An empty or below-MinTermLength term returns an empty page, not an error.
Notes
unaccentfallback. Both search-support columns are accent-folded via theunaccentextension:TsVectorExpression/TriggerFunctionwrap each field inunaccent(coalesce(...))beforeto_tsvector, and the trigger does the same before lowercasing intosearch_text. If a target environment cannot installunaccent(e.g. a managed Postgres without superuser access), drop theCREATE EXTENSION IF NOT EXISTS unaccent;line and neutralize bothunaccent(...)calls in the generated trigger body (the one inside thesearch_vectorassignment and the one inside thesearch_textassignment) before running it — the trigger then falls back to plainlower()/no accent-folding on both columns.SearchTerm.Normalizestill strips diacritics from the query term in C#, so unaccented data keeps matching correctly either way; withoutunaccentin the DB, accented data only matches an unaccented query if the stored value happens to already be unaccented.- Maintenance MUST be a trigger, not a
GENERATED ... STOREDcolumn.unaccent()isSTABLE, notIMMUTABLE, and PostgreSQL only allowsIMMUTABLEfunctions in aGENERATED ALWAYS AS (...) STOREDcolumn expression (or a plain index expression). ABEFORE INSERT OR UPDATEtrigger has no such restriction, which is whySearchSql.TriggerFunctionbuilds aplpgsqltrigger body instead of a generated column — do not "simplify" the migration into a generated column, it will fail to create.
Tuning (SearchOptions)
| Property | Default | Meaning |
|---|---|---|
Config |
"simple" |
tsvector/tsquery text-search configuration (regconfig). The platform default is simple (no language stemmer) — multilingual content with no built-in Armenian stemmer. |
MinTermLength |
1 |
Minimum normalized-term length before a search runs at all. |
EnableFuzzy |
true |
Whether the pg_trgm word_similarity disjunct is OR-ed into the filter. Not a fallback gated on the structured match missing; its word_similarity is the score for fuzzy-only rows (structured @@ matches score in the higher 1.0 + ts_rank_cd tier). |
FuzzyMinTermLength |
3 |
Minimum normalized-term length before the fuzzy disjunct is added (e.g. GIS raises this to 7 to mirror the old fuzzy gate). Raising it does not cost in-word matching any more — that is the prefix disjunct's job, with its own floor. |
FuzzyThreshold |
0.3 |
Minimum pg_trgm word_similarity() score for a fuzzy match to count as a hit. Because word_similarity scores the best-matching word extent (not the whole string), 0.3 is a usable recall default even against long bodies. |
EnablePrefix |
true |
Whether the to_tsquery(config, '<token>:*') prefix disjunct is OR-ed into the filter. This is what makes a term match inside a word. Skipped automatically for a term using websearch operator syntax — a quoted phrase, a -negated word, a bare or — because rebuilding those from sanitised tokens would change what the caller asked for (napoleon -naples would turn an exclusion into a required term). |
PrefixMinTermLength |
2 |
Minimum normalized-term length before the prefix disjunct is added. Its own floor, independent of FuzzyMinTermLength: a one-character prefix matches most of a corpus, whereas the length at which guessing at a typo becomes worthwhile is a different, higher number. |
Choosing how widely to match — SearchMatchTier
EnablePrefix/EnableFuzzy are the mechanism; SearchMatchTier is the vocabulary a service should
expose and reason in. It names the widest tier a caller will accept, and the same three words label
what a row actually matched — so a request and a response can never describe matching differently.
var options = new SearchOptions { FuzzyMinTermLength = 7 };
options.ApplyQuality(quality); // Exact | Prefix | Fuzzy -> the two toggles
...
MatchType = r.Score >= SearchMatchTiers.ExactFloor ? SearchMatchTier.Exact
: r.Score >= SearchMatchTiers.PrefixFloor ? SearchMatchTier.Prefix
: SearchMatchTier.Fuzzy
Name the constants in the projection rather than writing 1.0/2.0: a const folds into the
expression tree and translates to a SQL CASE, while a method call over the projected row does not —
and a band restated in three service repositories drifts the first time the scoring is retuned.
The enum is ordered widest-last, so default(SearchMatchTier) is the strictest value: a forgotten
quality errs toward too few results rather than opening the search up.
Widening only ever adds rows. A row that matched exactly keeps its Exact label when fuzzy is
admitted alongside — the tiers are bands on one score, not alternative queries.
Build / test / pack
dotnet build Bg.Shared.slnx -c Release
dotnet test tests/Brainograph.Utilities.Search.Tests
dotnet pack src/Brainograph.Utilities.Search -c Release
| Product | Versions Compatible and additional computed target framework versions. |
|---|---|
| .NET | net10.0 is compatible. net10.0-android was computed. net10.0-browser was computed. net10.0-ios was computed. net10.0-maccatalyst was computed. net10.0-macos was computed. net10.0-tvos was computed. net10.0-windows was computed. |
-
net10.0
- Brainograph.Utilities (>= 6.20.0)
- Microsoft.AspNetCore.Authentication.JwtBearer (>= 10.0.5)
- Microsoft.EntityFrameworkCore (>= 10.0.5)
- Microsoft.Extensions.Configuration (>= 10.0.5)
- Microsoft.Extensions.Configuration.Abstractions (>= 10.0.5)
- Microsoft.Extensions.DependencyInjection (>= 10.0.5)
- Microsoft.Extensions.DependencyInjection.Abstractions (>= 10.0.5)
- Microsoft.Extensions.Http (>= 10.0.5)
- Microsoft.Extensions.Logging (>= 10.0.5)
- Microsoft.Extensions.Logging.Abstractions (>= 10.0.5)
- Microsoft.IdentityModel.Tokens (>= 8.16.0)
- Newtonsoft.Json (>= 13.0.4)
- Npgsql.EntityFrameworkCore.PostgreSQL (>= 10.0.2)
- System.IdentityModel.Tokens.Jwt (>= 8.16.0)
NuGet packages
This package is not used by any NuGet packages.
GitHub repositories
This package is not used by any popular GitHub repositories.
| Version | Downloads | Last Updated |
|---|