Bbieniek.Uax29
1.0.3
dotnet add package Bbieniek.Uax29 --version 1.0.3
NuGet\Install-Package Bbieniek.Uax29 -Version 1.0.3
<PackageReference Include="Bbieniek.Uax29" Version="1.0.3" />
<PackageVersion Include="Bbieniek.Uax29" Version="1.0.3" />
<PackageReference Include="Bbieniek.Uax29" />
paket add Bbieniek.Uax29 --version 1.0.3
#r "nuget: Bbieniek.Uax29, 1.0.3"
#:package Bbieniek.Uax29@1.0.3
#addin nuget:?package=Bbieniek.Uax29&version=1.0.3
#tool nuget:?package=Bbieniek.Uax29&version=1.0.3
Bbieniek.Uax29
A managed .NET word tokenizer implementing Unicode UAX #29 word boundary segmentation (Unicode 15.0). Splits text into words, numbers, punctuation, and whitespace tokens โ correctly handling multilingual text, emoji, contractions, and numeric formatting.
No native ICU required. Zero dependencies on netstandard2.0. SQL CLR compatible.
Why use this?
Splitting on spaces or regex gives inconsistent results with real-world text. Unicode UAX #29 is the standard used by ICU, Lucene, quanteda, and most NLP tooling. This package gives you the same tokenization in pure managed .NET:
"don't"stays as one token (apostrophe is MidLetter)"1,000,000.50"stays as one token (comma/period are MidNum)"self-aware"stays as one token (hyphen preservation)- Emoji ZWJ sequences like
๐ฉโโค๏ธโ๐โ๐จstay together - Works with Latin, Greek, Cyrillic, Hebrew, Katakana, CJK, and all Unicode scripts
Installation
dotnet add package Bbieniek.Uax29
Quick start
using Bbieniek.Uax29;
// Tokenize into strings
var tokens = WordBreakTokenizer.TokenizeToStrings("Hello, world! Price: $1,000.50");
// โ ["Hello", ",", " ", "world", "!", " ", "Price", ":", " ", "$", "1,000.50"]
// Tokenize into spans with metadata
var spans = WordBreakTokenizer.Tokenize("self-aware robot");
foreach (var span in spans)
{
var text = "self-aware robot".Substring(span.Start, span.Length);
Console.WriteLine($"{text} (IsWord: {span.IsWord})");
}
// โ self-aware (IsWord: True)
// โ (IsWord: False)
// โ robot (IsWord: True)
Zero-allocation API (net8.0+)
On .NET 8 and later, use the zero-allocation enumerator for high-throughput scenarios:
foreach (var token in WordBreakTokenizer.EnumerateWords("hello, world".AsSpan()))
{
// token.Span is ReadOnlySpan<char> โ no heap allocation
// token.IsWord, token.Start, token.Length also available
}
Quanteda/ICU-compatible mode
By default, colon is treated as MidLetter per the Unicode spec ("key:value" โ one token). Use WordBreakOptions.Quanteda to match ICU/quanteda behavior where colon splits words:
var tokens = WordBreakTokenizer.TokenizeToStrings("key:value", WordBreakOptions.Quanteda);
// โ ["key", ":", "value"]
API
| Method | Returns | Description |
|---|---|---|
Tokenize(string) |
List<TokenSpan> |
Word/separator spans with positions and IsWord flag |
Tokenize(string, WordBreakOptions) |
List<TokenSpan> |
Same, with custom options |
TokenizeToStrings(string) |
List<string> |
Token strings (convenience) |
TokenizeToStrings(string, WordBreakOptions) |
List<string> |
Same, with custom options |
EnumerateWords(ReadOnlySpan<char>) |
WordTokenEnumerator |
Zero-alloc enumerator (net8.0+) |
EnumerateWords(ReadOnlySpan<char>, WordBreakOptions) |
WordTokenEnumerator |
Same, with custom options |
TokenSpan
| Property | Type | Description |
|---|---|---|
Start |
int |
Start index in the original string |
Length |
int |
Number of characters in this token |
IsWord |
bool |
true for words/numbers, false for separators/punctuation |
WordBreakOptions
| Preset | Behavior |
|---|---|
WordBreakOptions.Default |
Strict Unicode UAX #29 (colon is MidLetter) |
WordBreakOptions.Quanteda |
ICU/quanteda-compatible (colon splits words) |
new WordBreakOptions(chars) |
Custom MidLetter exclusions |
UAX #29 rules implemented
| Rule | Behavior | Example |
|---|---|---|
| WB3 | Don't break within CRLF | \r\n stays together |
| WB3c | Emoji ZWJ sequences | ๐ฉโ๐ฉโ๐งโ๐ง stays together |
| WB3d | Group horizontal whitespace | "a b" โ ["a", " ", "b"] |
| WB4 | Ignore Extend/Format/ZWJ | Combining marks attach to base |
| WB5 | Don't break between letters | "hello" โ one token |
| WB6/7 | MidLetter keeps words | "e.g" โ one token |
| WB7a-c | Hebrew letter rules | Hebrew + quote combinations |
| WB8-10 | Numeric sequences | "12345" โ one token |
| WB11/12 | MidNum keeps numbers | "1,000,000.50" โ one token |
| WB13 | Katakana sequences | Adjacent katakana stay together |
| WB13a/b | ExtendNumLet (underscore) | "hello_world" โ one token |
| keep_hyphens | Infix hyphens preserved | "self-aware" โ one token |
Compatibility
Targets netstandard2.0 and net8.0:
- .NET 8, 9, 10+
- .NET 6, 7
- .NET Framework 4.6.1+
- SQL CLR (SAFE mode)
Performance
- Zero-allocation enumerator (net8.0+):
ref structviaEnumerateWords()โ no heap allocation per token - Bitwise property matching:
[Flags]enum with single-op combined checks - ASCII fast path: pre-computed lookup table for characters 0-127
- Latin-1 fast path: direct classification for U+00C0-U+024F
- Inlined hot paths:
AggressiveInliningon frequently called predicates
Unicode conformance
Validated against the official Unicode 15.0 WordBreakTest suite (1823 test cases), plus quanteda 4.3.1 and ICU4N verification.
Versioning
Package versions are generated from Git tags using MinVer. Release by pushing a tag like v1.2.3.
License
| Product | Versions Compatible and additional computed target framework versions. |
|---|---|
| .NET | net5.0 was computed. net5.0-windows was computed. net6.0 was computed. net6.0-android was computed. net6.0-ios was computed. net6.0-maccatalyst was computed. net6.0-macos was computed. net6.0-tvos was computed. net6.0-windows was computed. net7.0 was computed. net7.0-android was computed. net7.0-ios was computed. net7.0-maccatalyst was computed. net7.0-macos was computed. net7.0-tvos was computed. net7.0-windows was computed. net8.0 is compatible. net8.0-android was computed. net8.0-browser was computed. net8.0-ios was computed. net8.0-maccatalyst was computed. net8.0-macos was computed. net8.0-tvos was computed. net8.0-windows was computed. net9.0 was computed. net9.0-android was computed. net9.0-browser was computed. net9.0-ios was computed. net9.0-maccatalyst was computed. net9.0-macos was computed. net9.0-tvos was computed. net9.0-windows was computed. net10.0 was computed. net10.0-android was computed. net10.0-browser was computed. net10.0-ios was computed. net10.0-maccatalyst was computed. net10.0-macos was computed. net10.0-tvos was computed. net10.0-windows was computed. |
| .NET Core | netcoreapp2.0 was computed. netcoreapp2.1 was computed. netcoreapp2.2 was computed. netcoreapp3.0 was computed. netcoreapp3.1 was computed. |
| .NET Standard | netstandard2.0 is compatible. netstandard2.1 was computed. |
| .NET Framework | net461 was computed. net462 was computed. net463 was computed. net47 was computed. net471 was computed. net472 was computed. net48 was computed. net481 was computed. |
| MonoAndroid | monoandroid was computed. |
| MonoMac | monomac was computed. |
| MonoTouch | monotouch was computed. |
| Tizen | tizen40 was computed. tizen60 was computed. |
| Xamarin.iOS | xamarinios was computed. |
| Xamarin.Mac | xamarinmac was computed. |
| Xamarin.TVOS | xamarintvos was computed. |
| Xamarin.WatchOS | xamarinwatchos was computed. |
-
.NETStandard 2.0
- No dependencies.
-
net8.0
- No dependencies.
NuGet packages
This package is not used by any NuGet packages.
GitHub repositories
This package is not used by any popular GitHub repositories.