Femur.Html.Parser
0.0.31
dotnet add package Femur.Html.Parser --version 0.0.31
NuGet\Install-Package Femur.Html.Parser -Version 0.0.31
<PackageReference Include="Femur.Html.Parser" Version="0.0.31" />
<PackageVersion Include="Femur.Html.Parser" Version="0.0.31" />
<PackageReference Include="Femur.Html.Parser" />
paket add Femur.Html.Parser --version 0.0.31
#r "nuget: Femur.Html.Parser, 0.0.31"
#:package Femur.Html.Parser@0.0.31
#addin nuget:?package=Femur.Html.Parser&version=0.0.31
#tool nuget:?package=Femur.Html.Parser&version=0.0.31
HtmlParser
A streaming HTML 2.0 parser that reads from a Stream and builds an Abstract Syntax Tree (AST) of nodes.
Overview
HtmlParser extends StreamParser<DocumentNode> and implements a single-pass parser that processes HTML content character-by-character, building a tree structure of nodes as it reads.
Parsing Strategy
The parser uses a streaming approach with the following characteristics:
- Sliding Buffer: Reads stream in chunks (default 4KB) to handle large files efficiently
- Absolute Position Tracking: Maintains position across buffer boundaries for accurate location tracking
- Element Stack: Uses a stack to match opening/closing tags and maintain parent-child relationships
- Single Pass: Processes tokens in one pass without a separate tokenization phase
- Location Tracking: Every node includes
SourceLocationinformation (start position and length)
Architecture
Base Class: StreamParser
HtmlParser extends StreamParser<DocumentNode>, which provides:
- Buffer Management: Handles reading chunks from the stream
- Position Tracking: Manages buffer position and absolute byte position
- Template Method Pattern: Defines the parsing algorithm:
CreateDocument()- Creates the root document nodeInitializeParsing()- Sets up parsing stateProcessCharacter()- Processes each character (main parsing logic)Cleanup()- Cleans up resources
Parsing Flow
1. CreateDocument() → Creates DocumentNode
2. InitializeParsing() → Sets up stacks and state
3. Main Loop (for each character):
- ReadMore() → Ensures buffer has data
- ProcessCharacter() → Routes to appropriate handler
4. Cleanup() → Returns buffer to pool
Core Data Structures
State Variables
_document: Reference to the root document node_currentParent: Current container node where new children are added_elementStack: Stack ofElementNodeobjects for matching opening/closing tags_isInsideScriptOrStyle: Flag to handle script/style tags specially
Void Elements
HTML void elements (like <br>, <img>, <input>) cannot have children and don't need closing tags. These are tracked in a static HashSet and are never pushed onto the element stack.
Parsing Flow
Character Processing
The main entry point is ProcessCharacter(), which routes characters to appropriate handlers:
if (ch == '<')
ProcessTag() // Tag processing
else
ProcessTextContent() // Text content
Tag Processing
When encountering <, the parser examines the next character to determine tag type:
<!→ Special tag (comment, CDATA, DOCTYPE)</→ Closing tag- Otherwise → Opening tag
Opening Tags (ProcessOpeningTag)
- Peek ahead to check for SVG tags (special handling)
- Parse tag name and attributes (
ParseOpeningTag) - Create
ElementNodeand add to_currentParent.Children - If not void and not self-closing, push onto
_elementStack - Update
_currentParentto the new element - Track script/style tags to set
_isInsideScriptOrStyleflag
Closing Tags (ProcessClosingTag)
- Parse closing tag name (
ParseClosingTag) - Pop elements from
_elementStackuntil finding a matching tag name - Handle mismatched tags gracefully (similar to browser behavior)
- Update
_currentParentto the matched element's parent - Clear
_isInsideScriptOrStyleflag if exiting script/style tag
Special Tags (ProcessSpecialTag)
Handles three types of special tags:
- Comments: ``
- CDATA:
<![CDATA[...]]> - DOCTYPE:
<!DOCTYPE ...>
These don't affect the element hierarchy, so _currentParent remains unchanged.
Text Content Processing
ProcessTextContent() handles everything between tags:
- Reads all characters until encountering
<(start of next tag) - Inside script/style tags: Preserves ALL content including whitespace
- Outside script/style tags: Filters out pure whitespace text nodes
- Creates
TextNodewith location tracking
Script/Style Handling: Inside <script> and <style> tags, the parser:
- Preserves all whitespace and content exactly as written
- Only stops at
</script>or</style>closing tags - Treats other
<characters as literal text
Attribute Parsing
Attributes are parsed with support for:
- Quoted values:
attr="value"orattr='value' - Unquoted values:
attr=value(until whitespace or>) - Boolean attributes:
attr(no value, stored as empty string) - Self-closing tags:
<tag />or<tag/>
Special Features
SVG Handling
When encountering an <svg> tag, the parser:
- Rewinds to the opening
<svg>tag position - Creates a
SvgSubStreamwrapper that reads until</svg> - Delegates parsing to
XmlParserfor the SVG block - Adds the resulting
XmlElementNodeto the HTML AST - SVG elements don't go on the element stack (foreign elements)
Limitation: SVG blocks must fit within a single buffer (typically 4KB). Blocks spanning multiple buffers will throw an exception.
Location Tracking
Every node includes a Location property (SourceLocation) that tracks:
- Start Position: Absolute byte position in the stream
- Length: Number of bytes the node spans
This enables:
- Error reporting with exact positions
- Source mapping
- Round-trip editing
Example Flow
For HTML like:
<div>
<p>Hello</p>
<img src="test.jpg" />
</div>
The parsing flow:
<div>: CreateElementNode, push onto stack, set as_currentParent- Text " ": Filtered out (whitespace-only)
<p>: CreateElementNode, push onto stack, set as_currentParent- Text "Hello": Create
TextNode, add to<p>children </p>: Pop<p>from stack, restore_currentParentto<div>- Text " ": Filtered out
<img ... />: CreateElementNode(void element), add to<div>, don't push stack</div>: Pop<div>from stack, restore_currentParentto document
Error Handling
The parser handles malformed HTML gracefully:
- Mismatched closing tags: Pops up the stack until finding a match (browser-like behavior)
- Unclosed tags: Elements remain on stack (can be detected after parsing)
- Invalid characters: Handled according to HTML spec (generally ignored or treated as text)
Performance Considerations
- Streaming: Processes large files without loading entire content into memory
- Buffer Pooling: Uses
ArrayPool<byte>for efficient buffer management - Single Pass: No backtracking or multiple passes required
- Minimal Allocations: Reuses
StringBuilderand buffers where possible
Usage
// From stream
using var stream = new FileStream("page.html", FileMode.Open);
var parser = new HtmlParser(stream);
var document = parser.Parse();
// From string
var document = HtmlParser.Parse("<html>...</html>");
// From bytes
var bytes = Encoding.UTF8.GetBytes("<html>...</html>");
var document = HtmlParser.Parse(bytes);
Key Methods
ProcessCharacter(): Main character routing logicProcessTag(): Routes to opening/closing/special tag handlersProcessOpeningTag(): Handles opening tags and updates stackProcessClosingTag(): Handles closing tags and updates stackProcessTextContent(): Handles text between tagsParseOpeningTag(): Parses tag name, attributes, self-closing indicatorParseSpecialTag(): Parses comments, CDATA, DOCTYPEParseSvgAsXml(): Special handling for SVG blocks
| Product | Versions Compatible and additional computed target framework versions. |
|---|---|
| .NET | net5.0 was computed. net5.0-windows was computed. net6.0 was computed. net6.0-android was computed. net6.0-ios was computed. net6.0-maccatalyst was computed. net6.0-macos was computed. net6.0-tvos was computed. net6.0-windows was computed. net7.0 was computed. net7.0-android was computed. net7.0-ios was computed. net7.0-maccatalyst was computed. net7.0-macos was computed. net7.0-tvos was computed. net7.0-windows was computed. net8.0 is compatible. net8.0-android was computed. net8.0-browser was computed. net8.0-ios was computed. net8.0-maccatalyst was computed. net8.0-macos was computed. net8.0-tvos was computed. net8.0-windows was computed. net9.0 is compatible. net9.0-android was computed. net9.0-browser was computed. net9.0-ios was computed. net9.0-maccatalyst was computed. net9.0-macos was computed. net9.0-tvos was computed. net9.0-windows was computed. net10.0 is compatible. net10.0-android was computed. net10.0-browser was computed. net10.0-ios was computed. net10.0-maccatalyst was computed. net10.0-macos was computed. net10.0-tvos was computed. net10.0-windows was computed. |
| .NET Core | netcoreapp2.0 was computed. netcoreapp2.1 was computed. netcoreapp2.2 was computed. netcoreapp3.0 was computed. netcoreapp3.1 was computed. |
| .NET Standard | netstandard2.0 is compatible. netstandard2.1 was computed. |
| .NET Framework | net461 was computed. net462 was computed. net463 was computed. net47 was computed. net471 was computed. net472 was computed. net48 was computed. net481 was computed. |
| MonoAndroid | monoandroid was computed. |
| MonoMac | monomac was computed. |
| MonoTouch | monotouch was computed. |
| Tizen | tizen40 was computed. tizen60 was computed. |
| Xamarin.iOS | xamarinios was computed. |
| Xamarin.Mac | xamarinmac was computed. |
| Xamarin.TVOS | xamarintvos was computed. |
| Xamarin.WatchOS | xamarinwatchos was computed. |
-
.NETStandard 2.0
- Femur.Markup.Abstractions (>= 0.0.31)
- Femur.Parsing (>= 0.0.31)
- Femur.Xml.Abstractions (>= 0.0.31)
- Femur.Xml.Parser (>= 0.0.31)
- System.Memory (>= 4.6.3)
-
net10.0
- Femur.Markup.Abstractions (>= 0.0.31)
- Femur.Parsing (>= 0.0.31)
- Femur.Xml.Abstractions (>= 0.0.31)
- Femur.Xml.Parser (>= 0.0.31)
-
net8.0
- Femur.Markup.Abstractions (>= 0.0.31)
- Femur.Parsing (>= 0.0.31)
- Femur.Xml.Abstractions (>= 0.0.31)
- Femur.Xml.Parser (>= 0.0.31)
-
net9.0
- Femur.Markup.Abstractions (>= 0.0.31)
- Femur.Parsing (>= 0.0.31)
- Femur.Xml.Abstractions (>= 0.0.31)
- Femur.Xml.Parser (>= 0.0.31)
NuGet packages
This package is not used by any NuGet packages.
GitHub repositories
This package is not used by any popular GitHub repositories.