Femur.Xml.Parser
0.0.31
dotnet add package Femur.Xml.Parser --version 0.0.31
NuGet\Install-Package Femur.Xml.Parser -Version 0.0.31
<PackageReference Include="Femur.Xml.Parser" Version="0.0.31" />
<PackageVersion Include="Femur.Xml.Parser" Version="0.0.31" />
<PackageReference Include="Femur.Xml.Parser" />
paket add Femur.Xml.Parser --version 0.0.31
#r "nuget: Femur.Xml.Parser, 0.0.31"
#:package Femur.Xml.Parser@0.0.31
#addin nuget:?package=Femur.Xml.Parser&version=0.0.31
#tool nuget:?package=Femur.Xml.Parser&version=0.0.31
XmlParser
A streaming XML parser that reads from a Stream and builds an Abstract Syntax Tree (AST) of nodes.
Overview
XmlParser extends StreamParser<XmlDocumentNode> and implements a single-pass parser that processes XML content character-by-character, building a tree structure of nodes. XML parsing is stricter than HTML - all tags must be closed, attributes must be quoted, and tag names are case-sensitive.
Parsing Strategy
The parser uses a streaming approach with the following characteristics:
- Sliding Buffer: Reads stream in chunks (default 4KB) to handle large files efficiently
- Absolute Position Tracking: Maintains position across buffer boundaries for accurate location tracking
- Element Stack: Uses a stack to match opening/closing tags with case-sensitive matching
- Single Pass: Processes tokens in one pass without a separate tokenization phase
- XML-Specific Features: Handles processing instructions, namespaces, and strict tag matching
Architecture
Base Class: StreamParser
XmlParser extends StreamParser<XmlDocumentNode>, which provides:
- Buffer Management: Handles reading chunks from the stream
- Position Tracking: Manages buffer position and absolute byte position
- Template Method Pattern: Defines the parsing algorithm:
CreateDocument()- Creates the root XML document nodeInitializeParsing()- Sets up parsing stateProcessCharacter()- Processes each character (main parsing logic)Cleanup()- Cleans up resources
Parsing Flow
1. CreateDocument() → Creates XmlDocumentNode
2. InitializeParsing() → Sets up stacks and state
3. Main Loop (for each character):
- ReadMore() → Ensures buffer has data
- ProcessCharacter() → Routes to appropriate handler
4. Cleanup() → Returns buffer to pool
Core Data Structures
State Variables
_document: Reference to the root XML document node_currentParent: Current container node where new children are added_elementStack: Stack ofXmlElementNodeobjects for matching opening/closing tags
XML-Specific Features
Unlike HTML, XML has:
- No void elements: All tags must be closed (either
<tag></tag>or<tag />) - Case-sensitive: Tag names must match exactly (case-sensitive)
- Quoted attributes: All attribute values must be quoted (single or double quotes)
- Namespaces: Supports namespace prefixes (
prefix:localname) - Processing Instructions: Supports
<?target data?>syntax
Parsing Flow
Character Processing
The main entry point is ProcessCharacter(), which routes characters to appropriate handlers:
if (ch == '<')
ProcessTag() // Tag processing
else
ProcessTextContent() // Text content
Tag Processing
When encountering <, the parser examines the next character to determine tag type:
<!→ Special tag (comment, CDATA)<?→ Processing instruction (<?xml version="1.0"?>)</→ Closing tag- Otherwise → Opening tag
Opening Tags (ProcessOpeningTag)
- Parse tag name and attributes (
ParseOpeningTag) - Extract namespace prefix if present (
prefix:localname) - Create
XmlElementNodeand add to_currentParent.Children - If not self-closing, push onto
_elementStack - Update
_currentParentto the new element - Handle namespace declarations (
xmlnsandxmlns:prefix)
Closing Tags (ProcessClosingTag)
- Parse closing tag name (case-sensitive)
- Pop elements from
_elementStackuntil finding an exact match (case-sensitive) - Update
_currentParentto the matched element's parent - If no match found and stack is empty, restore to document root
Key Difference from HTML: XML requires exact case-sensitive matching. <Tag> and </tag> would be considered mismatched.
Processing Instructions (ProcessProcessingInstruction)
Handles XML processing instructions like <?xml version="1.0"?>:
- Parse target name (e.g., "xml")
- Read content until
?> - Create
ProcessingInstructionNode - If target is "xml", store in
_document.XmlDeclaration
Special Tags (ProcessSpecialTag)
Handles two types of special tags:
- Comments: ``
- CDATA:
<![CDATA[...]]>
These don't affect the element hierarchy, so _currentParent remains unchanged.
Text Content Processing
ProcessTextContent() handles everything between tags:
- Reads all characters until encountering
<(start of next tag) - Filters out pure whitespace text nodes (XML preserves whitespace, but parser filters for efficiency)
- Creates
TextNodewith location tracking
Note: XML preserves whitespace by default, but this parser filters whitespace-only text nodes for efficiency. Full whitespace preservation can be enabled if needed.
Attribute Parsing
Attributes are parsed with strict XML rules:
- Must be quoted:
attr="value"orattr='value'(required) - Escaped quotes: Supports
\"and\'escaping within attribute values - Namespace declarations:
xmlns="uri"→ Sets default namespace URIxmlns:prefix="uri"→ Sets namespace URI for prefix
Key Difference from HTML: XML requires all attribute values to be quoted. Unquoted attributes are not valid XML.
Namespace Handling
XML supports namespaces via prefixes:
- Tag names:
prefix:localname(e.g.,svg:circle) - Namespace prefix: Extracted and stored in
XmlElementNode.NamespacePrefix - Namespace URI: Extracted from
xmlnsorxmlns:prefixattributes
The parser extracts namespace information but doesn't fully resolve namespace URIs (that would require maintaining a namespace context stack).
Special Features
XML Declaration
The XML declaration (<?xml version="1.0"?>) is:
- Parsed as a
ProcessingInstructionNode - Stored in
XmlDocumentNode.XmlDeclarationfor easy access - Preserved in the document structure
Location Tracking
Every node includes a Location property (SourceLocation) that tracks:
- Start Position: Absolute byte position in the stream
- Length: Number of bytes the node spans
This enables:
- Error reporting with exact positions
- Source mapping
- Round-trip editing
Example Flow
For XML like:
<?xml version="1.0"?>
<root>
<child attr="value">Text</child>
<self-closing />
</root>
The parsing flow:
<?xml version="1.0"?>: CreateProcessingInstructionNode, store in_document.XmlDeclaration<root>: CreateXmlElementNode, push onto stack, set as_currentParent- Text " ": Filtered out (whitespace-only)
<child attr="value">: CreateXmlElementNodewith attribute, push onto stack, set as_currentParent- Text "Text": Create
TextNode, add to<child>children </child>: Pop<child>from stack (case-sensitive match), restore_currentParentto<root>- Text " ": Filtered out
<self-closing />: CreateXmlElementNode(self-closing), add to<root>, don't push stack</root>: Pop<root>from stack, restore_currentParentto document
Error Handling
The parser handles XML with strict rules:
- Case-sensitive matching:
<Tag>and</tag>are considered mismatched - Unclosed tags: Elements remain on stack (can be detected after parsing)
- Unquoted attributes: Parsed but may not be valid XML
- Invalid characters: Generally treated as text content
Performance Considerations
- Streaming: Processes large files without loading entire content into memory
- Buffer Pooling: Uses
ArrayPool<byte>for efficient buffer management - Single Pass: No backtracking or multiple passes required
- Minimal Allocations: Reuses
StringBuilderand buffers where possible
Usage
// From stream
using var stream = new FileStream("data.xml", FileMode.Open);
var parser = new XmlParser(stream);
var document = parser.Parse();
// From string
var document = XmlParser.Parse("<root>...</root>");
// From bytes
var bytes = Encoding.UTF8.GetBytes("<root>...</root>");
var document = XmlParser.Parse(bytes);
Key Methods
ProcessCharacter(): Main character routing logicProcessTag(): Routes to opening/closing/special/processing instruction handlersProcessOpeningTag(): Handles opening tags and updates stackProcessClosingTag(): Handles closing tags with case-sensitive matchingProcessProcessingInstruction(): Handles<?target data?>syntaxProcessSpecialTag(): Parses comments and CDATAProcessTextContent(): Handles text between tagsParseOpeningTag(): Parses tag name, attributes, namespaces, self-closing indicator
Differences from HTML Parser
| Feature | HTML | XML |
|---|---|---|
| Case sensitivity | Case-insensitive | Case-sensitive |
| Void elements | Yes (<br>, <img>) |
No (all must be closed) |
| Attribute quotes | Optional | Required |
| Unquoted attributes | Allowed | Not valid |
| Self-closing | <tag /> or <tag/> |
<tag /> only |
| Processing instructions | No | Yes (<?target?> |
| Namespaces | No | Yes (prefix:name) |
| Product | Versions Compatible and additional computed target framework versions. |
|---|---|
| .NET | net5.0 was computed. net5.0-windows was computed. net6.0 was computed. net6.0-android was computed. net6.0-ios was computed. net6.0-maccatalyst was computed. net6.0-macos was computed. net6.0-tvos was computed. net6.0-windows was computed. net7.0 was computed. net7.0-android was computed. net7.0-ios was computed. net7.0-maccatalyst was computed. net7.0-macos was computed. net7.0-tvos was computed. net7.0-windows was computed. net8.0 is compatible. net8.0-android was computed. net8.0-browser was computed. net8.0-ios was computed. net8.0-maccatalyst was computed. net8.0-macos was computed. net8.0-tvos was computed. net8.0-windows was computed. net9.0 is compatible. net9.0-android was computed. net9.0-browser was computed. net9.0-ios was computed. net9.0-maccatalyst was computed. net9.0-macos was computed. net9.0-tvos was computed. net9.0-windows was computed. net10.0 is compatible. net10.0-android was computed. net10.0-browser was computed. net10.0-ios was computed. net10.0-maccatalyst was computed. net10.0-macos was computed. net10.0-tvos was computed. net10.0-windows was computed. |
| .NET Core | netcoreapp2.0 was computed. netcoreapp2.1 was computed. netcoreapp2.2 was computed. netcoreapp3.0 was computed. netcoreapp3.1 was computed. |
| .NET Standard | netstandard2.0 is compatible. netstandard2.1 was computed. |
| .NET Framework | net461 was computed. net462 was computed. net463 was computed. net47 was computed. net471 was computed. net472 was computed. net48 was computed. net481 was computed. |
| MonoAndroid | monoandroid was computed. |
| MonoMac | monomac was computed. |
| MonoTouch | monotouch was computed. |
| Tizen | tizen40 was computed. tizen60 was computed. |
| Xamarin.iOS | xamarinios was computed. |
| Xamarin.Mac | xamarinmac was computed. |
| Xamarin.TVOS | xamarintvos was computed. |
| Xamarin.WatchOS | xamarinwatchos was computed. |
-
.NETStandard 2.0
- Femur.Markup.Abstractions (>= 0.0.31)
- Femur.Parsing (>= 0.0.31)
- Femur.Xml.Abstractions (>= 0.0.31)
- System.Memory (>= 4.6.3)
-
net10.0
- Femur.Markup.Abstractions (>= 0.0.31)
- Femur.Parsing (>= 0.0.31)
- Femur.Xml.Abstractions (>= 0.0.31)
-
net8.0
- Femur.Markup.Abstractions (>= 0.0.31)
- Femur.Parsing (>= 0.0.31)
- Femur.Xml.Abstractions (>= 0.0.31)
-
net9.0
- Femur.Markup.Abstractions (>= 0.0.31)
- Femur.Parsing (>= 0.0.31)
- Femur.Xml.Abstractions (>= 0.0.31)
NuGet packages (1)
Showing the top 1 NuGet packages that depend on Femur.Xml.Parser:
| Package | Downloads |
|---|---|
|
Femur.Html.Parser
Package Description |
GitHub repositories
This package is not used by any popular GitHub repositories.