Net4x.Html2Xhtml
1.2.0.26249
dotnet add package Net4x.Html2Xhtml --version 1.2.0.26249
NuGet\Install-Package Net4x.Html2Xhtml -Version 1.2.0.26249
<PackageReference Include="Net4x.Html2Xhtml" Version="1.2.0.26249" />
<PackageVersion Include="Net4x.Html2Xhtml" Version="1.2.0.26249" />
<PackageReference Include="Net4x.Html2Xhtml" />
paket add Net4x.Html2Xhtml --version 1.2.0.26249
#r "nuget: Net4x.Html2Xhtml, 1.2.0.26249"
#:package Net4x.Html2Xhtml@1.2.0.26249
#addin nuget:?package=Net4x.Html2Xhtml&version=1.2.0.26249
#tool nuget:?package=Net4x.Html2Xhtml&version=1.2.0.26249
Net4x.Html2Xhtml
A .NET wrapper around html2xhtml, the stream filter by Jesús Arias Fisteus that turns HTML into XHTML and repairs, on the way, most of what real world HTML gets wrong: unclosed elements, mis-nested elements, unquoted attribute values, uppercase tag names, bare ampersands and void elements without a closing slash.
The package carries html2xhtml.exe and libiconv-2.dll with it and copies both next to your
application, so there is nothing to install on the machine that runs your code.
Install
dotnet add package Net4x.Html2Xhtml
Target frameworks: net35, net40, net6.0-windows, net8.0-windows, net10.0-windows.
Converting a document
Html2Xhtml.RunAsFilter writes to the standard input of the converter and hands you back its
standard output and its standard error:
using System.IO;
using Html2Xhtml.Extensions;
StreamReader standardError;
var output = Html2Xhtml.Html2Xhtml.RunAsFilter(
stdin => stdin.Write("<HTML><body><p class=lead>First<br><p>Second & last</body></HTML>"),
out standardError);
string xhtml = output.ReadToEnd();
string problems = standardError.ReadToEnd(); // empty when nothing had to be reported
Read the output before the standard error: the converter writes its diagnostics while it works, and draining the error stream first leaves it waiting for its output to be consumed.
The result is a complete XHTML document:
<?xml version="1.0" encoding="utf-8"?>
<!DOCTYPE html
PUBLIC "-//W3C//DTD XHTML 1.0 Transitional//EN"
"http://www.w3.org/TR/xhtml1/DTD/xhtml1-transitional.dtd" >
<html xmlns="http://www.w3.org/1999/xhtml">
<head>
<title>
****
</title>
</head>
<body>
<p class="lead">
First<br />
</p>
<p>
Second & last
</p>
</body>
</html>
Straight into an XDocument
ReadToXDocument reads the output and parses it. It converts the named entities of XHTML into
numeric ones first, because a plain XML parser knows only five of them, and it removes the XHTML
namespace so that elements can be reached by their plain name:
using Html2Xhtml.Extensions;
StreamReader standardError;
var document = Html2Xhtml.Html2Xhtml
.RunAsFilter(stdin => stdin.Write(html), out standardError)
.ReadToXDocument();
string title = document.Root.Element("head").Element("title").Value.Trim();
Pass keepXhtmlNamespace: true to keep the namespace, and address elements with
XNamespace.Get("http://www.w3.org/1999/xhtml") as usual. "...".ToXDocument() does the same for
a string you already have.
Options
Every option of RunAsFilter is optional and named after what it does.
| Option | Default | What it does |
|---|---|---|
outputDocType |
DocType.Auto |
The doctype to write. Auto picks XHTML 1.0 Transitional, or Frameset for a frameset document. |
e |
false |
Propagate input the converter cannot adapt to the output doctype, at the cost of a possibly non-valid document. |
inputEncodingName |
"utf-8" |
Encoding of the HTML you write. null leaves the converter to detect it. |
outputEncodingName |
"utf-8" |
Encoding of the XHTML you get back. null uses the encoding of the input. |
lineLength |
int.MaxValue |
Characters per line. Values below 40 are ignored; null leaves the converter to choose. |
tabLength |
0 |
Indentation per level, 0 to 16 characters. 0 writes no indentation at all. |
preserveWhitespaceInComments |
true |
Keep spaces, tabs and ends of line inside comments instead of rearranging them. |
noProtectCData |
false |
Enclose CDATA in script and style as <![CDATA[, instead of the browser-safe //<![CDATA[. |
compactBlockElements |
false |
Write <p>x</p> on one line instead of breaking the line after the start tag. |
emptyElementTagsAlways |
false |
Write <div /> for any empty element, not only for the ones the DTD declares empty. |
compactEmptyElementTags |
false |
Write <br/> instead of the recommended <br />. |
dosEndOfLine |
false |
Write CRLF instead of LF. |
Doctypes
DocType.Auto, Xhtml_1_0_Transitional, Xhtml_1_0_Strict, Xhtml_1_0_Frameset,
Xhtml_Basic_1_0, Xhtml_Basic_1_1, Xhtml_1_1, Xhtml_MobileProfile_1_0, Xhtml_Print_1_0.
Content the requested doctype does not allow is dropped, and what was dropped is reported on the
standard error. A strict document, for example, keeps the text of a font element but not the
element itself.
Encodings
Both .NET and the libiconv the converter is built on have to know an encoding under the same name. Ask before you convert:
Html2Xhtml.Html2Xhtml.SupportedEncodings.Contains("iso-8859-15"); // true
foreach (var name in Html2Xhtml.Html2Xhtml.SupportedEncodings.Names) { /* ... */ }
Ends of line
The output has the end of line style you asked for on every platform: LF, or CRLF with
dosEndOfLine: true. The converter writes through the standard streams of the operating system,
which rewrite ends of line on their own on Windows, so the wrapper normalizes them back to what was
requested.
This is done on the bytes of the output, which is only safe in an encoding where a line feed and a carriage return are single bytes — UTF-8, US-ASCII, the ISO 8859 family, the Windows code pages. Ask for a UTF-16 or UTF-32 output and the bytes are handed over exactly as the converter wrote them; note that the converter does not write those encodings through the standard output of Windows unharmed, so prefer UTF-8 and re-encode the result yourself.
Finding the converter
html2xhtml.exe is looked for in this order, and the first one that has it wins:
- the directory the library was deployed to,
- the base directory of the application, and its private search path (
binunder ASP.NET), - the current directory,
- whatever
PATHresolves, which is what happens when none of the above has it.
So the converter is found whatever the current directory of the host is — a test runner, a Windows
service and a web application all have one of their own. When it is found nowhere,
RunAsFilter throws Html2Xhtml.Exceptions.CommandNotFoundException, with the failure to start it
as the inner exception. Anything else that goes wrong — an exception out of your own input action,
for instance — surfaces unchanged.
Command.ResolveCommandPath("html2xhtml") answers where it would be started from, which is the
first thing worth printing when a deployment does not work.
Running many conversions
Each call starts a converter of its own, and calls from several threads do not interfere with each
other. Every converter that is still running when the application ends is stopped;
set Html2Xhtml.Diagnostics.Command.KillDangling to false to leave them alone.
Building and testing
dotnet build Html2Xhtml/Html2Xhtml.csproj
dotnet test Html2Xhtml.Tests/Html2Xhtml.Tests.csproj
The assembly is strong named, and code coverage instrumentation breaks that signature, so a coverage run needs an unsigned rebuild:
dotnet build Html2Xhtml.Tests/Html2Xhtml.Tests.csproj -t:Rebuild -p:SignAssembly=false
dotnet test Html2Xhtml.Tests/Html2Xhtml.Tests.csproj --no-build --collect:"XPlat Code Coverage"
Licence and credits
Distributed under the GNU General Public License, version 2, the licence of the html2xhtml program this package carries.
Copyright (c) Cetin Sert 2010; Jesús Arias Fisteus 2001-2010; Rebeca Díaz Redondo, Ana Fernández Vilas 2001. Company: CORSIS. The original program and its documentation are at http://www.it.uc3m.es/jaf/html2xhtml/.
| Product | Versions Compatible and additional computed target framework versions. |
|---|---|
| .NET | net6.0-windows7.0 is compatible. net7.0-windows was computed. net8.0-windows was computed. net8.0-windows7.0 is compatible. net9.0-windows was computed. net10.0-windows was computed. net10.0-windows7.0 is compatible. |
| .NET Framework | net35 is compatible. net40 is compatible. net403 was computed. net45 was computed. net451 was computed. net452 was computed. net46 was computed. net461 was computed. net462 was computed. net463 was computed. net47 was computed. net471 was computed. net472 was computed. net48 was computed. net481 was computed. |
-
.NETFramework 3.5
- Net4x.AsyncBridge (>= 1.5.0.26243)
-
.NETFramework 4.0
- Net4x.AsyncBridge (>= 1.5.0.26243)
-
net10.0-windows7.0
- Net4x.AsyncBridge (>= 1.5.0.26243)
-
net6.0-windows7.0
- Net4x.AsyncBridge (>= 1.5.0.26243)
-
net8.0-windows7.0
- Net4x.AsyncBridge (>= 1.5.0.26243)
NuGet packages
This package is not used by any NuGet packages.
GitHub repositories
This package is not used by any popular GitHub repositories.
| Version | Downloads | Last Updated |
|---|---|---|
| 1.2.0.26249 | 105 | 9/6/2026 |
| 1.2.0 | 277 | 4/4/2025 |
| 1.1.4 | 269 | 8/27/2023 |