DocSharpParsers 0.0.1
dotnet add package DocSharpParsers --version 0.0.1
NuGet\Install-Package DocSharpParsers -Version 0.0.1
<PackageReference Include="DocSharpParsers" Version="0.0.1" />
<PackageVersion Include="DocSharpParsers" Version="0.0.1" />
<PackageReference Include="DocSharpParsers" />
paket add DocSharpParsers --version 0.0.1
#r "nuget: DocSharpParsers, 0.0.1"
#:package DocSharpParsers@0.0.1
#addin nuget:?package=DocSharpParsers&version=0.0.1
#tool nuget:?package=DocSharpParsers&version=0.0.1
DocSharpParsers
A fork of DocSharp that adds a small, unified façade over its binary Office parsers.
Open legacy binary Microsoft Office files — .doc, .ppt, .xls — and:
- browse their internal OLE / compound-file structure,
- extract text and images,
- convert them to modern OOXML (.docx / .pptx / .xlsx).
Everything is lazy: opening only reads the compound-file directory, the format model is parsed on first text/convert access, and image bytes are decoded only when requested.
Quick start
using DocSharpParsers;
using var doc = OfficeDocument.Open("presentation.ppt");
Console.WriteLine(doc.Format); // Word / PowerPoint / Excel
// 1. Browse the OLE compound-file structure (lazy).
foreach (var e in doc.Root.DescendantsAndSelf())
Console.WriteLine($"{e.Kind,-7} {e.Path} ({e.Size} bytes)");
// Read a raw stream's bytes:
byte[] raw = doc.EnumerateStreams().First().ReadBytes();
// 2. Extract text.
string text = doc.GetText();
// 3. Extract images (bytes decoded on demand).
foreach (var img in doc.GetImages())
File.WriteAllBytes($"img.{img.Extension}", img.GetBytes());
// 4. Convert to OOXML.
doc.ConvertToFile("presentation.pptx");
// or: doc.ConvertTo(someStream);
Opening from a stream
using var fs = File.OpenRead("report.xls");
using var doc = OfficeDocument.Open(fs, fs.Name); // doc closes fs on dispose
// Keep ownership of the stream (it is NOT closed with the document):
using var doc2 = OfficeDocument.Open(fs, fs.Name, leaveOpen: true);
The stream must stay open and seekable for the lifetime of the OfficeDocument
(its contents are read on demand). Pass the file name — or just its extension — to
aid format detection when the content check is inconclusive.
Notes
- Format is auto-detected from the streams inside the OLE container (falling back to the file extension).
- Images are extracted for all three formats (Word, PowerPoint and Excel).
OfficeImage.Format/Extensionreport the encoding (PNG, JPEG, GIF, TIFF, BMP, EMF, WMF, PICT); metafiles are decompressed and headerless DIB blips are wrapped into standalone BMPs.
License
MIT — see LICENSE. This project is a fork of DocSharp (© Manfredi Marceca), also MIT-licensed.
| Product | Versions Compatible and additional computed target framework versions. |
|---|---|
| .NET | net10.0 is compatible. net10.0-android was computed. net10.0-browser was computed. net10.0-ios was computed. net10.0-maccatalyst was computed. net10.0-macos was computed. net10.0-tvos was computed. net10.0-windows was computed. |
-
net10.0
- No dependencies.
NuGet packages
This package is not used by any NuGet packages.
GitHub repositories
This package is not used by any popular GitHub repositories.
| Version | Downloads | Last Updated |
|---|---|---|
| 0.0.1 | 134 | 7/10/2026 |