Rag.NET.Parsers.Pdf
1.0.0
dotnet add package Rag.NET.Parsers.Pdf --version 1.0.0
NuGet\Install-Package Rag.NET.Parsers.Pdf -Version 1.0.0
<PackageReference Include="Rag.NET.Parsers.Pdf" Version="1.0.0" />
<PackageVersion Include="Rag.NET.Parsers.Pdf" Version="1.0.0" />
<PackageReference Include="Rag.NET.Parsers.Pdf" />
paket add Rag.NET.Parsers.Pdf --version 1.0.0
#r "nuget: Rag.NET.Parsers.Pdf, 1.0.0"
#:package Rag.NET.Parsers.Pdf@1.0.0
#addin nuget:?package=Rag.NET.Parsers.Pdf&version=1.0.0
#tool nuget:?package=Rag.NET.Parsers.Pdf&version=1.0.0
Rag.NET.Parsers.Pdf
PDF parser for the Rag.NET ingestion pipeline: PdfPig-based text extraction with table detection, plus an opt-in OCR fallback for scanned pages.
Install
dotnet add package Rag.NET.Parsers.Pdf
Install alongside the core pipeline package (dotnet add package Rag.NET), which supplies
the AddRagNet(...) builder the parser registers into.
Setup
Inside your AddRagNet(...) builder callback:
using Rag.NET.Parsers.Pdf;
rag.AddPdfParser();
Example
Table extraction is on by default; OCR is opt-in because it needs an engine:
using Rag.NET.Parsers.Pdf;
rag.AddPdfParser(options =>
{
options.ExtractTables = true; // default: tables become Markdown in the chunk text
options.MinTableRows = 3; // default
options.UseOcrFallback = true; // pages under OcrMinCharacters go through OCR
options.OcrMinCharacters = 50; // default: the OCR trigger threshold
});
The built-in fallback uses Tesseract (compile-time opt-in). For managed document-level
OCR, chain UseAzureDocumentIntelligenceOcr from the
Rag.NET.Parsers.Pdf.AzureDocumentIntelligence package instead — AddPdfParser dispatches
to whichever IDocumentOcrEngine is registered.
Full guide
| Product | Versions Compatible and additional computed target framework versions. |
|---|---|
| .NET | net10.0 is compatible. net10.0-android was computed. net10.0-browser was computed. net10.0-ios was computed. net10.0-maccatalyst was computed. net10.0-macos was computed. net10.0-tvos was computed. net10.0-windows was computed. |
-
net10.0
- Microsoft.Extensions.Logging.Abstractions (>= 10.0.12)
- PdfPig (>= 0.1.16)
- Rag.NET.Abstractions (>= 1.0.0)
NuGet packages (1)
Showing the top 1 NuGet packages that depend on Rag.NET.Parsers.Pdf:
| Package | Downloads |
|---|---|
|
Rag.NET.Parsers.Pdf.AzureDocumentIntelligence
Azure Document Intelligence document-level OCR engine for Rag.NET's PDF parser |
GitHub repositories
This package is not used by any popular GitHub repositories.