MarkdownIndexer 1.2.0

dotnet add package MarkdownIndexer --version 1.2.0
                    
NuGet\Install-Package MarkdownIndexer -Version 1.2.0
                    
This command is intended to be used within the Package Manager Console in Visual Studio, as it uses the NuGet module's version of Install-Package.
<PackageReference Include="MarkdownIndexer" Version="1.2.0" />
                    
For projects that support PackageReference, copy this XML node into the project file to reference the package.
<PackageVersion Include="MarkdownIndexer" Version="1.2.0" />
                    
Directory.Packages.props
<PackageReference Include="MarkdownIndexer" />
                    
Project file
For projects that support Central Package Management (CPM), copy this XML node into the solution Directory.Packages.props file to version the package.
paket add MarkdownIndexer --version 1.2.0
                    
#r "nuget: MarkdownIndexer, 1.2.0"
                    
#r directive can be used in F# Interactive and Polyglot Notebooks. Copy this into the interactive tool or source code of the script to reference the package.
#:package MarkdownIndexer@1.2.0
                    
#:package directive can be used in C# file-based apps starting in .NET 10 preview 4. Copy this into a .cs file before any lines of code to reference the package.
#addin nuget:?package=MarkdownIndexer&version=1.2.0
                    
Install as a Cake Addin
#tool nuget:?package=MarkdownIndexer&version=1.2.0
                    
Install as a Cake Tool

MarkdownIndexer

NuGet

A C# library for parsing markdown into structured JSON trees with optional LLM enrichment. Zero NuGet dependencies — BCL only.

Designed to match the output of PageIndex's markdown pipeline.

Install

dotnet add package MarkdownIndexer

Or vendor the two source files directly — see Architecture below.

Two files you drop into any .NET project:

File Responsibility External deps
MarkdownIndexer.cs Parse markdown → tree + JSON None (BCL only)
MarkdownIndexer.Enrichment.cs LLM summaries + doc description None (BCL only)

All external concerns are injected as delegates:

  • Token counting → Func<string, int>
  • LLM calls → Func<string, CancellationToken, Task<string>>

You wire these up from whatever libraries your project already uses.

Pipeline

MARKDOWN STRING
      │
      ▼
┌─────────────────────────────────────────┐
│ EXTRACT HEADERS                         │
│   Regex ^(#{1,6})\s+(.+)$,              │
│   skips ``` code blocks                 │
│   → [{title, line_num, level}, ...]    │
└──────────────────┬──────────────────────┘
                   │
                   ▼
┌─────────────────────────────────────────┐
│ EXTRACT TEXT PER SECTION                │
│   Slice from header line to next        │
│   header (or EOF)                       │
│   → [{title, line_num, level, text}]   │
└──────────────────┬──────────────────────┘
                   │
          ┌────────┴────────┐
          │ if thinning     │
          ▼                 │
┌──────────────────┐        │
│ THIN TREE        │        │
│ Bottom-up token  │        │
│ count. Merge     │        │
│ small subtrees   │        │
│ into parents.    │        │
└────────┬─────────┘        │
         │                  │
         └────────┬─────────┘
                   │
                   ▼
┌─────────────────────────────────────────┐
│ STACK-BASED TREE BUILDING               │
│   Push/pop by header level.             │
│   Children of nearest lower-level       │
│   ancestor.                             │
└──────────────────┬──────────────────────┘
                   │
                   ▼
┌─────────────────────────────────────────┐
│ DEPTH-FIRST NODE IDS + CLEANUP          │
│   "0001", "0002", ...                   │
│   Remove empty nodes arrays             │
└──────────────────┬──────────────────────┘
                   │
          ┌────────┴────────┐
          │ if enrichment   │
          ▼                 │
┌──────────────────┐        │
│ LLM SUMMARIZATION│        │
│ Per node (parallel│       │
│ + throttled).    │        │
│ Leaf → summary   │        │
│ Parent → prefix_ │        │
│          summary │        │
│                  │        │
│ DOC DESCRIPTION  │        │
│ One-sentence     │        │
│ (single LLM call)│       │
│                  │        │
│ STRIP TEXT       │        │
│ (optional)       │        │
└────────┬─────────┘        │
         │                  │
         └────────┬─────────┘
                   │
                   ▼
┌─────────────────────────────────────────┐
│ JSON OUTPUT                             │
│   { doc_name, line_count,               │
│     structure: [{title, node_id,         │
│       line_num, summary?,              │
│       prefix_summary?, text?,           │
│       nodes?}, ...] }                   │
└─────────────────────────────────────────┘

Quick Start

Add the two .cs files to your project. Install your choice of tokenizer and AI packages:

dotnet add package Microsoft.ML.Tokenizers
dotnet add package Microsoft.Extensions.AI.OpenAI

Parse only (no LLM)

using MarkdownIndexer;
using Microsoft.ML.Tokenizers;

var md = File.ReadAllText("doc.md");
var tokenizer = TiktokenTokenizer.CreateForModel("gpt-4o");

var result = MarkdownIndexer.Index(md, "doc", new IndexerOptions
{
    AddNodeId = true,
    AddNodeText = true,
    Thinning = new(Threshold: 5000)       // optional
}, text => tokenizer.CountTokens(text));

string json = MarkdownIndexer.ToJson(result);

Parse + LLM enrichment

var chatClient = new OpenAIClient(apiKey)
    .AsChatClient("gpt-4o-mini");

Func<string, CancellationToken, Task<string>> llm =
    async (prompt, ct) =>
    {
        var response = await chatClient.GetResponseAsync(prompt, cancellationToken: ct);
        return response.Text;
    };

var result = await MarkdownIndexer.IndexWithEnrichmentAsync(
    md, "doc",
    text => tokenizer.CountTokens(text),
    llm,
    enrichOptions: new EnrichmentOptions
    {
        SummaryTokenThreshold = 200,
        MaxConcurrency = 5,
        AddDocDescription = true
    });

string json = MarkdownIndexer.ToJson(result);

Using any LLM provider

The delegate pattern works with any client library:

// Anthropic
Func<string, CancellationToken, Task<string>> llm = async (prompt, ct) =>
{
    var msg = await anthropic.Messages.GetResponseAsync(new()
    {
        Model = "claude-sonnet-4-20250514",
        Messages = [new() { Content = prompt }]
    }, ct);
    return msg.Content.First().Text;
};

// Ollama (local)
Func<string, CancellationToken, Task<string>> llm = async (prompt, ct) =>
{
    var response = await httpClient.PostAsJsonAsync(
        "http://localhost:11434/api/generate",
        new { model = "llama3", prompt, stream = false }, ct);
    var json = await response.Content.ReadFromJsonAsync<JsonElement>(ct);
    return json.GetProperty("response").GetString()!;
};

Output Shape

{
  "doc_name": "my-document",
  "line_count": 42,
  "structure_quality": "good",
  "doc_description": "A document covering markdown parsing techniques.",
  "structure": [
    {
      "title": "My Document",
      "node_id": "0001",
      "line_num": 1,
      "end_char": 450,
      "prefix_summary": "The document introduces...",
      "nodes": [
        {
          "title": "Section A",
          "node_id": "0002",
          "line_num": 5,
          "start_char": 150,
          "end_char": 350,
          "summary": "Section A details the core algorithm."
        }
      ]
    }
  ]
}

Key conventions (matching Python PageIndex):

  • summary — leaf nodes (no children)
  • prefix_summary — parent nodes (signpost before descending into children)
  • text — excluded when AddNodeText is false or KeepText is false (after summarization)
  • start_char / end_char — character offsets (0-indexed, half-open) of the section's full text in the original markdown string. Omitted from JSON when zero (e.g., root node at offset 0).
  • EndLineNum — in-memory convenience property (not serialized to JSON) giving the last line number of the section
  • Empty nodes arrays are omitted from JSON
  • structure_quality — automated assessment of the TOC, always present (see below)

Structure Quality

Every IndexResult includes a structure_quality field that classifies the document's heading hierarchy into one of three levels:

Quality Meaning When
good Well-structured TOC — suitable for tree-index retrieval as-is Good coverage, decent nesting depth, branching, and descriptive headings
degraded Has issues that may reduce retrieval accuracy At least 2 of: too many vague headings, no branching, low heading coverage
flat No usable structure — treat as a plain document No headings, only one heading level, or most text comes before any heading

The assessment runs automatically during Index() and checks four signals:

  • Coverage — fraction of the document that falls under the heading hierarchy (not "orphan" text before the first heading)
  • Depth — maximum heading nesting level (e.g., ### = level 3)
  • Branching — whether any heading level has 2+ siblings (a chain of single-child headings has no branching)
  • Vague headings — generic titles like "Introduction", "Overview", "Section 1", or very short titles (< 4 chars) at top levels

Use this to triage documents before spending LLM tokens on enrichment or retrieval:

var result = MarkdownIndexer.Index(md, "doc");

if (result.StructureQuality == DocumentStructureQuality.Flat)
{
    // Skip enrichment — treat as a plain document
    Console.WriteLine("No meaningful structure, using fallback strategy.");
    return;
}

// Proceed with enrichment or tree-index retrieval
await MarkdownIndexer.EnrichAsync(result, tokenCounter, llm);

You can also call AssessStructure directly if you already have a tree:

var quality = MarkdownIndexer.AssessStructure(tree, fullMarkdownText);

Cross-Library Integration

When using start_char / end_char alongside other text-processing libraries (e.g., ApproximateSpanMatching), note that MarkdownIndexer does not apply Unicode NFC normalization to the input text. If the other library does (e.g., ApproximateSpanMatching normalizes before tokenizing), character offsets may diverge for text containing combining characters. For most markdown (typically ASCII), this is not a concern. To be safe, apply NFC normalization upstream in your pipeline:

var normalized = markdownContent.Normalize(System.Text.NormalizationForm.FormC);
var result = MarkdownIndexer.Index(normalized, docName);

Options

IndexerOptions

Option Default
AddNodeId true Assign depth-first "0001", "0002", ...
Thinning null new(Threshold: 5000) to merge small subtrees
AddNodeText false Include text in output

EnrichmentOptions

Option Default
SummaryTokenThreshold 200 Skip LLM for sections below this token count
MaxConcurrency 0 (unlimited) Max parallel LLM calls
AddDocDescription false Generate one-sentence document description
KeepText false Retain text after summarization

Customization

Both files are designed for copy-paste vendoring. Edit prompts directly:

// In MarkdownIndexer.Enrichment.cs
private const string SummaryPromptTemplate = """
    You are given...
    Partial Document Text: {0}
    Directly return the description...
    """;

Or customize the pipeline by calling individual methods (GenerateSummariesAsync, GenerateDocDescriptionAsync, StripText) instead of using the convenience orchestrators.

When to Use This

Tree-indexing is reasoning-based retrieval — the LLM reads a structured table of contents and navigates to the right sections, like a human expert scanning a document.

🟢 Strong fit

Document type Why tree-index beats chunking Typical size
Documentation sites (mdBook, Docusaurus) Already have #/## hierarchy; chunks destroy cross-section context 50–500 sections
Technical specs / RFCs Requirements cross-reference each other; need section-level context 20–200 pages
Internal wikis / knowledge bases Naturally hierarchical; vector search returns similar-but-wrong pages 100–1000+ pages
Research notes / lab notebooks Methods → results → conclusions form a reasoning chain 5–50 pages
Regulatory filings, financial reports (PDF→MD converted) Quarterly data lives in specific sections; chunking mixes Q1 and Q3 100–300 pages
Legal contracts, policy documents Meaning depends on which section/clause; definitions matter 20–200 pages

🟡 Moderate fit

API docs with clear hierarchy, blog archives, textbooks — benefits from structure but vector search may be sufficient for simple factoid queries.

🔴 Poor fit

  • FAQ / Q&A collections — flat by nature, vector search is faster
  • Chat logs, support tickets — no hierarchy
  • Product reviews, social media — each item is independent
  • Documents shorter than ~5K tokens — just feed to the LLM directly

Size sweet spot

┌────────────┐   ┌───────────────┐   ┌─────────────────────────┐
│ < 5K tokens │   │ 5K – 100K     │   │ 100K+ tokens             │
├─────────────┤   ├───────────────┤   ├─────────────────────────┤
│ No RAG      │   │ Tree indexing  │   │ Tree indexing is          │
│ needed —    │   │ SHINES         │   │ THE ONLY option           │
│ just feed    │   │                │   │                            │
│ to the LLM   │   │ Vector RAG     │   │ Vector RAG fails: chunks   │
│             │   │ also works     │   │ lose cross-section context │
└────────────┘   └───────────────┘   └─────────────────────────┘

vs. Vector RAG

Vector RAG Tree-Index RAG
How Chunk → embed → similarity search Parse → build TOC → LLM navigates structure
Accuracy ~30-50% on complex docs (FinanceBench) ~98.7% (PageIndex on FinanceBench)
Cost Cheap (no extra LLM calls) More expensive (LLM reads summaries, navigates)
Latency Fast (milliseconds) Slower (sequential LLM reasoning steps)
Traceability "Which chunks were similar" "Reasoning path through the document tree"
Best for Factoid queries, conceptual similarity Multi-hop reasoning, context-dependent answers

Core insight: similarity ≠ relevance. A paragraph can be semantically similar to a query but completely wrong for the answer. Tree-indexing trades speed for precision — worth it when getting the right answer is non-negotiable.

How Retrieval Works

This library builds the index. The retrieval is a separate concern — it's how a downstream LLM agent uses the tree to answer questions.

The key design: prefix_summary vs summary

Root (prefix_summary: "FY2024 report covering revenue, risk, guidance")
├── Section 3 (prefix_summary: "Revenue breakdown by segment")     ← signpost
│   ├── 3.1 (summary: "Enterprise $2.1B, +15% YoY")               ← leaf
│   └── 3.2 (summary: "Consumer $1.8B, flat")                     ← leaf
├── Section 4 (prefix_summary: "Risk factors and forward guidance") ← signpost
│   └── ...
└── Section 5 (summary: "Conclusion: strong year, cautious outlook") ← leaf

prefix_summary acts like a signpost — "this branch is about X, descend or skip?"
summary acts like a destination — "this leaf contains Y, grab it or pass?"

The LLM reads the tree and decides which nodes to fetch text for.

Strategy 1: One-shot tree navigation

Feed the entire tree (without text) to the LLM. The tree is compact — just titles + summaries — even a 300-section document is ~5-15K tokens. The LLM selects relevant node_ids in one pass.

┌──────────────┐     ┌───────────────────┐     ┌──────────────────┐
│ Tree (no     │────▶│ LLM reads all     │────▶│ Fetch text for   │
│ text, just   │     │ summaries, picks  │     │ selected nodes   │
│ summaries)   │     │ relevant node IDs │     │ → generate answer│
└──────────────┘     └───────────────────┘     └──────────────────┘

Strategy 2: Agentic iterative navigation

An LLM agent with three tools navigates step by step, never loading the whole tree or document:

get_document_structure()  →  "I see 'Financial Results' section. Let me check it."
                           ↓
get_page_content("15-18") →  "Q3 revenue $4.2B. Need MD&A for Q2 comparison."
                           ↓
get_page_content("8-10")  →  "Q2 was $3.8B. Q3 grew 10.5%." → answer

This handles multi-hop questions ("compare Q3 to Q2") where the answer spans multiple sections.

Complete retrieval example

// ── Build phase (offline) ──
// Generate tree WITH summaries but WITHOUT text (compact for context)
var tocTree = MarkdownIndexer.Index(md, "report", new IndexerOptions
{
    AddNodeId = true,
    AddNodeText = false   // text would bloat the context
});

// Generate summaries (or use cached)
await MarkdownIndexer.EnrichAsync(tocTree, tokenCounter, llm, new EnrichmentOptions
{
    SummaryTokenThreshold = 200,
    MaxConcurrency = 5
});
string tocJson = MarkdownIndexer.ToJson(tocTree);

// ── Query phase (online) ──
string query = "What was Q3 enterprise revenue?";

// Step 1: LLM reads TOC, selects relevant nodes
string selectionPrompt = $"""
    Question: {query}
    Document tree structure:
    {tocJson}

    Return JSON with node_ids of relevant sections:
    {{"node_ids": ["0003", "0007"]}}
    """;

string selection = await chatClient.GetResponseAsync(selectionPrompt);
var ids = JsonSerializer.Deserialize<SelectionResult>(selection).NodeIds;

// Step 2: Fetch full text only for selected nodes
var fullTree = MarkdownIndexer.Index(md, "report", new IndexerOptions
{
    AddNodeId = true,
    AddNodeText = true   // now with text
});
var allNodes = MarkdownIndexer.FlattenTree(fullTree.Structure);
var context = string.Join("\n\n",
    allNodes.Where(n => ids.Contains(n.NodeId)).Select(n => n.Text));

// Step 3: Generate answer
string answerPrompt = $"Context:\n{context}\n\nQuestion: {query}";
string answer = await chatClient.GetResponseAsync(answerPrompt);

Tip: Cache the tree-with-summaries. Regenerate only when the source document changes. Summaries are the expensive part (LLM calls); the parse itself is deterministic and fast.

Customization

Both files are designed for copy-paste vendoring. Edit prompts directly:

// In MarkdownIndexer.Enrichment.cs
private const string SummaryPromptTemplate = """
    You are given...
    Partial Document Text: {0}
    Directly return the description...
    """;

Or customize the pipeline by calling individual methods (GenerateSummariesAsync, GenerateDocDescriptionAsync, StripText) instead of using the convenience orchestrators.

Publishing

Releases are automated via GitHub Actions with NuGet Trusted Publishing — no long-lived API keys.

Release process

git tag v1.0.0
git push origin v1.0.0

Pushing a v* tag triggers the workflow which builds, packs, and pushes to NuGet.org.

One-time setup

  1. NuGet.org: Create a Trusted Publishing policy:

    • Owner: ypyl
    • Repository: MarkdownIndexer
    • Workflow: publish.yml
  2. GitHub: Add a repository secret NUGET_USERNAME with your nuget.org username.

Product Compatible and additional computed target framework versions.
.NET net10.0 is compatible.  net10.0-android was computed.  net10.0-browser was computed.  net10.0-ios was computed.  net10.0-maccatalyst was computed.  net10.0-macos was computed.  net10.0-tvos was computed.  net10.0-windows was computed. 
Compatible target framework(s)
Included target framework(s) (in package)
Learn more about Target Frameworks and .NET Standard.
  • net10.0

    • No dependencies.

NuGet packages

This package is not used by any NuGet packages.

GitHub repositories

This package is not used by any popular GitHub repositories.

Version Downloads Last Updated
1.2.0 125 7/16/2026
1.1.0 122 7/9/2026
1.0.0 112 7/9/2026