Omne.TikaContentDetector 3.3.1.25

dotnet add package Omne.TikaContentDetector --version 3.3.1.25
                    
NuGet\Install-Package Omne.TikaContentDetector -Version 3.3.1.25
                    
This command is intended to be used within the Package Manager Console in Visual Studio, as it uses the NuGet module's version of Install-Package.
<PackageReference Include="Omne.TikaContentDetector" Version="3.3.1.25" />
                    
For projects that support PackageReference, copy this XML node into the project file to reference the package.
<PackageVersion Include="Omne.TikaContentDetector" Version="3.3.1.25" />
                    
Directory.Packages.props
<PackageReference Include="Omne.TikaContentDetector" />
                    
Project file
For projects that support Central Package Management (CPM), copy this XML node into the solution Directory.Packages.props file to version the package.
paket add Omne.TikaContentDetector --version 3.3.1.25
                    
#r "nuget: Omne.TikaContentDetector, 3.3.1.25"
                    
#r directive can be used in F# Interactive and Polyglot Notebooks. Copy this into the interactive tool or source code of the script to reference the package.
#:package Omne.TikaContentDetector@3.3.1.25
                    
#:package directive can be used in C# file-based apps starting in .NET 10 preview 4. Copy this into a .cs file before any lines of code to reference the package.
#addin nuget:?package=Omne.TikaContentDetector&version=3.3.1.25
                    
Install as a Cake Addin
#tool nuget:?package=Omne.TikaContentDetector&version=3.3.1.25
                    
Install as a Cake Tool

TikaContentDetector for .NET

A .NET Core library that provides content type detection functionality similar to Apache Tika's DefaultDetector. It analyzes file content and metadata to determine MIME types, suitable file extensions, and other content metadata.

Features

  • ๐Ÿ” Content-based detection using magic number patterns
  • ๐Ÿ“ Filename-based detection using glob patterns
  • ๐Ÿ“ฆ Container format awareness (DOCX, JAR, EPUB, etc.)
  • ๐Ÿš€ Stream-based processing for memory efficiency
  • ๐ŸŽฏ High accuracy using Apache Tika's MIME database
  • โšก Fast detection with optimized pattern matching
  • ๐Ÿงช Comprehensive test coverage with real file samples

Installation

dotnet add package Omne.TikaContentDetector

Quick Start

using Omne.TikaContentDetector;

// Create detector instance
var detector = new ContentDetector();
detector.LoadMimeTypes();

// Detect from file
var result = detector.DetectFile("document.pdf");

Console.WriteLine($"MIME Type: {result.MimeType}");
Console.WriteLine($"Extensions: {string.Join(", ", result.SuggestedExtensions)}");
Console.WriteLine($"Confidence: {result.Confidence:P0}");
// Output:
// MIME Type: application/pdf
// Extensions: pdf
// Confidence: 95%

Usage Examples

Basic Detection

// Detect from file path
var result = detector.DetectFile("document.pdf");

// Detect from stream
using var stream = File.OpenRead("mystery.file");
var streamResult = detector.Detect(stream);

// Detect with filename hint for better accuracy
var resultWithHint = detector.Detect(stream, "document.docx");

// Detect from byte array
byte[] fileBytes = File.ReadAllBytes("image.jpg");
var byteResult = detector.Detect(fileBytes, "image.jpg");

Working with Different File Types

// Image detection
var imageResult = detector.Detect(imageStream, "photo.jpg");
// Returns: image/jpeg

// Office document detection (detects specific format, not just ZIP)
var docResult = detector.Detect(docStream, "report.docx");
// Returns: application/vnd.openxmlformats-officedocument.wordprocessingml.document

// Archive detection (including TAR)
var zipResult = detector.Detect(zipStream, "files.zip");
// Returns: application/zip

var tarResult = detector.Detect(tarStream, "archive.tar");
// Returns: application/x-tar

Handling Edge Cases

// Empty file
var emptyResult = detector.Detect(emptyStream);
// Returns: application/octet-stream

// Misnamed file (JPEG with .png extension)
var mismatchResult = detector.Detect(jpegStream, "image.png");
// Returns: image/jpeg (content takes precedence)

// No extension
var noExtResult = detector.Detect(stream, "datafile");
// Uses content detection only

// Large files are handled efficiently
using var largeStream = File.OpenRead("huge-video.mp4");
var result = detector.Detect(largeStream); // Only reads what's needed

Metadata Lookup

You can also query MIME type metadata directly:

// Get MIME type repository
var repository = detector.GetMimeTypesRepository();

// Look up by MIME type or alias
var mimeType = repository.GetMimeTypeByTypeOrAlias("application/pdf");
// or by alias
var mimeType2 = repository.GetMimeTypeByTypeOrAlias("application/x-pdf");

// Find all MIME types for an extension
var mimeTypes = repository.GetMimeTypesByExtension("xml");
foreach (var mt in mimeTypes)
{
    Console.WriteLine($"{mt.Type}: {mt.Description}");
}

CLI Usage

The package includes a command-line interface that can be used in two ways:

For Development
# Use the provided shell script (Linux/macOS)
./tika-detector document.pdf

# Or on Windows
tika-detector.cmd document.pdf

# Or run directly with dotnet
dotnet run --project src/Omne.TikaContentDetector.Cli -- document.pdf
For Installation (Production)
# Install the CLI tool globally
dotnet tool install -g Omne.TikaContentDetector.Cli

# Then use it anywhere
tika-detector document.pdf

# Detect multiple files with JSON output
tika-detector detect *.jpg --output json --pretty

# Include verbose metadata
tika-detector detect archive.zip --verbose

# Text output format
tika-detector detect image.png --output text

# Look up MIME type information by type or alias
tika-detector mime-type application/pdf
tika-detector mime-type application/x-pdf  # Works with aliases too

# Find MIME types by file extension
tika-detector extension pdf
tika-detector extension .pdf  # Leading dot is optional

# Get human-readable output
tika-detector mime-type text/plain --output text
tika-detector extension jpg --output text

Design & Architecture

Overview

TikaContentDetector uses the same tika-mimetypes.xml database as Apache Tika, ensuring compatibility and accuracy. The library implements a multi-strategy detection approach:

  1. Magic Number Detection: Examines file content for characteristic byte patterns
  2. Container Inspection: Looks inside ZIP/TAR files to identify specific formats
  3. Filename Matching: Uses file extensions as hints or fallbacks
  4. Priority Resolution: Resolves conflicts using Tika's priority system

Key Components

  • ContentDetector: Main entry point, orchestrates detection strategies
  • MagicMatcher: Performs binary pattern matching with offset and mask support
  • GlobMatcher: Handles filename pattern matching
  • ContainerInspector: Identifies specific ZIP-based formats (DOCX, JAR, etc.)
  • MimeTypesRepository: Manages the MIME type database

Detection Pipeline

Input Stream โ†’ Magic Detection โ†’ Container Inspection โ†’ Filename Matching โ†’ Result
                     โ†“                    โ†“                    โ†“
                Priority-based conflict resolution โ† โ† โ† โ† โ† โ† โ†

Performance Considerations

  • Stream-based processing (never loads entire file)
  • Configurable buffer sizes for large files
  • Lazy loading of MIME database
  • Efficient pattern matching algorithms

Supported Formats

The library supports all formats defined in Apache Tika's MIME database, including:

Documents

  • PDF, DOCX, XLSX, PPTX (Office Open XML)
  • DOC, XLS, PPT (Legacy Office)
  • ODT, ODS, ODP (OpenDocument)
  • RTF, TXT, HTML, XML, Markdown

Images

  • JPEG, PNG, GIF, BMP
  • WebP, TIFF, ICO, SVG
  • HEIF, AVIF, RAW formats

Archives

  • ZIP, 7Z, RAR, TAR
  • GZIP, BZIP2, XZ
  • JAR, WAR (Java archives)
  • EPUB (e-books)

Media

  • MP3, MP4, AVI, MKV
  • WAV, OGG, FLAC
  • WebM, MOV, WMV

Data

  • JSON, CSV, XML
  • SQLite databases
  • YAML, TOML

More Examples

The Examples/ directory contains comprehensive examples:

  • BasicUsage - Simple file detection examples
  • StreamDetection - Working with various stream types
  • BatchProcessing - Processing multiple files efficiently
  • WebApiExample - ASP.NET Core Web API integration

Run any example:

cd Examples/BasicUsage
dotnet run

Building from Source

# Clone the repository
git clone https://github.com/omnesoft/tika-content-detector-dotnet.git
cd tika-content-detector-dotnet

# Build the solution
dotnet build

# Run tests
dotnet test

Testing

The library includes comprehensive tests with real file samples:

# Run all tests
dotnet test

# Run specific test categories
dotnet test --filter Category=Images
dotnet test --filter Category=Documents
dotnet test --filter Category=Containers

Contributing

Contributions are welcome! Please read our Contributing Guide for details on our code of conduct and the process for submitting pull requests.

Development Setup

  1. Fork the repository
  2. Create a feature branch (git checkout -b feature/amazing-feature)
  3. Commit your changes (git commit -m 'Add amazing feature')
  4. Push to the branch (git push origin feature/amazing-feature)
  5. Open a Pull Request

License

This project is licensed under the Apache License 2.0 - see the LICENSE file for details.

The project includes the tika-mimetypes.xml file from Apache Tika, which is also licensed under the Apache License 2.0.

Acknowledgments

  • Apache Tika - This library uses Tika's MIME type database and follows its detection algorithms
  • FreeDesktop.org - For the shared MIME info specification
  • Contributors - Thanks to all who have contributed to this project

Sources and References

FAQ

Q: How accurate is the detection?
A: Very accurate for common formats. The library uses the same detection rules as Apache Tika, which is widely used and well-tested.

Q: Can it detect encrypted or password-protected files?
A: It can identify the format (e.g., encrypted ZIP), but cannot read the contents without the password.

Q: Does it support custom MIME types?
A: Yes, you can provide a custom MIME types XML file following the Tika format.

Q: How does it handle large files?
A: Efficiently! It only reads the necessary bytes (typically the first 8-64KB) for detection.

Q: Is it thread-safe?
A: Yes, the ContentDetector instance can be safely shared across threads.

Product Compatible and additional computed target framework versions.
.NET net10.0 is compatible.  net10.0-android was computed.  net10.0-browser was computed.  net10.0-ios was computed.  net10.0-maccatalyst was computed.  net10.0-macos was computed.  net10.0-tvos was computed.  net10.0-windows was computed. 
Compatible target framework(s)
Included target framework(s) (in package)
Learn more about Target Frameworks and .NET Standard.
  • net10.0

    • No dependencies.

NuGet packages

This package is not used by any NuGet packages.

GitHub repositories

This package is not used by any popular GitHub repositories.

Version Downloads Last Updated
3.3.1.25 762 7/28/2026

MIME database synced to Apache Tika 3.3.1; added model/step (ISO-10303-21) detection for .step / .stp files.