Omne.TikaContentDetector
3.3.1.25
dotnet add package Omne.TikaContentDetector --version 3.3.1.25
NuGet\Install-Package Omne.TikaContentDetector -Version 3.3.1.25
<PackageReference Include="Omne.TikaContentDetector" Version="3.3.1.25" />
<PackageVersion Include="Omne.TikaContentDetector" Version="3.3.1.25" />
<PackageReference Include="Omne.TikaContentDetector" />
paket add Omne.TikaContentDetector --version 3.3.1.25
#r "nuget: Omne.TikaContentDetector, 3.3.1.25"
#:package Omne.TikaContentDetector@3.3.1.25
#addin nuget:?package=Omne.TikaContentDetector&version=3.3.1.25
#tool nuget:?package=Omne.TikaContentDetector&version=3.3.1.25
TikaContentDetector for .NET
A .NET Core library that provides content type detection functionality similar to Apache Tika's DefaultDetector. It analyzes file content and metadata to determine MIME types, suitable file extensions, and other content metadata.
Features
- ๐ Content-based detection using magic number patterns
- ๐ Filename-based detection using glob patterns
- ๐ฆ Container format awareness (DOCX, JAR, EPUB, etc.)
- ๐ Stream-based processing for memory efficiency
- ๐ฏ High accuracy using Apache Tika's MIME database
- โก Fast detection with optimized pattern matching
- ๐งช Comprehensive test coverage with real file samples
Installation
dotnet add package Omne.TikaContentDetector
Quick Start
using Omne.TikaContentDetector;
// Create detector instance
var detector = new ContentDetector();
detector.LoadMimeTypes();
// Detect from file
var result = detector.DetectFile("document.pdf");
Console.WriteLine($"MIME Type: {result.MimeType}");
Console.WriteLine($"Extensions: {string.Join(", ", result.SuggestedExtensions)}");
Console.WriteLine($"Confidence: {result.Confidence:P0}");
// Output:
// MIME Type: application/pdf
// Extensions: pdf
// Confidence: 95%
Usage Examples
Basic Detection
// Detect from file path
var result = detector.DetectFile("document.pdf");
// Detect from stream
using var stream = File.OpenRead("mystery.file");
var streamResult = detector.Detect(stream);
// Detect with filename hint for better accuracy
var resultWithHint = detector.Detect(stream, "document.docx");
// Detect from byte array
byte[] fileBytes = File.ReadAllBytes("image.jpg");
var byteResult = detector.Detect(fileBytes, "image.jpg");
Working with Different File Types
// Image detection
var imageResult = detector.Detect(imageStream, "photo.jpg");
// Returns: image/jpeg
// Office document detection (detects specific format, not just ZIP)
var docResult = detector.Detect(docStream, "report.docx");
// Returns: application/vnd.openxmlformats-officedocument.wordprocessingml.document
// Archive detection (including TAR)
var zipResult = detector.Detect(zipStream, "files.zip");
// Returns: application/zip
var tarResult = detector.Detect(tarStream, "archive.tar");
// Returns: application/x-tar
Handling Edge Cases
// Empty file
var emptyResult = detector.Detect(emptyStream);
// Returns: application/octet-stream
// Misnamed file (JPEG with .png extension)
var mismatchResult = detector.Detect(jpegStream, "image.png");
// Returns: image/jpeg (content takes precedence)
// No extension
var noExtResult = detector.Detect(stream, "datafile");
// Uses content detection only
// Large files are handled efficiently
using var largeStream = File.OpenRead("huge-video.mp4");
var result = detector.Detect(largeStream); // Only reads what's needed
Metadata Lookup
You can also query MIME type metadata directly:
// Get MIME type repository
var repository = detector.GetMimeTypesRepository();
// Look up by MIME type or alias
var mimeType = repository.GetMimeTypeByTypeOrAlias("application/pdf");
// or by alias
var mimeType2 = repository.GetMimeTypeByTypeOrAlias("application/x-pdf");
// Find all MIME types for an extension
var mimeTypes = repository.GetMimeTypesByExtension("xml");
foreach (var mt in mimeTypes)
{
Console.WriteLine($"{mt.Type}: {mt.Description}");
}
CLI Usage
The package includes a command-line interface that can be used in two ways:
For Development
# Use the provided shell script (Linux/macOS)
./tika-detector document.pdf
# Or on Windows
tika-detector.cmd document.pdf
# Or run directly with dotnet
dotnet run --project src/Omne.TikaContentDetector.Cli -- document.pdf
For Installation (Production)
# Install the CLI tool globally
dotnet tool install -g Omne.TikaContentDetector.Cli
# Then use it anywhere
tika-detector document.pdf
# Detect multiple files with JSON output
tika-detector detect *.jpg --output json --pretty
# Include verbose metadata
tika-detector detect archive.zip --verbose
# Text output format
tika-detector detect image.png --output text
# Look up MIME type information by type or alias
tika-detector mime-type application/pdf
tika-detector mime-type application/x-pdf # Works with aliases too
# Find MIME types by file extension
tika-detector extension pdf
tika-detector extension .pdf # Leading dot is optional
# Get human-readable output
tika-detector mime-type text/plain --output text
tika-detector extension jpg --output text
Design & Architecture
Overview
TikaContentDetector uses the same tika-mimetypes.xml database as Apache Tika, ensuring compatibility and accuracy. The library implements a multi-strategy detection approach:
- Magic Number Detection: Examines file content for characteristic byte patterns
- Container Inspection: Looks inside ZIP/TAR files to identify specific formats
- Filename Matching: Uses file extensions as hints or fallbacks
- Priority Resolution: Resolves conflicts using Tika's priority system
Key Components
- ContentDetector: Main entry point, orchestrates detection strategies
- MagicMatcher: Performs binary pattern matching with offset and mask support
- GlobMatcher: Handles filename pattern matching
- ContainerInspector: Identifies specific ZIP-based formats (DOCX, JAR, etc.)
- MimeTypesRepository: Manages the MIME type database
Detection Pipeline
Input Stream โ Magic Detection โ Container Inspection โ Filename Matching โ Result
โ โ โ
Priority-based conflict resolution โ โ โ โ โ โ โ
Performance Considerations
- Stream-based processing (never loads entire file)
- Configurable buffer sizes for large files
- Lazy loading of MIME database
- Efficient pattern matching algorithms
Supported Formats
The library supports all formats defined in Apache Tika's MIME database, including:
Documents
- PDF, DOCX, XLSX, PPTX (Office Open XML)
- DOC, XLS, PPT (Legacy Office)
- ODT, ODS, ODP (OpenDocument)
- RTF, TXT, HTML, XML, Markdown
Images
- JPEG, PNG, GIF, BMP
- WebP, TIFF, ICO, SVG
- HEIF, AVIF, RAW formats
Archives
- ZIP, 7Z, RAR, TAR
- GZIP, BZIP2, XZ
- JAR, WAR (Java archives)
- EPUB (e-books)
Media
- MP3, MP4, AVI, MKV
- WAV, OGG, FLAC
- WebM, MOV, WMV
Data
- JSON, CSV, XML
- SQLite databases
- YAML, TOML
More Examples
The Examples/ directory contains comprehensive examples:
- BasicUsage - Simple file detection examples
- StreamDetection - Working with various stream types
- BatchProcessing - Processing multiple files efficiently
- WebApiExample - ASP.NET Core Web API integration
Run any example:
cd Examples/BasicUsage
dotnet run
Building from Source
# Clone the repository
git clone https://github.com/omnesoft/tika-content-detector-dotnet.git
cd tika-content-detector-dotnet
# Build the solution
dotnet build
# Run tests
dotnet test
Testing
The library includes comprehensive tests with real file samples:
# Run all tests
dotnet test
# Run specific test categories
dotnet test --filter Category=Images
dotnet test --filter Category=Documents
dotnet test --filter Category=Containers
Contributing
Contributions are welcome! Please read our Contributing Guide for details on our code of conduct and the process for submitting pull requests.
Development Setup
- Fork the repository
- Create a feature branch (
git checkout -b feature/amazing-feature) - Commit your changes (
git commit -m 'Add amazing feature') - Push to the branch (
git push origin feature/amazing-feature) - Open a Pull Request
License
This project is licensed under the Apache License 2.0 - see the LICENSE file for details.
The project includes the tika-mimetypes.xml file from Apache Tika, which is also licensed under the Apache License 2.0.
Acknowledgments
- Apache Tika - This library uses Tika's MIME type database and follows its detection algorithms
- FreeDesktop.org - For the shared MIME info specification
- Contributors - Thanks to all who have contributed to this project
Sources and References
- Apache Tika - The inspiration and source of the MIME database
- Shared MIME Info Spec - The specification Tika's format is based on
- Tika MIME Database - The source XML file
FAQ
Q: How accurate is the detection?
A: Very accurate for common formats. The library uses the same detection rules as Apache Tika, which is widely used and well-tested.
Q: Can it detect encrypted or password-protected files?
A: It can identify the format (e.g., encrypted ZIP), but cannot read the contents without the password.
Q: Does it support custom MIME types?
A: Yes, you can provide a custom MIME types XML file following the Tika format.
Q: How does it handle large files?
A: Efficiently! It only reads the necessary bytes (typically the first 8-64KB) for detection.
Q: Is it thread-safe?
A: Yes, the ContentDetector instance can be safely shared across threads.
| Product | Versions Compatible and additional computed target framework versions. |
|---|---|
| .NET | net10.0 is compatible. net10.0-android was computed. net10.0-browser was computed. net10.0-ios was computed. net10.0-maccatalyst was computed. net10.0-macos was computed. net10.0-tvos was computed. net10.0-windows was computed. |
-
net10.0
- No dependencies.
NuGet packages
This package is not used by any NuGet packages.
GitHub repositories
This package is not used by any popular GitHub repositories.
| Version | Downloads | Last Updated |
|---|---|---|
| 3.3.1.25 | 762 | 7/28/2026 |
MIME database synced to Apache Tika 3.3.1; added model/step (ISO-10303-21) detection for .step / .stp files.