KlutzyNinja.Filtrace
0.9.0
Prefix Reserved
dotnet tool install --global KlutzyNinja.Filtrace --version 0.9.0
dotnet new tool-manifest
dotnet tool install --local KlutzyNinja.Filtrace --version 0.9.0
#tool dotnet:?package=KlutzyNinja.Filtrace&version=0.9.0
nuke :add-package KlutzyNinja.Filtrace --version 0.9.0
filtrace
A small, agent-shaped CLI and MCP server for analyzing .NET CPU, allocation,
blocking, and wall-clock traces. Built on the
Microsoft.Diagnostics.Tracing.TraceEvent library; reads EventPipe
(.nettrace / .speedscope.json) and ETW (.etl) captures from both .NET and
.NET Framework runs.
Install
filtrace targets .NET 10. Both heads are published on NuGet.org:
KlutzyNinja.Filtrace
(the filtrace CLI) and
KlutzyNinja.Filtrace.Mcp
(the MCP server).
CLI (global tool)
Installing the CLI as a .NET global tool needs the .NET 10 SDK (dotnet tool
ships with it):
dotnet tool install --global KlutzyNinja.Filtrace
filtrace rank app.nettrace --metric cpu
Update or remove it later with dotnet tool update --global KlutzyNinja.Filtrace
or dotnet tool uninstall --global KlutzyNinja.Filtrace.
MCP server
The MCP server runs on demand - there is no install step. Add the stdio server to
your agent's MCP config and dnx (bundled with the .NET 10 SDK) fetches
KlutzyNinja.Filtrace.Mcp and launches it. See
Using filtrace from an AI agent for the exact
config block and the tool workflow.
From source
Activate the current checkout for one Git repository without changing a global Filtrace installation:
./tools/Use-LocalFiltrace.ps1 -Action Install -TargetRepository ../consumer -Configuration Release
Run Install again to refresh the private CLI, MCP server, and skill from the
current source. -TargetRepository defaults to the current directory. Restore
the target's recorded MCP and skill baseline and remove the private CLI with:
./tools/Use-LocalFiltrace.ps1 -Action Restore -TargetRepository ../consumer -Configuration Release
The command requires PowerShell 5.1 or 7, Git, and the .NET 10 SDK. It does not need elevation, change a global tool installation, or upload the prepared package.
Using filtrace
Every analysis command takes a trace path and prints a dense text report (or compact
JSON with --format json); collect instead launches the executable it records.
The canonical investigation is orient → rank → drill →
compare: inspect the capture, rank the matching metric, drill an unwindowed CPU
ranking when needed, then diff comparable traces or capture manifests against a baseline.
# Workflow: orient, rank the hottest frames, drill into one, then diff two runs.
filtrace info app.nettrace # 0. orient: format, providers, event counts, symbol rate
filtrace rank app.nettrace --metric cpu # 1. what's hot (self weight: milliseconds or raw samples)
filtrace callers app.nettrace MyApp.Parse # 2. who calls the hot frame
filtrace source app.nettrace --view lines --symbols bin/Release/net10.0 # 3. hot source lines
filtrace diff before.nettrace after.nettrace # 4. what changed between runs
filtrace batch BenchmarkDotNet.Artifacts/filtrace-runs/run/manifest.json
The same analysis core is exposed as a stdio MCP server: every analysis command has a
matching trace_* tool (eighteen in all - info → trace_info, rank →
trace_rank, callers → trace_callers, and so on), returning the same envelope
shape and results. The capture and housekeeping commands (collect, cache) are
CLI-only, as is widening an ETW analysis with --all-processes; MCP
supports automatic or named-process scope. See
Using filtrace from an AI agent for the client
config and tool workflow.
Commands
Orient - see what a capture holds before ranking (the CLI counterpart of the
trace_info tool):
| Command | Purpose | Example |
|---|---|---|
info |
Format, sample count, lost-event and symbol-quality warnings, supported analyses, and per-analysis capture/event state | filtrace info app.nettrace |
Ranking - rank stacks by a metric:
| Command | What it ranks | Example |
|---|---|---|
rank |
Any metric (cpu, alloc, exceptions, threadtime, contention, wait, activity) |
filtrace rank app.nettrace --metric contention |
Implemented scope inventory:
- Named process: CLI
info,rank,source,callers,tree,classify,timeline,diff,batch,export, andreport(for--kind diskio); MCPtrace_info,trace_rank,trace_callers,trace_lines,trace_heatmap,trace_tree,trace_classify,trace_timeline,trace_diff,trace_batch, andtrace_export, plustrace_diskio. Every listed surface except the disk report auto-scopes a multi-process.etlto the busiest process tree. Runprocesses/trace_processesfirst to inspect the capture, then set--process <name>/processto override. CLI commands expose--all-processeswhere an aggregate is supported; MCPtrace_infoandtrace_rankexposeallProcesses. Capture-wideinfo/trace_infopreserves whole-capture totals but budget-limits its busiest-thread detail with an explicit truncation warning. Process inventory aggregates ETL and EventPipe CPU ownership without materializing frames; speedscope uses its full stack reader. Useinfo/trace_infofor frame-name and source/PDB quality. Other stack-backed MCP analyses have no all-process aggregate;trace_diskiois the exception and, like the CLI disk report, remains machine-wide by default. When scoped, disk reports correlateDiskIOInitissuer IRPs to completions because a completion PID may be System or Idle. This is direct issuer scope, not causal ownership: deferred file-system/cache write-back issued by System is included only when System is selected, while Idle-issued I/O appears only in an unscoped report. - Exact process ids: the same commands and tools accept
--pid <id>[,<id>](comma-separated, not repeated) /pidinstead of a name. A name substring is right for discovery, but a common host name such asdotnetmatches every unrelated instance in a machine-wide capture and ranks them together; an exact id set cannot. Prefer it for manifests and automation. The three selectors are mutually exclusive, an id reused by two processes in one trace is refused rather than merged, and an id that is not in the trace is reported. - Descendants: the same commands and tools accept
--children include|exclude/children. Both selectors follow descendants by default, because the common capture shapes put the measured work in a child the host launched. Passexcludeto separate a parent's own CPU from a child runtime's; without it a native host's own cost is blended with the CoreCLR frames of the child it launched. - Disk completion time: CLI
report --kind diskioand MCPtrace_diskioaccept--time <start>,<end>/time. The window filters completion timestamps after issuer-process IRP correlation; either bound may be open. - Invocation roots: CLI
lifecycleand MCPtrace_lifecycletake the same--process/--pidselectors, but each matched process instance is one invocation and descendants always follow, so neither takes--childrenor--all-processes. - Root subtree: CLI
rank,callers,tree,classify,diff,batch, andexport; MCPtrace_rank,trace_callers,trace_tree,trace_classify,trace_diff,trace_batch, andtrace_export. Set--root <frame>/rootto keep the subtree under a frame. Root filtering is stack ancestry, not causal correlation: stacks without the selected frame are excluded, including sibling workers. Root-aware structured results identifyrootKind: stackAncestryand report available versus retained weight and record counts; direct diffs report both sides, and manifest batch/diff report each case. Use an instrumented activity or validated time window for a parallel phase, and ETWthreadtimewhen sampled CPU does not explain elapsed time. - BenchmarkDotNet workload: CLI
rank,callers,tree,classify,diff,batch, andexportaccept--benchmark; MCPtrace_rank,trace_callers,trace_tree,trace_classify,trace_diff,trace_batch, andtrace_exportacceptbenchmark: true. The preset isolates theWorkloadActionsubtree from harness and overhead scaffolding; it is mutually exclusive with an explicit root.sourceviews are not root-aware, so narrow them by method/file and treat percentages as process-scoped whole-trace values.
The rank command adds two more scopes: --activity <name> (the CPU samples taken inside one
start-stop request/job) and --time <start>,<end> (milliseconds from the trace
start, either bound optional; any metric on .nettrace / .etl), to zoom in on one
request or the slice around a latency spike. Speedscope input is aggregate-only for
--time and warns that the window was ignored.
filtrace rank bdn.nettrace --metric cpu --benchmark # just the [Benchmark] code
filtrace rank bdn.nettrace --metric alloc --benchmark # allocations under the workload
filtrace processes machinewide.etl # list every process by weight
filtrace rank machinewide.etl --metric cpu --process MyApp # one process tree
filtrace rank machinewide.etl --metric cpu --pid 9144,40356 --children exclude # exactly those, parent-only
filtrace rank app.nettrace --time 1000,5000 # just the spike window
Native runtime symbols. Managed frames (including NGEN and ReadyToRun
framework methods) resolve for free from the trace's CLR rundown. The unmanaged
runtime frames - the GC, the JIT, memset / memcpy, write barriers - need PDBs
from the Microsoft public symbol server, which rank fetches only when you
opt in with --native-symbols (cached under --symbol-cache, default in the temp
path). It is off by default so analysis stays offline and deterministic; the first
run downloads, later runs hit the cache.
filtrace rank app.etl --metric cpu --process MyApp --native-symbols # name the GC/JIT/memcpy frames
CPU drill-down - follow an unwindowed CPU ranking into detail:
| Command | Purpose | Example |
|---|---|---|
callers |
Immediate CPU callers of a frame, or a caller/callee view with --callees |
filtrace callers app.nettrace MyApp.Parse --callees |
source |
Ranked source lines or per-file heat | filtrace source app.nettrace --view heatmap --file Parser.cs |
tree |
Top-down CPU call tree from the root | filtrace tree app.nettrace --max-depth 5 |
Inventory - see what a (possibly machine-wide) capture contains:
| Command | Purpose | Example |
|---|---|---|
processes |
List processes by CPU-sample weight, to pick a --process target |
filtrace processes machinewide.etl |
classify |
Summarize CPU weight by runtime work category (ms when established, otherwise samples) |
filtrace classify app.etl --native-symbols |
Temporal - see what happened when, then scope a ranking to the busy window:
| Command | Purpose | Example |
|---|---|---|
timeline |
Aligned activity buckets or one bounded cross-lane snapshot | filtrace timeline app.nettrace --mode snapshot --at 1500 |
Compare and export:
| Command | Purpose | Example |
|---|---|---|
diff |
Absolute/normalized CPU changes for traces or paired manifests | filtrace diff before.nettrace after.nettrace |
batch |
One compact ranking across every capture-manifest case | filtrace batch run/manifest.json |
export |
Write a flame graph (speedscope / chromium) | filtrace export app.nettrace --format speedscope -o app.json |
Structured reports:
| Command | Purpose | Example |
|---|---|---|
report |
GC, JIT, thread-pool, or physical disk-I/O report | filtrace report app.nettrace --kind gc |
lifecycle |
Per-invocation wall-clock phases: root lifetime, first child, child span, teardown (ETW) | filtrace lifecycle run.etl --process myapp --image hostfxr |
events |
Query raw events, filtered by name / payload / pid / tid, paged | filtrace events app.etl --payload ConnectionReset |
Capture (Windows, elevated) - record an ETW .etl yourself, no external recorder:
| Command | Purpose | Example |
|---|---|---|
collect |
Launch an executable and record a CPU, thread-time, startup, or physical disk-I/O .etl |
filtrace collect --launch bin/Release/net10.0/MyApp.exe --output myapp.etl --profile threadtime |
filtrace collect --launch bin/Release/net10.0/MyApp.exe --output myapp.etl # CPU
filtrace collect --launch dotnet --launch-args MyApp.dll --output tt.etl --profile threadtime
filtrace collect --launch MyApp.exe --output start.etl --profile startup # low perturbation
filtrace collect --launch MyApp.exe --output io.etl --profile diskio # physical disk/files
filtrace collect --launch cmd.exe --launch-args "/d /c build.cmd" --working-directory C:\src\app --output build.etl
filtrace collect --launch cmd.exe --launch-args "/d /c build.cmd" --output build.etl --rundown --rundown-pid 1234 # known persistent CLR server
filtrace collect --launch MyApp.exe --output ring.etl --max-size-mb 512 # bounded ring buffer
--working-directory sets and records the absolute directory inherited by every
subject launch. Omit it to inherit the collector's current directory.
--rundown appends and merges a separate minimal CLR naming rundown after the
launched command exits. Use it only when captured CPU belongs to managed servers
that were already running and remain alive, such as compiler/build servers. The
opt-in pass is machine-wide by default; pass up to 8 comma-separated exact ids
to --rundown-pid to filter CLR naming events when the target servers are already
known. The capture result records that filter. Rundown requests 512 MB of ETW
buffers, and TraceEvent's derived maximum-buffer count permits the pool to grow
to roughly 641 MiB. Provider activation can take up to 30 seconds, followed by up
to 30 seconds of quiescence polling; the subsequent ETL merge has no timeout.
Rundown can add hundreds of megabytes. It is not valid with --profile diskio or
--max-size-mb; inspect the capture and info lost-event warnings before
trusting resolved names.
A targeted process must be running before capture and remain the same process
through rundown; exit or PID reuse fails explicitly.
The diskio profile enables only physical DiskIO/DiskIOInit events, process/thread
attribution, and the DiskFileIO name rundown. It deliberately omits CPU sampling,
stacks, verbose FileIO, and CLR events. The capture and unscoped disk report are
machine-wide; write the trace to a different volume when recorder writes must not
contend with the workload volume. Scope a report to direct issuers with
--process or --pid, and to completion time with --time.
With --format json, stdout contains only the capture-result JSON; identified
subject stdout and stderr are forwarded to stderr. The command still exits successfully
when capture succeeds even if the subject fails, and reports that failure in
processExitCode and invocations. Text output continues to inherit the subject's streams.
For an EventPipe (.nettrace) capture - cross-platform, no elevation - use the
first-party dotnet-trace (dotnet tool install -g dotnet-trace, then
dotnet-trace collect -- <app>); collect is ETW-only.
File ops - manage the ETLX conversion cache TraceEvent keeps beside a trace:
| Command | Purpose | Example |
|---|---|---|
cache |
Build/reuse or remove the ETLX cache | filtrace cache app.nettrace --action convert |
ETLX conversion is coordinated per canonical trace path across threads and
processes, with unique temporary files and atomic publication. Same-trace MCP
queries may run in parallel; trace_info.etlxCacheState and cache --action convert report
hit, waited, converted, or recovered.
Preview alias migration
The previous command names remain callable during the current migration window and print their canonical replacement to stderr, but they are hidden from top-level help and are not used in examples or generated guidance. Removal requires the explicit VN5 migration policy; it is not tied automatically to the passage of one preview release:
| Previous names | Canonical command |
|---|---|
cpu, alloc, exceptions, threadtime |
rank --metric <name> |
lines, heatmap |
source --view <name> |
gcstats, jitstats, threadpool, diskio |
report --kind gc|jit|threadpool|diskio |
convert, clean |
cache --action convert|clean |
Run filtrace <command> --help for the full option set of any command.
Using filtrace from an AI agent
filtrace is built for an agent mid-investigation. Two ways to wire it in:
MCP server - add the stdio server so the agent calls the
trace_*tools directly:{ "servers": { "filtrace": { "type": "stdio", "command": "dnx", "args": ["KlutzyNinja.Filtrace.Mcp", "--yes"] } } }CLI - install the global tool (
dotnet tool install -g KlutzyNinja.Filtrace) and let the agent shell out tofiltrace <verb>.
Either way, the canonical loop is orient → rank → drill → compare: read
trace_info (CLI: filtrace info) first; when symbol resolution is below 0.8,
inspect its warning and unresolved rows. Treat that as frame-name quality; before
source-line analysis, inspect sourceResolution for exact matching PDB modules,
mapped sampled managed frames, searched directories, highest-unmapped modules, and
pdbIdentityMismatchModules. A mismatch means a same-named local PDB was found but
its GUID or age differs from the trace. Once the relevant module matches, use
sourceMappedManagedMethodCount versus sampledManagedMethodCount to confirm sampled
methods resolve sequence points; unmappedNamedManagedFrameCount and
highestUnmappedMethods expose the remaining <no source> impact.
Use the generated BenchmarkDotNet child output when the outer build PDB does not
match; use native symbols for CPU ETW runtime frames as applicable. Rank by the metric that matches the question (cpu, alloc, exceptions,
threadtime, contention, wait, activity); for an unwindowed CPU ranking, drill the
hot frame with callers / lines / tree; diff comparable CPU traces against a baseline.
Layout
| Path | Purpose |
|---|---|
src/Filtrace.Core/ |
Analysis core: trace readers, stack-source providers, the provider-agnostic question-service engine. The only place logic lives. |
src/Filtrace/ |
CLI host, packaged as the filtrace .NET global tool. |
src/Filtrace.Mcp/ |
Stdio MCP host, packaged separately for dnx KlutzyNinja.Filtrace.Mcp. |
benchmarks/Filtrace.Benchmarks/ |
BenchmarkDotNet performance harness for the analysis core. |
benchmarks/Filtrace.PerfWorkload/ |
Parameterized CPU/activity workload for reproducible Track D traces. |
tests/Filtrace.Core.Tests/ |
Unit + golden-file contract tests. |
tests/Filtrace.Parity.Tests/ |
Numeric parity against the frozen legacy oracles. |
eval/ |
Headless-agent eval harness, tasks, baselines. |
docs/ |
Design, roadmap, competitive analysis, and the single-source workflow text for the skill / README / help. |
.agents/skills/filtrace/ |
The shipped agent skill. |
Self-containment
filtrace carries its own Directory.Build.props, Directory.Build.targets,
Directory.Packages.props, global.json, and .editorconfig (root = true),
so the build is fully self-contained. Its only external dependency is the
published KlutzyNinja.Touki NuGet package; it references no other project.
Build and test (standalone)
cd filtrace
dotnet build filtrace.slnx
dotnet test filtrace.slnx
| Product | Versions Compatible and additional computed target framework versions. |
|---|---|
| .NET | net10.0 is compatible. net10.0-android was computed. net10.0-browser was computed. net10.0-ios was computed. net10.0-maccatalyst was computed. net10.0-macos was computed. net10.0-tvos was computed. net10.0-windows was computed. |
This package has no dependencies.