ai-raccoon.linux-arm64
1.57.5
Prefix Reserved
dotnet add package ai-raccoon.linux-arm64 --version 1.57.5
NuGet\Install-Package ai-raccoon.linux-arm64 -Version 1.57.5
<PackageReference Include="ai-raccoon.linux-arm64" Version="1.57.5" />
<PackageVersion Include="ai-raccoon.linux-arm64" Version="1.57.5" />
<PackageReference Include="ai-raccoon.linux-arm64" />
paket add ai-raccoon.linux-arm64 --version 1.57.5
#r "nuget: ai-raccoon.linux-arm64, 1.57.5"
#:package ai-raccoon.linux-arm64@1.57.5
#addin nuget:?package=ai-raccoon.linux-arm64&version=1.57.5
#tool nuget:?package=ai-raccoon.linux-arm64&version=1.57.5
AiRaccoon
An MCP server that gives AI agents persistent, project-scoped memory and a searchable index of the project's own code. It runs locally on .NET 10 and SQLite, with hybrid keyword and semantic search, workspace sandboxes, a shared cross-project tier, and optional cloud sync.
flowchart LR
Agent["MCP client<br/>Claude Code / Hermes / IDE"] <-->|JSON-RPC over stdio| Proxy["ai-raccoon<br/>(proxy)"]
Proxy <-->|authenticated loopback :7721| Server["ai-raccoon serve<br/>(HTTP backend)"]
Server <--> Store[("memory.db<br/>FTS5 + vec0<br/>memory + code")]
Server -.->|optional snapshot sync| Cloud[("S3 / Azure Blob")]
Quick Start
Install the global tool:
dotnet tool install -g ai-raccoon
Activate the bundled embedding engine. A fresh bank has no memory engine until you run this, and memory_search stays keyword-only (it says so in its warning):
ai-raccoon model embedding set local
Add AiRaccoon to your agent's .mcp.json:
{
"mcpServers": {
"ai-raccoon": { "command": "ai-raccoon" }
}
}
The full walkthrough, including indexing your code, is Get started with AiRaccoon.
What's new
ai-raccoon project id register | get | checkand the MCP toolproject_id_getfind or register the id a project already uses, before a new one is minted. (1.57.0) ADR-0127- On Apple silicon the bundled engine can also run on the Neural Engine, opt-in with
ai-raccoon settings model device coreml. (1.53.0) ADR-0118 · how-to - On Windows and Linux x64 the bundled engine runs on the GPU through WebGPU again, from ONNX Runtime's own WebGPU-enabled core that now ships inside the package; linux-arm64 stays on the CPU. (1.52.0) ADR-0115 · how-to
- The server went from up to a full core of CPU and a 6.8 GB footprint to about 2% of one core and 1.75 GB, and embeds on the GPU with 30-90x less CPU per embed, at unchanged search quality. (1.44.3-1.51.2) report
- On Windows and Linux x64 the bundled engine can run on CUDA, opt-in with
ai-raccoon settings model device cuda <path>, unmeasured; the WebGPU plugin shipped in 1.51.0 is off again in 1.51.2. (1.51.0) ADR-0112 · how-to - An encrypted bank tells a wrong key (exit
21) from a corrupt file (exit32), using a key-check file next to the bank that never holds the key. (1.50.0) ADR-0111
Older releases: What's new history. A release tag is not proof of a nuget.org package, see Releases and publishing.
Breaking changes
Upgrading from an older version? Check Breaking changes for what to do past each version. The latest one is 1.57.0.
Performance
We can only speak for one chip so far: an Apple M4 (Mac16,12). The numbers below come from scripts/device-benchmark.py, which drives the shipped product through the same corpus (225 docs files, 4914 chunks) on each device, three repeats each. Every other Mac, and every Windows or Linux machine, is unmeasured.
| device on an M4 | drain time | system energy | server CPU | p95 search latency |
|---|---|---|---|---|
coreml (Neural Engine, opt-in) |
33 s | 379 J | 26 CPU-s | 35-39 ms |
mlx (GPU, opt-in) |
49 s | 1146 J | 37 CPU-s | not measured |
auto (WebGPU, the default) |
70 s | 1389 J | 42 CPU-s | 51-88 ms |
Medians over three repeats; latency is the range of per-repeat p95s over 50 searches. Search quality is the same on every device.
On an M4, switch to the Neural Engine. It embeds twice as fast as the default, at about a quarter of the energy and 40% less server CPU, and searches come back faster:
ai-raccoon settings model device coreml
ai-raccoon serve --restart
The first start after switching compiles the model for the Neural Engine, about 40 s, while WebGPU keeps serving. The compiled cache takes about 808 MiB in the data root, and later starts load in about a second. Newer chips (M5, M6) have faster Neural Engines, so we expect the same advice to hold there, but we have not measured them.
Running AiRaccoon on a different Apple silicon chip? Please run the benchmark and share the result. You need the tool installed, a full clone of this repo (the script reads its corpus from git history), Python 3 with httpx, and AC power. From the clone:
python3 -m venv .venv && .venv/bin/pip install httpx
.venv/bin/python scripts/device-benchmark.py --devices auto,mlx,coreml --repeats 3
It asks for your password once, for powermetrics (add --no-power to skip it), and takes 10-15 minutes. Then open an issue and paste the result.md from the results folder it prints at the end. Full findings: device benchmark report and performance report.
What it does
| Feature | What you get | Docs |
|---|---|---|
| Project-scoped memory | Notes partitioned by project id, in ~/.ai-raccoon or <project>/.ai-raccoon |
Capabilities |
| Hybrid search | FTS5 keyword and vec0 vector search, fused with Reciprocal Rank Fusion, with relevance floors and per-hit evidence | Search pipeline · Parameters |
| Code corpus | A separate, never-synced index of 37 source extensions across C#, F#, Java, Kotlin, Scala, Python, Ruby, PHP, Lua, shell, TS/JS, Vue, Go, Rust, C/C++, Objective-C, Swift, HTML, CSS/SCSS, SQL, Terraform/HCL and Gherkin; memory_search kind=code or kind=both, code_get by hash |
How-to · Dossier |
| File watching | Watched directories re-ingest on change, honouring ai-raccoon.ignore (template) |
Dossier |
| Workspace sandboxes | Isolated in-progress notes that consolidate into the project or get discarded | Workspaces |
| Shared tier | Cross-project facts promoted by memory_share, exempt from decay |
Shared tier |
| Rating and decay | Retrieval raises a memory's rating; a background sweep expires what nobody uses | Degradation |
| Cloud sync | Optional S3 or Azure Blob snapshot sync with optimistic locking | Architecture |
| Encryption at rest | Page-level ChaCha20 via AIRACCOON_DB_PASSPHRASE, rekeyable |
Configure · Rekey |
| Telemetry | OpenTelemetry metrics and traces, OTLP export, memory_performance |
Monitor · Metrics |
Every MCP tool and its parameters are in the tool reference; every command, option and exit code is in the CLI reference.
Embedding models
Memory and code each have their own engine. Both default to the same bundled model, so nothing is downloaded.
| Engine | Model | Dimensions | Notes |
|---|---|---|---|
| Bundled (default) | granite-embedding-small-english-r2, fp16 |
384 | Memory and code. WebGPU on macOS, CPU elsewhere |
| Downloaded | Any Hugging Face model with a tokenizer.json or a WordPiece/SentencePiece vocab, via model download |
from the model | Memory or code |
| Remote | Any OpenAI-compatible /v1/embeddings endpoint |
from the endpoint | Memory only |
Against the previous defaults, granite scores memory nDCG@10 0.632 (MiniLM 0.605) and code MRR@10 0.788 over 12 languages (MiniLM 0.643, code-daemon-embed-v1 0.391). With identifier-aware code search on top, code nDCG@5 went from 0.501 to 0.824. Sources: ADR-0108, code retrieval eval. Setup and the full model table: Configure embedding engines · Benchmarks.
Running the server
ai-raccoon # proxy (what MCP clients run): attaches to the backend, starting it if needed
ai-raccoon serve # the HTTP backend: owns the bank, idle watchdog, token-authenticated loopback
ai-raccoon serve observability counters # live counters via dotnet-counters
Ports, environment variables, encryption and upgrades: Configure and run the server.
Project layout
src/AiRaccoon/ # CLI, proxy and MCP tool handlers (thin)
src/AiRaccoon.Core/ # Domain: memory, search fusion, chunking, workspaces, rating
src/AiRaccoon.Infrastructure/ # SQLite store, embeddings, sync, maintenance jobs
tests/AiRaccoon.Tests/ # xUnit v3 suite (about 2,600 test methods)
benchmarks/ # BenchmarkDotNet and retrieval benchmarks
scripts/ # Python tooling, see docs/how-to/run-the-python-scripts.md
Design background: Architecture and the ADRs.
Documentation
Documentation tree: tutorials, how-to, explanation, reference, ADRs.
Contributing & Security
- CLAUDE.md holds repo conventions, including the mandatory TDD workflow.
- Before packing the tool from source, run
python3 scripts/download-embedding-model.py. The bundled model's weights are downloaded, not committed, anddotnet packstops without them. - Report security issues privately per SECURITY.md.
License
MIT. Copyright (c) 2026 Rafał Araszkiewicz.
Learn more about Target Frameworks and .NET Standard.
This package has no dependencies.
NuGet packages
This package is not used by any NuGet packages.
GitHub repositories
This package is not used by any popular GitHub repositories.
| Version | Downloads | Last Updated |
|---|---|---|
| 1.57.5 | 0 | 10/7/2026 |
| 1.57.3 | 105 | 10/6/2026 |
| 1.57.2 | 77 | 10/5/2026 |
| 1.57.0 | 78 | 10/5/2026 |
| 1.56.1 | 84 | 10/3/2026 |
| 1.56.0 | 82 | 10/3/2026 |
| 1.54.0 | 100 | 9/29/2026 |
| 1.53.13 | 83 | 9/28/2026 |
| 1.53.12 | 82 | 9/28/2026 |
| 1.53.10 | 89 | 9/28/2026 |
| 1.53.8 | 81 | 9/27/2026 |
| 1.53.7 | 83 | 9/27/2026 |
| 1.53.6 | 83 | 9/27/2026 |
| 1.53.5 | 89 | 9/27/2026 |
| 1.53.1 | 85 | 9/26/2026 |
| 1.53.0 | 97 | 9/26/2026 |
| 1.52.3 | 87 | 9/25/2026 |
| 1.52.2 | 86 | 9/25/2026 |
| 1.52.0 | 83 | 9/25/2026 |
| 1.51.2 | 85 | 9/25/2026 |