Techxot.Umbraco.AiReadiness
18.2.4
dotnet add package Techxot.Umbraco.AiReadiness --version 18.2.4
NuGet\Install-Package Techxot.Umbraco.AiReadiness -Version 18.2.4
<PackageReference Include="Techxot.Umbraco.AiReadiness" Version="18.2.4" />
<PackageVersion Include="Techxot.Umbraco.AiReadiness" Version="18.2.4" />
<PackageReference Include="Techxot.Umbraco.AiReadiness" />
paket add Techxot.Umbraco.AiReadiness --version 18.2.4
#r "nuget: Techxot.Umbraco.AiReadiness, 18.2.4"
#:package Techxot.Umbraco.AiReadiness@18.2.4
#addin nuget:?package=Techxot.Umbraco.AiReadiness&version=18.2.4
#tool nuget:?package=Techxot.Umbraco.AiReadiness&version=18.2.4
Techxot AI Readiness for Umbraco
Scores how readable your pages are to AI answer engines, generates JSON-LD from a page's own content, and records a score on every publish so a regression is visible.
Measured on a stock Umbraco Starter Kit: site average 22/100. One page went 24 → 51 after applying generated JSON-LD.
Requires .NET 10 and Umbraco 17.5.3 or later. Pick the package major that matches your Umbraco
major — 17.x for Umbraco 17, 18.x for Umbraco 18. They are built from the same source; Umbraco
moves interface members between majors in ways that are source- but not binary-compatible, so one
assembly cannot serve both. NuGet will refuse the wrong pairing rather than fail at runtime.
What you get
| Feature | Where |
|---|---|
| Page audit — 7 weighted levers, fixes ranked by ROI | AI Readiness tab on every document |
| Generate JSON-LD from the page's own content | Button in that tab |
| Score on publish, with history and regression detection | Background worker |
| Site-wide ranking, weakest page first | AI Readiness dashboard in Content |
| Site signals — which AI crawlers robots.txt lets in, llms.txt health, sitemap | Top of that dashboard |
| llms.txt, llms-full.txt and a Markdown rendering of every page, generated from content | Served at the root on Razor sites; fetched by your front end on headless ones — see llms.txt and Markdown below |
Install
Four steps. Step 4 depends on how your site renders, and there is one extra step at the end that does too. Read step 4 even if you are only running locally — on a headless site it is the difference between scoring your pages and scoring your 404 page.
1. Add the package
# Umbraco 18
dotnet add package Techxot.Umbraco.AiReadiness --version 18.*
# Umbraco 17 LTS
dotnet add package Techxot.Umbraco.AiReadiness --version 17.*
No startup code is needed — the package registers itself.
On Umbraco Cloud: add the reference in your own source repository, then push to the Cloud deployment repository. Cloud git repos are deployment repos — they hold build output, not your source — and NuGet packages are restored during the Cloud build. Nothing else is required.
2. Configure
{
"Techxot": {
"AiReadiness": {
"BaseUrl": "https://inmotion.techxot.com",
"ApiKey": "<your key>",
"TimeoutSeconds": 60,
// Origin to fetch a page from when scoring it. Leave unset on a normal Razor site.
// REQUIRED if you are headless — see step 4.
"SiteUrlOverride": null,
// Property aliases searched, in order, for somewhere to store JSON-LD.
"JsonLdPropertyAliases": [ "structuredData", "jsonLd", "customSchema" ],
"ScoreOnPublish": {
"Enabled": false, // off by default
"DelaySeconds": 5, // let caches settle before fetching the page
"MaxQueueLength": 500, // work waiting to be scored; excess is dropped
"HistoryLimit": 50, // scores kept per page
"RegressionThreshold": 5, // point drop that counts as a regression
// Engine allowance — see "Keys and quota" below. The defaults are sensible.
"DeduplicationSeconds": 30, // do not score the same page twice for one publish
"MaxThrottleRetries": 2, // patience when the engine says "slow down"
"MaxThrottleWaitSeconds": 120, // longest single wait before standing down
"QuotaWarningThreshold": 0.1 // warn below 10% of the daily allowance
}
}
}
}
Put the API key in environment variables or a secret store, not in
appsettings.json. The key travels server-to-server only; this package exists partly to keep it out of the browser.
If you cannot follow that rule, cap the key instead. On Umbraco Cloud the key belongs in
Security → Secrets → your environment, as Techxot__AiReadiness__ApiKey — double underscores are
Cloud's convention for overriding an appsettings.json value, and environment variables outrank
JSON, so the file can keep an empty "ApiKey": "". But Cloud's Starter plan allows only five
secrets per environment, and a real project spends those on SMTP, the Delivery API, Forms and
reCAPTCHA before this package asks for one. Higher plans have no limit.
When there is genuinely no slot, putting the key in appsettings.{Environment}.json is a defensible
trade rather than a mistake — provided you do two things:
- Know what it costs. The key enters your repository's history permanently. Rotating it later does not remove it. Treat it as disclosed to anyone who has ever had repository access.
- Cap the blast radius. Ask for a
rateLimiton the key record. A key that can spend at most a few thousand calls a day is a bounded, visible problem; an uncapped one is not. A full scan of a 40-page site costs about 30 calls, so a modest cap is invisible in normal use.
What a leaked key can and cannot do is worth knowing either way: the engine grants keys read and score only. It cannot apply a fix, change content, or read anything back out. A leak costs you compute, not data.
Every setting is read live through IOptionsMonitor, so editing configuration takes effect without
a restart. That is deliberate: if the engine misbehaves, ScoreOnPublish.Enabled can be turned off
without taking a deployment.
3. Add somewhere to store the JSON-LD
The package never creates content-type properties — doing so breaks Umbraco Deploy, which would find a schema change it did not author and try to reconcile it.
Add one manually, once:
- Settings → Document Types → the type all your pages compose. In the Starter Kit that is the shared SEO composition; adding it there gives every page type the field in one edit.
- Add a property: Textarea, alias
structuredData, name Structured Data (JSON-LD). - Save.
If the alias does not match, the AI Readiness tab tells you which aliases it looked for — it will not fail silently.
4. Tell the package which host to score
Scoring fetches a page's own rendered HTML. So the package has to know where that HTML actually lives, and the answer differs by rendering model. Get this wrong and scoring still runs — it just scores the wrong thing, which is worse than failing.
If your site renders with Razor
Set the application URL:
"Umbraco": { "CMS": { "WebRouting": { "UmbracoApplicationUrl": "https://www.example.com/" } } }
Background scoring runs without an HTTP request, so IPublishedUrlProvider has no host to work from
and produces portless, wrong URLs (locally: http://localhost/about-us/). Scoring still functions,
but every URL it records is wrong.
On Cloud this matters more, not less — each environment has a different hostname, so this is a per-environment setting.
If your site is headless
Set SiteUrlOverride to your FRONT-END host — not the CMS host.
"Techxot": { "AiReadiness": { "SiteUrlOverride": "https://www.example.com" } }
On a headless site the CMS renders almost nothing. Typically one template answers / and every
other path 404s, because rendering lives in the front end. Point scoring at the CMS and it dutifully
fetches those 404 pages and scores them — producing a uniformly terrible score that describes your
error page, not your site. Nothing errors, which is exactly what makes it dangerous.
IPublishedUrlProvider still resolves the correct path, so this setting swaps only the origin
and keeps the path: /about-us/ is fetched from https://www.example.com/about-us/, which is what a
visitor and an AI crawler actually see. The URL recorded in history is the one that was fetched, so
the dashboard links somewhere real.
This is the correct production value for a headless site. The setting also has a development use — pointing a local site at its plain-HTTP binding, because an app calling itself over HTTPS has to validate its own dev certificate — but do not mistake it for a dev-only switch.
Two things to check on a headless site, both of which this package will surface for you:
- Pages the CMS publishes but the front end does not route appear in the log as
could not fetch '<page>' … returned HTTP 404, naming the URL. That is a real routing gap worth fixing, not a package error. - Settings and composition nodes — headers, footers, global settings — usually have URLs and are therefore scored. If your front end renders them as pages, they are also publicly indexable, which is worth knowing on its own. Either stop routing them or expect them in the ranking.
Keys and quota
Auditing and JSON-LD generation are performed by the InMotion engine, which meters and rate-limits per key. Without a key the backoffice explains what the product does rather than showing an error.
One key per environment
Umbraco Cloud gives every project a Live, a Staging and usually a Dev environment. Ask for one key per environment, not one per project. It costs nothing and buys three things:
- Blast radius. A leaked Dev key cannot touch Live's allowance.
- Attribution. You can see which environment is spending what.
- Billing. Non-production keys are not billable. A Staging environment running a site scan on every deploy should never reach an invoice, and with separate keys it cannot.
Non-production keys are also rate-limited harder than production ones. That is intentional and not something to work around — if Staging genuinely needs production-scale volume, ask for the limit to be raised on that key.
What happens at the limit
The engine advertises the remaining daily allowance on every response, and this package acts on it without any configuration:
| Situation | What the package does |
|---|---|
Below QuotaWarningThreshold (10% left) |
Logs a warning naming the number remaining and when it resets — before anything fails |
| Short refusal (per-minute burst, seconds) | Waits and retries, up to MaxThrottleRetries. A bulk publish recovers on its own |
| Long refusal (daily quota, an hour) | Stands down until the quota resets, logs it once, and resumes automatically |
The distinction is drawn from the length of the wait the engine asks for, not from guesswork. It matters because the two need opposite handling: waiting out a daily quota would stall the queue until it overflows and started dropping work, while giving up on a per-minute burst would lose history for exactly the bulk operations most worth recording.
You do not need to intervene. If the daily quota is being hit routinely, that is a plan conversation, not a configuration one.
Rendering the JSON-LD
The package stores JSON-LD on the document. Emitting it is your site's job — this is the one caveat in "does it just work?", and it differs by rendering model.
Non-headless (Razor)
One line in your layout's <head>:
@await Component.InvokeAsync("TechxotStructuredData")
That is the whole integration. The component reads the stored property, validates and escapes it, and renders nothing at all when there is nothing valid to emit.
Why a component rather than a
<script>tag you write yourself. With@addTagHelper *, Microsoft.AspNetCore.Mvc.TagHelpersin scope, Razor re-encodes the attributes of any<script>it can see and shipsapplication/ld+json. Browsers decode that, so the page looks perfect and nothing errors — but any consumer that string-matches the media type sees no structured data at all. Measured on one page: 24/100 encoded vs 51/100 decoded, identical content. The component returns raw content and never enters the tag-helper pipeline, so the<!script>opt-out cannot be forgotten in someone else's layout.
Headless
The package cannot render into your front end, so read structuredData from the Delivery API and
emit it yourself. Copy headless/structured-data.ts from this package — it is dependency-free
and framework-agnostic:
import { serializeJsonLd } from "@/lib/structured-data";
const jsonLd = serializeJsonLd(properties.structuredData, path);
return (
<>
{jsonLd ? (
<script type="application/ld+json" dangerouslySetInnerHTML={{ __html: jsonLd }} />
) : null}
{/* ...the rest of your page */}
</>
);
A native <script> is correct here — next/script is for executable JavaScript, and JSON-LD is
data. React writes the type attribute verbatim, so the headless path does not suffer the Razor
tag-helper problem described above.
The caching trap — check this before you suspect the code. If JSON-LD is live in the CMS but missing on the front end, your front end is almost certainly serving a cached Delivery API response. In the pilot a
revalidateof 3600 meant newly published JSON-LD did not appear for a full hour, which looks exactly like a broken integration. Wire an Umbraco publish webhook to a revalidation endpoint (revalidateTag/revalidatePathin Next.js) so publishing invalidates the page. Debugging the emit code when the real problem is the cache costs an afternoon.
llms.txt and Markdown
AI crawlers read better from Markdown than from your rendered HTML, and /llms.txt is the
convention for telling them what a site contains. The package generates three things from your
published content — no template work, and it works on a site it has never seen, because it walks
property editors (text, rich text, Block List, Block Grid), not your document types:
| Document | What it is |
|---|---|
/llms.txt |
Site name, a one-paragraph summary, then every indexable page as a link with a one-line description, grouped by section |
/llms-full.txt |
Every indexable page rendered in full, one document — not the same bytes as llms.txt, which is a common mistake |
/{page-path}.md |
One page as Markdown; also served for Accept: text/markdown on the page URL |
Generation is always on. Where they are served depends on how your site renders, and the default is the safe one for both:
Razor (pages render on this host): set Techxot:AiReadiness:Emission:ServeAtRoot to true.
The package answers the three routes above at this host's root and adds
Link: <…/page.md>; rel="alternate"; type="text/markdown" to every HTML page so crawlers find the
alternate without reading the markup.
Headless (pages render on another host): leave ServeAtRoot at false — serving at the CMS
root would publish the files at a host no crawler visits — and have your front end serve them.
Copy headless/emission.ts (and headless/middleware.example.ts) from this package. It fetches
the documents from /umbraco/techxot/ai-readiness/public/{llms|llms-full|markdown} using the
Delivery API key you already have as Api-Key, forwards ETags so a crawler's If-None-Match
becomes a 304, and can fall back to a hand-written llms.txt if the CMS is unreachable:
// app/llms.txt/route.ts
import { proxyEmission } from "@/lib/emission";
export const GET = (req: Request) => proxyEmission("llms", req);
Links inside the generated files honour SiteUrlOverride, so they point at your front end, not
at the CMS.
Tuning, all under Techxot:AiReadiness:
| Setting | Default | Purpose |
|---|---|---|
ExcludedContentTypes |
[] |
Document type aliases that are not pages (settings, shared header/footer). Kept out of llms.txt, the ranking and the scan |
Emission:NoIndexPropertyAliases |
noIndex, hideFromSearchEngines, robotsNoIndex, seoNoIndex, excludeFromSitemap |
A truthy value keeps the page out of llms.txt and llms-full.txt |
Emission:DescriptionPropertyAliases |
metaDescription, seoDescription, description, summary, excerpt |
Where a page's one-line summary comes from; falls back to its first sentence |
Emission:ExcludedPropertyAliases |
SEO/meta aliases | Properties never rendered into Markdown |
Emission:SiteDescription |
— | The blockquote under the llms.txt title; falls back to the home page's description |
Emission:MaxFullTextPages |
200 |
Cap on pages in llms-full.txt |
The readiness score rewards all of this: a page that advertises a Markdown alternate scores higher
on retrievability, and the dashboard's Site signals strip shows whether /llms.txt exists and
is well-formed at your public host.
Who publishes this site — Organization and site-level JSON-LD
An answer engine that cannot tell who is speaking will not cite you. Every page therefore carries,
alongside whatever JSON-LD its author stored, three site-level nodes in one @graph:
| Node | Source |
|---|---|
WebSite |
The site name and origin, with a SearchAction if you give it a search URL |
Organization |
Techxot:AiReadiness:Organization — only when Name is set |
BreadcrumbList |
The page's ancestors, from the content tree |
"Techxot": {
"AiReadiness": {
"Organization": {
"Name": "Your Organisation",
"LegalName": "Your Organisation Ltd",
"Url": "https://www.example.com/",
"Logo": "https://www.example.com/logo.png",
"Description": "One paragraph on what you are.",
"SameAs": ["https://www.linkedin.com/company/your-org", "https://www.wikidata.org/wiki/Q…"],
"KnowsAbout": ["the topics you want to be cited for"],
"SearchUrlTemplate": "https://www.example.com/search?q={search_term_string}"
}
}
}
Configuration, not content: nothing is added to your document types, and every default is empty — a package that shipped an example organisation would be shipping somebody's identity to every site that forgot to change it. The dashboard shows what is configured (read-only) and says so plainly when nothing is.
The author wins. If a page's own JSON-LD already contains an Organization or a
BreadcrumbList, the site-level one of that type is dropped for that page — a hand-written node
on the About page is more specific than the configured one.
- Razor: nothing to do; the
TechxotStructuredDatacomponent you already invoke now emits the graph. - Headless:
structured-data.tsgainscomposeJsonLd(raw, siteNodes); fetch the nodes from/umbraco/techxot/ai-readiness/public/site-schema?id=…with the Delivery API key (fetchSiteSchemainemission.tsdoes exactly that) and emit the result as before.
The same identity is sent to the engine, so JSON-LD it generates names your organisation as publisher instead of guessing from the page title, and the audit credits the site-level types as provided by the platform when scoring HTML the editor posts from the workspace.
AI traffic — who crawls you, and who arrives from an assistant
The dashboard's AI traffic panel shows, for the last 7 or 30 days, visits referred by an AI
assistant (chatgpt.com, claude.ai, perplexity.ai, Copilot, Gemini…) with their landing pages, and
which AI crawlers (GPTBot, ClaudeBot, PerplexityBot, Google-Extended…) read what — HTML, your
Markdown, or llms.txt.
The data comes from a collector on the host that sees the traffic, posting to the engine in
batches. On a headless site that is your front end, not the CMS — which is why middleware-based
crawler analytics on a headless CMS show nothing. Copy headless/aeo-hits-collector.ts into
your front end and call recordHit(request) from your middleware; set INMOTION_API_KEY (and
optionally INMOTION_BASE_URL) in the front end's environment, using a key issued for that
environment.
What is sent: timestamp, path, host, User-Agent, Referer host, utm_source, country. What is
never sent: the visitor's IP, cookies, other query strings. The engine keeps only bot hits, AI
referrals and unknown automated agents; ordinary page views are dropped on arrival. Reporting does
not count against your audit quota.
Security — both paths
The property is a free-text textarea, so it is untrusted regardless of who has edit rights.
Anyone able to edit content could otherwise close the tag with </script> and inject executable
markup — stored XSS. Both reference implementations parse before emitting, reject anything that is
not an object or array, escape <, and log rejections rather than dropping them silently:
| Path | Implementation |
|---|---|
| Razor | StructuredDataHtml.Serialize / .ScriptTag (in this package) |
| Headless | headless/structured-data.ts (ships in this package) |
Both are unit-tested over the same cases, including a </script> breakout attempt. If you write
your own, test it against those cases too.
Verifying the install
Work down the list. Each row tells you which step failed if it does.
| Check | Expected | If it fails |
|---|---|---|
| AI Readiness tab on a document | Present | Step 1 — package not restored |
| AI Readiness dashboard in Content | Present | Step 1 |
| Open the tab on a published page, run an audit | A score, 7 levers, a ranked fix list | Step 2 — check BaseUrl and ApiKey |
| The tab on a page whose type has no matching property | Names the aliases it looked for | Step 3 — expected until you add the property |
| Generate JSON-LD, save, publish | Value stored on the page | Step 3 |
curl the published page, grep application/ld+json |
Present, and the + is literal — not + |
Rendering step; for headless suspect the cache first |
Set ScoreOnPublish.Enabled: true, publish |
A row appears within ~10 seconds | Step 2, and check the log |
| Run a site scan from the dashboard | Every published page scored, weakest first | — |
| Open a scored page's recorded URL | It resolves, and is the page a visitor sees | Step 4 — on headless, SiteUrlOverride is unset or points at the CMS |
| Scores are wrong but nothing errors | Real, varied scores — not the same low score everywhere | Step 4 — an identical score on unrelated pages means you are scoring one error page |
| Publish a parent with children | Only published pages are queued | — |
Reading the log
Scoring is a background worker, so the log is where it tells you things. Every line starts with
AEO, so filtering on that gives you the whole story.
| Line | Meaning |
|---|---|
'<page>' scored 51 (was 47) on Single |
Normal. The server role is recorded so scaled setups stay checkable |
AEO REGRESSION: '<page>' dropped 12 points |
The line the feature exists for |
skipped N of M entities in this publish |
Normal on a branch publish — unpublished descendants are not scored |
'<page>' has no routable URL, so it was not scored |
Page has no template or no published URL |
engine is rate limiting; waiting Ns |
Transient. It will recover on its own |
the InMotion key is over its quota |
Automatic scoring is paused until the stated reset |
the quota window has reset; automatic scoring has resumed |
Recovered |
scoring queue is full (500); dropped '<page>' |
A very large bulk publish. The next publish of that page records it |
AEO not configured (BaseUrl set: …, ApiKey set: …) |
Step 2. Usually an environment variable that died with its terminal |
Load balancing
Safe on scaled and Cloud plans. ContentPublishedNotification fires only on the instance that
performed the publish, so one publish produces one score and one history row regardless of instance
count — verified on Umbraco 18.0.2 across two instances sharing a database. Each score records the
server role that produced it, so this stays checkable with a query.
There is deliberately no IServerRoleAccessor gate. A publish request can land on any instance
and the work queue is per-instance, so suppressing the listener on subscribers would silently drop
every score for publishes that happened to land on one.
Licence
MIT.
| Product | Versions Compatible and additional computed target framework versions. |
|---|---|
| .NET | net10.0 is compatible. net10.0-android was computed. net10.0-browser was computed. net10.0-ios was computed. net10.0-maccatalyst was computed. net10.0-macos was computed. net10.0-tvos was computed. net10.0-windows was computed. |
-
net10.0
- Umbraco.Cms.Web.Common (>= 18.0.2 && < 19.0.0)
NuGet packages
This package is not used by any NuGet packages.
GitHub repositories
This package is not used by any popular GitHub repositories.
| Version | Downloads | Last Updated |
|---|---|---|
| 18.2.4 | 149 | 8/23/2026 |
| 18.2.3 | 106 | 8/23/2026 |
| 18.2.2 | 102 | 8/23/2026 |
| 18.2.1 | 117 | 8/22/2026 |
| 18.2.0 | 110 | 8/22/2026 |
| 18.1.0 | 106 | 8/21/2026 |
| 18.0.1 | 101 | 8/18/2026 |
| 18.0.0 | 110 | 8/18/2026 |
| 17.2.4 | 116 | 8/23/2026 |
| 17.2.3 | 111 | 8/23/2026 |
| 17.2.2 | 107 | 8/23/2026 |
| 17.2.1 | 114 | 8/22/2026 |
| 17.2.0 | 109 | 8/22/2026 |
| 17.1.0 | 103 | 8/21/2026 |
| 17.0.1 | 94 | 8/18/2026 |
| 17.0.0 | 115 | 8/18/2026 |
x.1.0: generates llms.txt, llms-full.txt and a Markdown rendering of every page from content,
with a Link rel=alternate header on HTML pages. Served at the root on Razor sites
(Emission:ServeAtRoot); fetched by your front end through a Delivery-API-key-guarded endpoint
on headless sites - headless/emission.ts ships in the package. New Site signals strip on the
dashboard: which AI crawlers robots.txt lets in, llms.txt health, sitemap. ExcludedContentTypes
keeps settings nodes out of the ranking, the scan and llms.txt.
Scores pages against seven weighted AI-readiness levers, records a score on every publish so a
regression is visible, generates JSON-LD from a page's own content, and ranks a whole site
weakest-first.
Supports Razor and headless sites. On a headless site set SiteUrlOverride to your front-end
host, or scoring will fetch and score the CMS's 404 page - see the README, step 4.
Requires an InMotion subscription for scoring. Package major tracks the Umbraco major: 17.x for
Umbraco 17, 18.x for Umbraco 18.