// AGENT SEO · FIRST-PARTY DATA

LLM SEO: what the checklist did for us, in numbers

LLM SEO is shaping a site so large language models — ChatGPT, Perplexity, Google AI Overviews, Copilot — can retrieve it and cite it. We ran the technical half of the standard checklist on clize.ai and deliberately skipped the half that needs other people: JSON-LD on every page, an llms.txt, a sitemap with hreflang, visible dates, a robots.txt that names eleven AI crawlers one by one — and not a single outreach email sent, not one directory pull request submitted. Here is the reading in that state, from Google Search Console for 2026-08-04 to 08-31, read on 2026-09-04: 1,415 impressions, 0 clicks, and 0 visits from any AI-engine referrer. Below: the numbers, what they do and do not prove, and the commands to get the same reading for your own site.

First-party data1,415 impressions · 0 clicks0 AI referralsRead 2026-09-04

What LLM SEO actually is

Two routes get your writing into an AI answer, and they behave nothing alike. The first is the training corpus: text that was already on the open web when a model was trained, which you cannot edit after the fact and cannot measure at all. The second is live retrieval — the search index the assistant queries while it answers you, which for ChatGPT and Copilot leans on Bing, for AI Overviews on Google, and for Perplexity on its own crawl plus partners. Only the second route responds to anything you do this month.

That split is why LLM SEO reads as a restatement of ordinary SEO with new vocabulary. It largely is. If the retrieval layer is a search index, then being retrievable means being crawlable, being indexed, and ranking well enough on the underlying query to make the shortlist the model reads. The genuinely new parts are narrow: answers get assembled from passages rather than pages, so a paragraph that stands alone travels further than one that depends on the three above it; a citation is a mention with a link attached, so being written about matters more than being linked to; and the click is optional, which means your analytics will understate the whole thing.

Everything downstream of that follows. The checklists you find for this term are, almost without exception, the same eight to ten actions in a different order. The question nobody answers is what happens after you do them, which is the only question this page is about.

The eight things every guide tells you to do

The page ranking second for llm seo as we write this is a 3,026-word guide built on eight steps. The one ranking third is a 4,103-word vendor guide with seven best practices and four advanced ones. Strip the wording and they are the same list, and it is a reasonable list. What none of them carries is a number produced by following it.

Here it is, with the only column that matters added: can you verify the step was done, without trusting anyone?

The stepVerifiable by you?Did we do it?
Set up Bing Webmaster ToolsYes — you either have the account or you do notNo. Never verified the property
Put schema markup on the pages that matterYes — view source, or run any validatorYes, on every page
Write for Bing as well as GooglePartly — indexation is checkable, intent is notPartly. IndexNow key is live; the Bing account is not
Answer real questions in plain, liftable proseNo — this is a judgement callWe think so, which is worth nothing
Keep publish and update dates current and visibleYes — the date is on the page or it is notYes, visible date plus dateModified
Do not use AI to generate the contentNo — unfalsifiable from the outsideWe cannot claim this one. See below
Earn mentions on other sites and publicationsYes — a link exists or it does notNo. Zero sent.
Grow branded search volumeYes — Search Console shows brand queriesNo deliberate work

Three of those eight cannot be audited by anyone, including the person who wrote the advice. Five can. We did four of the five and skipped the one that requires talking to strangers — which, as the numbers below suggest, may be the one that was carrying the weight.

The sixth row deserves its own sentence, because the honest answer is uncomfortable. This site is written by an AI agent under human review. We cannot tell you we followed the rule against machine-written content, so we do not. If that step is the load-bearing one, this whole page is a report from the wrong side of it — and you should read the numbers with that in mind rather than take our word for anything.

What we actually shipped, and what we never sent

The technical column is not a claim, it is a set of files you can open right now.

  • Structured data on every page. Each page carries JSON-LD — TechArticle or WebApplication, plus BreadcrumbList, plus FAQPage where there is an FAQ. The FAQ blocks are generated from one array, so the visible questions and the markup are the same strings by construction and cannot drift apart.
  • An llms.txt. clize.ai/llms.txt is generated from the same manifest that generates the sitemap, so an English page cannot ship without entering the LLM-facing index. Whether that file does anything is a separate question we answer honestly on the llms.txt page: no engine has confirmed it as a ranking input.
  • A sitemap with hreflang. Generated from the same manifest, one entry per public page, each carrying its full language cluster rather than a lone self-reference.
  • A robots.txt that names the crawlers. Twelve user-agent blocks: the wildcard plus eleven named AI and answer-engine agents — GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, anthropic-ai, Claude-Web, PerplexityBot, Perplexity-User, Google-Extended, Applebot-Extended, CCBot — each explicitly allowed. Intent stated, not merely implied by the wildcard.
  • IndexNow. The key file is live and returns 200, so Bing and its partners get pinged on deploy.

Now the other column, and this is the part most pages like this one leave out. Our own placement queue lists the outreach: a pitch letter to the one independent publisher ranking on our main cluster, drafted 2026-09-02, never sent. A pull request adding us to a 16k-star awesome list, patch prepared against a named upstream commit, never submitted. Two directory claims, not done. Bing Webmaster Tools, never verified. Indexing requests for the seven tool pages published on 2026-09-02, not submitted — our Search Console service account has restricted permissions and cannot request indexing through the API, so that one waits on a human clicking a button.

So the experiment, unintentionally, is clean: on-page maximal, off-page zero. That is not a strategy we recommend. It is the state the numbers below describe.

The reading in that state

Everything here comes from Google Search Console for the 28 days from 2026-08-04 to 2026-08-31, read on 2026-09-04, plus referrer data from our own analytics for the trailing 30 days. Nothing is modelled, projected or averaged across sites.

ReadingValue
Impressions1,415
Clicks0
Visits from any AI engine0
New queries discovered this windowNone
Biggest page/best-mcp-servers/ — 866 impressions, average position 70.6, 0 clicks (61% of all impressions)
The next five pageshome 142 @ 10.6 · /fr/mcp-server/ 59 @ 69.7 · /es/mcp-server/ 58 @ 64.0 · /ai-agent-email/ 54 @ 8.8 · /seo/geo-tools/ 45 @ 33.4
Top four non-brand queries219 @ 68.7 · 167 @ 68.8 · 156 @ 72.0 · 54 @ 66.4 — four wordings of one cluster, all 0 clicks
Pages published 2026-09-02All still reported as URL is unknown to Google at the 09-04 read
Trailing 30-day pageviews (sampled)2,290 total: 1,690 direct, ~200 from Bing, ~100 from Google, 0 from AI engines

Read the second row before anything else. Zero clicks is not the interesting number; average position 70 is. Position 70 is the seventh page of results. Nobody has ever been to the seventh page of results. The zero is arithmetic, not mystery — and any guide that promises a checklist will fix visibility owes you a sentence about what average position it expects you to reach, because below roughly the top ten the click-through rate of a perfect page and a terrible page are the same number: zero.

The shape is worth naming, because it decides what to do next. Google has clearly associated our pages with the queries — 1,415 impressions means the association exists, 219 of them on a single wording. What it withholds is the position. More pages on the same topic do not move that; it is not a content deficit. It is the missing half of the checklist showing up in the data exactly where the theory said it would.

What this proves, and what it does not

A single site is one data point, and a data point with an obvious confounder is worth less than a good argument. So, precisely:

What it supports. Doing the technical half of an LLM SEO checklist thoroughly, on a site with almost no inbound links, produced impressions and produced nothing else in this window. Markup, an llms.txt, a clean sitemap and an explicitly welcoming robots.txt did not, on their own, buy position, clicks, or a single AI referral. Anyone selling the technical half as sufficient is selling something our numbers do not support.

What it does not support. It does not show the checklist is wrong. Every item on it is cheap and none of it hurt. It does not show llms.txt is useless — you cannot prove that from a site that never ranked in the first place. And it says nothing whatsoever about the seven tool pages we published on 2026-09-02, because this window closed on 08-31, before they existed. Their first honest reading arrives in the window that starts after they were crawled, and we will publish it here with the same discipline: a window, a read date, and the number as it came out.

The confounder, stated plainly. This site is young and has almost no links. You could argue the whole result is authority, not LLM SEO, and you would be largely right — that is the point. The checklist is priced as though it were the variable, and on this site it was not the variable. Two of the eight items — earn mentions, grow brand search — are the ones that move authority, and those are the two we did not do. The reading is what "on-page complete, off-page zero" looks like from the inside.

What we will not do is publish a follow-up in six weeks claiming the numbers went up because of one thing we changed. With one site, no control, and Google shipping changes we cannot see, that claim is not available to us. It is not available to the guides either; they just make it anyway.

How to get this reading for your own site

The measurements above are not a product feature, they are Search Console and referrer logs. You can reproduce all of it for your own domain in a couple of minutes, and you should trust your own numbers more than ours.

$ npm i -g @clize/clize && clize login
$ clize seo check --domain example.com

The first run explains what it is missing. To read Search Console it needs a property: you add a service-account address as a Restricted user in Search Console — the command prints the exact address — and re-run. There is no OAuth consent screen and no review; restricted permission is enough to read search analytics. Pass --gsc none if you would rather it never asks again.

What comes back is four faces of the same site, printed as facts with the raw figures attached rather than as advice:

  • Queries. Impressions, clicks and impression-weighted average position for this window against the previous one, with the demand-side volume for the same word beside it — so "what the keyword tool claims" and "what Google actually gave you" sit on one line.
  • Pages. Your biggest pages with the previous window for comparison, movers whose impressions halved or doubled, a live HTTP status for each, and Google's own index verdict from URL Inspection, quoted rather than paraphrased: Submitted and indexed, Crawled — currently not indexed, URL is unknown to Google. That last string is why we know our new pages are not in the index rather than merely unranked. The two are not the same problem and no amount of writing fixes the first.
  • Unseen pages. URLs in your sitemap with zero impressions this window, newest first. On a site that just shipped a batch, this is the batch.
  • Traffic by source. Referrers split into search engines, AI engines and direct — AI listed separately because that is the number this whole field is about.

Positions and impressions come from Search Console and cost nothing. The only part of the command that costs money is pricing new keywords against a paid metrics source, and those metrics are cached for a month, so re-running to watch a trend is free. If you want to know what a query's results page actually looks like — who occupies it, whether there is an AI Overview on it — that is a separate, explicit call: clize seo serp <keyword>. Every charge is itemised by clize seo spend, so nobody has to hand-tally anything.

What we could not verify

A page about measurement should be exact about where its instruments stop.

AI referrals undercount, by design. Our AI figure is referrer-based: a visit arrives with chatgpt.com in the referrer, we count it. The list we match against has seventeen domains, including chatgpt.com, chat.openai.com, perplexity.ai, gemini.google.com, copilot.microsoft.com, claude.ai, you.com, phind.com, grok.com and several non-Western engines. Everything that method misses stays missed: an answer that cites you without the reader clicking; a desktop app or mobile client that sends no referrer; an engine not on the list; a reader who reads your paragraph inside the answer and never comes at all. So 0 AI referrals means no measured click-throughs. It does not mean nobody was ever cited. We cannot tell the difference, and neither can anyone else without a monitoring product — which we do not sell.

Our own analytics are sampled. The traffic figures come from adaptive sampling: the true count is the sampled rows times the average sample interval, which is why the numbers are round. They are for comparing against themselves over time, never for quoting as exact counts. Same-host referrers are internal navigation and are excluded, which sounds obvious and is the single biggest reason naive referrer tables are wrong.

Two instruments, one disagreement. Search Console reports 0 clicks for the window. Our referrer data reports roughly 100 pageviews with a Google referrer over 30 days. Both readings are ours, both are honest, and they do not reconcile. The windows differ, the sampling is coarse, and a Google referrer is not only web search. We are not going to resolve it with a paragraph — we are telling you it is there, because a page that showed you only the number that suited its argument would deserve exactly as much trust as the guides that show you no numbers at all.

Window shift matters more than you would think. Read on 2026-09-02, the 28 days ending 08-30 gave 1,342 impressions. Read on 2026-09-04, the 28 days ending 08-31 gave 1,415. Two overlapping windows, two days apart, 73 impressions of difference. Neither is wrong. This is why every figure on this page carries a window and a read date, and why a screenshot with no dates on it is not evidence.

A GEO checklist you can finish today

Most published GEO checklists are either gated behind an email form or written so abstractly that no item has a finish line. Here is the short one, where every line is a thing you can complete in an afternoon and then verify with something free.

  1. Confirm the AI crawlers are allowed. Open your robots.txt and look for GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Google-Extended. A blanket Allow: / covers them, but naming them states intent and survives the next person who edits the file.
  2. Confirm your sitemap parses. Not "exists" — parses, with valid lastmod dates and no oversized file. Paste it into the sitemap validator; it runs in the page and uploads nothing.
  3. Confirm your structured data matches what a reader sees. Markup that claims content the page does not show is the fastest way to lose rich results. If it is an FAQ, generate both halves from one source with the FAQ schema generator, which emits the JSON-LD and the matching visible HTML together.
  4. Confirm your language versions point at each other. One-way hreflang annotations are discarded silently. The hreflang tag generator builds the reciprocal set with x-default; we found 164 one-way pairs on our own site this way, which is how we know the failure is common.
  5. Check your llms.txt if you have one. The llms.txt validator checks it against the published spec, line by line. Treat it as cheap hygiene, not as a ranking lever — see the honest version of that argument on our llms.txt page.
  6. Then get your baseline. Run clize seo check once and write down three numbers: impressions, average position, AI referrals. Today's numbers are worthless. The comparison in six weeks is the whole point.

Notice what is not on this list: word count targets, entity density, prompt-stuffing, and every other proxy metric that exists because the real one is hard to measure. If an item cannot be finished and verified, it is not a checklist item, it is a mood.

ChatGPT SEO, specifically

Searches for chatgpt seo usually mean one of two different things, and mixing them up wastes a lot of effort.

The first is using ChatGPT as an SEO assistant — drafting briefs, clustering keywords, writing outlines. That is a workflow question, and we sort the tools for it by whether a human or a program can drive them on the AI SEO tools page.

The second is getting ChatGPT to cite your page, which is a retrieval question. When ChatGPT browses, it resolves queries against a search index rather than recalling your site from memory, and that index has historically leaned on Bing. The practical consequences are unglamorous: be indexed in Bing, not only Google; allow OAI-SearchBot and ChatGPT-User in robots.txt, which are the fetchers that matter for browsing and are distinct from GPTBot; and keep answers in self-contained passages, because what gets lifted is a passage, not a page.

Then check the result the only way available to you: ask the questions your customers ask, in a fresh session with no memory or personalisation, and write down whether you appear and who appears instead. Doing that for fifteen minutes every two weeks is a worse instrument than a monitoring dashboard and a much better one than nothing — and it is the instrument we actually use, because we decided a scheduled bot querying assistants on a cron was producing motion rather than decisions.

How to get cited: the part we skipped

Here is our best guess at what would move our own numbers, offered as a hypothesis rather than a lesson, because we have not run it yet.

Assistants cite what the underlying index ranks, and at the top of most commercial queries that means independent roundups, forum threads, documentation sites and curated lists — not vendor pages. Our own results pages say so plainly: on the cluster carrying 61% of our impressions, the top result is a forum thread. A vendor page competing with a forum thread on "best X" loses, because the reader asked for a comparison and one of those two is not in a position to give one.

Which makes the move obvious and unautomatable: be in other people's lists. Pitch the independent publisher who already ranks. Open the pull request against the curated repository your category lives in. Answer the actual question in the actual thread, as a participant. Get the directory entry claimed so the listing is accurate.

All four of those are sitting in our queue, drafted, addressed, and unsent — which is why this page can honestly say what a checklist with the last two items missing looks like in Search Console, and cannot yet tell you what it looks like with them done. When it does, the number will appear here with its window and its read date, whichever direction it goes.

If you want the same discipline applied to a query before you write anything: clize seo serp "your query" returns who currently occupies the results page, whether an AI Overview sits on top of it, and which of those occupants accept contributions — which is a faster way to find out whether a page is even the right instrument than writing one and waiting a month.

// FAQ

What is LLM SEO?

LLM SEO is optimising a site so large language models retrieve and cite it when they answer a question. In practice it splits in two: the training corpus, which you cannot edit or measure after the fact, and live retrieval, where the assistant queries a search index while answering. Only live retrieval responds to work you do this month, which is why most LLM SEO advice is ordinary technical SEO plus passage-level writing plus being mentioned on other sites.

Does LLM SEO actually work?

Our own reading says the technical half alone did not, on a site with no links. Doing schema, an llms.txt, a clean sitemap and an explicitly crawler-friendly robots.txt produced 1,415 Search Console impressions and 0 clicks in the window 2026-08-04 to 08-31, with an average position around 70 on the biggest page, and 0 measured visits from any AI engine. That is one site with an obvious confounder, so treat it as a data point rather than a verdict.

How is LLM SEO different from GEO or AEO?

Mostly in vocabulary. GEO (generative engine optimization), AEO (answer engine optimization) and LLM SEO describe the same work: be crawlable, be indexed, rank well enough to be in the shortlist a model reads, write passages that stand alone, and be mentioned by sources the index already trusts. The distinctions matter for naming a service, not for deciding what to do on Monday.

Is llms.txt worth adding?

It is cheap, so probably, but do not expect it to rank you. No search or answer engine has confirmed llms.txt as an input, and our own site has had one throughout a window that produced zero AI referrals — which proves nothing on its own, because a site that never ranked cannot demonstrate the file is useless either. Add it as hygiene and keep your expectations where the evidence is.

How do I measure whether AI engines send me traffic?

Split your referrers and look for the AI engine hosts specifically: chatgpt.com, chat.openai.com, perplexity.ai, gemini.google.com, copilot.microsoft.com, claude.ai and the rest. Clize does this in one command with clize seo check, which reports traffic by source with AI engines listed separately alongside your Search Console numbers. Remember the method undercounts: citations without a click, clients that send no referrer, and engines outside your list are all invisible to it.

How do I get cited by ChatGPT?

Be retrievable and be in other people's lists. Retrievable means indexed in Bing as well as Google, with OAI-SearchBot and ChatGPT-User allowed in robots.txt, and answers written as self-contained passages that survive being lifted out of context. Being in other people's lists means roundups, curated repositories, documentation and forum threads, because on commercial queries those are what the underlying index ranks above vendor pages. We have done the first half and none of the second, which is what this page reports.

Does Clize monitor whether AI answers mention my brand?

No, and we would rather say so than imply otherwise. Clize reads your own site: Search Console impressions, clicks, positions and index status, plus traffic by source with AI referrers listed separately, and it can tell you who occupies a given results page. It does not query assistants on your behalf or track your share of voice in AI answers. For that you need a monitoring product, and we do not sell one.

clize seo check — your own numbers, not ours

Get your baseline before you write anything.

One command reads your Search Console positions, your index status page by page, and your traffic split with AI engines listed separately. Today's numbers are worthless; the comparison in six weeks is the point.

$ npm i -g @clize/clize && clize login
$ clize seo check --domain example.com
[ Agent SEO → ]