Most GEO checklists start with crawlers. That is the wrong end of the problem, and it produces audits that are technically correct and commercially useless.
Here is the list I work through instead. It starts with the questions a company needs to appear for, and only then looks at whether systems can reach the pages, what they receive, how clearly the company is described, which sources outside your control are being cited, and how any of it gets measured. Technical access is necessary throughout. It cannot make generic information worth citing.
This is an audit, so it tells you what is true about the site. Deciding what to do about it is a separate job and a harder one.
One note on evidence. Google publishes far more about its own AI features than anyone else, so it is cited several times below. That documentation is authoritative for Google’s surfaces and for nothing else. ChatGPT, Claude and Perplexity retrieve content differently, say almost nothing publicly about how, and account for a growing share of the answers buyers see. Where Google has said something useful I have used it, and I have tried to be clear about where its answer stops.
1. What should this company be found for?
Before anything technical, there has to be an answer to this: which questions, from which people, at which point in their decision, does this company need to appear in?
- Collect the questions from sales and support as well as from keyword data. Keyword tools tell you what people search for and how much demand exists, which is real information you need. They are weaker at the long, conversational questions people put to a model, which often carry context a search box never had. Use both sources.
- Check which of those questions the site answers, and how honestly. Most sites cover the flattering questions thoroughly and skip the awkward ones. The awkward ones are where a buyer compares you to someone else.
- Check that the topics connect to something commercial. Visibility on questions that lead nowhere is easy to produce and easy to mistake for progress.
- Check whether the answer exists somewhere unpublished. It often does, in a sales deck, a spreadsheet, or someone’s head.
That last point is worth dwelling on. In one client project the information customers were looking for already existed, inside an interactive tool rather than on an indexable page. I identified the gap and the existing material was published as an ordinary web page. During the first 3 months that page recorded about 38,000 Google impressions, roughly 6,000 of them in AI Overviews. No new source material had to be produced. What that demonstrates is discoverability. It does not show that the visibility produced revenue, and I have written up the method and the limits separately. Read more about this GEO case.
Google’s own guidance puts this first too. Its guide to generative AI features says that creating unique, non-commodity content will likely influence your visibility more than anything else in the document (Google Search Central, May 2026).
2. Crawler access
- Read robots.txt. Look at the retrieval and user-triggered agents, not only the training crawlers. They do different jobs.
- Check the CDN or WAF settings. Cloudflare has blocked AI crawlers by default on new domains since July 2025, and from 15 September 2026 its defaults block training and agent crawlers on ad-carrying pages for new customers, new sites and all free-tier accounts (TechCrunch, July 2026). Plenty of sites now have a permissive robots.txt and a block one layer up.
- Read the logs at every layer. A request stopped at the CDN or WAF may never appear in the origin log at all. CDN, WAF and origin logs together show how far a request got and what response it received.
- For Google specifically, check eligibility. A page has to be indexed and eligible to show with a snippet, and the site has to be included in generative AI features in Search Console.
Crawler names and their functions change several times a year, so check each provider’s current documentation before you write a rule.
3. Rendering
- Fetch a page raw and search the HTML for a sentence you can see on screen. If it is not there, treat it as missing.
- Test the interactive elements rather than assuming. Content hidden behind an accordion or a tab is usually still in the HTML and perfectly readable. Content that is fetched when you click is not. The difference is in the markup, not in how it looks.
- Check whether facts live only in images. A number in a chart, a price in a graphic, a comparison in a screenshot.
A Vercel and MERJ log study published in December 2024 found that none of the major indexing crawlers executed JavaScript, including GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot and PerplexityBot (Vercel). Google’s generative features are the exception, because they run on the same infrastructure as Googlebot, which does render. A client-side rendered page can therefore perform in AI Overviews and return almost nothing to a crawler that does not render.
Two caveats on that study. It is a snapshot from late 2024 and worth re-testing on your own logs rather than assuming it still holds. And it covers indexing crawlers, not browser agents, which are a separate category: those can load a page in a browser and read the rendered DOM and the accessibility tree, which is why Google now documents them separately.
4. Content structure
Here is where a lot of GEO advice goes wrong. You will read that content should be written in self-contained chunks, shaped around how retrieval systems split documents. Google says the opposite in writing: there is no requirement to break content into small pieces, its systems handle multiple topics on a page, and you should make pages for your audience rather than for generative AI search. The same document says you do not need to rewrite content specifically for AI systems, because they understand synonyms and general meaning.
That is Google’s position on Google’s own surfaces, and no other provider has published anything equivalent. It is strong evidence against writing for a chunker, not proof of how ChatGPT or Perplexity assemble an answer. The reason I am comfortable following it anyway is that the alternative asks you to rewrite for a system nobody has documented, at a cost you can see, for a benefit nobody can demonstrate.
What is left is ordinary editorial work.
- Check that the answer comes early. Long build-up before the point is a habit from a time when the reader was already on the page and had nowhere else to go.
- Check the page is organised in a way a reader can follow. Clear sections, clear headings, paragraphs that hold one idea. Google’s guidance asks for exactly this and nothing more exotic.
- Check page scope. One page covering nine loosely related topics serves nobody. There is no ideal length, but there is usually an obvious point where a page is trying to be two pages.
- Check that comparable information is in real tables and lists, as markup rather than as a picture.
- Check the information architecture. Whether the structure follows how customers think about the problem or how the company is organised internally. It is very often the second one.
5. Entity clarity
- Check there is one canonical description of the company, used consistently, rather than five different positioning statements across five templates.
- Check naming consistency across the site, the structured data and the profiles you control elsewhere.
- Check the links outward to LinkedIn, Wikidata and industry registries, which let a system connect your site to an entity it already knows.
- Ask the models directly. Run “what is X”, “what does X sell” and “how much does X cost” in the tools your audience uses, and record the answers.
People react to that last check faster than to anything else on the list, because the finding is usually not invisibility. It is being described inaccurately. A model that recommends you warmly while getting your pricing wrong sends people to your sales team already misinformed.
6. Structured data
- Check the core types exist and validate: Organization, Person for authors, Article, Product where relevant.
- Check the markup matches what is visible on the page.
Be honest about the ceiling here. Google states plainly that structured data is not required for generative AI search and that there is no special markup to add, while still recommending it as part of normal SEO because it makes you eligible for rich results. The same guidance says Google Search ignores llms.txt entirely. I have not found public confirmation from any of the major providers that either one is used as a retrieval input, and none of them has ruled it out either. That is the honest state of the evidence.
Costs differ more than the advice usually admits. llms.txt takes twenty minutes. Correct, maintained structured data across a large site is a real engineering commitment. Neither is a GEO strategy.
7. Sources and authorship
- Check bylines. Named people with relevant expertise, linked to their profiles.
- Check that claims have sources, and that original data comes with its method and its limits. That is what makes a number usable by someone else.
- Look at the sources you do not own. Run your brand questions and record which domains get cited: comparison sites, trade press, forums, video, community threads.
That last check often moves the whole project somewhere else. An answer will cite a trade publication, a review site or a forum thread as readily as your own pages, and which kinds of source a system reaches for differs between tools in ways nobody has documented. You find out by running your questions in each of them and writing down what comes back. Where visibility is being decided is frequently somewhere the marketing team has no edit access.
One caution before anyone acts on that. Google specifically warns against chasing inauthentic mentions. This is an argument for real PR, community and partnership work, not for buying placements.
8. Internal linking
- Check for orphan pages and pages buried too deep to be reached from anywhere people actually land.
- Check that navigation uses real links rather than interactions, search boxes or filters.
- Check whether related pages are connected. Twelve articles on one subject that never link to each other read as twelve unrelated pages, and feel like a dead end to someone who wanted more.
9. Measurement
Set this up before making changes, so there is something to compare against.
- Use the Generative AI performance report in Search Console. Google added it in 2026 and it covers AI Overviews and AI Mode. It is the only platform-provided impression data currently available to site owners, and it covers Google alone.
- Fix a question set for everything else. 20 to 50 real questions, the ones from section 1, run on a schedule. Changing the questions between rounds means comparing two different things.
- Repeat each question. AI answers vary between runs on the same day, so one check is an anecdote. How many repeats you need depends on the variance you see, which is the part of the method that matters more than the tool.
- Track share of voice, mentions, citations, and factual accuracy. The last one is the one that costs money when it is wrong.
- Check referral tracking and consent settings. Referrals from AI tools are real but small, and a consent banner can reduce them to something that looks like zero.
- Agree what success means before starting. If traffic is the only scoreboard, a working programme will look like a failure.
Be sceptical of tools that claim access to internal ranking signals. Google states directly that no third-party tool has access to its internal ranking or AI systems. It can only speak for itself, but the same question is worth putting to any vendor selling numbers about the other platforms.
What this list does not cover
It does not tell you whether your content deserves to be cited. Every item above is about access, clarity and measurement. Those are necessary and none of them is sufficient. A perfectly accessible page that says what ten thousand other pages already say will not be used, and no technical work changes that.
It also does not decide what to fix. That depends on the company, the category and what the measurement says once it exists.
This is what I check in September 2026. Some of it will be out of date within a year, including at least one thing I was confident about six months ago.
I do this as an engagement: an audit of where a company stands in AI answers, plus a prioritised plan for the next three to six months. The list above is the audit half, in public. If you want to discuss what that could look like for your company, get in touch.
If you work through it and find something missing, tell me. That is how this version gets to a second one.
Read more:
Can AI Read Your Website? Why Rendering Matters for AI Search
How I Track AI Search Visibility Without Paid Tools: A Repeatable Method
