Auditing a site for AI visibility, not just SEO
Answer engines read a different set of signals than Google. Here is what we check, and why llms.txt is the least interesting part of it.
Every agency has a story about the client who ranked well on Google and still lost enquiries. There is a new version of that story now: the business that ranks fine and never gets cited by an assistant. When somebody asks ChatGPT or Perplexity for "a good emergency plumber in Leeds", the answer is assembled from what those systems can fetch, parse and trust, and that is a measurably different set of signals from the ones a classic SEO audit looks at.
Access comes first, and it is where most sites fail
Before any of the content questions matter, an answer engine has to be allowed in. A surprising number of sites block the very crawlers they would benefit from, sometimes because a security plugin added a blanket rule, sometimes because somebody pasted a robots.txt from a site that had made a deliberate choice to opt out.
So the first thing we check is plain: is GPTBot allowed? ClaudeBot? PerplexityBot? Google-Extended? A site can be immaculate in every other respect and be invisible because of four lines in a text file. That is also the easiest finding in the world to convert into a conversation, because it is unambiguous and it takes five minutes to fix.
llms.txt is the least interesting part
llms.txt gets the attention because it is new and easy to write about. We do check for it, but honestly, it is a weak signal. It is a proposal, adoption is uneven, and no assistant currently treats it as load-bearing. Selling a client on an llms.txt file as an AI-visibility strategy is selling them a deliverable, not an outcome.
What actually moves the needle is duller: structured data that is present and valid, an entity the model can resolve (this business, this location, this service, stated unambiguously in text), content that answers questions in the form people ask them, and consistency between what the site claims and what third parties say about it.
Answerability is a content property, not a markup property
Assistants quote passages. A page that buries "we open at 7am on Saturdays" inside a paragraph about the team's values is harder to quote than a page with a heading that asks the question and a sentence that answers it. This is why FAQ content keeps outperforming its reputation, not because FAQPage schema is magic, but because the format forces a question-and-answer shape that is trivially extractable.
What we score, and what we deliberately do not
Our AI-visibility category combines crawler access, structured-data presence and validity, entity definition, answerable content patterns, and citation signals. What it does not do is claim to measure your actual share of AI answers. Nobody can measure that reliably at the moment: the outputs are non-deterministic, personalised and unlogged, and a score that pretends otherwise would be the kind of number that falls apart the first time a client checks it.
That distinction matters commercially. "Your site blocks three of the four major AI crawlers, here are the lines that do it" is a claim you can defend in a meeting. "Your AI visibility score is 42" is not, unless you can say exactly what produced the 42, which is why every finding in our reports carries its measurement, and why the methodology is public.
Written by
SiteAssay