Not a survey. Not estimates. Eleven real businesses, audited one at a time between 9 and 28 July 2026 — what AI engines could actually find, read and repeat about each one.
The denominators differ because not every check was possible on every site. One business's website was offline when we audited it, and on two others we could not confirm the robots.txt state. We are reporting what we verified, not what we assumed. Where a number is out of 8 or out of 7, that is how many sites we could actually check that item on.
An AI engine answers "who should I hire for X" in three steps: it decides what the question is about, it retrieves sources, and it summarises across them. A business gets recommended when it appears in the sources the engine retrieves. So we checked the things that decide whether a website can be a source at all.
| Check | What it decides | Result across the sample |
|---|---|---|
| robots.txt | Whether AI crawlers are allowed to fetch the site | 2 of 7 were blocking citation crawlers |
| llms.txt | A plain-text summary of what the business does | 7 of 8 had none |
| Structured data | Whether facts are machine-readable, not just prose | 3 of 11 had none at all; most others were partial |
| Quotable answers (FAQ) | Whether the site answers buyer questions in a form AI can lift | 7 of 11 had no FAQ schema anywhere |
| AI visibility index | Mentions and cited pages across four engines | 4 of 11 scored zero |
Two of the seven sites whose robots.txt we could read were blocking the crawlers that allow an AI to cite them. In both cases it looked deliberate and was almost certainly inherited — a blanket "keep AI off our content" rule added before anyone distinguished between the two kinds of crawler.
They are not the same thing. GPTBot collects data to train models — blocking it keeps your content out of training and costs you nothing in visibility. OAI-SearchBot and ChatGPT-User fetch a page so ChatGPT can cite it to a customer. Blocking those removes the business from the answer.
One site in our sample went further and blocked every crawler except Googlebot. Its owner had no idea; the rule had been in place for years and predated ChatGPT having a search index at all.
The most striking result in the sample was not a broken website. It was a very good one.
One business had built 128 pages, each answering a single, narrow buyer question — the kind of architecture consultants spend months on. AI engines had cited its pages 76 times. They had named the company 26 times. On ChatGPT the ratio was nearly ten to one: nineteen pages read, the company named twice.
We then asked ChatGPT the question its buyers actually ask. It named four competitors, put this business fourth, and left it out of the closing recommendation entirely. The sources it cited were a competitor's own website, an industry directory and a marketplace.
The content was being used as a reference library. The recommendation went somewhere else. Crawlability gets you read. Structured data, third-party sources and reviews get you named.
Four of the eleven had zero cited pages across all four engines. That does not mean AI had never heard of them — two of the four were mentioned in answers. It means every word an AI had ever said about them came from somewhere else: a map listing, a directory entry, a forum thread, a review count.
In one case a shop's own homepage carried the exact claim that won the answer for its competitor — a fast turnaround promise, stated plainly on the front page. The competitor's version was read off their website and quoted. This shop's was never opened, so it never entered the answer. The two finished a tenth of a star apart.
Notably, llms.txt is on that list but not at the top. It is cheap, harmless and worth publishing, but no engine will recommend a business because it has one. The crawler settings and the cited sources are what move the needle.
Eleven businesses, audited individually between 9 and 28 July 2026. They were not selected for this study — each had requested a free AI visibility audit, and this analysis aggregates what those audits found. The sample spans insurance, e-commerce, marine repair, phone repair, scaffolding software, healthcare, consultancy and professional services, across the United States, United Kingdom and Australia.
For each site we fetched robots.txt, llms.txt and sitemap.xml directly, inspected rendered HTML and JSON-LD structured data, pulled AI visibility figures from Semrush's AI Search index, and ran live buyer-intent queries against ChatGPT in a clean session, recording which businesses were named and which sources were cited.
No business is identified in this page, and no figures from any individual audit are attributed. The audits themselves are private to the businesses that requested them.
No. GPTBot is a training crawler. Blocking it keeps your content out of model training and costs you nothing in whether ChatGPT recommends you. It is OAI-SearchBot and ChatGPT-User you must allow — those are the ones that fetch a page so it can be cited.
It is cheap and harmless, so publish one. But in our sample it was not the thing separating visible businesses from invisible ones. Crawler access, structured data and third-party sources mattered more.
Not automatically. One business in our sample was strong on Google's AI surfaces — hundreds of mentions — while sitting near zero on ChatGPT, because its robots.txt allowed Googlebot and blocked the rest. Google visibility and AI visibility are related but they are not the same channel.
Open yoursite.com/robots.txt and look for OAI-SearchBot, ChatGPT-User, PerplexityBot or ClaudeBot with a Disallow: / under them. Then ask ChatGPT the question your customers ask, and follow it with "what sources did you use for that?" If your own website is not in the list, that is the finding.
We run the same audit on request — five findings, the evidence behind each, and a 90-day fix plan. No charge, and the findings are yours to keep either way.