pilcrow
← All research Pilly stands on a stool and slides one glowing block into the single empty gap on a long shelf of blocks, with two AI company symbols beside the shelf.

If you opened after 2023, half the models don't know you exist

Five of the ten systems people ask for recommendations answer from memory alone. Their memory was written before you existed. Here is the sixty-second test, and the half of the market you can actually buy.

vendor documentation + our own runs · August 2026 · August 6, 2026 · 5 min read


There are two kinds of AI system taking questions about your industry right now, and the difference between them decides whether a business founded in 2024 is invisible or merely unranked.

The first kind searches the web before it answers. You ask it for the best roofer in Tulsa, it runs a search, reads some pages, writes a paragraph. If your website exists and is readable, you are eligible.

The second kind does not search. It answers from what it absorbed during training, and training finished at a fixed date. You ask the same question and it names four companies it learned about years ago. If you were not in that corpus, you do not exist to it. No amount of publishing this week changes what it already learned.

Which is which, as of August 2026:

Searches the web before answeringAnswers from training only
ChatGPT (search enabled)Gemini (in the API, without grounding)
PerplexityDeepSeek
Google AI OverviewsMistral
Google AI ModeQwen
Claude (with web search on)Most self-hosted open models

The split is not stable and it is not clean. Gemini in the consumer app grounds against Search; Gemini called through the API without grounding does not. ChatGPT will answer from memory alone if it judges a search unnecessary — which it frequently does for questions it thinks it already knows. Treat the table as two behaviours, not two products.

Why this hits recent businesses hardest.

An analysis published in June 2026, covering 91 brands across three categories, found the training-only half of the market accounted for between 35% and 65% of total visibility share depending on the category. Newer entrants in that dataset scored almost entirely on the web-search side: their numbers on training-only models rounded to nothing.

That is a structural disadvantage and it has nothing to do with quality. A company that opened in 2019 and stopped trying is inside the weights. A company that opened in 2024 and is better at the job is not. The models are not ranking you badly. They are not ranking you at all.

The sixty-second test

Do this before you believe anyone, including us.

One. Open ChatGPT and turn web search off. Ask: “Who are the best [your category] in [your city]?” Write down the names.

Two. Turn web search on. Ask the identical question. Write down the names.

Three. Compare the two lists.

There are three possible outcomes and each one means something different.

You appear in both. You are in the training data and you are findable live. This is the strong position and it is rarer than you’d expect for a business under five years old. Your work is defensive: stay in the answer.

You appear only with search on. This is the common case for anything founded after 2023, and it is the one worth understanding properly. You are reachable — but only by the half of the market that looks. Every improvement you make lands on that half and lands quickly, sometimes within days of a page being crawled. Nothing you do lands on the other half until the next training cycle, and nobody can tell you when that is or guarantee you’re in it.

You appear in neither. Then the question isn’t training data, it’s the first gate. Something is stopping the search-enabled models from reading you at all, and that is fixable this week. Run the check in blocking the wrong robot before you conclude anything about AI having an opinion of you.

What this means for what you buy

Here is the part that most articles on this subject skip, because it is inconvenient for the people selling.

You cannot buy your way into a training set. There is no submission form, no partner programme, no vendor with a relationship. The corpora are assembled from crawls of the open web at a moment nobody publishes in advance. Anyone who implies otherwise is selling you a lottery ticket and calling it a strategy.

What is true — and this is the mechanism, not a promise — is that training corpora are built from the same open web that search-enabled models read today. Being widely written about now is the only known input to being known later. That’s a reason to build presence, not a reason to believe a timeline. If a vendor gives you a date, that date is invented.

So the honest framing is: buy the half you can affect.

And it happens to be the better half. The search-enabled systems are the ones with the users who are shopping. ChatGPT with search, Perplexity, AI Overviews and AI Mode sit in front of people mid-decision. The training-only models are disproportionately API traffic, developer tooling, and self-hosted deployments — real, but not where your customer is standing when they ask who to call.

The four things that move the half you can affect

Not a list of everything. The four with a documented mechanism.

One — be readable at all. The retrieval crawlers must be able to fetch you, and the user-triggered fetchers must not be blocked at the firewall. This is gate one, it is usually broken, and it is usually broken for a reason nobody remembers.

Two — answer the question in the first sentence of a passage. Search-enabled models quote paragraphs, not pages. A paragraph, not a page has the shape that gets quoted and the median length that gets reused.

Three — put the facts on the page in plain text. Price, service area, hours, what you actually do. A model that has to guess will guess, and what these businesses charge is about how often it guesses wrong.

Four — exist somewhere other than your own website. Independent pages are what the models read when they decide whether you’re real. Written by somebody else has the citation split; the strongest signal is YouTube has the single strongest correlate anyone has measured.

None of this is exotic. All of it lands on the half of the market that can see you, and it lands on a timescale you can observe.

Related

Free scan, no card required

Now find out where your own business stands.

That was the research. The free report runs the buying questions from your industry through the AI systems and shows you who gets named instead of you.

Your services page rather than your homepage, if you have one. That's the page an AI reads when someone asks what you sell.

We'll email you the report. No card, no call required to get it.