The one-sentence version

ChatGPT runs a three-stage pipeline — retrieval, passage selection, citation decision — that operates only on the ~40% of answers where SearchGPT is triggered; the other ~60% come from the model’s innate training-data knowledge and require an entirely different playbook of earned media, entity development, and Wikipedia/Wikidata presence (Seer Interactive, humanswith.ai).

The 40/60 split that changes everything

According to Seer Interactive’s testing:

  • ~40% of ChatGPT answers trigger SearchGPT to pull live pages from Bing.
  • ~60% of ChatGPT answers are generated from the model’s innate knowledge — pre-training data plus post-training fine-tuning — with no live retrieval.

Which means every retrieval-based tactic applies to at most 40% of the conversations your buyers are having about your category. The other 60% is a training-data problem, and it moves on multi-year timescales, not quarterly.

You need both tracks in parallel.

The three-stage pipeline (the 40%)

Per humanswith.ai’s citation-signal analysis:

Stage 1 — Retrieval

Candidate pages are found from Bing’s index. Unlike Perplexity’s blended index, ChatGPT Search retrieves exclusively from Bing (Machine Relations). What matters here: Bing indexation, crawlability by OAI-SearchBot / GPTBot / ChatGPT-User, server-rendered HTML.

Stage 2 — Passage selection

Specific snippets or passages are chosen from within retrieved pages. This is where most pages lose. ChatGPT tends to reference passages that are direct (claim stated, not implied), self-contained (makes sense standalone), easy to fetch and parse (semantic HTML, no accordions), and credible enough for the specific claim.

Stage 3 — Citation decision

The answer model picks a small number of sources — averaging just 3.4 per multi-constraint query. This concentration effect means being a candidate isn’t enough; you have to be the best candidate. Trust signals — editorial coverage, government sources, named-expert bylines — get disproportionate weight at this stage.

The 60% — training-data optimization

The model’s innate knowledge is built from pre-training data (a corpus assembled before a fixed cutoff date) and post-training fine-tuning. OpenAI has not disclosed the weighting or full inclusions; Stanford’s Foundational Model Transparency Index rated ChatGPT 4 at 0/10 for transparency (Seer Interactive).

Omniscient Digital’s analysis of 23,387 unique citation sources across 240 branded queries found that only 23% of citations come from the brand’s own website (Outpace SEO summary):

Source category Share of citations
Earned media (combined) 48%
Commercial content from third-party publishers 30%
Owned brand content 23%

Getting represented in training data means getting third-party corroboration of your brand’s identity across the pages the training crawl saw. Owned content alone doesn’t do it.

What actually moves the needle

For the 40% (live retrieval)

  • Verify Bing indexation of every priority URL in Bing Webmaster Tools.
  • Allow OAI-SearchBot, GPTBot, ChatGPT-User in robots.txt.
  • Server-render or pre-render critical content — don’t rely on client-side JS.
  • Front-load direct answers in the first 100 words of every page.
  • Break content into self-contained passages with H2/H3 sub-questions.
  • Include statistics with clear attribution — +37% citation probability lift.
  • Cite authoritative sources yourself — +40% citation probability lift.
  • Use named-expert bylines with Person schema linked to authoritative profiles.

For the 60% (training data)

  • Earn independent editorial coverage in outlets your buyers actually read.
  • Get named in industry roundups and review sites — consistency across sources is what teaches the model where you belong.
  • Build a Wikipedia article if you meet notability, otherwise a well-formed Wikidata item.
  • Homepage schema.org Organization with a sameAs array of authoritative profile URLs (LinkedIn, Crunchbase, GitHub, X, Wikipedia, Wikidata).
  • Consistent brand description across every appearance — inconsistent definitions split your entity in the model’s memory.
  • Long horizons. Brands seeking training-data inclusion should expect to wait months or years for a new frontier-model snapshot.

The ChatGPT checklist

  • [ ] All priority URLs indexed in Bing (verified in Bing Webmaster Tools).
  • [ ] robots.txt allows OAI-SearchBot, GPTBot, ChatGPT-User.
  • [ ] Critical content is server-rendered (not client-side JS-only).
  • [ ] Direct-answer paragraph (60–100 words) at the top of every page.
  • [ ] Self-contained passages under H2/H3 sub-questions.
  • [ ] Statistics inline with named sources; own content cites authoritative sources.
  • [ ] Named-expert bylines with Person JSON-LD linked to LinkedIn / ORCID / other authoritative profiles.
  • [ ] Homepage Organization schema with sameAs array of 4+ authoritative profiles.
  • [ ] Wikidata item exists and is filled out (minimum: instance of, industry, country, official website, founded).
  • [ ] Wikipedia article exists if you meet notability thresholds.
  • [ ] Canonical brand description used verbatim across homepage, LinkedIn, Crunchbase, bylines, and Wikipedia.
  • [ ] Active independent editorial coverage — at least monthly presence in named outlets.

Deep dives

How ChatGPT fits the wider GEO picture

ChatGPT is the hardest AI engine to win on, and also the most valuable citation once you do. Its 3.4-average concentration means each citation contributes more of the answer’s actual language and structure than a Perplexity citation would (Machine Relations).

ChatGPT rewards patient, multi-year entity work — the kind of coverage, positioning and consistency that also compounds elsewhere. Start with the GEO Guide for the universal foundations, cross-reference with Perplexity and Google AI, and treat ChatGPT as a two-track investment: retrieval optimization for the 40%, entity development for the 60%.