Skip to content

Guide · GEO

Where does AI get its information? What ChatGPT, Perplexity and Gemini actually cite

Line illustration of a web page, an open book, a document and a globe feeding one funnel into an AI brain icon, which returns a chat answer bubble

Where does AI get its information? From the live web, reached through a search index. Overwhelmingly from company-owned pages.

Citation data: 5–11 September 2026.

We tracked 14,503 citation mentions across ChatGPT, Perplexity and Gemini in that week (Amadora AI project-wide tracking). Corporate sites dominated every other source type in that data (blog posts, product pages and listicles, mostly).

An AI citation is the linked or named source an engine’s answer draws from. A brand mention is different. It can happen with no link and no source attached at all.

AI engines cite corporate web pages far more than anything else: 88.7% of 14,503 tracked citation mentions in the week of 5–11 September 2026, against 0.7% for academic sources and 2.3% for community forums. Blog articles alone took 50.1%. The engines’ own documentation says the entry condition is ordinary indexation, not AI-specific markup.

But the corporate pages engines cite mostly belong to other companies. In B2B software, almost none of the brand mentions inside AI citations come from the brand’s own domain. So “corporate sources win” isn’t permission to publish more on your own site.

The source mix behind AI answers

Corporate sites dominate, and the gap isn’t close.

Source type is how we classify the origin of a cited document (the domain’s character, not the page’s format): corporate, technical, community, commercial, editorial, social, academic, news, government, informational.

Source typeCitation mentionsShare
Corporate12,86388.7%
Technical3922.7%
Community3332.3%
Commercial2811.9%
Editorial2271.6%
Social1731.2%
Academic990.7%
News960.7%
Informational210.1%
Government170.1%

Amadora AI project-wide tracking, 5–11 September 2026.

Do AI engines prefer editorial or vendor sources? Vendor, by more than fifty to one in this window (corporate against editorial). If you are deciding where to place a story, that ratio is the whole answer.

Vendor sources ahead of editorial sources by more than 50 to 1 in AI citations, a long violet bar beside a single dot

That answer holds beyond one week. Across a larger sample, the same corporate-source pattern appears in a 10,563-citation cross-project dataset, where corporate sites hold 79–85% of the citation surface and editorial publications 0.2% to 2.6%. The single-week share sits just above that range (fewer prompts, less averaging).

But which domains do AI engines cite most? Other companies’ corporate domains first, and among publications the smaller ones. Ranking 71 publications by citation volume, tier-1 outlets (TechCrunch, VentureBeat, Forbes, Wired, Inc., Fast Company, The Verge) delivered cumulatively roughly zero AI citations. Medium delivered 36 and TechRadar 31.

Smaller and more specialized beat bigger. That is the figure that should move your press budget.

Two limits before you reuse these numbers. The table counts mention volume, not unique URLs. It covers one tracked prompt set in one category. Treat the decimals as a shape rather than a benchmark for your own category.

Which page formats get cited

Format tracks source type closely, which narrows what you should build. Blog articles took 50.1% of mentions, website pages 25.0% and listicles 17.9%. Documentation took 2.1% and forum threads 1.3% (same window).

Blog articles and listicles together reached 68.0%. That sits inside the 60–72% band the cross-project dataset records.

Research papers took 0.6%. An engine answering a technical question reaches for vendor documentation before it reaches for a study. So your docs can earn citations your blog cannot.

Do AI engines cite Reddit and other forums?

Rarely in aggregate, and very unevenly. Forum threads by themselves took just 1.3% of mentions in the tracked week. That sits inside the small community share in the table above.

But the aggregate hides the mechanism, because community citations concentrate instead of spreading. In one project, all 84 community citations came from a single Reddit thread. In another, 417 came from a cluster of fewer than ten threads. Quora, Stack Exchange, Indie Hackers, Hacker News and Slack groups produced near-zero citations across the projects we track (same cross-project sample).

So a share that small isn’t a reason to skip forums. It’s a reason to stop treating them as a channel.

One cited thread can often carry your whole category. Working a subreddit for months may produce nothing. A single answer inside the thread an engine already pulls from changes what the engine repeats.

Find the thread first. Search your own tracked prompts for community URLs, read which threads appear, and post where the citation already exists.

What makes a page eligible to be cited

Both major engines document their entry conditions.

  • Google requires ordinary indexation. Its Search Central documentation states that “To be eligible to be shown as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to be shown in Google Search with a snippet”. The same page adds that “You don’t need to create new machine readable files, AI text files, or markup to appear in these features. There’s also no special schema.org structured data that you need to add.” It also describes a query fan-out technique, issuing multiple related searches across subtopics to build one response (documentation last updated 10 December 2025).
  • ChatGPT uses a separate search crawler. OpenAI’s bot documentation states that “OAI-SearchBot is used to surface websites in search results in ChatGPT’s search features,” and that sites blocking it “will not be shown in ChatGPT search answers.” GPTBot is a different agent, used “to crawl content that may be used in training our generative AI foundation models.”

AI-specific markup is not the entry condition. Indexation is.

That gives you a free check to run today. Open your robots.txt and confirm OAI-SearchBot is allowed (GPTBot is a separate decision, about training rather than citation). Then confirm your target pages are indexed and snippet-eligible.

One Amadora AI client site had disallowed the ChatGPT, Perplexity and Claude crawlers outright. It sat at zero AI visibility while remaining a plausible answer to the query (one client site, August 2026).

But eligibility isn’t selection. Passing both checks only puts your page in the pool the engine chooses from.

How citation sources differ by engine

Yes, AI citations differ by engine. What changes is which sources each favors, and how widely each spreads.

EngineHow it reaches the web (documented or observed)Observed source tendency
ChatGPTDedicated search crawler, separate from the training crawlerCites no video at all; fewer sources per answer
PerplexityLive retrieval on every querySurfaces Reddit more readily than the others
GeminiGrounded in Google’s search indexGoogle’s cited domains lead with YouTube
Google AI OverviewsSame index, with query fan-out across subtopicsWider and more diverse link set per response

Reach mechanics from Google and OpenAI documentation. Source tendencies from Amadora AI citation data filtered by engine, June 2026, 1,904 cited URLs, where Reddit accounted for about 2.4% of cited domains. Per-answer breadth from the arXiv measurement framework below.

Read that table as directional for your own category. It comes from one tracked project in one category over one month. Per-engine shares move with the prompt set.

Being cited is also not the same as being used. A measurement framework published on arXiv in April 2026, From Citation Selection to Citation Absorption, separated the two across 602 controlled prompts, 21,143 search-layer citations and 18,151 fetched pages. Perplexity and Google cited more sources broadly (breadth). ChatGPT cited fewer sources but showed substantially higher average citation influence among the pages it fetched (depth).

The pages with the most influence were longer, better structured, more semantically aligned with the query and richer in evidence.

Why this matters: your citation count can rise while your contribution to the answer stays flat. Breadth and depth need separate tracking.

How to track which sources win citations

Where does AI get its information for your own prompts? Only a dataset can tell you. Most published AI source analysis omits the fields that make its numbers usable.

In practice, that usually means demanding four disclosures before you trust a figure:

  1. The window. A seven-day share and a ninety-day share describe different things, and neither is evergreen, so pin the window before you quote it.
  2. The unit. Mention counts, unique URLs and unique domains produce very different percentages from identical data (the same corpus can look twice as concentrated).
  3. The taxonomy. Ask how a source became “editorial” or “corporate” (taxonomies differ between vendors), because the classification decides the headline.
  4. The collection method. Consumer interface or API, logged in or logged out, daily or monthly. We query the public ChatGPT, Perplexity and Gemini interfaces daily with no account, history or memory, deliberately not through an API, because replies and cited sources diverge between the two surfaces. We then extract brand mentions inside cited documents with full-text search rather than model calls.

But check your source mix per topic rather than per site. Listicles took 17.9% of mentions project-wide in the tracked week, but 76% of citations inside a single “AI visibility tools for agencies” topic (one tracked topic, May 2026). A site-wide average would’ve hidden that entirely.

You can see your own brand’s source-type breakdown across your tracked prompts and engines in Amadora AI.

How to get into the sources AI engines trust

Start by giving up the obvious reading of the corporate share. In B2B software, roughly 1–2% of the brand mentions found inside AI citations come from the brand’s own domain. Three tracked competitors: 2 owned of 186, 0 of 135, and 3 of 621 (Amadora AI citation tracking, April to August 2026). Other niches record a 30–50% owned share.

Owned-domain share of AI citation mentions: 1–2% in B2B software, 2/186, 0/135 and 3/621 for three brands, 30–50% elsewhere

GEO, or generative engine optimization, is the work of getting retrieved and cited by AI answer engines. Ranking in a results page is a different job. Most of that work happens off your own site (which is why an owned-content plan alone stalls). Four actions, in the order they pay:

  1. Clear the eligibility checks first. Robots and indexation cost you nothing to fix and gate everything else.
  2. Get named on other companies’ pages, in the two formats that get cited. Blog articles and listicles took 68.0% of all mentions in the tracked week. Inclusion in a “best tools” listicle or a comparison article on somebody else’s domain is often worth more citation volume than another post on yours.
  3. Find the one community thread that’s already cited, then answer in it. Thread-level concentration makes your channel-wide engagement poor value.
  4. Pick review platforms from your own citation data rather than from brand recognition. In one B2B SaaS project, G2 accounted for 9 citations, fewer than Slashdot, and the same analysis surfaced 61 relevant review platforms.

But how far any of this transfers depends on your category. Owned share is a category pattern rather than a constant: the low single digits in B2B software against 30–50% elsewhere. The only way to know which pattern you are in is to measure your own prompts.

Common questions

Does the source mix change from month to month?

Yes. Shares move as the underlying index and retrieval behavior change, which is why the figures here are stamped to the week of 5–11 September 2026 rather than presented as stable. Re-measure on a fixed window before you compare two periods.

How many sources does one AI answer draw on?

Roughly 50 to 100 analyzed pages sit behind a single reply, and engines avoid pulling repeatedly from the same source (observed across tracked projects, May to July 2026). Your brand therefore contributes one or two pages at most to any given answer, however much you publish.