Skip to content

How do AI citations work?

Short answer

An AI citation is a link an engine shows to a page behind its answer. The engine searches, retrieves pages, writes a response, and links its sources. Citations are not always accurate, so a link does not prove a page supports the claim.

Key takeaways
  • A citation is a link to a page the engine used for its answer, and each engine shows and counts them in its own way.
  • Google, Perplexity, and ChatGPT search each fetch pages, through fan-out searches, user-triggered visits, or a search crawler.
  • In a controlled test of 252,000 trials, topical relevance and list position were the strongest predictors of which source got cited, and formatting changes had little effect.
  • Studies find citations can mislead, with one measuring 30.6% that distort their source and another finding 3% to 13% of cited URLs hallucinated.
  • Citation choices are not stable, since one audit saw about 15% of binary decisions change on a fresh run.

When an AI engine answers a question, it often shows links beside the answer. Those links are citations. This page explains how engines produce them and how far you can trust them.

What is an AI citation?

It is a link or reference to a page that sits behind part of an answer. Google says its generative AI features show "prominent, clickable links to relevant web pages that support the information in the response" [1].

In the UK, the Competition and Markets Authority (CMA) now requires Google to attribute publisher content with clear links in AI-generated search results, and gave it nine months from 3 June 2026 to put the changes in place [9].

Engines also count citations. Bing's AI Performance report includes a total citations figure, which counts how often a site's content is referenced in AI answers [2].

A citation is not the same as a mention. An answer can name your business without linking to you, and a link does not mean the engine recommends you.

How does an engine get the pages it cites?

It searches, fetches pages, and then writes the answer. The three engines this page draws on do it differently.

  • Google. Its AI features use query fan-out, which its guide defines as "a set of concurrent, related queries generated by the model to request more information and fetch additional relevant search results" [1].
  • Perplexity. When a user asks a question, its Perplexity-User agent might visit a web page to help provide an accurate answer and include a link to the page in its response [3].
  • ChatGPT. OpenAI says OAI-SearchBot is used to surface websites in ChatGPT's search features [4].

Bing shows the retrieval side to site owners. Its report lists grounding queries, which are the key phrases the AI used when retrieving content [2].

What decides which pages get cited?

Research points to relevance and position more than formatting. The research below is controlled testing, not a measurement of every live engine, so read it as evidence of tendencies.

A May 2026 study put exactly two candidate sources into the model's context and ran 252,000 trials across six language models, varying 18 content factors. It found that topical relevance and list position were the strongest predictors of citation selection. Price information and recent timestamps also raised the chance of a citation, and formatting changes had minimal impact [5].

A September 2026 audit of 103 pairs of sources that support the same fact found a 42.3 percentage point gap in citation rate between the top and fifth-ranked document. A structured layout raised the target citation count by 0.50 per answer, but it moved existing citations around instead of adding new ones [6].

Are citation choices stable?

No. The same audit found that about 15 percent of binary decisions change when the model decodes again [6]. Asking one question once tells you little, so ask it several times.

Are AI citations accurate?

Not always. A cited page can be real and still not support the claim next to it.

One study of 112,000 responses from ten models across five providers found that 30.6 percent of citations distort their source and 27.1 percent come from domain-inappropriate sources [7].

Another study of commercial models and deep research agents found that 3 to 13 percent of citation URLs were hallucinated, and 5 to 18 percent did not resolve [8]. Open the link before you trust it.

Do reviews, directories, and forums get cited?

The sources used on this page do not say. None of them ranks review sites, directories, or forums by how often engines cite them. Treat any claim that a certain kind of site is trusted most as unproven until it shows its method.

What does this mean for your site?

Make each page easy to retrieve and easy to quote. Answer the question on the page in the first sentence, keep the page about one topic, and make sure engines can fetch it. Then check what each engine cites for your questions, several times, and open the links.

What AI answered

We asked 11 buyer questions, 3 times on each engine.

ChatGPT
39
different sites cited in 33 runs
Perplexity
155
different sites cited in 33 runs
Gemini
None
no source list returned
22
different agencies or tools named in 99 runs
developers.google.com
most cited source, in 16 of 99 runs

C8 check, October 2026. ChatGPT (chatgpt.com, logged out, model not shown), Perplexity (Perplexity API (sonar)), Gemini (Flash-Lite). Each prompt asked 3 times on each engine, logged out, one fresh session per run.

Sources

  1. Optimizing your website for generative AI features on Google Search, Google Search Central. Accessed 2 Oct 2026.
  2. Introducing AI Performance in Bing Webmaster Tools Public Preview, Bing Webmaster Blog. Accessed 2 Oct 2026.
  3. Perplexity crawlers, Perplexity. Accessed 2 Oct 2026.
  4. Overview of OpenAI crawlers, OpenAI. Accessed 2 Oct 2026.
  5. What Gets Cited: Competitive GEO in AI Answer Engines, Vishwakarma, Kumar, and Jamidar, arXiv preprint, May 2026. Accessed 2 Oct 2026.
  6. CITECHOICE: A Causal Audit of How Document Presentation Redistributes Citation Credit in Agentic Search, Selvam and Ghosh, arXiv preprint, September 2026. Accessed 2 Oct 2026.
  7. Verified Misguidance: Measuring Structural Citation Failures in Search-Augmented LLMs, Seo et al., arXiv preprint, May 2026. Accessed 2 Oct 2026.
  8. Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents, Rao, Wong, and Callison-Burch, arXiv preprint, April 2026. Accessed 2 Oct 2026.
  9. CMA secures fairer deal for publishers and improves Google search services in UK, Competition and Markets Authority, GOV.UK, 3 June 2026. Accessed 4 Oct 2026.

First published


Pervaj Ahmed Tuhin is the Co-Founder and Chief Technology Officer of Ravenence, leading how the sites are built and structured.

More from Pervaj →

Related questions

View all answers →

Want your business named in answers like this one?

We rebuild your website around the questions customers ask Google and AI. Done for you, in 7 days.