How Can You (Really) Get Cited by ChatGPT?

Backlinks: The New Holy Grail for SEO Professionals

For several weeks now, I've been asking myself a simple question that's on everyone's lips: How can I get cited by ChatGPT?

Because today, the challenge is no longer limited to traditional SEO to appear in Google’s search results. It’s about understanding how a brand, a web page, or a website can appear in the responses generated by AI. In short, how to be mentioned by ChatGPT.

To get a clearer picture, I conducted my own empirical tests. Together with friends located on different continents— in the Americas, Africa, Asia, and Europe—we submitted exactly the same queries to ChatGPT on topics related to health, sports, daily tasks, and e-commerce, varying the search intent: informational, commercial, and transactional.

The results I observed are described below:

Information Queries (Know / Know Simple)

Example: How do you change a car tire?

When searching for this query in Google’s SERP, sites like Midas, Ornikar, and Europcar appear among the top results. However, when the same question is asked of ChatGPT without enabling web search, the AI generates a complete answer without citing any sources, regardless of the continent from which the prompt is sent.

But once web search is enabled, the AI starts suggesting links to Feu Vert, TotalEnergies, Allopneus, VroomVroom, and even a YouTube video from France Casse.

The same is true for a medical query: without a web search, there are no results.

This behavior makes sense. The AI generates its responses based on the knowledge it acquired during training on a vast dataset (that’s what we’ve been told since the advent of generative models). It doesn’t automatically search the Internet for every question. Instead, it constructs a response by predicting the most likely continuation of a text based on the context.

Transactional/Action-Oriented Queries (Do)

Example: I have a hike next week. What are the best shoes for hiking?

When it comes to transactional queries, the behavior changes significantly.

For queries with a transactional or action-oriented intent, ChatGPT may automatically trigger a web search, even if the user has not explicitly requested it.

The AI then displays links to online stores, product listings, comparison charts, images, and, in some cases, links that allow users to directly schedule an appointment or make a purchase.

In other words, the likelihood of being cited depends heavily on the search intent.

So, what should we take away from this?

ChatGPT cannot cite all websites under the same conditions.

When web search is not enabled and the query is purely informational, ChatGPT generally does not cite any sources. To date, OpenAI has not published any statistics indicating the breakdown among ChatGPT’s various operating modes (responses based on its training, web search, in-depth research, etc.).

On the other hand, when ChatGPT decides to conduct a web search, several factors appear to influence the sources it selects: domain authority, content relevance, the quality of the information provided, and, above all, search intent.

Does ChatGPT only read snippets, or does it actually crawl the pages?

This raised another question: When ChatGPT performs a web search, does it simply read the snippets displayed in the SERP, or is it actually able to explore the pages?

To answer this question, I ran several tests by modifying the file robots.txt from various websites and by observing the behavior of several AI systems.

The results show that a distinction must be made between traditional indexing crawlers and the content retrieval systems used by certain AI systems.

Block Googlebot in the file robots.txt often serves as a signal not to crawl. This directive is also followed by many other indexing or training crawlers, such as Bingbot and GPTBot, which may cause them to ignore the resource.

On the other hand, certain content search or retrieval systems that feed LLMs—such as those used by Perplexity, Mistral, or Bing Copilot—may, in some cases, retrieve the page using a rendering engine, execute the JavaScript, and reconstruct the final DOM, even when the URL is not crawled by traditional crawlers.

However, this behavior depends on their data retrieval pipeline and their data sources (search index, cache, headless browser, etc.) and is likely to change over time.

In other words, preventing a crawler from crawling a page does not necessarily mean that a chatbot will be unable to use its content.

These experiments show that visibility rules are changing rapidly.

Ranking well in Google’s SERP remains important, but it doesn’t guarantee that you’ll be cited by an AI. Conversely, certain queries automatically trigger a web search and open up the possibility of appearing in the generated answers.

Search engine optimization is gradually entering a new phase. It is no longer just about optimizing one’s search engine rankings, but also about understanding how AI systems select their sources, crawl pages, and determine which content is reliable enough to be cited.

This is likely one of the new challenges facing SEO in the age of language models.

Back to top