I finally know how to be cited by ChatGPT

Mask vs. Machine

This is an issue that is becoming increasingly important for website publishers, SEO specialists, and content creators. If an AI reads only a few lines of an excerpt, the implications are not the same as if it analyzes the entire page.

On August 22, 2025 (about 4 months ago), I decided to test an idea that had been on my mind for a while: Does ChatGPT-5 rely solely on the snippets displayed in search results, or is it capable of crawling an entire page to construct its response?

When a prompt is entered, ChatGPT first determines the intent of the search. Depending on the type of query, it may respond using only its own knowledge or trigger a web search. When this in-depth search is used, it scans the Internet, selects multiple sources, and retrieves content from the pages it deems relevant.

So I wanted to see what was really going on.

My test focused on a subject that has fascinated me for a long time: African art.

In 2019, I published an article about the Wanyugo mask, an article that is, incidentally, cited by Wikipedia:

WANYUGO: THE SYMBOLISM OF KNOWLEDGE
Read the article on Wanyugo by clicking on this clickable text

I then conducted two (2) experiments (yes, I know, I don't get much sleep).

  • Test #1 (automatic mode, around midnight): no sources displayed, no citations. That was exactly what I expected.
  • Test #2 (12:25 a.m., with advanced search enabled): This time, ChatGPT took about eight minutes to compile a much more comprehensive report.

The next day, I checked the logs on my OVH server to figure out exactly what had happened. Here’s what I found:

Bots detected:

  • 00:05 → GPTBot checks the robots.txt file
  • 00:28:32 → ChatGPT-User views the article
  • 00:36:58 → ChatGPT-User visits again (probably to complete their analysis)
  • 00:38:30 to 00:38:33 → downloading the images in the article

In summary, the robots accessed:

  • the robots.txt file
  • a content page
  • two (2) images

Based on these observations, several lessons emerge.

First, GPTBot appears to be querying the site as part of a training or model preparation process (the frequency of these queries has yet to be confirmed).

Furthermore, ChatGPT-User clearly doesn't settle for just a simple excerpt. It crawls the entire page, analyzes its content, extracts the relevant information, and then rephrases it to generate its response.

For this single query alone, ChatGPT aggregated 24 different sources and generated 57 citations.

Of these, 14 came directly from my article, accounting for about 23% of the citations used.

In other words, yes, ChatGPT quotes me… but it rewrites the content extensively rather than reproducing it verbatim.

Let's take a very simplified example.

ChatGPT's response:

Wanyugo: a Senufo helmet-mask from northern Côte d’Ivoire, part of an initiation rite associated with the Poro.

In my article, I explained this in more detail:

the origins of the Wanyugo mask, its role in Poro society, and the various stages of the initiation process (ages 7, 14, and 21).

The same ideas are there, but they’ve been summarized and rephrased. I’m sure this comes from my article because I’ve read all the articles that ranked for the keyword “Wanyugo,” and as I said, my article is cited as a source by Wikipedia.

My conclusion, therefore, is as follows:

No, ChatGPT-5 doesn't stop at snippets when web search is enabled.

It crawls web pages, analyzes them, compiles them, cites them in some cases, and rewrites the information to generate its own response.

For website publishers, this distinction is essential.

Your content can be used to generate an AI-powered response that goes far beyond the simple snippet displayed on a results page, as some SEO experts claim.

Another interesting observation: ranking first on Google isn't a prerequisite for being cited by ChatGPT.

My article is cited by Wikipedia, but it doesn't appear on the first page of Google search results for this query. Despite this, ChatGPT picked it up and cited it several times.

This confirms a hunch.

Editorial quality remains the foundation, but visibility on Google also depends on the link graph and proximity to “seed” sites (the Nearest-Seed PageRank algorithm): the more a page is linked to high-trust nodes (institutions, museums, journals, media outlets), the more it stands out.

For your information, a “seed” site is a site considered an extremely reliable starting point in link-graph ranking models (such as PageRank). These are generally institutional, academic, or editorial sites with very high authority (universities, museums, public agencies, major media outlets). These sites serve as anchor points: the closer a page is—in terms of the number of links—to these “root” sources, the more reliable it is considered to be in propagating the reputation of a piece of content.

This is consistent with the logic of Nearest-Seed PageRank.

In my view, the objective thus becomes a two-part equation:

  • produce high-quality content that precisely addresses a search intent;
  • strengthen ties with leading websites to enhance its credibility.

There are a few simple steps that can be taken right away.

  • Proximity to “seed” sites: Obtain editorial links from museums, universities, specialized journals, or reputable media outlets, while strengthening internal linking from the most powerful pages on your site.
  • Facilitating Aggregation by AI : Add a summary box at the beginning of the article containing 5 to 10 factual points (with figures, dates, and sources). The goal is to enable AI to perform fast and reliable extraction:
    • 5 to 10 Key Facts
    • 100% in numerical terms as soon as possible (dates, durations, volumes, percentages)
    • explicit sources or direct links
    • short formulation (chunk size to be determined)

This type of structure greatly improves the readability of content for AI systems whose job is to aggregate content. Of course, these are not general, absolute, or definitive truths. They are merely observations drawn from a specific case. I will continue my testing…

Back to top