Skip to main content

How AI assistants pick sources: what is known, and what is guesswork

What can actually be established about how AI assistants choose which sites to cite, what is reasonable inference, and what is confident speculation sold as fact.

By Kaivalya Deshpande, Founder, RankBrain AI·Published ·5 min read·4 sources cited

The short version

  • Providers do not publish a complete source-selection formula. Anyone presenting a ranked list of "AI ranking factors" is inferring, not reporting.
  • What is well established: assistants with live retrieval mostly draw on indexed, crawlable web content. Being findable remains the prerequisite.
  • Reasonable inference from observation: clearly stated, attributable claims from recognisable sources get quoted more than vague ones from unknown sites.
  • Treat this whole area as an engineering problem with an unknown spec. Optimise for being unambiguous and verifiable, because those help under every plausible mechanism.

This is the question everyone in the category wants answered and nobody can answer properly. The starting position is that the selection process is not documented, varies between products, changes without notice, and is not directly observable from outside. That has not stopped a substantial industry from publishing confident factor lists.

What follows separates three things that usually get blended: what is established, what is reasonable inference, and what is speculation. The practical advice at the end holds under all of them, which is the main reason to trust it.

What is actually established

  • Assistants with browsing or retrieval use web content. They fetch pages, and pages they cannot fetch cannot be used.
  • Crawler access is controllable by you. Your robots.txt governs which agents may fetch your pages, and blocking a compliant agent prevents that crawl, not every possible source of an AI answer.
  • Some assistants surface citations with links. Which means the source set is at least partially inspectable for a given answer.
  • Training data and live retrieval are different mechanisms. A model may "know" about you from training without ever fetching your site, and that knowledge has a cutoff you cannot influence.

That last distinction matters more than it gets credit for. If an assistant describes your company from training data, on-site changes do not directly update the model weights, though retrieval or other sources may change the answer. If it retrieves live, your current pages matter immediately, and the crawlers that do the retrieving are documented: OpenAI lists its crawlers, Perplexity lists its own, and Google explains how its AI features use your pages. Most products now do some of both, which is why answers about you can be simultaneously current and out of date.

What is reasonable inference

Observed patterns and how much weight to put on them
PatternConfidenceBasis
Cited pages tend to already rank wellHighConsistently observed; retrieval commonly uses a search index
Clearly stated facts get quoted more than implied onesMedium-highExtraction favours self-contained statements
Recognisable, established sources are preferredMedium-highConsistent with how these systems are evaluated
Consistency across the web increases confidenceMediumPlausible mechanism; widely observed
Valid structured data helps extractionMediumStrong rationale, no published confirmation
Recency matters on fast-moving topicsMediumObserved, varies by product
A specific word count is optimalVery lowNo evidence whatsoever
A specific "AI readability score" mattersVery lowVendor-invented metric

Confidence ratings are our editorial judgement from observation and mechanism, not measured findings. We include the bottom two rows because they circulate widely and deserve naming.

What is speculation sold as fact

A number of claims circulate with a confidence their evidence does not support: that assistants weight particular schema types heavily, that an llms.txt file materially affects citation, that there is an optimal answer length, or that a proprietary score predicts citation likelihood. None of these has published support from the people who build these systems.

Some may well turn out to be true. The problem is not the hypotheses, it is presenting them as findings, particularly when the presenter sells a product that measures the thing they have declared important. Apply the same scepticism you would to any vendor-supplied ranking factor list. See llms.txt vs robots.txt.

What to do given the uncertainty

The useful move when a specification is unknown is to optimise for properties that help under every plausible mechanism. Fortunately those are also just good practice, though implementation has costs and needs testing.

  1. Be retrievable. Check your robots.txt permits the agents you want, and that pages render without requiring JavaScript execution to produce their main content.
  2. Rank well. Under every observed pattern, ranking correlates with being in the candidate set. Ordinary SEO is not superseded here.
  3. State facts explicitly, with attribution and dates. "X costs $129/mo as of September 2026, per the vendor pricing page" survives extraction; "X is competitively priced" does not.
  4. Describe yourself in one plain sentence and use it consistently everywhere you appear.
  5. Keep structured data valid, checked in the Rich Results Test. The rationale is strong even without published confirmation, and it costs nothing.
  6. Be corroborated. Being described the same way in several credible places is a plausible confidence signal under any reasonable design.

Applying this standard to ourselves

This page could have been a ranked list of twelve AI ranking factors with confident percentages. Its percentages would be fabricated; its ranking is unknown. We have instead said what is known, what is inferred and what is invented, which is the same standard we apply to the pricing tables elsewhere on this blog, and the reason any of it is worth reading.

Frequently asked questions

How do AI assistants choose which sites to cite?

Providers publish some guidance, not a complete selection formula. Observation suggests cited pages usually rank well already, come from recognisable sources, and state relevant facts clearly enough to extract. Anything more specific than that is inference.

Is there an AI ranking factor list?

Not a real one. Published lists are inferred from observation, often by companies selling tools that measure the factors they have declared important. Treat them accordingly.

Does schema markup affect AI citations?

The rationale is strong (valid markup makes facts unambiguous to parse) but no major assistant has confirmed weighting it. It is good practice with a plausible mechanism, not a proven lever.

Can I submit my site to an AI assistant?

There is no submission process comparable to a sitemap. The closest equivalent is making sure you are crawlable, well-indexed and clearly written.

Why does an assistant describe my company incorrectly?

Possible causes include outdated training data, retrieval errors, conflicting sources and ambiguous descriptions. Correct your public information and use available feedback channels; model updates do not guarantee a correction.

Sources

Every figure on this page traces to one of these. Dates are when we last read each page: prices and features change, so check anything older than a few months. The last entry is a product page, listed so you can find the tool, not as evidence.

  1. [1]
    Introduction to robots.txt

    Google Search Central · developers.google.com · Official documentation · read 2026-09-18

  2. [2]
    Overview of OpenAI Crawlers

    OpenAI · platform.openai.com · Official documentation · read 2026-10-07

  3. [3]
    Perplexity Crawlers

    Perplexity · docs.perplexity.ai · Official documentation · read 2026-10-07

  4. [4]
    AI Features and Your Website

    Google Search Central · developers.google.com · Official documentation · read 2026-10-07

  5. [5]
    Rich Results Test

    Google · search.google.com · Product page · read 2026-09-18

Keep reading

This is the question everyone in the category wants answered and nobody can answer properly. The starting position is that the selection process is not documented, varies between products, changes wi…