{
  "@context": "https://schema.org",
  "@type": "Article",
  "schema_version": "1.1",
  "content_item_id": "guide-how-ai-search-finds-sources.en",
  "translation_group_id": "guide-how-ai-search-finds-sources",
  "locale": "en",
  "type": "guide",
  "section": "knowledge",
  "slug": "how-ai-search-finds-sources",
  "title": "How AI search finds and selects sources",
  "description": "An AI answer is not a transparent search log. We can observe a bot request, a published answer, and displayed links, but one citation cannot reveal the complete retrieval trace, source weights, or internal reasoning. Google describes query fan-out as a possible way to split a complex question into additional searches; that is a documented possibility for one system, not a universal rule for every answer engine.",
  "direct_answer": "An AI answer is not a transparent search log. We can observe a bot request, a published answer, and displayed links, but one citation cannot reveal the complete retrieval trace, source weights, or internal reasoning. Google describes query fan-out as a possible way to split a complex question into additional searches; that is a documented possibility for one system, not a universal rule for every answer engine.",
  "sections": [
    {
      "heading": "1. Seven different stages",
      "paragraphs": [
        "A URL may be discovered, fetched, processed, indexed, retrieved for a question, used in synthesis, and displayed as a citation. This is not one funnel with guaranteed progression. When a report collapses stages into one score, it loses diagnostic value: we cannot tell whether the issue is access, index coverage, or selection for a particular answer."
      ],
      "source_ids": [
        "google-ai-search"
      ]
    },
    {
      "heading": "2. Why a passage must stand alone",
      "paragraphs": [
        "A passage should name its subject, action, period, and claim boundary. “The platform uses sources” is weaker than “OpenAI documents a separate search crawler; the document does not say that every fetch appears in an answer.” The second can be checked, separates official fact from assumption, and does not require five screens of context."
      ],
      "source_ids": [
        "openai-bots",
        "perplexity-bots"
      ]
    },
    {
      "heading": "3. Scenario: a complex question",
      "paragraphs": [
        "“How should we choose an AI-search platform?” contains several intents: define criteria, compare features, and check freshness. A system may search documentation for each platform separately. The editor should not expect one “main” source. Build a claim map, connect every criterion to a primary document, and label which parts are the newsroom’s comparison."
      ],
      "source_ids": [
        "google-ai-search"
      ]
    },
    {
      "heading": "4. A reproducible test",
      "paragraphs": [
        "Save exact wording, locale, country, device, mode, timestamp, the full answer, and every link, including the absence of links. Run the same panel on at least two dates. A mode, language, or interface change creates a new slice. Compare only compatible slices and show absolute counts beside rates."
      ],
      "source_ids": [
        "perplexity-bots"
      ]
    },
    {
      "heading": "5. How to verify a citation",
      "paragraphs": [
        "Open the source and locate the claim the link supposedly supports. Record semantic match, date, version, and context. If the URL merely accompanies the answer, call it a displayed link, not evidence. If the source contradicts the answer, record a quality finding; it is not proof that the platform always errs."
      ],
      "source_ids": [
        "google-ai-search"
      ]
    },
    {
      "heading": "6. Limits of inference",
      "paragraphs": [
        "Operator documentation and interface observation have different evidentiary strength. Do not claim a page was used in hidden retrieval when all you saw was a bot visit. Do not generalize one answer to every language, region, and date. Say: “in this panel, on this surface, at this time, we observed…”"
      ],
      "source_ids": [
        "google-ai-search"
      ]
    },
    {
      "heading": "7. Publication checklist",
      "paragraphs": [
        "Before release check: one question per direct answer; every factual claim has primary evidence; the source is relevant to the nearby claim; observation/inference/hypothesis are labeled; panel and period are recorded; no citation promises are made; answer version is saved. Only then is the piece research rather than an impression from one chat."
      ],
      "source_ids": [
        "google-ai-search"
      ]
    },
    {
      "heading": "8. An evidence map for one answer",
      "paragraphs": [
        "Take one question, such as “how do I check a page’s technical availability for AI search,” and split the result into four fields. Field one is the URL and its fetch; field two is the visible answer passage; field three is the link and the exact paragraph it supports; field four is the referral, or its absence. If a link opens a page but does not support the nearby claim, call it a displayed link, not evidence. If a bot arrived but no link appeared, record fetch without citation. If the answer named a brand without a URL, record mention without citation. The map cannot reconstruct a hidden retrieval trace, but it stops observable events being conflated and gives the editor a concrete next-check list."
      ],
      "source_ids": [
        "google-ai-search",
        "openai-bots",
        "perplexity-bots"
      ]
    },
    {
      "heading": "9. Choosing the next change",
      "paragraphs": [
        "Use a simple matrix: an access problem calls for a technical fix; an understanding problem calls for a clear heading, definition, and standalone passage; an evidence problem calls for a primary source and data period; a citation problem calls for another surface test, not a promised new tag. Score each candidate for cost, reversibility, and observability. Start with a change that can be rolled back and checked on a small control set. Do not call a variant a winner after one favorable answer: it may reflect a temporary index or an interface change."
      ],
      "source_ids": [
        "google-ai-search",
        "perplexity-bots"
      ]
    },
    {
      "heading": "10. What to archive",
      "paragraphs": [
        "A minimal run archive is more than the final link. Save the prompt, mode and locale, timestamp, page version, full answer, every displayed URL, and a source snapshot where the license permits it. Record which passage supports the claim and which part is the newsroom’s inference. On a repeat, do not overwrite the old slice: a new date or changed surface must remain distinguishable. This archive makes a disputed observation checkable and lets the editor say honestly when the original answer is no longer available."
      ],
      "source_ids": [
        "google-ai-search",
        "openai-bots",
        "perplexity-bots"
      ]
    },
    {
      "heading": "11. Operator-neutral caveat",
      "paragraphs": [
        "A site owner controls publication, access, facts, and measurement. The owner does not control the operator’s index, hidden retrieval signals, answer composition, or whether a referral appears. A strong article therefore describes the interface and test conditions instead of assigning universal behavior to a platform. Phrases such as “we observed” and “in this slice” are more accurate than “the model always chooses.” This boundary remains even when technical guidance is followed completely."
      ],
      "source_ids": [
        "google-ai-search",
        "openai-bots"
      ]
    }
  ],
  "published_at": "2026-09-11",
  "modified_at": "2026-09-11",
  "data_through": "2026-09-11",
  "next_review_at": "2026-10-11",
  "author": "GeoAeoAle Editorial",
  "origin": "editorial",
  "publisher": "GeoAeoAle Editorial",
  "license": "https://creativecommons.org/licenses/by/4.0/",
  "canonical_url": "https://geoaeoale.com/en/knowledge/how-ai-search-finds-sources/",
  "claims": [
    {
      "claim_id": "source-query-fanout",
      "text": "Google describes query fan-out as a possible way to break a complex question into additional searches.",
      "status": "observed",
      "confidence": "high",
      "source_ids": [
        "google-ai-search"
      ],
      "publication_status": "public",
      "confidentiality": "public"
    },
    {
      "claim_id": "source-citation-not-trace",
      "text": "A displayed citation is observable but does not reveal the complete source-selection path.",
      "status": "inference",
      "confidence": "high",
      "source_ids": [
        "google-ai-search",
        "perplexity-bots"
      ],
      "publication_status": "public",
      "confidentiality": "public"
    }
  ],
  "sources": [
    {
      "source_id": "google-ai-search",
      "canonical_url": "https://developers.google.com/search/docs/appearance/ai-features",
      "title": "Top ways to ensure your content performs well in Google's AI experiences",
      "publisher": "Google Search Central",
      "source_type": "official",
      "locale": "en",
      "published_at": null,
      "checked_at": "2026-09-11",
      "sha256": "c6b267ed42c26ee63151c87d27de45d8882a5b7dc77c3d3dece510b21d63c300",
      "license": "Source terms apply",
      "visibility": "public"
    },
    {
      "source_id": "openai-bots",
      "canonical_url": "https://developers.openai.com/api/docs/bots",
      "title": "Overview of OpenAI crawlers",
      "publisher": "OpenAI Developers",
      "source_type": "official",
      "locale": "en",
      "published_at": null,
      "checked_at": "2026-09-11",
      "sha256": "ccbdef3018bd08dceaacb7fe0ea07a2020d25e201ab84aa44625827aac925440",
      "license": "Source terms apply",
      "visibility": "public"
    },
    {
      "source_id": "perplexity-bots",
      "canonical_url": "https://docs.perplexity.ai/docs/resources/perplexity-crawlers",
      "title": "Perplexity crawlers",
      "publisher": "Perplexity Docs",
      "source_type": "official",
      "locale": "en",
      "published_at": null,
      "checked_at": "2026-09-11",
      "sha256": "75930d803650ae046c37ed3529840f7cbf1a6599c1bf2e41cc70bf83cc5b6470",
      "license": "Source terms apply",
      "visibility": "public"
    }
  ],
  "related_slugs": [
    "retrieval",
    "rag",
    "grounding",
    "query-fan-out",
    "source-attribution"
  ],
  "limitations": [
    "One answer and one fetch do not reveal universal architecture or the complete set of sources used.",
    "Trace one page end to end: check raw HTML, status 200, canonical, robots, and sitemap; then ask the same question on the selected surface. Save the full answer, visible links, exact supporting passage, locale, time, and page version. Repeat after the change.",
    "A mention without a URL, a link without a relevant passage, and a crawler visit are different events. Count a citation only when the verifiable claim matches the link; without a saved answer the observation cannot be reproduced."
  ],
  "corrections": []
}