Seven different stages

A URL may be discovered, fetched, processed, indexed, retrieved for a question, used in synthesis, and displayed as a citation. This is not one funnel with guaranteed progression. When a report collapses stages into one score, it loses diagnostic value: we cannot tell whether the issue is access, index coverage, or selection for a particular answer.

Section sources:[1] Google Search Central

Why a passage must stand alone

A passage should name its subject, action, period, and claim boundary. “The platform uses sources” is weaker than “OpenAI documents a separate search crawler; the document does not say that every fetch appears in an answer.” The second can be checked, separates official fact from assumption, and does not require five screens of context.

Section sources:[2] OpenAI Developers[3] Perplexity Docs

Scenario: a complex question

“How should we choose an AI-search platform?” contains several intents: define criteria, compare features, and check freshness. A system may search documentation for each platform separately. The editor should not expect one “main” source. Build a claim map, connect every criterion to a primary document, and label which parts are the newsroom’s comparison.

Section sources:[1] Google Search Central

A reproducible test

Save exact wording, locale, country, device, mode, timestamp, the full answer, and every link, including the absence of links. Run the same panel on at least two dates. A mode, language, or interface change creates a new slice. Compare only compatible slices and show absolute counts beside rates.

Section sources:[3] Perplexity Docs

How to verify a citation

Open the source and locate the claim the link supposedly supports. Record semantic match, date, version, and context. If the URL merely accompanies the answer, call it a displayed link, not evidence. If the source contradicts the answer, record a quality finding; it is not proof that the platform always errs.

Section sources:[1] Google Search Central

Limits of inference

Operator documentation and interface observation have different evidentiary strength. Do not claim a page was used in hidden retrieval when all you saw was a bot visit. Do not generalize one answer to every language, region, and date. Say: “in this panel, on this surface, at this time, we observed…”

Section sources:[1] Google Search Central

Publication checklist

Before release check: one question per direct answer; every factual claim has primary evidence; the source is relevant to the nearby claim; observation/inference/hypothesis are labeled; panel and period are recorded; no citation promises are made; answer version is saved. Only then is the piece research rather than an impression from one chat.

Section sources:[1] Google Search Central

An evidence map for one answer

Take one question, such as “how do I check a page’s technical availability for AI search,” and split the result into four fields. Field one is the URL and its fetch; field two is the visible answer passage; field three is the link and the exact paragraph it supports; field four is the referral, or its absence. If a link opens a page but does not support the nearby claim, call it a displayed link, not evidence. If a bot arrived but no link appeared, record fetch without citation. If the answer named a brand without a URL, record mention without citation. The map cannot reconstruct a hidden retrieval trace, but it stops observable events being conflated and gives the editor a concrete next-check list.

Section sources:[1] Google Search Central[2] OpenAI Developers[3] Perplexity Docs

Choosing the next change

Use a simple matrix: an access problem calls for a technical fix; an understanding problem calls for a clear heading, definition, and standalone passage; an evidence problem calls for a primary source and data period; a citation problem calls for another surface test, not a promised new tag. Score each candidate for cost, reversibility, and observability. Start with a change that can be rolled back and checked on a small control set. Do not call a variant a winner after one favorable answer: it may reflect a temporary index or an interface change.

Section sources:[1] Google Search Central[3] Perplexity Docs

What to archive

A minimal run archive is more than the final link. Save the prompt, mode and locale, timestamp, page version, full answer, every displayed URL, and a source snapshot where the license permits it. Record which passage supports the claim and which part is the newsroom’s inference. On a repeat, do not overwrite the old slice: a new date or changed surface must remain distinguishable. This archive makes a disputed observation checkable and lets the editor say honestly when the original answer is no longer available.

Section sources:[1] Google Search Central[2] OpenAI Developers[3] Perplexity Docs

Operator-neutral caveat

A site owner controls publication, access, facts, and measurement. The owner does not control the operator’s index, hidden retrieval signals, answer composition, or whether a referral appears. A strong article therefore describes the interface and test conditions instead of assigning universal behavior to a platform. Phrases such as “we observed” and “in this slice” are more accurate than “the model always chooses.” This boundary remains even when technical guidance is followed completely.

Section sources:[1] Google Search Central[2] OpenAI Developers

What this does not prove

  • One answer and one fetch do not reveal universal architecture or the complete set of sources used.
  • Trace one page end to end: check raw HTML, status 200, canonical, robots, and sitemap; then ask the same question on the selected surface. Save the full answer, visible links, exact supporting passage, locale, time, and page version. Repeat after the change.
  • A mention without a URL, a link without a relevant passage, and a crawler visit are different events. Count a citation only when the verifiable claim matches the link; without a saved answer the observation cannot be reproduced.

Sources

  1. 2
    Overview of OpenAI crawlersOpenAI Developers · official · 11 Sept 2026
  2. 3
    Perplexity crawlersPerplexity Docs · official · 11 Sept 2026

Correction history

No material corrections have been published.