{
  "@context": "https://schema.org",
  "@type": "Report",
  "schema_version": "1.1",
  "content_item_id": "weekly-research-2026-08-17.en",
  "translation_group_id": "weekly-research-2026-08-17",
  "locale": "en",
  "type": "research",
  "section": "research",
  "slug": "weekly-research-2026-08-17",
  "title": "Weekly GEO research: four boundaries of measuring AI visibility",
  "description": "A synthesis of four primary studies on context, retrieval, citation quality, and web-content provenance.",
  "direct_answer": "The week did not reveal a universal growth factor, but it established four practical limits: conversation context changes answers, one run does not cover all sources, more citations can reduce evidential quality, and page provenance is not a recipe for generating content.",
  "sections": [
    {
      "heading": "Research question and weekly scope",
      "paragraphs": [
        "This week asks which fresh publications actually change how AI visibility should be measured and how answer changes should be explained. The material combines four primary studies published or updated in August 2026. It is not a platform ranking and does not promise citations. The unit of analysis is a specific experiment, its sample, and its artifact. We separate fetch, retrieval, mention, citation, and referral: temporal alignment does not prove causation. This frame matters because the word visibility often mixes technical access, source selection, and human behavior."
      ]
    },
    {
      "heading": "Conversation context is part of the test condition",
      "paragraphs": [
        "The August 3 study compared multi-turn conversations while keeping the final user message identical. The model received either the full conversation, only the final question, or a short reconstruction of history. In 44.7% of observations, the full conversation and the question without history produced materially different answers; a judge considered some differences capable of changing user action. This does not show that longer history is always better and did not measure web search. It does make comparisons invalid when one run continues a dialogue and another is a clean question. A test record should store conversation_state and the history-transfer method."
      ]
    },
    {
      "heading": "Repeats expose the incompleteness of one answer",
      "paragraphs": [
        "In the August 13 medical sample, three production chatbots selected primary studies for 20 clinical questions and were compared with an expert Cochrane set. One answer found 39.2% of included works on average; combining four repeats found 74.2% at least once. A researcher-framed role produced broader coverage than a patient role. These percentages belong only to twenty medical questions and cannot be transferred to brands or commercial search. The transferable methodological lesson is to count the stable core, cumulative coverage, and long tail separately. Missing from one answer does not prove that a system lacks the source."
      ]
    },
    {
      "heading": "Citation frequency without evidence quality is risky",
      "paragraphs": [
        "An August 11 laboratory study repeatedly rewrote already retrieved documents to make a generative system cite them more often. Under repeated interventions, the texts accumulated unsupported claims; a mechanism rewarding verifiable content produced a better balance on the benchmark. This was not a production audit of ChatGPT, Google, or Perplexity and does not reveal their algorithms. An editorial metric should separate candidate-set presence, displayed links, support for a particular claim, and absence of new errors. A higher citation share is not success if the source becomes harder to verify."
      ]
    },
    {
      "heading": "What the AI-generated provenance observation means",
      "paragraphs": [
        "A Google AI Overview audit from May 25–29 across 2,597 YMYL queries found that pages labeled AI-generated by a commercial detector were cited more often in the observed sample. The observation does not establish causation: the detector is contestable, quality and confounders were limited, and an author’s relationship with the detector company calls for caution. The gap mainly appeared among URLs outside the collected organic top 100. Provenance may therefore be studied as an experimental feature, but it cannot justify recommending generated text for citations."
      ]
    },
    {
      "heading": "Limitations and next check",
      "paragraphs": [
        "S01 measured generation without web search; a search-enabled replication is needed to test whether history changes query rewriting and link sets. S02 concerns medical questions, S03 is a simulation, and S04 is an observational audit with sensitive classification. None reveals a universal ranking factor. The next reproducible test should preregister clean versus continued questions, run at least four repeats, preserve full answers, and separately verify citations and supporting passages. Results should include absolute counts, period, language, sample, and confidence level."
      ]
    },
    {
      "heading": "A minimum editorial protocol",
      "paragraphs": [
        "Before every comparison, record the exact question, language, region, surface, mode, and conversation state. Preserve the full answer, every link, timestamp, and page version. Give each result a separate verdict: fetch, mention, citation, supporting evidence, or referral. If sources differ, do not call it a ranking change without a matched panel. Readers need the denominator, exclusions, and boundary of applicability, not only an attractive number. This protocol turns a weekly note into a cumulative database: a new run can be compared with an old one, and a correction does not erase the earlier observation."
      ]
    },
    {
      "heading": "Editorial conclusion",
      "paragraphs": [
        "The week’s main editorial error is treating a single observation as system knowledge. A useful publication shows the question, condition, source, and limitation. If a conversation continued, that is part of the test; if an answer was repeated, that is part of the method; if a link appeared, check which claim it supports. This format lets readers separate new knowledge from an attractive assumption. It also preserves history: the next cycle can confirm, refine, or challenge the conclusion without rewriting the old date or hiding disagreement."
      ]
    },
    {
      "heading": "The observation unit: from URL to action",
      "paragraphs": [
        "For monitoring, an answer should be decomposed into candidate, retrieval, mention, link, claim support, and user action. These events diverge: a page may enter the candidate set without being cited; a link may point to a domain without supporting the sentence. The report should preserve an event matrix, full answer, conversation conditions, and evidence passage rather than one visibility score. A repeatable protocol turns a weekly note into a cumulative database: the next run can be compared with the earlier observation without rewriting it."
      ]
    }
  ],
  "published_at": "2026-08-17",
  "modified_at": "2026-08-17",
  "data_through": "2026-08-17",
  "next_review_at": "2026-10-11",
  "author": "GeoAeoAle Editorial",
  "origin": "editorial",
  "topics": [
    "weekly-research"
  ],
  "publisher": "GeoAeoAle Editorial",
  "license": "https://creativecommons.org/licenses/by/4.0/",
  "canonical_url": "https://geoaeoale.com/en/research/weekly-research-2026-08-17/",
  "claims": [
    {
      "claim_id": "weekly-0817-context",
      "text": "In a controlled sample, the full conversation and the same final question without history produced different answers in 44.7% of cases.",
      "status": "observed",
      "confidence": "medium",
      "source_ids": [
        "weekly-0817-s01"
      ],
      "publication_status": "public",
      "confidentiality": "public"
    },
    {
      "claim_id": "weekly-0817-repeats",
      "text": "One answer found 39.2% of an expert set, while four repeats found 74.2% at least once.",
      "status": "observed",
      "confidence": "medium",
      "source_ids": [
        "weekly-0817-s02"
      ],
      "publication_status": "public",
      "confidentiality": "public"
    },
    {
      "claim_id": "weekly-0817-citation",
      "text": "Optimizing only for citation frequency accumulated unsupported claims in a laboratory simulation.",
      "status": "observed",
      "confidence": "medium",
      "source_ids": [
        "weekly-0817-s03"
      ],
      "publication_status": "public",
      "confidentiality": "public"
    }
  ],
  "sources": [
    {
      "source_id": "weekly-0817-s01",
      "canonical_url": "https://arxiv.org/abs/2608.02556",
      "title": "Conversation history changes model answers",
      "publisher": "Primary research",
      "source_type": "primary-research",
      "locale": "en",
      "published_at": null,
      "checked_at": "2026-09-11",
      "sha256": "055a85b7e9eab3d66ce51d71d68b604cd74e6a8aad56a849ee08cabf96e8a1dc",
      "license": "Source terms apply",
      "visibility": "public"
    },
    {
      "source_id": "weekly-0817-s02",
      "canonical_url": "https://arxiv.org/abs/2608.13786",
      "title": "Chatbots retrieve different subsets of evidence",
      "publisher": "Primary research",
      "source_type": "primary-research",
      "locale": "en",
      "published_at": null,
      "checked_at": "2026-09-11",
      "sha256": "e0c2be30ab5d803f1aa70b88dcdf736530c823bf66705f98c68e2bee5ee6dfa0",
      "license": "Source terms apply",
      "visibility": "public"
    },
    {
      "source_id": "weekly-0817-s03",
      "canonical_url": "https://arxiv.org/abs/2608.11390",
      "title": "Citation optimization and evidence quality",
      "publisher": "Primary research",
      "source_type": "primary-research",
      "locale": "en",
      "published_at": null,
      "checked_at": "2026-09-11",
      "sha256": "46949ff2bbd6d7271202eadcceb5d7a363a979992d97885e882dae0fbf54fdc3",
      "license": "Source terms apply",
      "visibility": "public"
    },
    {
      "source_id": "weekly-0817-s04",
      "canonical_url": "https://proceedings.mlr.press/v318/kakimov26a.html",
      "title": "Auditing AI-generated provenance and citation quality",
      "publisher": "Proceedings of Machine Learning Research",
      "source_type": "primary-research",
      "locale": "en",
      "published_at": null,
      "checked_at": "2026-09-11",
      "sha256": "e6d3f42f3fde295adab962829b9a9dd1fefb857dde55ac85240beaa6c787df0e",
      "license": "Source terms apply",
      "visibility": "public"
    }
  ],
  "related_slugs": [
    "measuring-ai-visibility",
    "ai-citation",
    "retrieval"
  ],
  "limitations": [
    "Samples, models, languages, and surfaces differ; observations cannot be generalized to the whole web or treated as causal effects."
  ],
  "corrections": []
}