AI search systems do not read a web page the way a person does. When a user asks a question, the system retrieves candidate documents, divides them into passages, scores those passages against the question and its related sub-questions, and assembles an answer from the segments that fit best. Google describes AI Overviews and AI Mode as using a “query fan-out” technique, which issues multiple related searches across subtopics and data sources to develop a response. A page is therefore assessed as a collection of extractable units, and the way those units are arranged influences which of them, if any, reach the final answer.
This article reviews what published research says about content structure and AI citations. It separates findings drawn from large datasets from vendor-reported figures with unpublished methods, and it sets both beside what Google itself states about optimizing for its AI features. Our reading of the evidence, which we label as our own interpretation wherever it appears, is that structure governs how easily a passage can be extracted once a page is already eligible for retrieval, and that it has little capacity to replace the authority signals that determine eligibility. The article continues a series that has examined schema markup, publisher mentions, entity clarity, original research and backlinks.
How AI Systems Retrieve and Extract Passages
Many AI answer systems that cite sources follow a retrieval-augmented pattern, in which a model is supplied with text retrieved from the web or an index and writes its response from that text. The retrieval step operates on passages rather than whole pages, because language models work within limited context windows and because a long page usually contains material that is irrelevant to any single question. A passage that states its subject, makes a complete claim and can be understood without the surrounding text is easier to retrieve and to quote than one that depends on earlier paragraphs for its meaning.
Query fan-out adds a second layer. A single question, such as which payment processors a fintech company should consider, may be expanded into several narrower searches covering pricing, regulation, integration and reliability. Each of those searches can retrieve a different passage from a different page, and the answer draws on the results of all of them. Our reading is that a page competes in many small contests at once under this arrangement, and that a page whose sections each answer a distinct sub-question has more chances of being selected than one that treats a topic as a single continuous narrative.
Google’s documentation adds a practical point. It states that the best practices for SEO remain relevant for AI features in Google Search, and it advises ensuring that important content is available in textual form. Material that exists only inside images, interactive elements or scripts that crawlers do not render may never enter the pool of passages from which an answer is built, so crawlable text is the first structural requirement, ahead of any decision about headings or formatting.
What the Available Research Does and Does Not Show
Four sources bear directly on the question, and they differ considerably in quality. The table sets out what each one measured and how far its findings can be relied upon.
| Source | Sample | Reported finding | Quality note |
|---|---|---|---|
| Kevin Indig, via Search Engine Land (February 2026) | 3 million ChatGPT responses and 30 million citations; 18,012 verified citations analyzed | 44.2% of citations came from the first 30% of a page, 31.1% from the middle and 24.7% from the final third | Independent analyst; method described (sentence embeddings matched responses to source sentences); ChatGPT only |
| Same study, passage traits | Same sample | Cited text was nearly twice as likely to use clear definitions; heavily cited text averaged 20.6% proper nouns against a typical 5 to 8 percent | As above |
| AirOps | More than 12,000 URLs | 68.7% of ChatGPT-cited pages used a sequential heading structure, against under 25% of Google page-one results; nearly 80% of cited URLs contained a structured list, against fewer than one-third of Google’s top results | Vendor report; methodology not published on the page |
| Aggarwal et al., GEO (KDD 2024) | GEO-bench, a benchmark of queries across domains | Optimization methods raised visibility in generative engine responses by up to 40%; effectiveness varied by domain | Peer-reviewed venue; results come from a benchmark |
| Google Search Central | Official documentation | No additional requirements or special optimizations for AI Overviews or AI Mode, and no special schema.org markup | Primary source for Google only |
The Indig figures are the most useful because the dataset is large and the method is described. In plain terms, roughly 44 of every 100 cited passages came from the top 30% of the page they were taken from, with a sharp drop toward the footer. The same analysis found that 53% of cited sentences came from the middle of a paragraph, 24.5% from first sentences and 22.5% from last sentences. More than half of the cited material therefore sat inside a paragraph, which complicates advice to place the answer in the opening sentence of every paragraph.
Three limits apply to this evidence. The Indig study measured ChatGPT, so it does not establish that Google’s AI features or other assistants behave in the same way. It records where cited text was found and which traits it shared, and it cannot show that moving a passage higher on a page causes it to be cited, since pages that state their main claims early may also be better written or more authoritative in other respects. The AirOps figures compare ChatGPT-cited pages with Google page-one results, which shows that cited pages tend to carry more headings and lists, while the page leaves the construction of the sample unexplained. The GEO paper offers peer-reviewed support for the general proposition that content changes can alter visibility in generated answers, though the effect it reports varies by domain and the paper does not establish what the effect would be for a particular publisher.
Why Early, Self-Contained Passages Are Cited More Often
There are plausible mechanisms for the position pattern, and we offer them as our interpretation and as no part of the study’s findings. Retrieval systems often work with a limited amount of text per page, which would favor material that appears early. Pages that open with a direct statement of their subject also tend to supply the definitions and summary claims that answer the questions users actually ask, so a concentration of citations near the top may reflect both truncation and the quality of the writing. The two explanations point to the same practical conclusion, which is that the principal claim of a page belongs in its opening third.
The passage-level traits point in a consistent direction. Cited text was nearly twice as likely to use clear definitions, and heavily cited text averaged 20.6% proper nouns, compared with a typical 5 to 8 percent. A passage dense in named entities, such as companies, regulations, products and people, gives a retrieval system more to match against a query and gives a model more specific material to attribute. This connects directly to the work described in our post on entity clarity, since a passage that names its subject explicitly is easier to connect to a query and easier to attribute to a source.
The same dataset found that cited content was twice as likely to contain a question mark, and that 78.4% of the citations tied to questions came from headings. A heading that restates a reader’s question in plain terms marks the start of a passage that answers it, which helps a system pair a sub-question with the right segment. For finance and crypto publishers, where readers ask precise questions about regulation, custody, taxation and risk, this favors sections that state the question in the heading, answer it in the opening sentences and place qualifications in the sentences that follow.
Which Structural Practices Carry Over and Which Do Not
Heading hierarchy and lists have the most consistent support in the available data, though the support is correlational and comes largely from a vendor report. AirOps found that most ChatGPT-cited pages followed a sequential heading structure and that nearly 80% of cited URLs contained at least one structured list, and it reports a 2.8x citation lift for pages with question-based headings and FAQ sections. None of these figures comes with a published method, so they are best treated as directional. Clear headings and well-formed lists are worth adopting because they improve readability and give extraction systems clean boundaries, while the evidence that they cause additional citations is weaker than the evidence that they accompany them.
Schema markup occupies a narrower position than many guides suggest. Google states that no special schema.org structured data is needed to appear in AI Overviews or AI Mode. Structured data retains value for entity disambiguation and for conventional search features, a subject covered in our post on schema markup, and it does not substitute for passages that are clear when read alone.
Passage length is a further area where precision exceeds the evidence. AirOps recommends paragraphs of roughly 100 to 300 tokens and cautions that much longer ones risk truncation, which is sensible guidance in the form of a recommendation, and no independent test of a threshold has come to our attention. The more defensible position is that each paragraph should carry one complete idea, since the Indig data shows cited sentences coming from within paragraphs more often than from a first or last sentence.
The broader limit concerns what formatting cannot supply. A well-structured page on a domain with no third-party corroboration still lacks the signals that determine whether it is retrieved, and the GEO benchmark shows that the effects of content changes depend on the domain in which they are made. Structure is best regarded as hygiene that prevents a strong page from being passed over, and as a weak lever for a page that is not otherwise eligible.
Auditing a Page for Extractability
A structural audit can be run page by page, and it benefits from being applied first to the pages that carry the most commercial weight, such as service pages, comparison pages and the articles that already attract organic traffic. The following sequence reflects the evidence reviewed above and our own interpretation of it.
- Confirm that the answer exists as crawlable text. Check the rendered HTML to verify that core claims appear as text and are not confined to images, tabs that load on interaction or scripts that crawlers may not execute.
- Place the principal claim and the key definition early. Within the first third of the page, state what the page covers, define its central terms and give the main conclusion, since that portion supplied the largest share of citations in the Indig dataset.
- Make each section self-contained. Name the subject explicitly in each section and avoid references such as “this” or “the above” that depend on earlier text, so that a passage can be understood when it is read alone.
- Name entities precisely and attribute figures. Use the full names of companies, regulations, products and people, and attribute each statistic to its source in the same sentence that presents it.
- Align headings with reader questions. Where a section answers a real question, phrase the heading as that question or a close statement of it, and answer it directly in the opening sentences.
Where This Fits Into a Broader Authority Strategy
The evidence reviewed here supports a sequence in which eligibility comes first and extractability second. Whether a page is retrieved at all depends on signals that sit largely outside the page, including publisher mentions, the backlink profile, the clarity of the underlying entity and the presence of original research that others have reason to cite. Structure then determines how much of an eligible page can be used. Our interpretation is that structure acts as a multiplier on authority, which would be consistent with the domain-dependent effects reported in the GEO benchmark.
The same logic applies to third-party pages. An editorial placement is itself a page that AI systems may retrieve and quote, and a placement whose early paragraphs name the company, define what it does and attribute its claims offers a retrieval system a clean statement from an independent source. Publisher placements are therefore worth planning with the passage in mind, which includes the opening paragraphs of the piece and the manner in which the company is described, in addition to the publication and the link.
Organizations deciding where to start can use the AI Authority Audit to review content, backlink and publisher-mention signals together, so that structural fixes are prioritized against the authority gaps that sit outside the page. Where the audit identifies missing media signals, the PR Marketplace offers editorial and sponsored placements across publishers relevant to finance, technology, AI and crypto, which can be matched to the gaps the audit surfaces.