
How Generative Engine Optimization Actually Works
Generative engine optimization improves a website's eligibility and usefulness as a source for AI answers. The practical work is crawler access, clear answers backed by original evidence, consistent business information, and repeated measurement. Crawling, retrieval, grounding, corroboration, and attribution are a useful model for diagnosing gaps, not a universal sequence every engine follows. You control what you publish and how your site responds. Each engine decides what to retrieve, cite, and say.
Generative Engine Optimization (GEO) work is easier to evaluate when you separate the parts you can improve from the decisions an engine makes. The five stages below are our diagnostic model. Operators document different retrieval systems, and their internal processes are not fully public. Use the model to find missing evidence or access problems, not to promise a citation.
What actually decides whether you get cited
- Crawling, indexing, citations, visits, and enquiries are different outcomes. A successful fetch proves access; it does not prove selection or demand.
- Historical Cloudflare counts on our site show requests claiming AI-fetcher identities. Those counts need operator verification before they can establish which bots made the requests.
- Retrieval is not your ranking. Late-2025 analyses by BrightEdge and Originality.AI put the share of AI Overview citations drawn from the organic top 10 between roughly 17% and 52%, so being cited is a separate outcome from being ranked.
- Publish evidence a buyer can check: useful comparisons, original observations, and sources for material claims. Experimental visibility results do not supply a universal citation uplift.
- Independent sources can help buyers and search systems evaluate a claim. Repeating an assertion across several sites does not make it true.
The pipeline, and where the levers are
| Stage | What the engine does | What you control | Strongest lever |
|---|---|---|---|
| 1. Crawl | May fetch your page or use indexed content | Site response, not crawl timing | Crawler access, render speed, static HTML |
| 2. Retrieve | Breaks the question into sub-questions, pulls candidate passages | Partly | Topic coverage, heading relevance to the sub-question, freshness |
| 3. Ground | Uses selected sources to support an answer | Published evidence, not source selection | Clear answers, context, linked sources |
| 4. Corroborate | May compare a claim with other evidence | Evidence you provide, not the engine’s judgment | Independent support and original data |
| 5. Attribute | Decides whether and how to name a source | Barely | Entity clarity, being the origin of the fact |
Everything below is a walk through that table.
Stage 1: the crawl, and the three bots that are not the same bot
Search crawlers, training crawlers, and user-requested fetchers have different jobs. Blocking direct access can limit how your pages are retrieved, but a brand may still appear through other sources or existing index data.
Using OpenAI’s fleet as the example, documented in OpenAI’s bots reference, the roles split cleanly:
- GPTBot concerns content that may be used for model training. Its robots policy is independent of SearchBot.
- ChatGPT Search crawler supports ChatGPT Search. Allowing it and its verified network addresses supports source access; it does not guarantee inclusion.
- ChatGPT-User retrieves content for user actions. A matching user-agent string alone proves neither operator identity nor a qualified visitor.
Check each operator’s documentation. PerplexityBot supports search; ClaudeBot concerns training, while Claude-SearchBot supports search. Google-Extended is a robots.txt policy token, not a separate HTTP crawler. Google Search inclusion is a separate control.
Our historical week-to-July-27, 2026 log report counted 2,370 requests claiming ChatGPT-User and 1,390 claiming Googlebot. Those totals were classified by user agent. They cannot establish verified crawler volume or human demand without additional identity evidence. Keep claimed requests separate from verified bots, referrals, and leads.
Two checks follow. Important text and links should be available in the returned HTML, so retrieval does not depend on a crawler executing your scripts. Google can render JavaScript, but rendering support differs across systems. Also compare access logs with analytics without treating them as the same population: a bot fetch is not a browser session.

Stage 2: retrieval, or why your ranking stopped predicting your citations
Some AI search features use query fan-out: they search related questions to gather sources for a response. Google documents this behavior for its AI features. The exact queries and retrieval process vary by product and request, so this is not a universal sequence you can reverse-engineer from one answer.
That decomposition is why ranking and citation have come apart. You are not competing for one query anymore, you are competing for a spray of sub-questions you never see, and the candidate set for each is drawn more broadly than the top ten links. The published late-2025 measurements disagree on magnitude but not direction: BrightEdge’s rank-overlap tracking (September 2025, 16.7% of citations from top-10 results) and Originality.AI’s citation study (November 2025, 52%) bracket the share of AI Overview citations drawn from the organic top 10 between roughly 17% and 52%.
What you control here is coverage and shape, not the fan-out itself:
- Cover useful follow-up questions. Keep related information together when that helps the reader. Separate pages should serve distinct needs, not merely target minor query variations. Google’s AI optimization guidance warns against generating such pages primarily to manipulate rankings or AI responses.
- Use headings that describe the answer. A question-and-answer format can help when readers have specific questions. It is not a required AI format or a reason to turn every article into an FAQ.
- Stay fresh with real updates. Recency carries unusual weight in these systems. A real update changes the content; a date bump changes nothing and the engines increasingly see through it.
Stage 3: grounding, and the finding that named the field
Grounding connects an answer to selected evidence. You can improve your page’s clarity and support for a claim, but the engine controls which sources and passages it uses.
The GEO benchmark study (Aggarwal et al., the 2024 ACM SIGKDD Conference on Knowledge Discovery and Data Mining) tested nine content modifications against roughly 10,000 queries. The winners were adding verifiable statistics, incorporating credible quotations, and citing reliable sources. Improving fluency helped less. Keyword stuffing produced negligible or negative effects.
Be careful with the headline number from that paper, because it is quoted badly everywhere including by agencies selling against it. The often-repeated “40% lift” is a relative gain over a 19.8% baseline on a position-adjusted word-count metric, measured among five sources that were already injected into the model’s context. It is a share-of-answer result inside a simulator, not a measure of whether you get discovered. A 45-study scoping review published in July 2026 makes the point directly, and C-SEO Bench (NeurIPS 2025) found most such methods are ineffective or actively negative on ranking, with traditional SEO outperforming them, and gains shrinking as more people adopt them.
The practical lesson is to make evidence useful and verifiable. Do not turn a benchmark’s visibility score into a forecast for your site. Its models, source set, prompts, and scoring method differ from a live buyer’s search.
Translated into editing rules:
- Answer the question clearly. Put the important point early and preserve the context needed to interpret it. There is no required answer length.
- Use numbers when they help the decision. Give the source, period, units, and limits. Do not add statistics solely to make a passage look authoritative.
- Link material claims to evidence. Original sources let readers verify a claim and assess whether it applies to them.
- Use real tables. Semantic HTML tables of options, costs, or benchmarks are dense, unambiguous, and easy to extract. Comparison content is disproportionately retrieved because comparison is what buyers ask for.
- Match certainty to evidence. Write directly, but keep limitations that would change a buyer’s decision. An uncertain claim does not become stronger when its caveat is removed.

Stage 4: corroboration, the stage nobody works on
Independent evidence can make a claim easier to assess. That does not mean every engine runs a separate corroboration step, or that repeating a claim across listicles makes it a fact. Seek accurate coverage and useful participation; avoid manufactured mentions or copied assertions presented as independent proof.
This is why GEO stops being a website project. The surfaces that do the corroborating are ones you do not own:
- Round-ups and “best X” lists, because they pre-package exactly the comparison an engine is trying to build. Being in someone else’s list beats publishing your own.
- Communities and forums, which are heavily represented in both training data and live retrieval. Genuine participation compounds; spam gets you filtered.
- Earned coverage and third-party mentions, including unlinked ones. An unlinked mention still teaches the model an association between your brand and a topic, and associations are what get recalled.
- Original data that other people quote. This is the move with the largest payoff, because it inverts the relationship: instead of chasing mentions, you publish the number everyone else has to cite. Every citation of your data is a corroboration event you did not have to negotiate.
Stage 5: attribution, and why you get named or skipped
The last stage decides whether the engine says “according to you” or just says the thing. You have the least control here, and two factors do most of the work.
Entity clarity. The model needs to know who you are with enough confidence to name you. Inconsistent company names, missing or contradictory profiles across LinkedIn, Crunchbase, and review platforms, and vague descriptions of what you do all make attribution risky, and engines skip risky attributions. Fixing this is unglamorous data hygiene and it is load-bearing.
Being the origin. Engines attribute to the source of a fact more readily than to a site that repeated it. If the number originated with you, the attribution follows the number. If you summarized someone else’s research, the citation usually goes to them. This is the strongest argument for original data that exists, and it is not a GEO argument at all, just being worth citing.
What you cannot control, and should stop trying to
Three things are outside the pipeline you can influence, and pretending otherwise burns quarters.
The fan-out. Engines choose related searches, and a sampled answer does not expose every retrieval decision. Cover real reader needs. Generating slight query variants primarily to manipulate search visibility is the problem Google warns against.
The wording of the answer. The model paraphrases. You can influence what it grounds on, never how it phrases the result.
Whether the user clicks. Citation, visit, and purchase are separate decisions. Track visibility alongside qualified visits, enquiries, and revenue. A prompt panel is a sample of answers, not a conversion rate or proof that a specific edit caused a gain.
What the pipeline produced on our own site
We ran this on strataigize.com before selling it to anyone, which is the only reason we are willing to publish specifics.
A late-2025 push on the previous site, published 2025-12-10, took ChatGPT-attributed sessions from 38 to 474 (engagement rate 44.7% to 53.8%) and Perplexity from 21 to 76, with 117 Semrush AI citations as of 2025-12-09. Leads arrived attributed directly to both engines. The full before-and-after tables are in the AI SEO case study. AI search was an emerging channel when we started; the reason we test those early is in why we test emerging channels.
The revenue picture, stated with its caveat attached: both of the only two clients our website has ever produced came through AI answers, worth $215,603 together, or 29.7% of our lifetime closed-won revenue. Roughly 95% of that is a single deal. That proves the channel closes at our deal size and proves nothing about a rate. We keep the caveat attached because a number this good is exactly the kind that gets misused.
One structural gotcha worth stealing: the larger deal was tagged “Organic Search” in our CRM, and the client identified the AI-assisted path. Google now documents a separate Generative AI performance report for AI Overviews and AI Mode impressions. Those impressions do not identify which lead clicked an AI answer. Keep first-party visibility reports and the buyer’s account of how they found you as separate evidence.
The agent-readable layer, as it actually works
We also run a public Model Context Protocol (MCP) server at mcp.strataigize.com, an agent-to-agent endpoint, and published Agent Skills. Agents can read our current published content and discover the consultation request flow. That makes the site easier to use through those interfaces; it does not establish a search advantage.
What it is not, today, is a citation lever. Machine-readable endpoints and agent manifests have no demonstrated effect on whether ChatGPT names you in an answer, and we are not going to pretend otherwise while the evidence is absent. Build the five stages above first. Treat the agent layer as positioning and future-proofing.
Where to go next
- Start here if the terms are the problem: answer engine optimization (AEO) vs GEO vs SEO explained.
- The plain-language introduction: what generative engine optimization is.
- Buying rather than building: what AI SEO services include and how to vet an agency.
- Tooling for each stage: the best AI SEO tools.
- See where you currently stand: the AI visibility checker, or our managed generative engine optimization service.
Frequently asked questions
Should I block AI crawlers? Decide separately for search, training, and user-requested retrieval. Blocking a search crawler can limit direct access to your content; it does not erase every mention of your brand. Verify operator identity and the effective robots and firewall policies before changing access. For Google, also check the property’s Search generative AI inclusion setting.
Does llms.txt help me get cited? There is no evidence that it does. It costs almost nothing to publish and it is reasonable hygiene for agent consumers, but treat any agency selling it as a citation strategy with suspicion. The levers are content, corroboration, and entity clarity.
How often do AI engines re-check a page? There is no universal schedule. Search indexing, cached content, and user-requested retrieval follow different paths. Use verified logs and the operator’s reports to see what happened on your site. Update material facts when they change rather than bumping dates to suggest freshness.
Why do I show up in one engine and not another? Engines use different retrieval systems, product settings, and source selections. Results can also vary by prompt, time, location, and user context. Record those conditions and repeat measurements before treating one missing citation as a site defect.
Is any of this different for B2B? The same access and evidence checks apply, but buying needs vary. For B2B readers, publish the information they need to evaluate a supplier: scope, proof, pricing context, implementation requirements, and a useful next step. Measure qualified enquiries as well as visibility.
Talk to us