Most of the articles you've read about Google's AI search guidance cite the wrong document.
They link to Google's ranking systems guide, which is a useful page about BERT and PageRank and the site diversity system. It says almost nothing about AI Overviews. The guidance everybody is actually summarizing lives somewhere else, at Optimizing your website for generative AI features on Google Search, and it was last updated on July 10, 2026.
That's a small error. It's also a useful tell. When a writer cites the wrong source for a document they claim to have read, they didn't read it. They read someone else's summary of it, rewrote the summary, and shipped. Which happens to be the exact behavior Google's guidance is written to punish.
So let's do this properly. Here's what Google published, what it means for the way you write and build pages - whether you run an HVAC company, a plumbing business, or any other service business - and the three areas where following Google's advice to the letter still leaves you invisible.
Google's position in one sentence
AI search runs on the same ranking systems as regular search, so the work is the same work.
Google is explicit about this. Its generative AI features are "rooted in our core Search ranking and quality systems." Two techniques sit on top of that foundation:
Retrieval-augmented generation, which Google also calls grounding. The model doesn't answer from memory. It pulls current pages out of the Search index, reads them, and generates a response with clickable links back to the pages that support it.
Query fan-out, where the system spawns a set of related searches alongside the original one. Google's own example: someone asks how to fix a lawn full of weeds, and the system also runs "best herbicides for lawns," "remove weeds without chemicals," and "how to prevent weeds in lawn."
That second mechanism is quietly good news if you're small. Fan-out means the system is generating long-tail queries on the user's behalf, and pulling links from a wider pool than a classic ten-blue-links page would. Google says this openly: the fan-out process lets it "display a wider and more diverse set of helpful links" than classic web search. More slots, more variety, more chances for a page that would never have cracked the top three.
For an HVAC company, this means a query like "AC not cooling" can fan out into "common causes of AC not cooling," "how much does AC repair cost," and "best HVAC company near me" - and your page about the specific repair you performed last August has a shot at being pulled into the answer even if it doesn't rank in the top three for the head term. For a plumbing company, "water heater leaking" fans out into "water heater replacement cost," "emergency plumber near me," and "signs your water heater is failing" - and your case study about the tankless install you did in a 1998 Chesterfield ranch has a shot at being cited.
On the terminology fight, Google picked a side. Answer engine optimization and generative engine optimization, AEO and GEO, are terms Google acknowledges exist and then declines to adopt. "From Google Search's perspective, optimizing for generative AI search is optimizing for the search experience, and thus still SEO."
Fine. For Google, it's SEO. Hold that thought, because it stops being true the moment you leave Google's property.
The originality bar just moved, and this is the part that matters
Everything else in Google's guide is housekeeping. This section is the whole game.
Google draws a line between commodity and non-commodity content, and it gives two examples that are worth reading twice. Commodity content: "7 Tips for First-Time Homebuyers." Non-commodity content: "Why We Waived the Inspection & Saved Money: A Look Inside the Sewer Line."
Look at what separates them. The first could have been written by anyone, about anything, in any market, without leaving the house. The second contains a decision, a consequence, and a specific thing somebody actually looked at. It carries risk. Somebody could disagree with it.
For an HVAC company, the commodity version is "5 Signs Your AC Needs Repair." The non-commodity version is "Why We Replaced a 12-Year-Old Condenser Coil Instead of Patching It: A St. Louis August Case Study." For a plumbing company, commodity is "How to Prevent Frozen Pipes." Non-commodity is "The Slab Leak We Found Under a 1998 Chesterfield Ranch - and Why We Told the Owner to Wait." The pattern is the same in every trade: a specific decision, a specific property, a specific outcome, and a reason someone could push back on.
Google's instruction is direct: "Don't just recycle what others on the internet have already said, or could easily be produced by a generative AI model."
The reason this bites harder in AI search than it did in classic search comes down to how retrieval works. Grounding pulls multiple candidate sources for the same question and compares them. If your page says the same eight things as the other fourteen pages in the candidate set, there is no reason for the system to reach for yours. Interchangeable is worse than mediocre. Mediocre and distinctive still gets picked sometimes.
Google also runs original content systems built specifically to surface original reporting "ahead of those who merely cite it." That system predates AI Overviews. It didn't go away.
Before you write anything, answer one question: what's in this piece that could only come from us? Our data. Our client outcomes. A method we tested and can describe. A position we'll defend. If the honest answer is nothing, the angle is wrong and no amount of drafting fixes it.
Nobody is talking about the plagiarism problem, so let's talk about it
Here's the failure mode we see most often in agency content right now, and it's rarely called what it is.
A writer opens the three articles currently ranking for the target keyword. They pull the outline from the best one. They feed the text to a model with an instruction to rewrite it in a different voice. They swap synonyms, reorder two sections, add an intro, and publish it as original work.
That's plagiarism. The words changed. The reporting, the structure, the sequence of ideas, and the actual intellectual work did not. Nothing in that page was learned, tested, or verified by the person whose name is on it.
It's also a policy problem. Google's spam policies define scaled content abuse, and Google's guidance on generative AI content states that "using generative AI tools or other similar tools to generate many pages without adding value for users may violate Google's spam policy on scaled content abuse." Google points site owners at two specific sections of the Search Quality Rater Guidelines: section 4.6.5 on scaled content abuse, and section 4.6.6 on main content created "with little to no effort, little to no originality, and little to no added value."
Read 4.6.6 with the paraphrase workflow in mind. Little effort, little originality, little added value describes it exactly.
Google's position on AI as a tool is not prohibitionist, and we shouldn't misrepresent it. The guidance says generative AI "can be particularly useful when researching a topic, and to add structure to original content." Using a model to organize your thinking is fine. Using one to launder somebody else's article is not.
The rules we work to:
Quote or write, nothing in between. Anything word-for-word goes in quotation marks with a link at the point of use. Everything else gets composed from source material you've actually read, not reworded from a page you've skimmed.
Rebuild the outline. If your section order matches a competitor's section order, you copied their thinking even if you didn't copy their words. Start the outline from your research notes instead.
Source every number. Every statistic, study, price, date, and policy quote gets a link. If a claim can't be sourced, it comes out or it gets labeled as our own observation.
Never invent. No fabricated statistics, no invented client results, no made-up quotes, no hallucinated URLs. Models produce plausible-looking citations that lead nowhere. Load every reference link before publishing. A dead source at the bottom of an article does more damage than a typo, because it proves nobody checked.
Check images too. Licensed, owned, or created. If you're publishing AI-generated imagery on ecommerce pages, Google Merchant Center requires IPTC DigitalSourceType metadata set to TrainedAlgorithmicMedia, and AI-generated product titles and descriptions have to be labeled as such.
There's a business argument here beyond the policy one. If you're producing content that any competitor could produce by running the same three URLs through the same model, you are not building an asset. You're renting a position until somebody with a bigger budget runs the same play.
The technical floor is lower than the industry pretends
Google states the requirement plainly. To be eligible to appear as a supporting link in AI Overviews or AI Mode, a page has to be indexed and eligible to be shown in Google Search with a snippet. "There are no additional technical requirements."
That's it. No special file, no special markup, no separate qualification process. The site does need to be included in Search generative AI features in Search Console, and indexing is never guaranteed, but there's no secondary technical bar to clear.
Which means the technical work is the technical work you should already be doing:
- Crawling allowed in robots.txt, and not blocked by your CDN or host
- Important content present as text, not trapped in images or JavaScript that fails to render
- JavaScript following Google's JS SEO basics
- A decent page experience across devices
- Duplicate content reduced and canonicals pointed correctly
- Internal links that make your pages findable
- Structured data that matches the text visible on the page
- Site verified in Search Console
One line in the guide deserves attention because it contradicts a lot of paid advice: perfect semantic HTML is not required. Google's words are that "the web in general is not valid HTML, and Google can understand it." Use semantic markup because screen readers and browser agents benefit from it, not because you think the ranking system rewards it.
For measurement, Search Console now carries a Generative AI performance report covering generative AI features in Search and Discover. Google also gets pointed about third-party tools: "No third-party tool has access to our internal ranking or AI systems." Anything marketed as exposing internal Google metrics is selling you a guess with a confident interface.
The myth list, and why it's an integrity issue
Google published a section on what you can ignore. It reads like a list of things being sold right now as AI SEO packages.
For Google Search, you don't need:
- llms.txt or similar AI text files. Google Search "doesn't use them" and "ignores them." Keeping one neither helps nor hurts.
- Special AI markup or AI-only schema. Structured data isn't required for generative AI features. Use it for rich results.
- Chunking content into small blocks. Google's systems handle multiple topics on a page and surface the relevant part. There is no ideal page length.
- Rewriting content specifically for AI systems. The systems understand synonyms and intent. You don't need to capture every phrasing variant.
- Separate pages for every query variation or fan-out query. Google names this one directly as a scaled content abuse risk.
- Manufactured brand mentions. Google calls these "inauthentic mentions" and says chasing them "isn't as helpful as it might seem."
We think this list should end some conversations. If a vendor's AI search proposal is built on llms.txt files, AI schema, and chunking, that proposal has no mechanism behind it for Google, by Google's own documentation. Selling it after this guidance was published is not an honest mistake anymore.
Which brings us to the part Google's guidance doesn't cover, and where most of the actual work now lives.
Google's advice is correct and insufficient
Google says optimizing for AI search is still SEO. For Google's surfaces, that's accurate. Then people apply it to ChatGPT, Perplexity, Copilot, and Claude, where it stops being true, because those systems don't retrieve from Google's index and don't obey Google's crawl controls.
Every one of them uses separate crawlers with separate names, and every one of them is controlled separately in your robots.txt. If you've never looked, your site may already be invisible to some of them, and Search Console will never tell you.
OpenAI runs three bots. GPTBot collects content that may be used to train foundation models. OAI-SearchBot is the one that indexes pages so ChatGPT can retrieve and cite them in search. ChatGPT-User handles fetches triggered by a user action. Those are three separate decisions with three separate robots.txt tokens. A site that blanket-blocked "OpenAI bots" in 2024 to stay out of training data may have also blocked itself out of ChatGPT's search citations, which is a very different trade. Allowing OAI-SearchBot while disallowing GPTBot is a supported configuration.
Perplexity runs two, and documents them at Perplexity's crawler documentation. PerplexityBot is the search crawler, and Perplexity states it "is not used to crawl content for AI foundation models." Perplexity-User handles user-initiated fetches, and their documentation notes that because a person requested it, that fetcher "generally ignores robots.txt rules."
Perplexity's page also carries an operational detail that catches more sites than robots.txt does. If you run a WAF, you may be blocking these bots at the edge regardless of what robots.txt says. Perplexity publishes current IP ranges at perplexity.com/perplexitybot.json and perplexity.com/perplexity-user.json and walks through Cloudflare and AWS WAF allow rules. We've stopped assuming this is configured correctly on any site we inherit. Check the logs.
Anthropic runs three. ClaudeBot for training, Claude-User for fetches when someone asks Claude a question that needs a page, and Claude-SearchBot for indexing content that feeds Claude's search results. Anthropic states all three honor robots.txt, and blocking the training crawler doesn't block the search crawler. Same structure as OpenAI, same trap for anyone who blocked by vibes. (Search Engine Land's coverage walks through the granularity.)
Microsoft feeds Copilot from Bing's index, which makes Bing Webmaster Tools a real channel rather than a legacy one. Microsoft rewrote its webmaster guidelines to address Copilot grounding and citations directly, and in February 2026 it shipped an AI Performance report that tracks how often Copilot and Bing's AI answers cite your pages. There's no Google equivalent for cross-platform citation data at that granularity. IndexNow matters more here than it does for Google, because Bing crawls smaller sites less aggressively, and a page that isn't in the index can't be cited by anything downstream of it. (Search Engine Journal's coverage of the Bing guideline rewrite is worth reading in full.)
Now the detail that makes the point better than any of the above.
Google says llms.txt is useless, and for Google that's true. Perplexity's own documentation serves an llms.txt file and opens the page with a note addressed to AI agents telling them to fetch it. So "llms.txt does nothing" is accurate about Google Search and wrong as a general statement about the web. Google's guidance even allows for this, noting it's "completely fine" to maintain such files "for other services or systems that use these files."
That's the shape of the whole problem. Google's documentation is authoritative about Google and silent about everything else, and the industry keeps quoting it as though it settles questions it never addressed.
What actually travels across all of them
Strip away the platform specifics and a short list survives everywhere.
Answer the question early, in plain text, in a self-contained paragraph. Retrieval systems lift passages, not pages. A page that states its answer in the first hundred words gets quoted. A page that arrives at the answer on screen three gets skipped by every one of these systems, and by most humans.
Make claims checkable. Dated, attributed, sourced statements get reused more readily than confident assertions with nothing behind them. This is the same behavior that makes content trustworthy to a reader, which is not a coincidence.
Keep your entity data identical everywhere. Name, address, phone, and business description consistent across your site, Google Business Profile, Bing Places, directories, and social profiles. Inconsistent entity data is one of the most common reasons a real business fails to surface in AI answers, and it's boring enough that nobody wants to fix it.
Earn mentions on sources these systems actually retrieve from. Industry publications, review platforms, active forums, reference sites. Earn them. Google explicitly warns against manufacturing them, and the manufactured kind tend to look manufactured.
Decide your crawler policy on purpose. Sit down with your robots.txt and your WAF rules and make a deliberate call on each bot. Training access and search access are different questions with different business answers, and you're allowed to answer them differently.
Measure in more than one place. Search Console for Google, including the generative AI report. Bing Webmaster Tools for Copilot citations. Referral traffic segmented by assistant in analytics. Manual prompt testing against your target queries in each assistant, run on a schedule and logged, because it's the only way to see what these systems say about you when they don't send a click.
Where this leaves you
The uncomfortable read on Google's guidance is that it removed the shortcuts without lowering the difficulty. There's no file to add, no markup to deploy, no chunking scheme to implement. What's left is the part that was always hard: knowing something worth publishing and writing it down in your own words.
For an HVAC or plumbing company, the thing that could only come from you is the job you actually ran last week - the specific equipment, the specific neighborhood, the specific decision you made and the reason you made it. That's the content Google's retrieval system reaches for, and it's the content no competitor can produce by running your URL through a model. The same is true for every other service business: the electrician who can describe the panel upgrade they did on a 1960s ranch, the roofer who can explain why they recommended a full tear-off instead of an overlay, the landscaper who can walk through the drainage fix they designed for a yard that flooded every spring. Specific work, specific properties, specific decisions. That is the non-commodity content Google is asking for, and it is the content that travels across every AI system.
Go look at the last ten pieces published under your brand. For each one, name the thing in it that could only have come from you. If you can't do that for at least seven of them, the crawler configuration isn't your problem.
Want to see how your business scores across Google, Maps, and AI search? Run a free Search Authority Audit - it scans your site for answer readiness, entity clarity, local authority, structured data, and AI citation potential, then gives you a prioritized action plan. No obligation, and we will tell you if your service area is still available.
References
- Optimizing your website for generative AI features on Google Search - Google Search Central, last updated July 10, 2026
- AI features and your website - Google Search Central, last updated December 10, 2025
- Google Search's guidance on using generative AI content on your website - Google Search Central, last updated December 10, 2025
- A guide to Google Search ranking systems - Google Search Central, last updated December 10, 2025
- Google Search Essentials: Spam policies - Google Search Central
- Search Quality Rater Guidelines - Google, sections 4.6.5 and 4.6.6
- Creating helpful, reliable, people-first content - Google Search Central
- Guidance on third-party SEO tools and advice - Google Search Central
- Generative AI performance report - Google Search Console Help
- Perplexity Crawlers - Perplexity documentation
- Overview of OpenAI Crawlers - OpenAI Platform documentation
- Does Anthropic crawl data from the web, and how can site owners block the crawler? - Anthropic Support
- Anthropic's Claude Bots Make Robots.txt Decisions More Granular - Search Engine Land
- Bing Adds GEO To Official Guidelines, Expands AI Abuse Definitions - Search Engine Journal
- Bing Webmaster Guidelines - Microsoft Bing Webmaster Tools
- IndexNow - IndexNow protocol documentation
- Google Merchant Center policies for AI-generated content - Google Merchant Center Help