Generative engine optimization (GEO) is the work of making a business easy for AI search tools to find, understand and cite correctly. ChatGPT search, Perplexity, Gemini, Copilot and Google's AI Overviews look up web pages when someone asks a question, so we make your facts consistent, crawlable, specific and worth quoting.
How AI Search Tools Find and Cite Sources
An AI model can know about your business in two different ways, and the difference shapes everything that follows.
Training is what a model learned from a large snapshot of text gathered months or years before you use it. You cannot edit that snapshot, and it may describe your business as it was, not as it is. Retrieval happens at the moment someone asks a question: the tool runs a search, reads a handful of current pages, writes an answer and links to its sources. Google's guide to optimizing for generative AI features (opens in a new tab) calls this retrieval-augmented generation, or grounding. GEO is mostly about retrieval, because that is where today's facts and today's citations come from.
Each tool retrieves from its own index or from a search partner's, and each company documents its own crawlers:
- Google AI Overviews and AI Mode pull from Google's search index using its core ranking systems, plus related searches Google calls query fan-out. Google says a page must be indexed and eligible to show a snippet, with no special extra requirements.
- ChatGPT search relies on OpenAI's OAI-SearchBot crawler. OpenAI's crawler documentation (opens in a new tab) says sites that opt out of OAI-SearchBot will not be shown in ChatGPT search answers, though they can still appear as navigational links.
- Perplexity uses PerplexityBot to surface and link websites in its results, and Perplexity's crawler page (opens in a new tab) says that crawler is not used to train AI foundation models.
- Microsoft Copilot citations are reported through Bing. In February 2026 Microsoft opened an AI Performance report in Bing Webmaster Tools (opens in a new tab), in public preview, that shows how often a site is cited in Copilot and in Bing's AI-generated summaries.
- Claude from Anthropic uses Claude-SearchBot to improve search results, and Anthropic's crawler article (opens in a new tab) notes that blocking it may reduce your site's visibility and accuracy in search results.
- Gemini is governed partly by Google-Extended, a robots.txt token that controls whether content Google crawls may be used to train future Gemini models and to ground answers in Gemini Apps. Google's crawler documentation (opens in a new tab) states that Google-Extended does not affect inclusion or ranking in Google Search.
The common thread: an AI tool can only cite what it can reach, and it can only describe you correctly if the pages it reaches agree with each other.
Entity Clarity: One Consistent Story About Your Business
In search terms, an entity is a distinct, identifiable thing: your business, its owner, its location. AI tools assemble answers from many sources. If your website lists one phone number, your Business Profile another and an old directory a third, the tool may pick any of them, blend them or hedge. Entity clarity means making every source tell the same story.
- Identical core facts: the same business name, phone number and address or service area on your website, Google Business Profile, Bing Places, Apple Business Connect and the directories that matter in your industry.
- A real About page: who owns the business, when it started, where it operates and what it does and does not do, stated plainly enough to quote.
- Organization or LocalBusiness markup: Google's Organization structured data documentation (opens in a new tab) says this markup can help Google understand your administrative details and disambiguate your organization, and recommends the most specific LocalBusiness type for a local business.
- sameAs links: inside that markup, links to your official profiles on other sites, such as your social accounts and industry listings, so systems can connect them to the same business.
- One summary sentence: a single plain description of what you do and where, reused across your site and profiles instead of five different taglines.
- Clear service-area wording: towns and counties named in text, not only shown on a map image.
We practice this on our own site. Every page, our structured data and our footer describe A1A Web Design the same way: a remote studio of AldoMedia, LLC. in Buffalo, New York, serving five Florida coastal counties by phone, video and email. No tool has to guess whether we have a Florida office. We do not.
Crawler Access: Decide Who May Read Your Site
The robots.txt file at the root of your site tells crawlers what they may fetch. Most AI companies now run separate crawlers for search and for model training, so you can allow one and refuse the other. OpenAI states this directly: each of its settings is independent, so a site can allow OAI-SearchBot to appear in search results while disallowing GPTBot.
| Company | robots.txt token | Documented purpose | If you block it |
|---|---|---|---|
| OpenAI | OAI-SearchBot | Surfacing websites in ChatGPT search | Not shown in ChatGPT search answers, though navigational links may still appear |
| OpenAI | GPTBot | Crawling content that may be used to train generative AI foundation models | Opts out of that training crawl; OpenAI says it is independent of search |
| Anthropic | Claude-SearchBot | Improving search result quality for Claude users | May reduce your visibility and accuracy in Claude's search results |
| Anthropic | ClaudeBot | Collecting web content that could contribute to model training | Opts your site out of that collection |
| Perplexity | PerplexityBot | Surfacing and linking websites in Perplexity search; not used for foundation model training | Your pages are not surfaced by that crawler in Perplexity results |
Google-Extended | Training future Gemini models and grounding in Gemini Apps and Vertex AI | No effect on Google Search inclusion or ranking, according to Google |
There are also fetchers that visit a page only because a person asked about it, such as ChatGPT-User, Claude-User and Perplexity-User. Perplexity says its user-initiated fetcher generally ignores robots.txt, since a user requested the page. That is a reminder that robots.txt is a request, not a lock. Google's own introduction to robots.txt (opens in a new tab) notes that not every crawler obeys it and that it is not a way to keep a page private.
Our default for client sites is to allow the search and retrieval crawlers, so your business can be found and cited, and to let you decide about training crawlers once the trade-offs are explained. We also check server logs and firewall settings, because security tools on some hosts block these crawlers without anyone noticing. OpenAI notes that its search systems can take about 24 hours to adjust after a robots.txt change.
Google's AI features follow different controls. AI Overviews and AI Mode are part of Google Search, so they respect Googlebot rules and snippet controls such as nosnippet, and Search Console now offers a Search generative AI control (opens in a new tab), on by default, that can exclude a site from those features. We leave it on unless a client has a specific reason not to.
llms.txt and ai.txt: Emerging Conventions, Not Requirements
llms.txt (opens in a new tab) is a proposal published by Jeremy Howard in September 2024: a Markdown file at the root of a site with the site's name as a heading, a short summary and lists of links to the most useful pages, meant to help AI agents use a website. It is a proposal, not an official standard, and support varies from tool to tool. Google is explicit about its own position: its generative AI guide says Google Search does not use llms.txt files, and that keeping one will neither help nor hurt your visibility in Google Search, while noting that it is fine to maintain one for other services that use it.
ai.txt is a separate proposal, introduced by Spawning in 2023, that lets a site state whether its text and media may be used to train AI models. Spawning passes those preferences to the partners that use its service. It does not control search crawlers and has nothing to do with being cited.
We include both files on the sites we build because they take little time to write, they summarize your business accurately and they record your choices in a readable place. We tell clients plainly that neither one is a ranking switch or a citation lever.
Citable Facts and Original Information
An AI answer is built from statements the system can trust and attribute. Vague marketing copy gives it nothing to work with; specific, checkable statements do. Pages that get cited tend to carry the kind of detail a careful reporter would want:
- Facts with numbers and context: the year you opened, the towns you serve, starting prices or price ranges you will honor, lead times and capacity, such as a 40-seat dining room or a six-passenger boat.
- Policies in plain words: deposits, cancellations, pets, refunds and what happens when a storm forces a closure.
- Original information: your own photos, project write-ups, lessons from your own jobs and numbers you collected yourself.
- Sources for claims: a link to the official source whenever you cite a regulation, a standard or a statistic.
Research points the same way. The paper that introduced the term, GEO: Generative Engine Optimization (opens in a new tab) by Pranjal Aggarwal and colleagues, including researchers at Princeton University (first posted in November 2023 and presented at KDD 2024), tested ways of editing content against a benchmark of queries. The authors report that GEO methods could boost visibility in generative engine responses by up to 40 percent, with adding citations, quotations from relevant sources and statistics among the strongest methods, while keyword stuffing offered little to no improvement. Those are results on a research benchmark, not a promise for any single business, but they match Google's advice to create original, non-commodity content.
Google's guide also warns against overdoing it. Creating separate pages for every variation of a question, mainly to manipulate rankings or AI responses, violates Google's scaled content abuse policy. More pages do not make a site more trustworthy. Better facts do.
Mentions on Trusted Third-Party Sites
Google notes that its AI features can show what is being said about products and services across the web, including blogs, videos and forums, and that seeking inauthentic mentions is not as helpful as it might seem. Other AI tools draw on the open web in similar ways. The mentions worth having are earned:
- Coverage in local news, community calendars and visitor guides.
- Member listings with your chamber of commerce and trade associations.
- Sponsorships of local events, teams and nonprofits, with a link back.
- Interviews, podcasts and expert quotes in your field.
- Honest reviews on the platforms your customers already use.
We never plant fake forum posts or buy reviews. Beyond being useless for long-term trust, fake reviews are illegal under the FTC's 2024 rule. Our local SEO service covers listings and review routines in detail.
Measuring AI Referrals
AI visibility is harder to measure than rankings, and anyone who shows you a single tidy number is simplifying. We combine several signals:
- Analytics referrals: visits from AI tools usually arrive with the tool's domain as the referrer, such as chatgpt.com, perplexity.ai, copilot.microsoft.com or gemini.google.com. We set up a saved report so these are visible at a glance.
- Google Search Console: the Generative AI performance report (opens in a new tab) shows impressions for links to your site in AI Overviews and AI Mode.
- Bing Webmaster Tools: the AI Performance report shows Copilot and Bing citations and the grounding queries that led to them.
- Logged prompt checks: each month we ask the main AI tools the questions your customers ask and record whether and how you are mentioned. Answers vary by person, place and day, so we look for trends, not single screenshots.
- Your contact form: an "AI assistant" option in the how-did-you-hear question captures people who never clicked a link at all.
Our GEO Process
Baseline
We record what the major AI tools say about your business today and note every error, gap and outdated fact.
Fix the facts at the source
Your website, Business Profile and key listings are corrected so they agree on name, contact details, hours and service area.
Add entity markup and a strong About page
Organization or LocalBusiness structured data with sameAs links, plus an About page that states the facts plainly.
Set your crawler policy
robots.txt rules you have chosen, llms.txt and ai.txt files, and a check of server logs and firewall rules to confirm the crawlers you allow can actually get in.
Make key pages citable
We rewrite service and FAQ pages with specific facts, original information and sources, working alongside our answer engine optimization methods.
Earn real mentions
We help you identify the local and industry sites where a genuine mention makes sense.
Measure monthly
Referrals, Search Console and Bing reports and logged prompt checks, explained in plain English.
What we will not promise: no one controls what an AI assistant says. Google itself warns about tools that claim access to internal ranking or AI metrics; none has that access. We will make your business easier to find, understand and cite, show you what changed and tell you when something did not work.
Frequently Asked Questions
Should I block AI crawlers from my website?
Block training crawlers if you prefer, but think carefully before blocking search crawlers. OpenAI, Anthropic and Perplexity each run separate crawlers for search and for training. Blocking a search crawler such as OAI-SearchBot or PerplexityBot can keep your site out of that tool's search answers. We explain each trade-off, and you decide.
Why does an AI assistant give wrong information about my business?
Usually because sources disagree or are out of date: an old phone number in a directory, former hours on a profile, or a website that never states your service area clearly. Some answers also come from training data gathered long ago. Fixing the facts at the source and keeping them consistent everywhere is the most reliable cure.
Will an llms.txt file get my business into ChatGPT answers?
Nothing about it guarantees that. llms.txt is a proposed convention, not an official standard, and Google says Google Search does not use it at all. We add one because it takes little time and gives a tidy, accurate summary of your site, not because it promises citations.
What is the difference between GEO and AEO?
AEO focuses on making the direct answer on a page easy to lift into featured snippets, voice replies and AI Overviews. GEO looks wider: whether AI tools such as ChatGPT, Perplexity, Gemini and Copilot can reach your site, recognize your business as one consistent entity and find specific facts worth citing. Both rest on solid SEO.
How can I tell if AI tools are sending visitors to my website?
Check your analytics for referrals from domains such as chatgpt.com and perplexity.ai, review Search Console's Generative AI performance report for Google's AI Overviews and AI Mode, and check Bing Webmaster Tools for Copilot citations. Adding an AI assistant option to the how-did-you-hear question on your contact form helps too.
Can you guarantee that AI tools will recommend my business?
No. Each AI tool chooses its sources with its own systems, answers vary from one person and one day to the next, and no outside company controls them. We make your business easier to find, understand and cite, and we report honestly on what changes.
Sources and Further Reading
- Google Search Central, Optimizing your website for generative AI features on Google Search (opens in a new tab)
- OpenAI, Overview of OpenAI Crawlers (opens in a new tab)
- Anthropic, Does Anthropic crawl data from the web, and how can site owners block the crawler? (opens in a new tab)
- Perplexity, Perplexity Crawlers (opens in a new tab)
- Google Search Central, Google's common crawlers (including Google-Extended) (opens in a new tab)
- llmstxt.org, The /llms.txt file proposal (opens in a new tab)
- Aggarwal et al., GEO: Generative Engine Optimization (arXiv 2311.09735, KDD 2024) (opens in a new tab)
