Skip to content
Pylon Digital

Websites & AI visibility · Guide

How professional firms get recommended by ChatGPT and Google AI answers

Professional firms get recommended by ChatGPT and Google AI answers by doing search basics precisely: pages that can be crawled and indexed, plain answers near the top, consistent firm details across the web, and robots.txt rules that let AI search crawlers in. No special file or markup switches it on, and nobody can promise a citation.

Published
Last reviewed
Reading time
8 min read

How do ChatGPT and Google choose which firms to mention?

They mostly choose from what their search systems can already find, read and trust. Google says its AI Overviews and AI Mode are rooted in its core ranking and quality systems, and that a page must be indexed and eligible to show with a snippet before it can appear as a supporting link (Google: AI features and your website). ChatGPT, Claude and Perplexity each run their own search crawler, and a site that blocks it drops out of, or fades from, their answers.

That gives two gates. The first is access: can the engine’s crawler reach and read your pages? The second is usefulness: when someone asks “which Geelong accountants look after SMSFs?” or “what happens at a first family law appointment?”, is your page the clearest and most specific answer the engine can find? Most of the work is careful, ordinary SEO. The AI-specific layer is small: crawler access, pages shaped around answers, and consistent details about the firm.

What works, and what doesn’t?

What works is the same work that earns a place in ordinary search results, done precisely. What doesn’t work is most of what is sold as “AI optimisation”: special files, special markup and rewriting pages for machines. Google’s guide to optimising for generative AI features, last updated in July 2026, says so in plain terms.

TacticVerdict as at September 2026Why
Crawlable, indexable pages with the content in the HTMLRequiredGoogle needs pages indexed and snippet-eligible. Vercel’s December 2024 crawler analysis found the major AI crawlers it studied, apart from Google’s and Apple’s, fetched JavaScript but did not run it
Letting AI search crawlers in, in robots.txt and at your CDNRequiredOpenAI says sites opted out of OAI-SearchBot are not shown in ChatGPT search answers
Answer-first pages with firm-specific detailWorksGoogle asks for original, non-commodity content written for people
An accurate Google Business Profile and consistent firm detailsWorksGoogle says Business Profiles can help services appear in AI responses and other results
Bylines, review dates and article markupHelpsGoogle’s helpful content guidance asks whether authorship is self-evident
Genuine mentions: law society or CPA events, local press, association directoriesHelpsGoogle warns that chasing inauthentic mentions “isn’t as helpful as it might seem”
An llms.txt fileNo effect on GoogleGoogle Search ignores it
“Chunking” pages or rewriting them for AINot neededGoogle says there is no need to break content into tiny pieces or write a special way for AI
Special schema for AI answersNot neededGoogle says structured data isn’t required for generative AI search

What should your robots.txt allow?

Allow the search and user-fetch agents that put your pages in AI answers, and make the training decision separately. The agents do different jobs, and blocking one does not block the others. OpenAI’s crawler documentation says each of its settings is independent, so a firm can allow OAI-SearchBot for search while disallowing GPTBot for model training.

JobAgentsWhat blocking does
Indexing pages for AI answersOAI-SearchBot (ChatGPT), Claude-SearchBot (Claude), PerplexityBot (Perplexity), plus Googlebot, Bingbot and Applebot for ordinary searchRemoves or reduces your visibility in that product’s answers
Fetching a page when a user asksChatGPT-User, Claude-User, Perplexity-UserStops live look-ups. OpenAI says robots.txt rules may not apply to ChatGPT-User and Perplexity says Perplexity-User generally ignores them; Anthropic says Claude-User honours them
Model trainingGPTBot, ClaudeBot, Google-Extended, Applebot-ExtendedSignals that content should be left out of future training. Google-Extended also covers grounding in the Gemini app, but Google says it does not affect inclusion or ranking in Google Search

For most firms, the simplest correct file lets every crawler in and keeps private paths out:

User-agent: *
Disallow: /portal/

Sitemap: https://www.example.com.au/sitemap.xml

If the partners decide to keep public pages out of model training, add a separate group for each training agent:

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: Applebot-Extended
Disallow: /

One detail catches firms out. A crawler obeys only the group that names it and ignores the * group, as Google’s robots.txt rules and the robots.txt standard, RFC 9309, both set out. If you add User-agent: OAI-SearchBot with Allow: /, that crawler stops reading the Disallow: /portal/ line under *. Either repeat private-path rules inside every named group or leave search crawlers to the * group, as above.

Three more checks:

  1. Your CDN or firewall. Edge settings act before robots.txt is read. Cloudflare, for example, lets site owners block AI crawlers by category: search, training or agent. Its training setting also blocks mixed-purpose crawlers that serve both search and training. Check the dashboard as well as the file.
  2. Timing. OpenAI says robots.txt changes take about 24 hours to reach ChatGPT search.
  3. Staging sites. Google says robots.txt is not a way to keep a page out of Google; use a password or noindex. Put staging behind a login.

If a firm ever wants its pages out of AI Overviews and AI Mode but still in ordinary results, Google now offers a Search generative AI control in Search Console, rolled out to all websites by 31 August 2026. The default is to allow. Snippet controls such as nosnippet also limit AI features, because a page must be snippet-eligible to appear.

Do you need an llms.txt file or special schema?

Not for Google, and it is optional elsewhere. llms.txt is a proposal, first published in September 2024, for a Markdown file at /llms.txt that points language models to a site’s key pages. Google says Search ignores these files and that they neither help nor harm visibility. The crawler documentation from OpenAI, Anthropic and Perplexity describes robots.txt controls and does not mention it. If you want one, it is an hour’s work: a one-paragraph summary of the firm and links to your main service pages. Our glossary has a plain-English definition of llms.txt.

Structured data is worth doing for ordinary search, not as an AI lever. Google says structured data isn’t required for generative AI search and there is no special schema.org markup to add. It still states facts without ambiguity:

One change to note: in June 2026 Google removed its FAQ rich result documentation because that result is no longer shown in Search (Search Central updates). FAQ markup will not earn extra space in results, so treat it as optional.

How should firm pages be written so engines can quote them?

Put the answer first and make it specific enough that no other firm could have written it. An engine quoting a page needs a sentence that stands alone and says who, what and where. Generic copy gives it nothing to use.

  1. Open every service page with a 40–60 word answer: what the service is, who it is for, and what happens next.
  2. Use headings phrased as the questions clients ask, then answer each one in the first sentence beneath it.
  3. Name the specifics: practice areas or tax work, locations, the systems you use, how fees are set and typical timeframes.
  4. Give each service its own page instead of listing ten services on one.
  5. Add a byline and a “last reviewed” date to guides, and change the date only after a real review.
  6. Keep the firm’s name, address, phone and a one-paragraph description identical on the site, Google Business Profile, LinkedIn and professional directories.

For example, picture a 20-lawyer firm whose estate planning page opens with “We provide estate planning services tailored to your needs.” An engine has nothing to quote. A stronger opening names the work and the reader: “We prepare wills, enduring powers of attorney and testamentary trusts for business owners and blended families in Melbourne’s eastern suburbs, with a fixed fee quoted after the first meeting.” The second sentence answers the question a client actually typed.

How do you know whether it is working?

Measure monthly with the same four checks, and judge the trend rather than any single result. AI answers vary from one run to the next, so one test proves little either way.

  1. Search Console. Google’s Generative AI performance report shows how often links to your site appeared in AI Overviews and AI Mode. It reports impressions, not clicks.
  2. A fixed prompt set. Write 20 questions your clients ask, run them monthly in ChatGPT, Claude, Perplexity, Gemini and Google AI Mode, and log whether the firm is named or linked.
  3. Server or CDN logs. Filter for OAI-SearchBot, Claude-SearchBot and PerplexityBot to confirm they are reaching your pages.
  4. Referrals. Track visits from chatgpt.com, perplexity.ai, claude.ai and gemini.google.com in your analytics.

Pylon Digital’s websites are built with these settings in place from launch, and Website Care includes ongoing AI-visibility monitoring.

What should a firm do in the next 30 days?

Fix access first, then firm details, then content, then set a baseline. Each week’s work is small, and the order matters: rewriting pages achieves little if crawlers cannot reach them.

WeekTaskOwner
1Check robots.txt and CDN bot settings; verify Google Search Console and Bing Webmaster Tools; submit the XML sitemapWeb developer
2Make the firm’s name, address, phone and description identical everywhere; complete Google Business Profile; add organisation markupPractice manager
3Rewrite the five most important service pages answer-first, with bylines and review datesPartner plus writer
4Record a baseline: prompt set, crawler log check and the Search Console reportWeb developer

The detail differs by profession. Our law firm website checklist covers the advertising rules lawyers must work within, and our guide to accounting firm website essentials covers registration details and client document handling.

Questions

Frequently asked questions

Will blocking GPTBot remove our firm from ChatGPT search?

No. GPTBot collects content that may be used to train OpenAI's models, while OAI-SearchBot finds pages for ChatGPT search. OpenAI's crawler documentation says each setting is independent, so you can disallow GPTBot and still allow OAI-SearchBot. Blocking OAI-SearchBot is what removes a site from ChatGPT search answers, although it can still appear as a navigational link.

Should our firm block AI training crawlers?

It is a partner decision, not a technical one. Google says Google-Extended does not affect Google Search, and OpenAI says GPTBot is separate from ChatGPT search. Google-Extended also covers grounding in the Gemini app, so blocking it can reduce your presence there. Either way, robots.txt only governs public pages; confidential material should never be on the public site at all.

Does an llms.txt file help a firm appear in AI answers?

Not in Google. Google's guidance says Search ignores llms.txt files, so they neither help nor harm visibility in AI Overviews, AI Mode or ordinary results. The crawler documentation from OpenAI, Anthropic and Perplexity describes robots.txt controls and does not mention llms.txt. A short file costs about an hour, so treat it as optional housekeeping, not a strategy.

How long does it take to start appearing in AI answers?

There is no fixed timeframe. Access fixes act quickly: OpenAI says robots.txt changes reach ChatGPT search in about 24 hours. Being chosen as a source depends on crawling, indexing and how many other pages answer the same question well. Measure monthly with the same set of prompts and judge the trend over several months rather than one test.

Can a web agency promise our firm will be cited by ChatGPT or Google?

No. Answer engines choose sources for each question, and their systems change without notice. Google says no special markup or file is needed to appear in its AI features. What an agency can do is remove every avoidable reason not to be chosen: access, indexing, speed, clear answers and consistent firm details, then measure the result honestly.

Secure by design. Set up correctly. Fully managed.

Talk to us before you commit to anything

Start with a free 45-minute discovery call. We look at your systems and priorities, then recommend a first step with a fixed scope, or tell you if we are not the right fit.

Book a free 45-minute discovery call