Websites & AI visibility · Guide
How professional firms get recommended by ChatGPT and Google AI answers
Professional firms get recommended by ChatGPT and Google AI answers by doing search basics precisely: pages that can be crawled and indexed, plain answers near the top, consistent firm details across the web, and robots.txt rules that let AI search crawlers in. No special file or markup switches it on, and nobody can promise a citation.
- By
- Peter
- Published
- Last reviewed
- Reading time
- 8 min read
How do ChatGPT and Google choose which firms to mention?
They mostly choose from what their search systems can already find, read and trust. Google says its AI Overviews and AI Mode are rooted in its core ranking and quality systems, and that a page must be indexed and eligible to show with a snippet before it can appear as a supporting link (Google: AI features and your website). ChatGPT, Claude and Perplexity each run their own search crawler, and a site that blocks it drops out of, or fades from, their answers.
That gives two gates. The first is access: can the engine’s crawler reach and read your pages? The second is usefulness: when someone asks “which Geelong accountants look after SMSFs?” or “what happens at a first family law appointment?”, is your page the clearest and most specific answer the engine can find? Most of the work is careful, ordinary SEO. The AI-specific layer is small: crawler access, pages shaped around answers, and consistent details about the firm.
What works, and what doesn’t?
What works is the same work that earns a place in ordinary search results, done precisely. What doesn’t work is most of what is sold as “AI optimisation”: special files, special markup and rewriting pages for machines. Google’s guide to optimising for generative AI features, last updated in July 2026, says so in plain terms.
| Tactic | Verdict as at September 2026 | Why |
|---|---|---|
| Crawlable, indexable pages with the content in the HTML | Required | Google needs pages indexed and snippet-eligible. Vercel’s December 2024 crawler analysis found the major AI crawlers it studied, apart from Google’s and Apple’s, fetched JavaScript but did not run it |
| Letting AI search crawlers in, in robots.txt and at your CDN | Required | OpenAI says sites opted out of OAI-SearchBot are not shown in ChatGPT search answers |
| Answer-first pages with firm-specific detail | Works | Google asks for original, non-commodity content written for people |
| An accurate Google Business Profile and consistent firm details | Works | Google says Business Profiles can help services appear in AI responses and other results |
| Bylines, review dates and article markup | Helps | Google’s helpful content guidance asks whether authorship is self-evident |
| Genuine mentions: law society or CPA events, local press, association directories | Helps | Google warns that chasing inauthentic mentions “isn’t as helpful as it might seem” |
| An llms.txt file | No effect on Google | Google Search ignores it |
| “Chunking” pages or rewriting them for AI | Not needed | Google says there is no need to break content into tiny pieces or write a special way for AI |
| Special schema for AI answers | Not needed | Google says structured data isn’t required for generative AI search |
What should your robots.txt allow?
Allow the search and user-fetch agents that put your pages in AI answers, and make the training decision separately. The agents do different jobs, and blocking one does not block the others. OpenAI’s crawler documentation says each of its settings is independent, so a firm can allow OAI-SearchBot for search while disallowing GPTBot for model training.
| Job | Agents | What blocking does |
|---|---|---|
| Indexing pages for AI answers | OAI-SearchBot (ChatGPT), Claude-SearchBot (Claude), PerplexityBot (Perplexity), plus Googlebot, Bingbot and Applebot for ordinary search | Removes or reduces your visibility in that product’s answers |
| Fetching a page when a user asks | ChatGPT-User, Claude-User, Perplexity-User | Stops live look-ups. OpenAI says robots.txt rules may not apply to ChatGPT-User and Perplexity says Perplexity-User generally ignores them; Anthropic says Claude-User honours them |
| Model training | GPTBot, ClaudeBot, Google-Extended, Applebot-Extended | Signals that content should be left out of future training. Google-Extended also covers grounding in the Gemini app, but Google says it does not affect inclusion or ranking in Google Search |
For most firms, the simplest correct file lets every crawler in and keeps private paths out:
User-agent: *
Disallow: /portal/
Sitemap: https://www.example.com.au/sitemap.xml
If the partners decide to keep public pages out of model training, add a separate group for each training agent:
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: Applebot-Extended
Disallow: /
One detail catches firms out. A crawler obeys only the group that names it and ignores the * group, as Google’s robots.txt rules and the robots.txt standard, RFC 9309, both set out. If you add User-agent: OAI-SearchBot with Allow: /, that crawler stops reading the Disallow: /portal/ line under *. Either repeat private-path rules inside every named group or leave search crawlers to the * group, as above.
Three more checks:
- Your CDN or firewall. Edge settings act before robots.txt is read. Cloudflare, for example, lets site owners block AI crawlers by category: search, training or agent. Its training setting also blocks mixed-purpose crawlers that serve both search and training. Check the dashboard as well as the file.
- Timing. OpenAI says robots.txt changes take about 24 hours to reach ChatGPT search.
- Staging sites. Google says robots.txt is not a way to keep a page out of Google; use a password or
noindex. Put staging behind a login.
If a firm ever wants its pages out of AI Overviews and AI Mode but still in ordinary results, Google now offers a Search generative AI control in Search Console, rolled out to all websites by 31 August 2026. The default is to allow. Snippet controls such as nosnippet also limit AI features, because a page must be snippet-eligible to appear.
Do you need an llms.txt file or special schema?
Not for Google, and it is optional elsewhere. llms.txt is a proposal, first published in September 2024, for a Markdown file at /llms.txt that points language models to a site’s key pages. Google says Search ignores these files and that they neither help nor harm visibility. The crawler documentation from OpenAI, Anthropic and Perplexity describes robots.txt controls and does not mention it. If you want one, it is an hour’s work: a one-paragraph summary of the firm and links to your main service pages. Our glossary has a plain-English definition of llms.txt.
Structured data is worth doing for ordinary search, not as an AI lever. Google says structured data isn’t required for generative AI search and there is no special schema.org markup to add. It still states facts without ambiguity:
- organisation markup with your legal name, logo, address and profile links
- a business type such as LegalService or AccountingService, both LocalBusiness subtypes on schema.org
- article markup with each author’s name and the published and modified dates
One change to note: in June 2026 Google removed its FAQ rich result documentation because that result is no longer shown in Search (Search Central updates). FAQ markup will not earn extra space in results, so treat it as optional.
How should firm pages be written so engines can quote them?
Put the answer first and make it specific enough that no other firm could have written it. An engine quoting a page needs a sentence that stands alone and says who, what and where. Generic copy gives it nothing to use.
- Open every service page with a 40–60 word answer: what the service is, who it is for, and what happens next.
- Use headings phrased as the questions clients ask, then answer each one in the first sentence beneath it.
- Name the specifics: practice areas or tax work, locations, the systems you use, how fees are set and typical timeframes.
- Give each service its own page instead of listing ten services on one.
- Add a byline and a “last reviewed” date to guides, and change the date only after a real review.
- Keep the firm’s name, address, phone and a one-paragraph description identical on the site, Google Business Profile, LinkedIn and professional directories.
For example, picture a 20-lawyer firm whose estate planning page opens with “We provide estate planning services tailored to your needs.” An engine has nothing to quote. A stronger opening names the work and the reader: “We prepare wills, enduring powers of attorney and testamentary trusts for business owners and blended families in Melbourne’s eastern suburbs, with a fixed fee quoted after the first meeting.” The second sentence answers the question a client actually typed.
How do you know whether it is working?
Measure monthly with the same four checks, and judge the trend rather than any single result. AI answers vary from one run to the next, so one test proves little either way.
- Search Console. Google’s Generative AI performance report shows how often links to your site appeared in AI Overviews and AI Mode. It reports impressions, not clicks.
- A fixed prompt set. Write 20 questions your clients ask, run them monthly in ChatGPT, Claude, Perplexity, Gemini and Google AI Mode, and log whether the firm is named or linked.
- Server or CDN logs. Filter for OAI-SearchBot, Claude-SearchBot and PerplexityBot to confirm they are reaching your pages.
- Referrals. Track visits from chatgpt.com, perplexity.ai, claude.ai and gemini.google.com in your analytics.
Pylon Digital’s websites are built with these settings in place from launch, and Website Care includes ongoing AI-visibility monitoring.
What should a firm do in the next 30 days?
Fix access first, then firm details, then content, then set a baseline. Each week’s work is small, and the order matters: rewriting pages achieves little if crawlers cannot reach them.
| Week | Task | Owner |
|---|---|---|
| 1 | Check robots.txt and CDN bot settings; verify Google Search Console and Bing Webmaster Tools; submit the XML sitemap | Web developer |
| 2 | Make the firm’s name, address, phone and description identical everywhere; complete Google Business Profile; add organisation markup | Practice manager |
| 3 | Rewrite the five most important service pages answer-first, with bylines and review dates | Partner plus writer |
| 4 | Record a baseline: prompt set, crawler log check and the Search Console report | Web developer |
The detail differs by profession. Our law firm website checklist covers the advertising rules lawyers must work within, and our guide to accounting firm website essentials covers registration details and client document handling.
