Home / Blog / Technical SEO for AI Search: Crawlers, Rendering and HTML

Technical SEO for AI Search: Crawlers, Rendering and HTML

Technical SEO for AI search means making sure AI systems can reach your pages, read the text in the raw HTML, and trust which URL is the real one. For Google’s AI Overviews and AI Mode, that’s ordinary SEO: Google says a page must be…

Technical SEO for AI Search: Crawlers, Rendering and HTML

Technical SEO for AI search means making sure AI systems can reach your pages, read the text in the raw HTML, and trust which URL is the real one. For Google’s AI Overviews and AI Mode, that’s ordinary SEO: Google says a page must be indexed and eligible to show with a snippet, nothing more. For ChatGPT, Claude and Perplexity, two things break most often. Your robots.txt or CDN blocks their crawlers, or your content only appears after JavaScript runs, which most AI crawlers don’t do.

So the job is short but specific. Check access per crawler, check what the server sends before any script runs, and keep status codes and canonicals boring. I’ll walk through each as a set of if/then decisions, because the right answer depends on how your site is built.

The Condition Map: What Changes Technical SEO for AI Search on Your Site?

Four variables decide how much work this is. Most sites I look at pass two of them without even trying, and then fail one so badly that no amount of content work elsewhere can make up for it. Usually it’s access.

  1. Which AI surfaces you care about. Google’s AI features ride on normal Googlebot crawling. Other assistants use their own bots, each with its own robots.txt token.
  2. How your pages render. A WordPress theme that sends finished HTML is in good shape. A JavaScript app that builds the page in the browser may look empty to most AI crawlers.
  3. What sits in front of your server. A CDN, firewall or security plugin can block bots before robots.txt is ever read.
  4. Your stance on AI training. Blocking training crawlers is a business call. Blocking search crawlers by accident is a mistake.

Does Google Need Anything Special for AI Overviews?

No. Honestly, this is the part I wish more clients heard first. Google’s guide to optimizing for generative AI features opens by saying SEO best practices “continue to be relevant” because its AI features are rooted in its core ranking and quality systems. It says a page must be indexed and eligible to show with a snippet, meeting the normal Search technical requirements.

The same guide says Google can process JavaScript when it isn’t blocked, and asks you to follow its JavaScript SEO best practices. It also lists things you can skip, such as AI text files and chopping content into tiny chunks. I cover the file question in my post on llms.txt, and markup in schema markup for AI search, so I won’t repeat either here.

If Googlebot can crawl and index the page, then it’s technically eligible for AI Overviews. Nothing else to switch on.

If you want to limit what AI features quote, then use nosnippet, data-nosnippet or max-snippet. Google’s AI features documentation names these as the controls.

If you’re tempted to block Googlebot to stay out of AI Overviews, then stop. You’d leave Search entirely. Google-Extended is a separate token, and Google says it doesn’t affect inclusion or ranking in Search.

For the full crawl, index and render audit, my technical SEO checklist goes stage by stage. This post only covers what’s different for AI systems, so for classic blockers like stray noindex tags, canonical conflicts and redirect chains, see my 12 most common technical SEO issues.

Which AI Crawlers Should You Allow in robots.txt?

Here’s the part most guides flatten. Each AI vendor runs different bots for different jobs, and blocking one doesn’t block the others. I sort them into three roles, using each vendor’s own crawler page as of September 2026.

VendorTraining crawlerSearch crawlerUser-triggered fetcher
OpenAIGPTBotOAI-SearchBotChatGPT-User
AnthropicClaudeBotClaude-SearchBotClaude-User
Perplexitynone listedPerplexityBotPerplexity-User
GoogleGoogle-Extended (a control token, not a separate bot)Googlebotnot applicable

The vendors spell out the consequences. OpenAI says sites that opt out of OAI-SearchBot won’t be shown in ChatGPT search answers, while disallowing GPTBot only signals your content shouldn’t train its models. Perplexity says PerplexityBot surfaces sites in its results and isn’t used to crawl content for foundation models. Anthropic says blocking Claude-SearchBot reduces your visibility in search results.

User-triggered fetchers are the odd ones. OpenAI says robots.txt rules “may not apply” to ChatGPT-User, and Perplexity says Perplexity-User “generally ignores” them, because a person asked for the page. So robots.txt isn’t a full opt-out for those.

If you want AI search visibility but not training, then allow the search crawlers and disallow the training ones:

User-agent: OAI-SearchBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

If you’re happy to be used for training too, then you don’t need any of these lines. A site with no AI-specific rules is open to all of them by default.

If you run several subdomains, then repeat the rules on each one. Anthropic’s help page says so directly, and it’s true of robots.txt generally, since each host has its own file.

For Perplexity’s side of this in more depth, including the stealth crawling dispute, see my Perplexity SEO guide.

Branch 1: When a CDN or Firewall Sits in Front of Your Site

This one catches people. A lot. Your robots.txt says “welcome” and your CDN says “no” before the bot ever reads it. The crawler just gets a block page or a challenge.

Cloudflare is the common case because so many small sites use it. On July 1, 2025, Cloudflare announced it would block AI crawlers by default on new domains and ask owners upfront whether to allow them. Its current documentation describes new defaults from September 15, 2026 for new domains: training and agent bots blocked on pages that show ads, search bots still allowed, and mixed-purpose crawlers that do both search and training also blocked.

My read? The defaults are sensible for publishers who sell ads. They’re a problem if you never looked at them and assumed you were visible. Settings live under Security settings, then the AI bot policy options, per Cloudflare’s docs.

If you use Cloudflare or another CDN, then open its security event log, filter by the user agents in the table above, and look for blocks or challenges.

If you find blocks you didn’t choose, then change the AI bot policy or add an allow rule for the specific verified bots you want. Keep the rest of your bot protection on.

If you block by IP address instead, then know the trade-off. Anthropic warns that IP blocking may not work reliably as an opt-out, because it also stops the bot from reading your robots.txt.

Security plugins on WordPress can do the same thing quietly. When I audit a site that “never shows up in ChatGPT”, the firewall log is one of the first places I look.

Branch 2: When Your Content Depends on JavaScript

Do AI crawlers render JavaScript? Mostly no. And for once there’s a primary source rather than a hunch, which is rare in AI search advice, where a single screenshot on social media tends to become an industry rule within a week. Vercel’s December 2024 analysis, The Rise of the AI Crawler, looked at crawler traffic across its network and found that none of the major AI crawlers it tracked rendered JavaScript. ChatGPT’s and Claude’s crawlers did fetch JavaScript files. They just didn’t execute them.

The same analysis found two exceptions. Gemini uses Googlebot’s infrastructure, so it gets full rendering, and AppleBot renders through a browser-based crawler. That study is almost two years old, so treat it as the best public evidence rather than a permanent rule. Vendors don’t publish rendering specs, so I plan for the conservative case.

If your main content is in the server’s HTML response, then you’re fine. Most WordPress, Shopify and static sites land here.

If product details, prices, reviews or article text load through client-side scripts, then assume AI crawlers other than Gemini’s see an empty shell. Move that content to server-side rendering, static generation or prerendering.

If only extras load by script, such as chat widgets or related-post carousels, then leave them. Vercel’s own recommendation was to keep client-side rendering for non-essential enhancements.

Testing is quick. Fetch the raw HTML without a browser and search it for a sentence from the page:

curl -s -A "Mozilla/5.0 (compatible; test)" https://www.example.com/your-page/ | grep -c "a sentence from your page"

A count of zero means the text isn’t in the initial HTML. Google may still see it after rendering, but most AI crawlers won’t. Broader JavaScript SEO for Google is its own topic, so here I only care about this one question.

Do Status Codes and Canonicals Matter to AI Crawlers?

Honestly, the AI vendors don’t document how their crawlers treat canonicals or redirect chains. So I won’t pretend to know their rules. What I do know is that confused signals hurt Google’s AI features for certain, since those depend on normal indexing, and clean signals can’t hurt anyone else.

  • Real pages return 200. Error pages return a real 404 or 410, not a 200 with “not found” text. That pattern is a soft 404, and it wastes crawls.
  • Moved pages use a single 301. Chains of three or four redirects are fragile for any bot.
  • One URL per piece of content. Parameter and duplicate versions point a canonical at the main URL, and internal links use the main URL too.
  • Key pages aren’t behind a login or a cookie wall. If a bot can’t see it logged out, it can’t cite it.

What Does Clean, Text-Accessible HTML Look Like?

Crawlers read text. That’s the whole game. If your important facts live in images, PDF viewers or scripts, an AI system has less to work with. Here’s what I check.

  • Facts in text, not in images. Prices, hours, specs and comparison tables belong in HTML, not in a screenshot.
  • A heading structure that matches the content. One H1, then H2s and H3s that describe what each section answers.
  • Real HTML tables for tabular data. A table built from divs and scripts may not survive text extraction.
  • Transcripts for key videos and alt text for meaningful images. It’s the text version a crawler can actually read.
  • Content in the page, not only in tabs that load on click. Tabs are fine when the text is already in the HTML.

None of this is new. It’s the accessibility and SEO work that good sites already did, which is exactly why Google says fundamentals apply.

Should You Use IndexNow for AI Search?

It helps with Bing, and Bing matters here because Bing Webmaster Tools now reports citations of your pages in Microsoft Copilot and Bing’s AI summaries. IndexNow is a protocol that pings search engines when a URL is added, updated or deleted. The IndexNow project site lists Microsoft Bing, Naver, Seznam.cz, Yandex and Yep as participants. Google isn’t on the list.

Bing’s May 2025 post on IndexNow tied faster indexing to staying current in search shaped by “immediacy and AI”. That’s a reasonable claim, not a measured one. If your CMS or SEO plugin supports IndexNow, then turn it on, since it costs nothing. If it doesn’t, then I wouldn’t build a custom integration just for this.

Edge Cases I Watch For

  • Staging sites left open. An indexable staging copy can get crawled and cited in place of the live site. Protect it with a login.
  • Country or bot blocking at the host. Some hosts block traffic from data center IP ranges, which is where most crawlers live.
  • Paywalls. If you want citations, the bot needs enough visible text to understand the page.
  • Rate limits. Aggressive limits can return 429 errors to crawlers during a busy crawl. Check your logs for them.

The Decision Matrix

If your site…And…Then
Is indexed in GoogleYou only care about AI OverviewsNothing extra; keep normal SEO healthy
Uses Cloudflare or a WAFYou never reviewed AI bot settingsCheck the event log and AI bot policy today
Builds content with client-side JSYou want ChatGPT or Claude citationsMove key content to server-rendered HTML
Allows all botsYou don’t want training useDisallow GPTBot, ClaudeBot and Google-Extended only
Blocks search crawlersYou want AI search visibilityAllow OAI-SearchBot, Claude-SearchBot and PerplexityBot
Has soft 404s or redirect chainsAny AI goalFix them first; they affect Google’s AI features too

When Should You Hand This to a Developer?

Bring in a developer when the fix means changing how pages render, such as moving a React or Vue storefront to server-side rendering. The same goes for firewall rules on a site that handles payments. A wrong WAF rule can open the door to bad bots too.

If you’d rather have the whole thing reviewed, our technical SEO service covers crawl access, rendering and indexing together, or you can start with a free SEO audit. And if you want to know whether AI answers are costing you clicks at all, read my companion piece on AI Overviews traffic impact.

Frequently Asked Questions

Do AI Crawlers Follow robots.txt?

The training and search crawlers from OpenAI, Anthropic and Perplexity say they do. User-triggered fetchers are different. OpenAI says robots.txt rules may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores them, because a person requested the page.

Will Blocking GPTBot Remove My Site From ChatGPT Search?

No. OpenAI treats them separately. GPTBot covers training, and OAI-SearchBot controls whether your site can appear in ChatGPT search answers. Block GPTBot if you don’t want training use, and leave OAI-SearchBot allowed if you want to stay visible.

Does Google Need Server-Side Rendering for AI Overviews?

Not strictly. Google says it can process JavaScript content when it isn’t blocked, and its AI features use normal Search indexing. Server-rendered HTML still helps because it’s faster to process and it’s also readable by AI crawlers that don’t run scripts.

Is Technical SEO for AI Search Different From Regular Technical SEO?

Mostly no. The new parts are managing separate AI user agents, checking CDN bot settings, and making sure content works without JavaScript for crawlers that don’t render. Crawlability, status codes, canonicals and clean HTML are the same work you’d do for Google.

Last updated: September 2026 by Mizanur Rahman

Put this guide to work.

Want help applying it? Start with a free audit of your site. We’ll show you what to fix first.

Get a free SEO audit