All insights
PONTEM ONE · READOUT 07/16 1 JUL 2026
TechnicalPontem One · 5 min read · 1 Jul 2026

Your robots.txt is probably blocking the AI that should be citing you

GPTBot, ClaudeBot, PerplexityBot, Google-Extended, CCBot. Most CMS defaults disallow them, and nobody checks.

Listen to this piece
5:00
MP3 128k
#pod-article-robots-txt-ai
Article hero1600×900
Technical — article lead image 16:9 · WebP/AVIF · ≤180KB · #img-hero-robots-txt-ai

Every site has a robots.txt telling crawlers which paths they may fetch. For twenty years the only crawler that mattered commercially was Googlebot. That has not been true since 2023, and most healthcare sites have not been updated since.

The crawlers that now matter

GPTBot, ClaudeBot, PerplexityBot, Google-Extended and CCBot feed the systems currently answering patient questions. If robots.txt disallows them, your content is invisible to those systems no matter how well it ranks in classic search.

Why it is usually accidental

Several CMS platforms and security plugins ship with restrictive defaults disallowing any user-agent outside a small allowlist. Nobody chose it, nobody reviews it, and the effect is that a 200-page clinical library is unreadable to the exact systems that could be recommending the practice.

The fix, and the caveat

Check for blanket Disallow rules targeting AI user-agents and remove them. Add explicit Allow directives. Verify with a live fetch rather than assuming the file deployed. The caveat: this is a policy decision as much as a technical one — allowing these crawlers means allowing your content to inform models. Make it deliberately. Most practices, having made it, want to be visible.

Want this applied to your own site?

The audit covers the same ground on your domain: crawl, entity graph, and fifty live queries across four answer engines.

Request an audit