I made a spreadsheet of 180 B2B SaaS and fintech domains and checked three things: their robots.txt, whether GPTBot was explicitly blocked, and whether ChatGPT still cited them across 60 brand-name queries (company name + "review", "alternative", "vs competitor"). Ran all queries on ChatGPT-4o between June 15 and July 3.
The expectation was simple. Block GPTBot → don't get cited. That's how it's supposed to work.
31% of domains blocking GPTBot still showed up as citations. Direct source attribution, domain name, link and everything. Most had been blocking for 4+ months — I checked Wayback Machine timestamps on their robots.txt files. A few since late 2023.
So I went looking for where those citations were actually coming from.
The biggest leak was syndication. Close to half of the "blocked" citations traced back to sites that had republished or aggregated the original domain's content. Press release networks, industry roundups, those content syndication partners that nobody really tracks. The original site locked the door, but their content was already living on a dozen other sites with the door wide open.
Then I found the part that made me reconsider the whole approach. A huge chunk — maybe a third — came from Reddit threads, forum posts, and Q&A sites where users had quoted or paraphrased the blocked domain. Someone copies a paragraph from an article, or drops a specific stat into a comment, and ChatGPT picks up the forum post as the source. You can block every AI crawler on the planet and it doesn't matter if someone screenshots your chart and posts it to a subreddit.
The rest was murky. Some looked like cached versions on search engines. Some matched content structures from before the block was in place — old crawled data still influencing responses. I couldn't pin this down with certainty, but the pattern was consistent enough to be unsettling.
Here's what I keep turning over: the robots.txt approach to AI opt-out has a massive hole in it. You can control whether a crawler hits your server. You can't control whether your content has already been copied, quoted, summarized, or cached somewhere the crawler CAN reach.
For domains thinking blocking GPTBot means their content won't appear in ChatGPT — it doesn't. It just means the citation credit goes to whoever republished you. The aggregator gets the visibility. You get nothing.
I'm not sure what the fix is. Watermarking? Stricter syndication agreements? None of those scale well. The uncomfortable realization is that your content strategy in the GEO era isn't just about what you publish — it's about every surface where your content might be living without your knowledge.
Wondering if anyone here has found a practical way to monitor where your content is being reproduced. We've been doing manual searches and it feels like bailing out a boat with a spoon.