I'd been reading these AI visibility studies and something bugged me. They're all run in English, on US brands, and then we cite them like they apply to a business in Bogota or Valencia. I went looking for the Spanish-language version and found nothing, so we ran it ourselves.
I picked a market with plenty of real players, marketing agencies in Argentina, and wrote 100 searches the way an actual person would type them. I ran them across four engines, ChatGPT, Gemini, Google AI Overview and Perplexity. And here's the part almost nobody does, I repeated the same 100 searches three times, same day, without changing a single word. That's 1,200 queries in total.
The first thing that came out already felt big. Two AIs agree on average on just 9.6% of the brands they recommend. So out of every ten brands that show up, nine don't repeat in the other model. And when I pooled everything I ended up with 932 distinct agencies, of which 81% appeared in a single engine and in none of the other three.
But hold on, because that still wasn't the thing that broke my brain. Before comparing the models against each other I asked the obvious question. If I run the same search on ChatGPT three times in a row, same day, do I get the same brands?
And the answer is no. ChatGPT agrees with itself just 26% of the time. The most stable, AI Overview, reaches 39% and even that doesn't clear half. Meaning three out of four brands it names change between one answer and the next.
And that one finding breaks half the study, mine included. Because two models can't agree with each other more than each one agrees with itself. That 9.6% doesn't mean the AIs think differently, it means a big chunk of what looks like disagreement is actually noise.
One last thing I didn't see coming. Perplexity plays a different game, it recommends global networks like VML, Ogilvy or Dentsu, while the other three reward local agencies. One example, the same agency that shows up in 70 of the 100 searches for Gemini shows up in 9 for Perplexity.
What I take from it is simple. Measure all four engines separately because one doesn't predict the others, never trust a single query, and if your brand is local start with Google, since AI Overview and Gemini are the closest pair.
Required disclaimer, I work at CreceRank and we ran all of this on our own tool, so it could be biased. And the worst part is I've got nothing to compare it against, because there's nothing like this done for the Spanish-speaking market. If anyone has data that contradicts me or wants to rip the methodology apart, please do, that's the most useful thing for me.