Field note
Six engines, six different 'best' agencies
I asked six AI assistants to name the best specialist in my own field. They gave six different answers — no two the same.
I asked six AI assistants the same question: who is the best independent specialist for AI search visibility? I expected some overlap — a name or two that kept coming up. Instead I got six different answers, with no two engines naming the same top pick. Not six lists that mostly agreed. Six different worlds.
So I checked whether that was a fluke. I took seven buyer-style questions — variations of "best agency," "who offers audits," "recommend a consultant," "who should I hire" — and asked all six engines (ChatGPT, Perplexity, Claude, Gemini, Copilot, Google AI Mode) the same wording, with web search on. That's 42 answers. Then I counted who got recommended.
The disagreement was the rule, not the exception:
- On none of the seven questions did all six engines name the same provider. Not once.
- Across the seven questions, more than 100 distinct agencies, consultants, and tools were named — and roughly three-quarters appeared in only one engine's answer.
- A provider reached even a majority (four of six engines) only a handful of times. The closest thing to a consensus pick still showed up in just four of the six.
One nuance worth keeping: the more specific the question, the more the engines converged. "Recommend an AEO consultant for a B2B SaaS company" produced the most agreement of the seven — but still no name that all six shared. Narrow the intent and the rooms overlap a little more; they never become one room.
And here's the part that surprised me most: all six engines searched the live web for these answers, and still diverged this much. So this isn't one engine answering from stale memory while another looks things up. They each searched, and each built a different shortlist from what they found — which means the disagreement is baked into how each engine selects and trusts sources, not into whether it bothered to look.
The honest caveats, because they're the whole point of doing this properly: this is one field, one set of questions, on one day, with search on, scored once. Engines drift between sessions, so the exact names would shift on a re-run — the finding is the low agreement, not the specific roster. And being named by an engine says nothing about whether a provider is actually any good; it only shows what the engine chose to say.
But the headline holds, and it matters if you're a company trying to be recommended: there is no single "the AI recommends us." Winning one engine's shortlist tells you almost nothing about the other five. Each engine is its own room, with its own door — which is exactly why an audit has to check all of them, the same way, rather than asking one and assuming the rest agree.
See also do AI engines cite the same sources? and who AI cites.
New here? Start with the overview → · By Denes Csaszar · written in the open