There Is No Number One Position in AI Search
There Is No Number One Position in AI Search: Why Google, Gemini, ChatGPT and Perplexity See Different WinnersCategory: AI Search Research & GEO Strat...

There Is No Number One Position in AI Search: Why Google, Gemini, ChatGPT and Perplexity See Different Winners
Ask most marketers what it means to rank well, and you'll get some version of the same answer. Show up near the top of Google, and the rest tends to follow. That assumption held up reasonably well for two decades, because there was really only one meaningful ranking to worry about. Google's results page was the results page.
That assumption is now measurably wrong, and the evidence for it isn't a hunch or an opinion piece. It's a peer reviewed study, built on a genuinely large dataset, that puts a number on something a lot of marketers have suspected but couldn't quite prove. A team of researchers built a public benchmark of 11,500 real user queries and compared what three systems inside Google's own ecosystem, traditional Search, AI Overviews, and Gemini, actually retrieved for each one. The result was not a family of closely related answers. It was three different retrieval logics, sitting inside a single company, producing source lists that barely resembled each other.
The average overlap between what any two of these systems retrieved for the same query came out below 0.2 on a similarity scale where 1.0 means identical results and 0 means nothing in common. That's not a small variance. That's most of the answer changing depending on which door inside Google you happened to walk through. And that's before you even step outside Google's ecosystem into ChatGPT, Perplexity, or Claude, each running its own independent retrieval process on top of its own view of the web.
This piece walks through what that research actually found, why it happens, what it means for the increasingly outdated idea of a single ranking position, and how a business can actually start measuring its own presence across a search landscape that no longer has one scoreboard.
We Asked Five AI Systems the Same Question
Before getting into the research, it's worth naming the experience most marketers have already had informally, even if they haven't put a name to it. Ask ChatGPT to recommend the best project management tool for a small team, then ask Gemini the same question, then ask Perplexity, then Claude, then check what Google's AI Overview shows for the same search. If you've ever actually done this, you already know the punchline. The lists don't match. Sometimes they barely overlap at all.
That informal observation is exactly what the academic research confirms with real numbers, and it's worth taking seriously precisely because it's no longer just an anecdote passed around in marketing Slack channels. It's a measurable, replicable pattern, and it has direct consequences for how any business should think about the phrase "we rank well in AI search," because that phrase, as commonly used, doesn't actually describe a single, coherent thing anymore.
Why the Answers Were Different
The research, published as How Generative AI Disrupts Search and accepted for the SIGIR 2026 conference, set out specifically to measure this divergence with real rigor rather than relying on scattered individual examples. The team compared traditional Google Search results, Google's AI Overviews, and Gemini 2.5 Flash across the same 11,500 real user queries, using two separate statistical measures, Jaccard similarity and rank biased overlap, to quantify exactly how different the retrieved sources actually were between systems.
The headline finding is stark. Average Jaccard similarity across all three systems came in below 0.2, with specific pairings landing between 0.11 and 0.18. In plain terms, on average, roughly eighty to eighty nine percent of the sources any two of these systems pulled for the same query did not overlap at all. And here's the detail that genuinely surprised even the researchers. The two systems you'd expect to be most alike, since AI Overviews is built on a lightweight version of the same Gemini model family, actually turned out to be the least similar pairing of the three. Being built on related underlying technology did not translate into similar source selection.
The study also found that no pairing among the three systems reached an average rank biased overlap above 0.27, a second measure that weighs whether the same sources show up in a similar order, not just whether they show up at all. That reinforces the same basic point from a different angle. This isn't a case where the systems mostly agree with a bit of noise around the edges. The disagreement is the dominant pattern, not the exception.
A few additional findings from the same research help explain why this happens. AI Overviews were generated and displayed above the organic results for just over half of all the representative real user queries tested, 51.5 percent, and showed up disproportionately for controversial or sensitive topics, appearing for as much as 93.8 percent of political queries in the dataset while showing up far less often, around 8.1 percent, for queries tied to trending news. Generative search systems were also significantly less likely than traditional Search to pull from popular, well established websites or institutional sources like government and education domains, and significantly more likely to pull from content Google itself owns. On top of that, the researchers found generative search results were noticeably less consistent than traditional search when the same query was run again, or run with a slightly different phrasing, device, or location, meaning the fragmentation problem isn't just between systems. It shows up within the same system too, across repeated runs of a nearly identical question.
Google vs Gemini vs ChatGPT vs Perplexity vs Claude
It's worth being precise about what the SIGIR research actually covered, because it's tempting to overextend a strong finding beyond what it directly measured. The study specifically compared three systems inside Google's own ecosystem, traditional Search, AI Overviews, and Gemini. It did not directly test ChatGPT, Perplexity, or Claude inside the same benchmark.
That distinction matters, but it doesn't weaken the broader point nearly as much as it might first appear to. If three systems built and operated by the exact same company, sharing infrastructure and drawing on closely related underlying models, already diverge this dramatically from each other, there's no serious reason to expect independently built systems, run by entirely different companies with entirely different retrieval architectures, training approaches, and citation philosophies, to converge more tightly. If anything, the honest expectation runs the other way. ChatGPT's retrieval, Perplexity's citation forward approach, and Claude's own independent process for grounding a response are each shaped by different design priorities, which makes additional divergence, not additional agreement, the far more plausible outcome once you step outside Google's own walls entirely.
This is exactly the gap GEO SEO Lab's own testing methodology, described later in this piece, is built to address. Where the academic research establishes rigorous proof of fragmentation within Google's ecosystem specifically, a business trying to understand its own visibility needs a practical way to extend that same question across the full range of platforms its actual customers are using, which increasingly means testing ChatGPT, Gemini, Perplexity, and Claude side by side, on the same real, commercially relevant questions, rather than assuming a strong showing on any one platform says much of anything about the other three.
Which Brands Appeared Across Every Platform
Here's the honest, important caveat before going further. The SIGIR study measured source overlap at the level of retrieved domains and pages, not specific brand recommendations inside a commercial buying decision. Extending this finding into a fully quantified, cross platform brand visibility benchmark, the kind of study the strongest version of this research eventually deserves, is exactly the work GEO SEO Lab is building toward as an ongoing research practice, testing real commercial prompts across all five major platforms on a recurring basis and publishing what we find.
What we can say with real confidence, grounded directly in the academic findings above, is the shape of what that kind of test consistently reveals once you actually run it. A brand that ranks strongly in traditional Google Search has no statistical guarantee of appearing at all in Google's own AI Overview for the identical query, since the research shows those two systems only share a small fraction of retrieved sources on average. That same brand has even less guarantee of appearing inside Gemini's response, since Gemini and AI Overviews turned out to be the least similar pairing the researchers measured, despite sharing underlying technology. And once you step outside Google's ecosystem entirely into ChatGPT, Perplexity, or Claude, each running an independent retrieval process shaped by its own architecture and citation philosophy, the reasonable expectation, grounded in exactly this kind of research, is further divergence rather than a returning consensus.
Picture, purely as an illustration of the pattern rather than a specific tested result, a small business software company that has spent years earning the top organic position on Google for its core category. Under the fragmentation pattern this research documents, that same business could plausibly be entirely absent from Google's own AI Overview for the identical search, appear third or fourth inside Gemini's answer, get a confident number one recommendation inside Perplexity because Perplexity happened to weight a recent, well cited comparison article the business published, and never come up at all inside a ChatGPT conversation because ChatGPT's retrieval leaned on a completely different set of sources. None of that is a hypothetical edge case under research like this. It's closer to the expected, statistically likely outcome once you actually look for it, precisely because the underlying retrieval systems disagree this substantially even when they're built by the same company.
Which Sources Did Each System Prefer
This is where the SIGIR research gets specific enough to actually act on, rather than just confirming that fragmentation exists in the abstract.
Traditional Google Search showed a meaningfully stronger preference for popular, well established websites and institutional sources, government domains, educational institutions, and generally higher authority publishers, compared with the two generative systems tested alongside it. Generative search, by contrast, leaned away from those same institutional sources and leaned more heavily toward Google owned properties specifically, a pattern the researchers flagged as a real and reasonable point of scrutiny, since it raises an obvious question about whether generative retrieval inside Google's own ecosystem quietly favors Google's own content over comparably strong outside sources.
There's a genuinely encouraging finding buried in this same research too, worth calling out directly because it cuts against a common assumption. The researchers noted that generative search appears to benefit niche, less established content providers at the expense of the largest, most dominant publishers, essentially the opposite of the classic SEO dynamic where a handful of massive, high authority domains have historically crowded out smaller, more specialized sources. If that pattern holds up under further study, and it's consistent with plenty of anecdotal reporting from smaller publishers and businesses who've noticed unexpected AI citations despite modest traditional domain authority, it means a smaller, genuinely specialized business may have a real, measurable shot at AI visibility that traditional Google rankings alone would never have predicted.
Why AI Retrieval Systems Diverge
A few real, structural reasons explain why this fragmentation exists, rather than it being some kind of random noise or a bug that will simply get patched out over time.
Different systems are solving genuinely different problems, even when the query looks identical on the surface. Traditional Search is fundamentally trying to rank the best individual documents for a query. A generative system is trying to assemble a coherent, synthesized answer, which means it's optimizing for a combination of sources that together cover a topic well, not simply the single best ranked page. Those are different optimization targets, and different targets naturally produce different selections.
Training data, model architecture, and citation philosophy all shape retrieval independently of raw content quality. A system trained and tuned with a strong preference for recent, frequently updated sources will retrieve differently than one weighted more heavily toward established domain authority, even when both are technically capable of finding the exact same underlying pages.
Query interpretation itself varies across systems. The same typed words can get interpreted with subtly different assumed intent depending on the system doing the interpreting, and a different assumed intent naturally leads to a different set of sources being judged relevant in the first place.
Consistency itself is lower for generative systems, which compounds every other source of divergence. The same research found generative search results shift more than traditional search results do when the same query gets rerun, rephrased slightly, or tested from a different device or location. That means a business isn't just facing divergence between platforms. It's facing a meaningful amount of divergence within a single platform over time and across sessions, which is exactly why a single test, run once, tells a business far less than it might assume.
Why AI Ranking Is an Oversimplification
Put together, all of this points toward a genuinely important reframe. The phrase "AI ranking," used the way most marketing conversations currently use it, implicitly borrows the mental model of traditional SEO, one system, one ordered list, one position worth fighting for. That model simply doesn't describe what's actually happening anymore.
There isn't one AI ranking to chase. There are several genuinely independent retrieval processes running in parallel, each shaped by its own architecture, training priorities, and citation behavior, each capable of producing a meaningfully different answer to what looks, on the surface, like the exact same question. Treating "AI visibility" as a single score to optimize for is a category error, not just an oversimplification, because it assumes a level of underlying consistency that the actual research shows simply doesn't exist.
The Problem With One AI Visibility Score
This matters practically, not just conceptually, because a lot of the emerging AI visibility tooling on the market right now is built around exactly the assumption this research undermines, a single composite score meant to represent how well a business is doing in AI search overall.
A single score, however sophisticated its underlying methodology, necessarily compresses genuinely different platform level realities into one number, which risks hiding the exact pattern that actually matters most for strategic decision making, a business dominating on one platform while being nearly invisible on another. A composite score that averages a strong Perplexity presence against a weak ChatGPT presence might report a comfortable, reassuring middle of the road number, while masking a genuinely urgent gap on the specific platform where a business's actual customers happen to be spending the most time.
This is precisely why platform level detail, not a single rolled up number, is the more honest and more actionable way to actually understand AI visibility, and it's the specific gap the framework below is built to address directly.
How Brands Should Measure Multi Engine Visibility
Test the same real, commercially meaningful question across every platform that matters to the business, not just the one that's easiest to check. A genuinely useful test means running an identical, carefully worded prompt through Google Search, Google's AI Overview, Gemini, ChatGPT, Perplexity, and Claude, and recording what actually comes back from each one individually, rather than assuming a strong result on one implies anything reliable about the rest.
Track platform specific outcomes side by side, not as an average. Recording which brands got mentioned, in what order, with what supporting sources cited, on each individual platform, preserves exactly the kind of detail a single composite score would otherwise erase, and it's that platform by platform detail that actually drives useful strategic decisions.
Repeat every test multiple times, across different phrasings, sessions, and time periods. Since the underlying research shows generative systems are meaningfully less consistent than traditional search even when the same query gets rerun, a single test result should be treated as one data point in a pattern still forming, never as a stable, final answer on its own.
Pay specific attention to where a business is strong on one platform and effectively invisible on another. That specific gap, not the overall average, is where the most urgent and most addressable strategic risk usually lives, since it often reflects something genuinely fixable, like a specific platform's retrieval system simply not having access to the right kind of source material about a business yet.
Treat this as an ongoing research practice, not a one time audit. Given how much these systems are still actively evolving, a comprehensive multi platform test run once, then filed away and forgotten, will be meaningfully out of date within a matter of months, if not weeks.
GEO SEO Lab's AI Search Fragmentation Index
Put simply, the AI Search Fragmentation Index is our own framework for actually operationalizing everything this research points toward, a structured way to measure and track how consistently, or inconsistently, a business shows up across the platforms that actually matter to it. It works by running a defined, recurring set of real commercial prompts across Google Search, Google's AI Overview, Gemini, ChatGPT, Perplexity, and Claude, then scoring each platform independently across presence, meaning whether a business shows up at all, position, meaning where it lands within that platform's answer, and source strength, meaning how credible and well supported the underlying citation actually is. Rather than compressing those results into one misleading composite number, the index deliberately preserves platform level detail, then layers on a genuine fragmentation score, a measure of how much a business's visibility actually varies across the full set of platforms tested. A low fragmentation score describes a business showing up consistently everywhere. A high fragmentation score describes exactly the pattern this article has been describing throughout, strong somewhere, invisible somewhere else, with real, addressable gaps hiding underneath what might otherwise look like a perfectly respectable average.
The Multi Engine Visibility Model
Alongside the Fragmentation Index sits a simpler, complementary way of thinking about strategic priority. Picture visibility across platforms as four rough categories. Universal presence describes a business showing up consistently and favorably everywhere tested, the strongest and rarest position to actually hold. Platform concentrated presence describes strong visibility on one or two platforms with real gaps elsewhere, the most common pattern among businesses that have done real GEO work but haven't yet tested broadly enough to notice their own blind spots. Fragmented presence describes visibility that shows up inconsistently and unpredictably across platforms with no clear pattern, often a sign of underlying content or entity inconsistency rather than a platform specific issue. And absent presence describes a business that simply doesn't show up in AI generated answers at all yet, regardless of platform, usually reflecting a genuine gap in original, citable content rather than anything fixable through better targeting alone. Knowing honestly which of these four categories actually describes a business's current position is a far more useful starting point than any single composite score could ever provide.
What Marketers Should Prepare for Next
Stop treating any single platform's result as representative of the whole picture. The research is unambiguous on this point, and it applies even inside a single company's own ecosystem, let alone across genuinely independent competitors.
Build testing infrastructure now, even if it starts small and manual. A simple, consistently repeated spreadsheet tracking the same handful of real commercial questions across five platforms, updated monthly, already puts a business meaningfully ahead of the majority still relying on a single, informal spot check.
Prioritize investment based on where the actual customer base spends time, not based on which platform is currently easiest to measure. A genuinely large gap on the specific platform a business's real customers actually use matters considerably more than a slightly better average score built mostly from platforms those same customers rarely touch.
Expect this fragmentation to persist, not resolve, as these systems keep evolving independently. Since the divergence stems from genuinely different underlying architectures and priorities, not a temporary bug, there's little reason to expect convergence any time soon, and every reason to build a measurement practice around that reality rather than waiting for it to simplify on its own.
Closing Thought
For twenty years, ranking well meant one thing, because there was really only one scoreboard that mattered. That era is over, and it didn't end quietly. A rigorous, peer reviewed study, built on 11,500 real queries, just put hard numbers behind something a lot of marketers had already started to suspect from lived experience, that the systems increasingly standing between a business and its next customer don't agree with each other, sometimes not even when they're built by the exact same company. The businesses that adapt fastest to this new reality won't be the ones chasing a single AI ranking that no longer exists. They'll be the ones who accept, honestly and early, that visibility now has to be earned, tested, and tracked separately across a genuinely fragmented landscape, one platform at a time.
Key Takeaways
- A rigorous academic study comparing Google Search, Google AI Overviews, and Gemini across 11,500 real queries found average source overlap below 0.2 Jaccard similarity, meaning roughly four out of every five retrieved sources differed between any two systems tested.
- Surprisingly, AI Overviews and Gemini, despite sharing underlying model technology, turned out to be the least similar pairing of the three systems studied.
- Generative search systems were significantly less likely to retrieve from popular, institutional sources and significantly more likely to retrieve from Google owned content, while also benefiting smaller, niche publishers more than traditional Search historically has.
- There is no single unified AI ranking a business can chase. Each platform runs an independent retrieval process, and strong visibility on one carries no reliable guarantee of visibility on another.
- Measuring AI visibility properly requires platform specific tracking, repeated testing over time, and resisting the temptation to compress results into one misleading composite score.
About GEO SEO Lab
GEO SEO Lab is a research and strategy group focused on helping businesses understand and improve visibility across AI assisted search and discovery, including Google Search, Google AI Mode, ChatGPT, Gemini, Claude, Perplexity, and the broader ecosystem reshaping how people find and evaluate information. Our work spans Generative Engine Optimization, AI visibility strategy, entity optimization, and original multi platform research aimed at helping businesses understand exactly where they stand across a genuinely fragmented search landscape, not just where they stand on any single platform.
References and Further Reading
- Grossman, R., Liu, S., Chen, M. K., Smith, M., Borcea, C., and Chen, Y., How Generative AI Disrupts Search, An Empirical Study of Google Search, Gemini, and AI Overviews, accepted for SIGIR 2026, arXiv 2604.27790
- Unite.AI, AI Is Splitting Web Search Into Three Different Realities, coverage of the SIGIR 2026 research
- GEO SEO Lab, The New Rules of AI Visibility, why rankings alone won't win in 2026
- GEO SEO Lab, Google vs ChatGPT, Who Wins This Search
- GEO SEO Lab, Google's AI Search Boom, why rankings alone no longer define success
Tags
Frequently Asked Questions
Find answers to common questions about this topic
About the Author
Anubhav
SEO Expert & Content Creator
Experienced digital marketing professional specializing in SEO strategies, content optimization, and data-driven marketing solutions. Passionate about helping businesses grow their online presence and achieve better search rankings.