If you ask one person who the best accountant in town is, you get an opinion. If you ask twenty people, you get a market. The difference is not that any one of the twenty is smarter than the first person. It is that twenty answers let you see which names keep coming up, which names come up once, and how confident you should be in either.
AI market intelligence has the same structure, and most of the industry ignores it. Tools that track "what AI says about your brand" typically query one model, or treat several models as interchangeable and pool the results. This article puts numbers on why that is a mistake. We looked at the most recent edition of every one of the 120 categories QuadrantX tracks, each scored independently by five production AI models: Claude Sonnet 4.6, GPT-5, Gemini 2.5 Flash, DeepSeek, and Perplexity's Sonar Pro. Then we asked a simple question. How often do they agree?
The Headline Number: 38%
All five models named the same leading vendor in 38% of categories. At least four of the five agreed in 62%. A simple majority agreed in 85%. In the remaining 15% of categories, no majority existed: the five models split at least three ways on who leads the market.
Put differently: if you rely on a single AI model to tell you who leads your market, you will get the "consensus" answer in roughly six categories out of ten and something else in the other four. You will have no way of knowing which situation you are in.
Multi-Model Consensus — A measurement approach that queries several independent AI models with the same prompt and treats agreement between them as the signal and disagreement as the uncertainty. A vendor named by every model has earned a stable position in AI's picture of the market. A vendor named by one model has earned a data point.
Below the Leader, Agreement Falls Apart
The leader is the easy case. The shortlist is where buyers actually make decisions, and the shortlist is where models diverge.
We compared each model's top five vendors in every category with every other model's top five. On average, two models' top-five lists overlapped by about half. Across all five models, a typical category produced ten distinct vendors in its combined top-five lists, twice what perfect agreement would produce. Fewer than two vendors per category, on average, appeared in every model's top five. In one category in five, no vendor did.
Go further down the list and the divergence becomes the norm. Across all 120 categories, 41% of every vendor mentioned was mentioned by only one model. Just 16% were mentioned by all five. The long tail of AI recommendations is not a shared long tail. It is five separate ones, and a buyer using one assistant sees only theirs.
Every Model Has a Personality
The disagreement is not random noise. Each model has consistent tendencies that show up across categories.
- GPT-5 is the generous lister. It named an average of 27 vendors per category, against roughly 20 to 22 for the others, and it produced the most vendors that no other model mentioned. Its shortlists are longer and its long tail is wider.
- Sonar Pro is the contrarian. Perplexity's model, which searches the live web before answering, named a leader that differed from the group's plurality pick in about a third of categories, more than any other model. Its picks are often more recent and more specialised.
- Claude is the consensus-seeker. Claude Sonnet 4.6 diverged from the plurality leader in only about one category in six, the lowest rate of any model. Gemini 2.5 Flash was close behind.
- Gemini and Sonar Pro are the optimists. Both scored vendors about four points higher on Sentiment, on average, than DeepSeek, the most reserved model. The same vendor, described by two models, can read as enthusiastically recommended or cautiously acknowledged.
None of these tendencies is a flaw. They are the reason multi-model analysis works. A model that lists more vendors surfaces challengers earlier. A model that retrieves the live web catches narrative shifts before training-data models do. A model that hews to consensus is a good baseline. Pooling them without knowing which is which throws that information away.
A single model tells you what one system believes. Five models tell you what is actually settled, what is still contested, and where the market is moving first.
Consensus Is Stable. Single Models Are Not.
QuadrantX regenerates every category on a rolling schedule, which means we can watch the same question asked dozens of times over months. That history reveals the second reason consensus matters: it is far more stable than any individual model's output.
In roughly half of our categories, the consensus leader has held its position in more than nine editions out of ten this year. SAP Concur has led expense management in 98% of editions. McKinsey has led management consulting in 98%. Salesforce has led CRM in every edition. Amazon Web Services, MongoDB, DocuSign, Shopify, and Tesla have never been displaced in theirs.
These are categories where the models agree with each other, and because they agree, the answer does not move. The unstable categories are the ones where they do not. In the AI visibility and observability category, the lead has changed hands in more than half of consecutive editions. In investment advisory, the lead has changed hands in roughly half of consecutive editions. A single-model reading of either category on any given day is a coin flip dressed up as a finding.
This is also why the Narrative Dominance score weights agreement explicitly. A vendor mentioned prominently by all five models earns a higher score than one mentioned prominently by one, even if the single mention is glowing. The score is designed to reward the stable signal, not the loud outlier.
What Consensus Actually Buys You
1. A way to separate signal from noise
When four models name a vendor and one does not, the omission is interesting but not alarming. When one model names a vendor and four do not, the mention is interesting but not evidence of market position. Without multiple models you cannot make this distinction, and you will over-react to both.
2. An early-warning system
Because the retrieval-based model responds to new content within days while training-data models respond within months, divergence between them is a leading indicator. When Sonar Pro starts naming a challenger the other four have never heard of, something has changed in the market's public narrative. That divergence is invisible in a single-model tool and invisible in a pooled score.
3. Protection against entity confusion
In the digital advertising software category, the five models named five different leaders: Google Ads, Google Display & Video 360, Google Marketing Platform, "Google," and The Trade Desk. Four of those are the same company described at different levels of granularity. A single model would have reported one of them as the definitive answer. Five models expose the ambiguity, which is itself a finding about how the category is understood.
4. Honest uncertainty
The most valuable output of multi-model analysis is not a better point estimate. It is a confidence level. A leader that is unanimous across models and stable across editions is a fact about the market. A leader that changes with the model is a hypothesis. Buyers, marketers, and analysts deserve to know which one they are looking at. See our earlier experiment, asking six AI models the same question, for a hands-on illustration.
When you evaluate any AI visibility measurement, ask three questions. How many models were queried? Are their results reported separately or pooled? And is disagreement between them treated as noise to be averaged away, or as information to be reported? If the answers are "one," "pooled," or "averaged," you are looking at an opinion, not a measurement.
The Case Against Averages
It is tempting to respond to model disagreement by averaging. Query five models, take the mean, report a single number. This is better than one model, but it discards the most useful part of the data. The variance between models is not measurement error. It is the shape of the market's uncertainty, and it tells you where a vendor's position is secure and where it is up for grabs.
That is why QuadrantX reports per-model scores alongside the consensus, and why our bias-reduction approach treats each model as a distinct witness rather than a redundant sensor. Five witnesses who agree are worth listening to. Five witnesses who disagree are worth listening to even more carefully.
The Bottom Line
AI recommendations are becoming the first shortlist most buyers see. Those recommendations agree on the leader in fewer than four categories out of ten, and on the shortlist far less often than that. Any measurement built on a single model is measuring one assistant's opinion and calling it the market. Multi-model consensus is not a methodological nicety. It is the minimum required to know what you are looking at.