SandsDX
Back to Perspectives
AI & Automation 11 min read

The AI Shortlist Report: What One Model Release Did to 21 B2B SaaS Categories

SandsDX original data: identical buyer questions through Claude Fable 5 and its predecessor across 21 B2B SaaS categories, 258 runs. Three category leaders flipped, 13% of top-five places changed hands, and the models know about rebrands they refuse to use.

Page Sands · · Updated

Fifteen years building B2B SaaS revenue systems: Microsoft agency side, Drift, Avalara, Blackbaud, ConnectWise. About →

A SandsDX original data report. 21 B2B SaaS categories, 2 frontier models, 258 API runs, July 2026. Methodology at the end; raw data available on request.

Update, July 7, 2026: Wave 2 retested this study inside a larger run across eight models from five AI labs. The membership findings below replicate. The individual number-one flips proved sensitive to how the question is asked, and from Wave 2 forward we report crown changes only when they survive multiple phrasings. Details in the Wave 2 report.

Executive summary

In July 2026 Anthropic shipped Claude Fable 5. We asked it the question a B2B software buyer asks, “who are the best vendors in this category,” across 21 categories, five times each, and compared its answers to Claude Opus 4.8, the model it replaced. Nothing happened in the market between the two sets of answers. Every difference comes from one thing: the model changed.

Shortlists that changed13 of 21Categories whose top five changed between the old model and the new one.
Places that moved13%Share of all top-five places that changed hands, with no market event in between.
The naming gap14 ptsThe models agree on which companies belong (79%) more than on what those companies are called (65%).
Young vs settled categories16% vs 10%Turnover in the eleven GTM-stack categories vs the ten biggest ones. The young stack moves harder.
  • 13 of 21 category shortlists changed. 14 vendors entered a top five, 13 exited. 13% of all top-five places changed hands.
  • Three category leaders lost the number one slot under this exact question: Drift to Intercom in conversational marketing, Demandbase to 6sense in ABM, Artisan to 11x in AI SDR. Four more categories flipped when a company’s product lines are counted together.
  • The naming gap: 79% vs 65%. The models agree on which companies belong far more than they agree on what those companies are called. The 14-point gap is rebrand lag, measured inside the models.
  • The GTM stack churned harder than the majors, 16% to 10%. The tools revenue teams buy sit in the categories that move the most.

If 81% of B2B buyers choose their vendor before ever talking to sales, these are market share events. They happened in a week, and almost nobody who gained or lost a slot knows.

The method, in one paragraph

Two groups of categories: the ten largest in B2B SaaS (CRM, marketing automation, HCM, help desk, project management, BI, cloud ERP, endpoint security, video conferencing, e-commerce) and eleven from the GTM stack (sales engagement, ABM, revenue intelligence, conversation intelligence, AI SDR, customer success, and five more). Both models got the identical buyer question five times per category through the API, with no web search attached. That last part is deliberate: we are measuring what the model itself believes, the thing that decides answers whenever live search is not invoked or comes back ambiguous. Each model’s five answers were combined into one ranking, and every result below held on the combined ranking, never on the strength of a single lucky answer.

Finding 1: Category leaders are not tenured

The number one slot is the whole game in an AI answer. It survives summarization, it gets read aloud, it anchors the shortlist. And it moved. Each row below shows a category’s top recommendation before and after the release.

CategoryOpus 4.8 saidFable 5 says
Conversational marketingDriftIntercom
ABM platformsDemandbase6sense
AI SDRArtisan11x

Source: SandsDX Wave 1, July 2026. Combined ranking across 5 answers per model per category.

The Drift result deserves a sentence, because I watched that category get created from the inside. Drift invented conversational marketing, named it, and held the AI crown for it through every model release since. This one demoted Drift to second and, in some runs, annotated it “Drift (Salesloft),” correctly processing the acquisition. The category creator lost the category’s top recommendation, in the same release that recorded its change of ownership.

Four more categories flipped at the company level, where a vendor’s product-line variants are combined: marketing automation (Adobe to HubSpot), cloud ERP (SAP to Oracle), identity security (Okta to Microsoft), and customer success. These are shifts in recommendation share across a company’s products rather than a clean crown change, so we report them separately. Either way you count, a third of the categories we measured have a different answer to “who should I look at first” than they did one model ago.

One honesty note, added after Wave 2: these single-slot flips are the most protocol-sensitive finding in the study. Asked for a top ten instead of a top five, some of them reverse. The membership churn, the naming gap, and the maturity gradient all replicate; treat the crown changes as true under this exact question, and see Wave 2 for the stricter test.

Finding 2: The models know your rebrand. They just don’t use it.

This is the finding we didn’t expect, and the one worth the most to anyone who has renamed anything.

We probed both models directly on eight publicly renamed products: what is Pardot called today, what happened to HelloSign, and so on. Asked directly, both models are nearly perfect: the old model scored 7 of 8, the new one 8 of 8. They know Pardot is Marketing Cloud Account Engagement. They know Azure AD is Microsoft Entra ID.

Then we looked at what the same models do in a buying context. The old model, which correctly answers the Pardot trivia question, recommends marketing automation products under retired names when a buyer asks for a shortlist: it offered “Salesforce Marketing Cloud” where a buyer would evaluate Account Engagement, and it listed Marketo twice in one consensus under two different names because it could not settle on what Adobe calls the product across runs.

Recall and usage are different behaviors. The models have learned your rebrand as a fact. Answering a buyer, they reach for the older habit. Your buyers never ask the trivia question. They ask the buying question, and the buying question surfaces whichever name the model saw most during training, correct or otherwise. If you renamed your company or product in the last two years, a direct-question audit will pass and mislead you. The test that matters is whether the new name shows up unprompted when a buyer asks who to consider.

The one direct-recall failure is instructive too: asked what Wingman is called now, the old model invented three different wrong answers across three runs, including products from entirely different companies. The new model answered Clari Copilot, correctly, all three times. Between those two models, a company’s answer to “what happened to that product” went from three hallucinations to a clean fact.

Finding 3: The maturity gradient

Read each row as: of that group’s categories, how many top fives changed, what share of the top-five places changed hands, and how closely the two models’ lists agree.

GroupShortlists that changedPlaces that movedAgreement between the models
The 10 biggest categories6 of 1010%83%
The 11 GTM-stack categories7 of 1116%76%

Source: SandsDX Wave 1, July 2026. 2 models, 5 answers per category each.

The biggest categories in B2B SaaS barely moved. CRM returned the identical five vendors in the identical order from both models, all ten runs. Video conferencing, identical. E-signature, identical. These shortlists are frozen: years of consistent corpus have locked them in, and a model release does not thaw them.

The GTM stack is the opposite. Sales engagement turned over two of its five places. Revenue intelligence turned over two of five. AI SDR reshuffled its entire order. Conversation intelligence was so scrambled by acquisition renaming that the old model listed the same product twice under two names.

The gradient cuts both ways. If you sell in a young category, your AI shortlist position is up for grabs at every release: the sources the models weigh are knowable, the incumbents are not entrenched, and the window between releases is when the work compounds. If you sell in a frozen category and you are not on the list, a model release will not save you; the only path in is changing what the corpus says between now and the next training cutoff, and that work takes exactly the time you think it does.

Finding 4: Acquisitions reach the models months after the press release

Catalyst merged with Totango well before either model shipped. On the old model, Catalyst still holds a customer success shortlist slot. On the new one, it is gone. Groove, acquired by Clari, appears on the new model annotated “Groove (Clari).” Drift picked up its “(Salesloft)” tag. The new model systematically carries ownership annotations the old one lacks.

For anyone integrating an acquisition, this is the timeline that matters: the market learns about your deal from a press release, but the machines advising your buyers learn about it from a training run, months later, all at once. Until then, buyers are being recommended a company that, in the form they will encounter it, no longer exists.

What this means if you run revenue

Three moves, in order of effort.

Baseline your category this week. Run your buying committee’s real questions through the assistants your buyers use. Record where you appear, at what rank, under which name, and who holds the places you don’t. Our free Am I on the List tool runs one buyer query for your category. Twenty minutes, and the next model release becomes a measured event instead of an invisible one.

If you repositioned, test the buying question. Asking the machines what your old name became proves little; they pass that test and mislead you. Ask the buying questions and see which name surfaces unprompted. The gap between those two answers is the size of your problem.

If you operate a portfolio, make this an index. For a PE operating partner, one portco’s shortlist position is a data point; the same measurement across ten portcos at every model release is an early-warning system for the exact revenue risk that never shows up in a board deck. The turnover we measured says the reading changes materially several times a year.

The shortlist is one output of the message system, and this month it moved for an eighth of the market we measured, without a single company in the dataset doing anything. The announcement ships in a day. The model catches up when it catches up. The companies that win the gap are the ones watching it.


Methodology: 21 B2B SaaS categories in two panels (10 majors, 11 GTM-stack), 2 models (Claude Fable 5 and Claude Opus 4.8), 5 runs per model per category, plus 8 rebrand probes at 3 runs per model, 258 total API runs in July 2026. Identical prompts, structured JSON output, no web grounding, zero refusals. Rankings scored by consensus across runs; overlap measured as Jaccard similarity of top-5 sets at two levels: company (product and naming variants collapsed to the owning company) and name (exactly what the model output). Number-one flips reported as outright only where the leading name changed at both levels. The buyer question, verbatim: “You are helping a B2B software buyer build a shortlist. What are the best [category] vendors right now? Return the top 5 vendors as a ranked list, best first. Vendor names only.” The rebrand probe, verbatim: “A colleague mentioned the software product ‘[old name]’. What is this product called today, and which company owns it? Give the current official product name.” Raw data available on request.

Frequently Asked Questions

Do AI model releases change which vendors get recommended?

Yes, measurably. In July 2026 SandsDX ran identical buyer questions through Claude Fable 5 and the previous flagship model across 21 B2B SaaS categories, five runs each, 258 total runs. 13 of 21 shortlists changed and 13% of all top-five places changed hands, findings that replicated when Wave 2 reran the study. Three category leaders also lost the number one recommendation under this exact question: Drift to Intercom in conversational marketing, Demandbase to 6sense in ABM, and Artisan to 11x in AI SDR.

Why do AI shortlists change when a new model ships?

Each model release is trained on a fresher snapshot of the world. Vendors that grew, rebranded, merged, or faded between training cutoffs get re-evaluated all at once. The shift lands on release day, not gradually, and nothing about your own marketing has to change for your position to move.

Do AI models know about company rebrands?

Asked directly, mostly yes: the older model mapped 7 of 8 renamed products to their current names, the newer model 8 of 8. But the same older model that correctly answers what Pardot is called today still recommends marketing automation products under retired names in a buying context. Knowing a rebrand and using it are different behaviors, and buyers only ever see the second one.

Are mature software categories safer from AI shortlist changes than new ones?

Much safer. In this study the ten largest B2B SaaS categories showed 10% shortlist turnover across the model release, while the eleven GTM-stack categories showed 16%. CRM, video conferencing, and e-signature shortlists were identical across both models. Young categories reshuffle at every release and consolidated ones barely move. Whether that is good news depends on whether you are already on the list.

Did the Wave 1 findings replicate?

The structural findings did. Wave 2 reran this comparison inside a larger cross-provider study and reproduced the membership churn: 13 of 21 categories changed and turnover measured 12%, against 13 of 21 and 13% published here. The individual number-one flips proved sensitive to how many vendors the question asks for, so crown changes are now reported only when they hold across a panel of phrasings.

How do I find out if my company appears in AI shortlists for my category?

Run your category's real buyer questions through the AI assistants your buyers use and record where you appear, under what name, and who holds the places you don't. The free Am I on the List tool runs one buyer query for your category. A Gap Map audits the full message system alongside the targeting, the motion, and the measurement.

Put this to work on your GTM

Thirty minutes. You bring the number and the story change; I'll tell you which of the four systems I'd look at first.

Schedule a Conversation

Not ready to talk? Run a free GTM diagnostic →