• Post a Project

What AI Actually Tells B2B Buyers (Three Engines)

Updated September 30, 2026

Anna Peck

by Anna Peck, Content Marketing Manager at Clutch

A B2B buyer asks ChatGPT, Gemini, and Claude the same question and gets three different conversations and three almost entirely different shortlists. We analyzed more than 69,000 of those answers to find out why.

For a growing share of B2B buyers, the search for a service provider no longer starts with a search bar. It starts with a question typed into the AI assistant of their choice: Who’s the best agency for this? In my market? For my kind of company?

Traditional search hasn’t gone anywhere. Buyers still click through and read reviews, but by the time they do, AI has already given them a first impression and a shortlist.  The AI answer is now often the starting point of the search, and whoever has that answer gets a head start on everyone who doesn’t.

Nearly one in four content marketers said LLMs are now the primary audience for the majority of their content, according to the 2026 State of Content Report from Clutch and Conductor. Proprietary research and original reports ranked as the top priority for earning visibility in AI-generated answers. A meaningful share of marketers are no longer writing for people who might ask an AI. They're writing for the AI itself.

Here’s the catch: there is no single “AI” to write for. Popular systems like ChatGPT, Gemini, and Claude behave like three different people, each with their own habits, own sources, and own idea of what a good answer looks like.

Every trust signal across the internet (verified reviews, case studies, third-party citations, etc.) is another piece of guidance pointing AI to the right recommendation. Our data shows that each AI assistant has a distinctive personality and a set of levers that buyers pull without realizing it. This article will dive into those findings and how parties can navigate a shifting landscape.

What We Asked

Every prompt in this dataset asks the question a buyer asks when they’re ready to hire (in various ways): Who are the best providers for this service?

The prompts, regardless of wording, share three traits:

  • One buyer. Every prompt is written from the perspective of the same persona: a strategic, growth-focused decision-maker.
  • One stage. Every prompt carries purchase intent. These are shortlist questions, not early research questions like "what should I look for in an SEO agency?", which may play out differently.
  • Many variations. The panel mixes bare "who are the best" questions with versions that ask for proof ("based on reviews," "proven track record") or name an industry. That mix is what lets us measure how phrasing changes the answer.

The prompts used for this report span over 20 B2B service categories, and run in both US and UK markets, adding over 69,000 AI responses from the pull in the beginning of September 2026.

The Three AI Engines at a Glance

Metric ChatGPT Gemini Claude
Personality The salesperson The analyst The editor
Average answer length ~484 words ~514 words ~416 words*
Length variation (coefficient of variation) 28.6% 9.5% 21.5%
Distinct providers mentioned per answer ~14 ~15 ~13
Providers elevated (bolded shortlist / formatted list) —** ~17 ~8–9
Offers to continue the conversation ~55% ~24% ~8%
Hedges ("it depends," "varies") ~39% ~17% ~9%
Uses comparison tables 14% 65% 1%
Links per answer ~7 ~0 since June ~11

ChatGPT - The Salesperson

ChatGPT is the only engine that treats an answer as the opening of a conversation rather than the end of one.

ChatGPT - The Salesperson

In about 55% of its answers, it closes by offering to keep going: compare the options head-to-head, narrow down that list by budget, etc., and then draft next steps. Gemini takes this approach in about 24% of its answers, and Claude does in about 8%.

If Gemini and Claude hand over a file, ChatGPT hands over a business card and asks when you’re free to continue the conversation.

That sales instinct shows up in tone, too. ChatGPT is all about hedging (using “it depends” and “pricing varies”) in nearly 39% of its answers, which is far higher than the others.

In terms of structure, it is the least structured of the three, with headers in about 16% of answers and tables in 14%. A typical answer (which reads more like a conversation than a report) runs about 484 words while mentioning 13-14 distinct providers.

Review and directory platforms are ChatGPT's #1 source type, and it's essentially the only engine that pulls from Wikipedia (11,935 Wikipedia citations in this dataset, versus 73 for Claude) and Reddit, which appears in roughly 7% of its answers.

It's also the most changeable engine month to month, and model updates are the likeliest reason. OpenAI ships model and product changes at a pace no search engine has ever done. 

The share of ChatGPT’s citations going to review platforms more than doubled over the summer of 2026, from 5.5% in May to 13.7% in early September, peaking at 25.9% in July along the way.

What providers can learn: ChatGPT rewards being present where a crowd would check: review platforms, community threads, and established reference sources. And because it keeps the conversation going, your first mention is the start of a buyer's evaluation, not the end. The follow-ups it offers, like "compare these on price" or "which have the best reviews?", are where having a robust presence on Clutch pays off for service providers.

Gemini - The Analyst

If ChatGPT is a conversationalist, Gemini acts as an analyst, using a templated response across most prompts.

Gemini - The Analyst

Gemini’s answers average about 514 words with a coefficient of variation of just 9.5%, so response length is very consistent from one conversation to the next. ChatGPT’s varies three times as much.

The format is just as consistent. Gemini uses headers in essentially every answer and comparison tables in about 65% of them, versus about 1% for Claude. It lays out the longest provider lists of any engine, roughly 16 to 17 names, and it almost never closes on a question (0.3% of answers).

It reads like a one-shot analyst brief: here's the landscape, here's the table, here's the list. Today, Gemini hands buyers a shortlist with essentially no visible citations, so the only way to track who it recommends is to read the names in the text.

What providers can learn: You can't count on a citation to carry you on Gemini. It builds its lists from what it already knows, so what gets you in is consistent, structured, corroborated information about what you do: clean, comparable data that slots neatly into a comparison table.

Claude - The Editor

Claude is the most selective of the three engines, not in how many providers it mentions, but in how many it elevates.

Claude - The Editor

It names about 13 distinct providers per answer, right in line with the others, but bolds a tight shortlist of just 8 or 9.

It’s also the most decisive – Claude hedges in only about 9% of its answers, the lowest of the three, and writes in structured prose: headers in nearly every answer, almost never a table.

It also shows its work more than anyone else: Claude cites about 11 links per answer – the most of any engine. The catch is what it cites: overwhelmingly, providers’ own websites and agency-published “top agencies” listicles.

Claude reads a lot, but it tends to take vendors at their word.

Claude has also been the steadiest of the three. A jump in answer length in June, to around 508 words, faded by August, and the length settled back to about 353. That was a temporary effect of a model update, not a new baseline.

What providers can learn: On Claude, the goal is to be bolded, not just named. Because it leans so heavily on providers' own sites, the substance on your website matters here more than anywhere: specific services, specific industries, specific results.

The Consensus Myth: Why "AI Visibility" Isn't One Number

In our analysis, consensus means a provider named by all three engines in response to the same prompt whereas self-consistency means the share of providers an engine names again when the same prompt is run again.

By that measure, agreement among the three personalities is rare. Across them, a single prompt surfaces about 70 distinct provider names, and only one or two of them appear in all three answers. Roughly 26% of prompts produce no consensus at all, meaning no single provider is named by all three engines. That's down from about 35% earlier in the year, so consensus is rising, but from a base so low that agreement is still the exception.

The engines don't even agree with themselves: run the same prompt again, and ChatGPT and Gemini repeat only about 30% of their providers. Claude, the steadiest, repeats about 63%.

That's a real break from how discovery used to work. Until a few years ago, search results were largely stable: the same query returned roughly the same 10 blue links for everyone, and a provider could earn a ranking and hold it. AI answers shift by engine, by phrasing, and from one run to the next.

A different game doesn’t mean starting from square one for companies, though. Much of what works in SEO still works for AI visibility. Some engines draw on search indexes when they retrieve sources, and the fundamentals (clear, specific content, authority, and citations from trusted third parties) carry over.

What changes is the goal – you’re no longer defending a position; you’re maximizing your company’s probability and potential. The aspiration is to become one of the one or two providers ever engine names.

It also changes how you measure. A provider can be a fixture in Claude's bolded shortlist and invisible in Gemini's tables, and an aggregate "AI visibility" score averages those realities into a number that describes none of them. Track your performance within each engine individually. Then prioritize: if you know which engine your buyers actually use, start there.

The Levers Buyers Control

Across three engines that agree on almost nothing, one thing holds constant: the buyer’s own prompt shapes the answer. It is a set of levers, not just one simple switch.

Evidence Framing Rotates the Proof

When a prompt asks for proof, the answer gets built from proof. Prompts worded around evidence and trust ("highest client satisfaction," "proven track record," "based on reviews") pull a review platform into roughly two-thirds (about 65%) of answers, on all three engines.

And the whole answer rotates, not just the citations. Answers that include numeric ratings jump from 2% of Gemini answers to 41%. Review counts and "verified reviews" language jump from 3–11% of answers to roughly a third on every engine. Awards and case studies rise. 
The engines don't do more research when asked for proof. They swap in a different kind of evidence.

Asking for proof also makes the engines more decisive, not more cautious. ChatGPT hedges in 51% of its bare "best" answers but only 28% of trust-worded ones, and it trims its shortlist from about 18 names to about 13.

Buyers who ask for evidence get a shorter, more committed answer.

Industry Specificity Helps

If evidence framing changes what kind of proof shows up, industry changes who shows up.

Adding an industry to the prompt changes roughly 70% of the companies AI recommends, swapping a general list for specialists in that field.

What AI Actually Tells B2B Buyers (Three Engines)

Compare that to the control – rewording the same question without changing its substance changes just 5% of the list.

The engines respond to substance, not phrasing. What a buyer is actually asking for moves the answer.

When no review platform shows up, each engine falls back on something different. ChatGPT leans on Wikipedia (45% of its no-platform answers) and media or analyst sources (26%). Claude takes vendors at their word: 52% of its no-platform answers cite only providers' own websites, and 42% cite agency-published listicles. Gemini falls back on nothing, and 72% of its no-platform answers cite zero sources.

In other words, when buyers don't ask for proof, the engines don't go get it. Unverified self-description wins by default, and verified evidence only enters the room when the buyer's wording invites it in.

What providers can learn: You can't control how buyers phrase the question, so you have to be present in both universes. The verified layer (review platforms, ratings, third-party citations) carries the evidence-seeking buyer. Your own site and the wider web carry everyone else. And specialize legibly: industry specificity is the single biggest factor in whether you're even in the candidate pool, so make the industries you serve unmistakable everywhere you show up.

Win Each Engine While the Ground is Still Moving

Every personality in this report is a moving target, whether they are aiming to close the deal or trim fluff content.

AI labs ship model and product updates quickly, and each one can change what an engine reads and who it recommends.

The engines aren’t just changing; they’re changing in different directions, as showcased in this data. But the landscape hasn’t settled yet, and that’s the argument for service providers to act now rather than wait for things to stabilize.

For providers, that means you can’t win “AI” as a monolith. You win each engine individually: by being present where it reads, specific about what you do and who you do it for, and verifiable through third-party sources it trusts.

Nearly one in four content marketers already say LLMs are the primary audience for most of their content.

The providers who make themselves legible to that audience now, while the engines are still deciding who to repeat, are the ones who'll keep showing up in the shortlist.

About the Author

Avatar
Anna Peck Content Marketing Manager at Clutch
Anna Peck is a content marketing manager at Clutch, where she crafts content on digital marketing, SEO, and public relations. Alongside editing and producing engaging B2B content, she plays a key role in Clutch's awards program and content initiatives. Originally joining Clutch on the reviews team, she now focuses on developing SEO-driven content strategies that deliver valuable insights to B2B buyers searching for the best service providers.
See full profile

Related Articles

Getting Ready for 2027: 5 Lead Generation Platforms to Find Better Prospects
Generative AI vs. Agentic AI: What's the Difference?
How to Build An AI Agent (No Coding Required)