8 min read

Analyst rankings can build your shortlist. They can’t make your decision. How to read analyst evaluations of B2B commerce platforms

Analyst rankings can build your shortlist. They can’t make your decision. How to read analyst evaluations of B2B commerce platforms
Analyst rankings can build your shortlist. They can’t make your decision. How to read analyst evaluations of B2B commerce platforms
13:04

An analyst rating shows how vendors compare in a market. It doesn’t show which one fits your business. This article walks through six questions to ask before a rating shapes your shortlist, and what to do once it has.

 

A disclosure and why this is worth knowing

Part of what I do here at Intershop involves working with analysts who are reviewing commerce platforms. Intershop appears in several of them. That is worth knowing before anyone reads further.

I will also admit that, every year, we look forward to talking to the analyst community and wait with much anticipation for the results following these intense evaluation cycles, which involve entire teams. Analyst evaluations are one of the few moments when a lot of quiet work across product, engineering, and the wider GTM teams gets put in front of an analyst with no particular reason to be generous about it. I don’t own these relationships here (high five to Tobias Giese and team), but I sit close to the process and have seen for years what goes into a submission. I also get to share the results when they land well (which is a good day for everyone!).

In this article, I am not going to discuss a specific report, and I am not going to quote scores. Our placements over the years can be found on our analyst page. What follows is a discussion about what these evaluations measure, what they structurally cannot, and how to go from a published rating to a purchase decision you can defend five years from now.

Analyst evaluations are valuable. They compress a crowded market into something comparable. They pressure-test vendor claims. They give a buying committee a shared starting point instead of a dozen competing opinions. The problem is that readers often expect these reports to provide a definitive answer to questions they were never designed to answer.

 

What analyst reports do and do not answer

A market evaluation is built to answer: how are these vendors performing relative to one another in this market right now, and how well positioned do they look for what comes next? 

What it’s not built to answer is which solution is right for your business.

These questions are not the same. The highest rated airline in the world is still the wrong airline if it does not fly where you need to go. Scale, financial health, and momentum are real information, and they matter for multi-year commitments. They say a vendor will still be there, still innovating, still shipping, still supporting. But what they do not say is whether the platform fits how a given business actually sells. For a manufacturer supplying spare parts to maintenance teams who search by machine compatibility rather than product name, or a distributor pricing every order against a contract negotiated 18 months ago, market momentum is not the variable that should point your team to a given solution.

Different analyst reports also draw the line between vendor and product differently. Some weight market position heavily. Others score product capability directly, alongside strategy and vision. It’s important to work out which one you are holding before reading a single graphic, because it determines what those results can honestly tell you.

Either way, the output is the same. A published analyst evaluation is a map; a compass. It tells you which direction the market is facing and produces a valuable vendor list. It does not show you the turn-by-turn to your destination.

 

Six questions for building a vendor shortlist

None of the following is exotic. Most experienced buyers already do some of it by instinct. Writing it down helps mainly because these reports tend to arrive in the middle of an already-crowded process, get skimmed once, and then quietly shape the conversation for months afterwards.

The first two questions ask what’s actually been measured, since the boundary drawn around a market and the choice between scoring vendors or products both decide what the numbers mean. The next two ask how it was weighted and what the evaluator could observe. The last two ask whether the result is usable at all, given how the field is distributed and what each vendor was allowed to bring.

Together, they give a sense of how much of a rating transfers to a particular business, and how much of it belongs to somebody else's.

 

1. What market is being evaluated and is it yours?

Every evaluation starts by drawing a boundary around a market, and that boundary shapes everything downstream.

A single report may combine vendors built for consumer retail alongside vendors built for contract-driven, account-based selling. That breadth is useful when showing an entire landscape at-a-glance, but it also means the criteria must be stretched to cover every vendor. As a basis for a B2B or a B2C platform decision, this is harder to defend because the moment one set of criteria has to serve both models, specific criteria get less room and the weighting starts to describe a compromise rather than business needs.

If your customers log in to see their custom negotiated pricing across twelve sites, and much of the vendor field is optimizing for anonymous checkout, that stretch works against you. Understand what market has been evaluated before you read the scores.

 

2. Is this scoring the vendor, the product, or both?

Vendor strength and product fit are separate things. A vendor can be commercially formidable, with a strong balance sheet and a large installed base, but still handle deep ERP integration or customer-specific catalogs poorly. The reverse happens too. Two analyst reports may weigh the two differently, which leaves the interpretation to whoever is reading it.

 

3. What do the weightings reward?

The weightings are usually just as revealing as the scores, because they tell you what the evaluation cares about, which is not necessarily what you care about.

Consumer-oriented weighting optimizes for acquisition and conversion: winning new buyers, increasing basket size, and removing checkout friction. Most B2B commerce programs are funded for different reasons: Taking cost out of order handling. Getting sales teams out of manual order entry. Protecting margin on contract pricing as volume grows. Those are efficiency and retention problems. An evaluation built around acquisition will not weight them properly.

The feature-level version of the same check is quicker. If, for example, storefront styling and campaign tooling carry the heaviest weights, while deep system integration and contract pricing logic sit near the bottom, that ranking speaks to a very specific buyer. What might be a five-star rating for one is a one-star rating for another.

 

4. How much was actually visible?

Evaluations are built from demos, briefings, and vendor-selected reference calls, usually inside a fixed and fairly short time window. That is a great way to survey a market. But it’s not the best way to judge how a platform holds up under real operational complexity.

The distinction matters more in B2B than in most categories, because the value here is realized over months and years of operation rather than at launch. What determines whether a platform succeeds in this market is rarely visible in the first hour: The 20th integration, not the first. The fifth country rollout. The unicorn customer whose approval chain does not fit the standard model. The catalog that doubles in size after an acquisition. A product demo shows the happy path and often the ‘art of the possible’, but B2B projects are as real as it gets.

Some of that gap is structural and not anyone's ‘fault’. Punchout into a customer's procurement system, search across a spare-parts catalog with hundreds of thousands of items unique to the installed base. In most cases, this sits behind a login, which makes it harder to assess than anything on a public storefront.

In B2B, the decisive capability is usually the one an outsider cannot easily see. In one evaluation, any feature/function that rests behind a firewall was completely discounted from the assessment. This level of disconnect from the market is, at the least, troubling.

 

5. How much does the scoring separate the field?

In a mature category, most of the field clusters near the top; and that can be an accurate reflection of the market. When that happens, the report is still useful, just for a narrower job. A tightly clustered field works well as a pass/fail filter, confirming which vendors are credible enough to be in the conversation at all.

But what looks like a compliment to all of the market is also a sign that the report has done as much as it can, and that the criteria are not discriminating between real differences any longer. A rating only helps if it would look meaningfully different for a weaker solution.

The differentiation then has to come from somewhere else, and the quickest source tends to be the vendors themselves: Where do they think their platform is a poor fit? Which projects has the vendor walked away from in the past? What has to be custom-built rather than configured? The ones with straight, honest answers are also telling you something useful about what the next three years of working with that vendor might look like. A vendor with no interest in qualifying you out is a vendor with no stake in project success. The best partnerships are selective on both sides.

 

6. Are you comparing platforms, or portfolios?

Vendors are sometimes evaluated as part of a ‘portfolio’ or a suite rather than a single product, and that introduces a lot of variability. It is a fair representation of what each vendor can bring (along with the higher portfolio price tag), but it also means that score reflects everything included in their submission, not the piece you may end up licensing. It is worth knowing each vendor's SKUs in scope and comparing them to what you truly need.

 

Your work starts where the reports stop

Once a report has informed your initial list, the question changes. It stops being ‘which vendors look strongest in the market’ and becomes ‘which one best solves our specific set of problems’. That evaluation usually is on you to run. No external report can run it for you because it depends on things only you can see: what your business needs today, what it will need next, and where the friction actually sits right now. And the method should not be complicated.

  1. List the capabilities that genuinely make or break the operation. Five to eight, not fifty. The list is specific to the business, and writing it down tends to surface disagreements inside the committee; cheaper to find out now rather than in a few months from project kickoff.

  2. Assign weightings. This is the step we see most committees skip, which is a pity because it is the one that matters most. A weighting is really a statement about what the business needs today and what it needs next. Those two things are worth establishing deliberately, rather than by instinct. Thankfully, a structured maturity assessment can help frame it.

  3. Score each vendor against those criteria using RFI responses and demos built around real-world use cases. Hand over actual catalog data, pricing rules, and your thorniest integration. Standard demos typically don’t solve for that.

  4. Pressure-test the two or three assumptions that would be expensive to get wrong. And remember, not every expense is a dollar. Reputation costs are real too! Not everything needs proving. Trying to prove everything is how a selection process loses momentum. Your genuinely vital requirements are the ones where being wrong is costly and reversal is painful. Pay attention to how far a vendor is willing to dive into the requirements before there is a signature. Vendors with genuine B2B depth tend to welcome that, because a difficult scenario is where they separate from a crowded field. The ones who can turn a messy requirement into a testable outcome during the evaluation are the same ones who can do so during delivery.

 

Three powerful takeaways

A rating describes a market, not a fit. Analyst reports narrow the field from a wide range of solutions to a list of viable vendors. Reading more than one is worth the time. And it is not unusual for two credible, independent assessments to reach different conclusions about the same platform in the same quarter. It usually points to different methodology, and the reasons behind it are worth understanding before either report drives a purchase decision.

If you’re curious how independent analysts have evaluated Intershop over time, here's what leading analysts are saying about us.

The weighting matters more than the position. Two vendors a short distance apart on a chart can be very far apart when it comes to addressing specific business needs. This depends on which criteria carried the most weight and whether those are the ones that matter for your project.

The report is the start of due diligence. It narrows a wide field to a serious shortlist and sharpens the questions worth putting to the vendors on it. What it cannot tell you is which criteria decide which commerce solution is best for your business, because that depends on where your business currently stands and how it envisions its future. A company whose integrations break with every release needs a different shortlist than one whose commerce engine runs fine but underdelivers on adoption, or one where growth has outpaced what the team can handle manually. Same market, same reports, three different decisions. Working out which situation applies to your unique needs is what should set the weightings, and the weightings are what turns a ranking into a decision. We wrote a short playbook on those situations, with checklists to learn which one fits and what tends to work from there. It is available here.

 

Read the reports. Read more than one. But know your own position first.