Listen to this article

The retail innovation blind spot 

Editor’s note: Marc Foreman is co-founder and co-CEO of Buzz 3D in the U.K. and has spent three decades developing interactive 3D technology. His work focuses on retail category management, shopper research and reducing the operational barriers to testing merchandising concepts. Find Foreman on LinkedIn. 

“Give me a faster horse!” Henry Ford could have said this. Instead, he discovered a new opportunity. 

Consumers had not asked for a car. They had not even conceived a car was what they needed. Everyone rode horses, so getting down to the shops and back without melting the ice cream would surely require a bigger, better, stronger, faster animal to get there.

And since strapping a saddle to a cheetah is notoriously difficult, a horse is what they wanted.

But not what they needed.

What, I hear you ask, does this have to do with the state of modern retail?

For years, retail innovation has been framed as an optimization challenge:

  • How do we generate better assortments?
  • How do we optimize shelf space more effectively?
  • How do we predict which merchandising concepts will perform best?

These are all important questions – retail is a highly competitive business, and getting something wrong in the rollout to stores can make all the difference between outstanding success and saying, “Let’s never speak of this again.” 

A complex retail ecosystem has built up to address these issues. Shelves are built as planograms. Sales data is hoarded and analyzed for high-value patterns. And testing! Lots and lots of testing. As much deep, tightly focused, intensive and well-researched, industry-established (expensive, time-consuming) testing as it’s possible to do with the very best enterprise-scale tools available.

Even AI has bought in on the game and is trained by a wealth of information from market leaders in shelf testing (giants such as NIQ Spaceman, Blue Yonder (formerly JDA), InContext Solutions, Decipher, Qualtrics and more). Never has a large language model been so sure of needing a cheetah.

Yes. The retail industry has a blind spot. 

It thinks that in order to get better concepts into stores, it must dig even deeper, perform more calculations, leverage the best technologies – such as AI or cloud-accelerated predictive algorithms – to try and glean the maximum benefit possible from every rollout strategy.

And Google also has this blind spot. Before I even thought about sitting down to write this article, I asked Google Gemini why category teams test so few concepts.

It answered exactly as the industry has for years:

  • “Testing is expensive, disruptive and resource intensive.”

So, I asked: 

  • “If technology reduced the cost of preparing concepts by an order of magnitude, would the optimal strategy still be to test only a handful?”

Google's answer changed:

  • “No. The economics change completely. When the cost of evaluation approaches zero, the rational strategy becomes evaluating many more concepts before committing to expensive downstream validation. The problem shifts from conserving resources to maximizing the chance of discovering exceptional ideas.”

I prompted: 

  • “So, the industry's assumption that only a few concepts should ever be tested isn't actually fundamental. It's a consequence of the historical cost of getting concepts into evaluation.”

Google’s answer:

  • “Correct. Lower the cost of evaluation sufficiently and the optimal workflow changes. Organizations should explore far more concepts before committing resources to deeper validation.”

I prompted: 

  • “Does this mean the industry may have been optimizing the wrong bottleneck all along?”

Google’s answer:

  • “Yes. If operational throughput is the real constraint, then reducing the cost of concept preparation and early evaluation changes the economics of retail innovation. The bottleneck moves downstream to physical execution, while upstream innovation becomes a high-throughput process.”

I prompted: 

  • “So, the question isn't, ‘How do we deeper-dive the few concepts we have?’ It's, ‘How do we make it economically practical to evaluate many more concepts before choosing which ones deserve deeper inspection?’”

Google’s answer:

  • “Exactly. Under that model, the objective shifts from selecting a small number of safe ideas to exploring a much broader pool of concepts, allowing evidence rather than intuition to identify the strongest candidates for downstream validation.”

Evidence rather than intuition. Determine the best candidates to put forward into the expensive and time-consuming depths of enterprise scale testing from a wider range of tested candidates. Test before testing. Pre-testing. Or, more accurately, pre-screening.

The bottleneck is therefore not creativity from our category teams but operational throughput and the sheer lack of time and resource to test every great idea that comes along.

And this, in my opinion folks, is where innovation goes to die. Every one of those discarded ideas might have been the next breakthrough.

Or not.

The point is that nobody knows, because most never failed. They simply never had the opportunity to succeed.

The hidden cost of evaluating ideas

Every merchandising concept has to travel through a surprisingly complex journey before anyone knows whether it deserves to reach a store.

A planogram has to be created or modified.

The concept must be communicated.

It needs to be represented realistically enough for meaningful evaluation.

It often has to be incorporated into a shopper study, deployed to respondents and analyzed across multiple stakeholders.

None of these activities is unreasonable in isolation. Collectively, however, they create enough operational friction that organizations naturally begin rationing the number of concepts they evaluate, unconsciously filtering every job sent through the system.

Faced with cost and time constraints, teams frequently find themselves selecting a few "best bets," while many other potentially valuable concepts never receive any meaningful evaluation at all.

When the economics change, the strategy should change

Economics, that exciting branch of bedtime reading, offers a useful way to think about this: Every concept carries two costs. The first is the cost of preparation and evaluation. The second is the cost of implementation.

Historically, the first cost has been high enough that organizations have concentrated their resources on only their highest-confidence concepts.

But what happens if the cost of preparing concepts falls dramatically? If, to sustain the analogy, some clever person like Henry Ford comes along and tells us they have something better than just that faster horse?

The economics change.

The objective shifts from selecting the safest idea to discovering the best idea.

A different workflow

Rather than treating deep validation as the starting point, organizations can introduce an intermediate stage focused on rapid exploration.

That workflow might be described as breadth before depth, and it might look something like this:

1. Generate broadly.

Encourage category, commercial and shopper teams to contribute a wide range of merchandising, pricing and assortment concepts without prematurely filtering them.

2. Evaluate rapidly.

Use low-friction virtual evaluation deliberately designed to avoid handover friction (likely through automation) and to focus exclusively on gathering empirical data. Its role is shallow: which concepts, to put it bluntly, made the most virtual cash?

Which were much-of-a-muchness and not worth writing home about. And of course, which were complete and total duds and could be chucked right out of consideration (and “never spoken of again”).

3. Validate deeply.

Now with your much smaller handful of empirically more successful concepts, go forth and enterprise away. At least then you’re giving creativity the breathing room it needs, and if an idea is any good, it’ll show up in the virtual cash register.

Expensive downstream resources are now focused where they deliver the greatest value.

4. Execute confidently.

Implementation becomes the final stage of a much broader evidence-based selection process rather than the outcome of a relatively narrow initial filter.

Breadth before depth

A mile wide and an inch deep isn’t great for swimming, but it sure does help the saplings grow.

Instead of asking: "Which three ideas should we test?"

Organizations can begin asking: "Which 30 ideas deserve the opportunity to compete?"

The difference is subtle, but strategically profound.

Many of tomorrow's strongest merchandising concepts may never have looked like obvious winners when they were first proposed.

They simply needed the opportunity to be evaluated.

AI will save, doom and/or meh us all

AI makes it easier than ever to generate content but the conversation I had with Google indicates that innovation does not live in AI. It lives in the hearts and minds of your team and as of right now, innovators have nowhere to go.

The organizations that gain competitive advantage over the coming years won’t be those that do the deepest validation. It will be those that implement a fast, shallow, frictionless testing ecosystem before their enterprise stack … right there. See it? No? Well … understandable, I guess. It’s bang in the middle of that blind spot.