Listen to this article

AI in market research: Are we measuring its impact the right way?

Editor’s note: Andy Sweet is a technology entrepreneur, and he currently serves as vice president of enterprise AI solutions at AnswerRocket. Previously, Sweet was co-founder and CEO of Cognitive Spark, an AI and management consulting firm acquired by AnswerRocket. Earlier in his career, Sweet co-founded and led Visual Software Integration, which was acquired by Compete. He spent over a decade in executive leadership at IBM and Daugherty Business Solutions. Find Sweet on LinkedIn.  

Insights professionals measure things for a living. You worry about sample, validity and whether a number means what people assume it means. It’s fair to turn that same scrutiny on your AI tools and ask how you’re measuring them. In most teams I talk to, the honest answer is hours saved, reports produced and how many people logged in. Those are activity metrics that tell you the tool gets used. They tell you almost nothing about whether the research got better.

The wrong yardstick

You would never field a tracker and report “we completed 2,000 interviews” as the finding. Completing the interviews is the activity. The finding is what the 2,000 people told you, and what the business should do about it. Grading AI on outputs produced makes that same mistake one level up. Speed and volume are easy to count, so they get counted. The questions that actually matter get skipped.

MIT’s NANDA initiative put a number on where this leads. Its 2025 study, “The GenAI Divide,” found that 95% of enterprise generative AI pilots delivered no measurable impact on the bottom line, despite tens of billions of dollars in spend. The headline reads like a verdict on the technology. It isn’t. The models work. The problem is that “we deployed a tool” and “we got value from a tool” keep getting treated as the same sentence. They are two very different claims, and most measurement programs only check for the first one.

AI doesn’t fail loudly

For researchers there’s a second problem sitting underneath the first, and it’s specific to your work. When AI is wrong, it rarely breaks. It produces a clean, confident, well-written output that looks right. An open-end summary that reads beautifully. A segmentation that holds together. A theme that sounds exactly like something a respondent would say. It passes the eyeball test, lands in the deck and shapes a recommendation. Nobody catches it because nothing looks broken.

That’s the failure mode worth fearing. A system that crashes tells you it failed. A system that hands you a confident wrong answer tells you nothing at all. You already know this risk from bad survey data and leading questions. AI raises the stakes. It generates plausible-looking output faster and at a far greater volume than any human could.

Synthetic respondents are the sharpest version of this. Simulated answers can fill a gap in a hard-to-reach segment, and they look like data. They carry decimal points and clean distributions and all the texture of the real thing. Whether they reflect how actual people think is a separate question, and you only answer it by validating against real respondents. The output looks like an answer either way. If your measurement can’t separate output that’s right from output that merely looks right, you don’t have a measurement program. You have a hope.

Start with the job, then measure against it

The fix starts before the tool. Define the job first. What decision is this research meant to inform? Measure the AI against that decision, not against the clock. Once the job is clear, three things are worth measuring, in roughly this order.

  1. Did the research get better? Did the insight change what the business did, and was that the right call? This is the hardest item on the list to track and the only one that maps to real impact. Start here anyway.
  2. Is the output accurate? Treat a new AI method the way you’d treat any new method. Validate it before you trust it. If the AI codes open-ends, hand code a few hundred yourself and check agreement. You’d do this for a new panel or a new coding vendor without thinking twice. Do it for the AI.
  3. How often is it wrong in a way that would have shipped? Track the catch rate. When someone reviews the AI’s work, how often do they find an error big enough to change the answer? If nobody reviews, that number is unknown and unknown is not the same as zero.

Efficiency still counts. Time saved and costs cut are real, and you should track them. Treat them as the floor. The ceiling is whether the work got better. Producing a wrong answer faster is negative value, and it happens to be the easiest kind of value to celebrate by accident.

What this looks like in practice

A few habits separate teams getting real value from teams reporting motion:

  • Validate before you scale. Prove the AI on a benchmark you trust before it touches a live study. A small hold-out you’ve coded by hand beats any vendor accuracy claim.
  • Keep a human in the loop where a confident error is expensive. Open-ends feeding a strategic decision need a reviewer. A first-pass summary for internal triage may not. Match the oversight to the stakes.
  • Measure again, then keep measuring. Models get updated, your categories shift, the market moves. A validation that passed in the spring may not hold in the fall. Build a recurring check rather than a one-time sign-off.
  • Start narrow. Give the AI the smallest, best-defined job first, prove it there, then expand. A tool that does one thing reliably is worth more than one that does 10 things you can’t vouch for.

Hold AI to your own standard

None of this asks for new methodology. It asks you to apply the methodology you already have to a tool you’ve been grading on a curve. You hold your studies to a standard. Hold your AI to the same one.

Before you report the hours your team saved this quarter, try to answer one question: Can you show that the research got better? If you can, you’re measuring impact. If you can’t yet, you’re measuring motion, and that gap is worth closing before the next budget review asks you to prove the difference.